# Parallel development with git worktrees Git worktrees let you check out several branches at once, each in its own directory, sharing one `.git`. This repo supports running multiple worktrees **simultaneously** by giving each worktree its own non-colliding block of ports. That applies to **dev servers**. **e2e suites are deliberately serialized**: one run at a time across the whole machine, everyone else waits in line. See [The e2e queue](#the-e2e-queue). The helper is `scripts/worktree.mjs`, exposed as `pnpm wt`. ## Port scheme Each worktree gets an offset of `index * 100`, where `index` is the worktree's position in `git worktree list`. The **main** worktree is always first, so it keeps the original defaults — nothing changes for the primary checkout. The port table itself is `common/lib/ports.mjs` — one copy, which `scripts/worktree.mjs` offsets and `pnpm archilyzer doctor` reports. Every port, its base and what uses it is in **[ENVIRONMENT.md → Ports](ENVIRONMENT.md#ports)** (generated from it). The injector also sets `PLAYWRIGHT_BASE_URL` (`http://localhost:`) for node-side fetches in specs. So worktree #1 runs editor on **3101**, test server on **3111**, export on **3110**, etc. The hundreds digit is the worktree index. The allocator never overrides a variable already set in the environment, so CI and manual overrides always win. ## Commands ```sh pnpm wt ports # print this worktree's assigned ports (table + KEY=VALUE) pnpm wt list # list all worktrees with their port blocks pnpm wt add # create a sibling worktree, seeded + ports printed pnpm wt add --from --share-data pnpm wt rm [--force] # remove a worktree (--force if it has local files, # e.g. a --share-data worktree's .worktree-env) pnpm wt run -- # run a command with this worktree's ports injected ``` `pnpm dev:editor`, `pnpm dev:export`, `pnpm start:export`, and `pnpm e2e` already run through `wt run`, so they pick up the right ports automatically — no manual setup needed. ### Load the ports into your shell (fish) ```fish node scripts/worktree.mjs ports --shell fish | source echo $PORT # -> e.g. 3111 in worktree #1 ``` For bash/zsh use `--shell bash` and `eval "$(node scripts/worktree.mjs ports --shell bash)"`. ## Typical workflow ```sh # from the main checkout pnpm wt add feature-x # creates ../feature-x, prints its ports, seeds settings.json # terminal A (main): pnpm dev:editor -> http://localhost:3001 # terminal B (../feature-x): pnpm dev:editor -> http://localhost:3101 # both run at once, no EADDRINUSE # e2e is queued, not parallel — start it anywhere, it waits its turn: pnpm e2e # main: runs now cd ../feature-x && pnpm e2e # feature-x: prints "waiting …", starts when main finishes ``` ## The e2e queue Every e2e entry point takes a machine-global lock before it runs, so **exactly one suite runs at a time**. Starting a second one is not an error — it prints who holds the queue and blocks until its turn: ``` queue-lock: waiting for the e2e queue — held by yt-dlp-transcript-browser (main, pid 2401494) for 4m12s (one e2e run at a time, machine-wide; E2E_QUEUE=0 to bypass) queue-lock: still waiting (5m00s) ``` **A long wait here is normal, not a hang** — the serial suite is ~24 minutes. Queued: `pnpm e2e`, `pnpm e2e:sharded` (including its `docker build`), export's `e2e`, `e2e:hub`, `e2e:2origin`, `e2e:report`, and homepage's and umtool's `e2e`. Wrapping is at the *package* level, so `pnpm --filter editor run e2e` is covered too, and a nested invocation passes through instead of deadlocking. Not queued on purpose: a raw `pnpm --filter editor exec playwright test`, the escape hatch for debugging a single spec, and the `e2e:ui` interactive sessions. ### Root `pnpm e2e` is the editor's, and only the editor's The root script is `pnpm --filter editor run e2e`, so a spec name appended to it goes to the **editor's** playwright: `pnpm run e2e clip-bench.spec.ts` matches nothing there, holds the machine-global lock for the whole ~24-minute editor suite, and never runs the spec you meant. Every other package's suite goes through its own filter: | Suite | Command | |---|---| | editor | `pnpm e2e [spec…]` | | export | `pnpm --filter export run e2e` (also `e2e:hub`, `e2e:2origin`, `e2e:report`) | | homepage | `pnpm --filter homepage run e2e` | | umtool | `pnpm --filter umtool run e2e [spec…]` | umtool's `e2e/fixtures/make-fixture.mjs` builds its fixture *from* the song project's bulk data — `umtool/song/paths.mjs`'s `SONG_DATA`, which is `SONG_DIR` if set and `~/reports/quartering-uh-song/data` otherwise. The suite itself never runs against that directory — the fixture is an empty-state copy with the heavy audio symlinked in — and a machine without the data builds an empty fixture whose song-data specs skip. The lock lives at `/e2e-queue.lock`, which resolves to the same file from every worktree. The kernel releases it when the holding process's fd closes, so a `kill -9` or a crashed run can never strand the queue — there is no stale-lock recovery to run. ### Aborting on a busy port After taking the lock, a run checks that the ports it is about to use are actually free. If one is not, it **aborts** rather than starting: ``` queue-lock: ABORTING — this suite's ports are already in use. PORT=3011 pid 2244791 cwd /path/to/checkout/editor next-server (v16.2.3) kill 2244791 ``` This closes a silent, destructive bug. The Playwright configs set `reuseExistingServer: !CI`, so a run that found port 3011 already bound used to attach to **whoever else's** editor test server was there — and the suite's 303 `resetData()` calls would then wipe that session's fixture data with no error at all. Because we hold the queue lock at that moment, no *queued* run can own those ports: what you are looking at is a leftover from a killed run, or a server someone started by hand. The trade-off: a hand-started `pnpm dev:test` on 3011 is no longer silently reused, so the ~30s server boot is no longer skippable. `E2E_PORT_CHECK=0` restores the old behavior — but a reused server must then be started with the test-only env that `editor/playwright.config.ts`'s `E2E_SERVER_ENV` gives the servers Playwright starts (`E2E_TEST_ROUTES=1 E2E_AUDIO_CHECK_INTERVAL_MS=300 E2E_AUDIO_CHECK_SIZE_GATE=4096 E2E_AUDIO_CHECK_INTERVAL_FLOOR_MS=50 E2E_AUDIO_CHECK_RECOVER_STEP_MS=100 E2E_AUDIO_CHECK_RECOVER_AFTER=2 pnpm dev:test`). Without it `/api/test/*` answers 404 and the audio checks run at production pace, which reads as real failures. ### Escape hatches | Variable | Effect | |---|---| | `E2E_QUEUE=0` | Skip the queue and the heavy slot (the port preflight and the memory floor still run) | | `E2E_PORT_CHECK=0` | Skip the port preflight, reusing whatever servers are up (a hand-started editor test server needs the config's `E2E_SERVER_ENV`, above) | | `E2E_QUEUE_TIMEOUT=` | Give up waiting after N seconds (default: wait forever) | Verify the queue with `pnpm test:scripts`. ## The heavy slot (`pnpm heavy`) An e2e suite, a `next build` and a video render each want several GB, and two of them at once is how this machine OOMed (taking the desktop session with it). So all three go through ONE more machine-global lock, the **heavy slot**, and start only once `/proc/meminfo`'s MemAvailable is at least a floor (6000 MB): ```sh pnpm heavy -- # any heavy command, by hand pnpm heavy -- node umtool/report-to-video/build-video.mjs … # a render ``` Who takes it, and in what order: - **Every e2e entry point** (the same ones the e2e queue covers): the heavy slot FIRST, then the e2e queue. One order everywhere, so nothing can deadlock — and a heavy command that starts another (`pnpm heavy -- pnpm e2e`, a build stage run by an e2e suite's editor) passes straight through (`HEAVY_HELD`), as a nested e2e run always has. - **The publish stages' `next build`** — a site's, the hub's, the homepage's (`common/publish/build.ts`, `heavyGated`). The wait shows in the stage's log, and a Cancel still stops the build (the gate forwards SIGTERM). In a docker-runner container the slot is the container's own; the floor still reads the host's memory, which throttles a fan-out when the host runs low. - **A render**, by hand, as above. The slot is taken before the floor is waited for, so nobody slips in while the holder waits for memory. A waiter is told what it waits behind: ``` queue-lock: waiting for the heavy slot — held by feature-x (feature-x, pid 31337) for 2m10s: playwright test heavy: waiting for memory — 4210 MB available, the floor is 6000 MB ``` The lock is `/heavy-queue.lock`, released by the kernel like the e2e one. A machine whose MemTotal is under the floor is told so and runs; with no usable `flock` the gate warns and runs on the floor alone — it is a safety net, not a correctness lock. | Variable | Effect | |---|---| | `HEAVY=0` | Skip the heavy slot AND the memory floor | | `HEAVY_MIN_FREE_MB=` | Move the floor (`0` turns it off) | | `HEAVY_TIMEOUT=` | Give up waiting (slot, then floor) after N seconds; an e2e run uses `E2E_QUEUE_TIMEOUT` | ### A render and the transcription lane A render competes with the transcription lane for memory and CPU, and the lane is not a heavy slot holder. Hold the lane for the render's duration, and release it whatever the render's outcome: ```sh pnpm ops lane --json '{"lane":"transcription","held":true}' pnpm heavy -- node umtool/report-to-video/build-video.mjs … ; \ pnpm ops lane --json '{"lane":"transcription","held":false}' ``` A hold stops new dispatches, not a transcription already running; and the release resumes the lane even if someone else held it for another reason — check `/operations` first. ## Data directories By default each worktree is **fully isolated**: `common/lib/paths.ts` resolves `transcripts/`, `.next/`, `.export-index/`, etc. relative to the worktree root, and the editor's `dev:test`/`start:test` scripts already point `TRANSCRIPTS_DIR` at the worktree's own `test-transcripts/`. This is safe for parallel runs — no shared LMDB lock. ### Sharing downloaded data (opt-in) Re-downloading channels into every worktree is wasteful. `pnpm wt add --share-data` writes a `.worktree-env` file pointing `TRANSCRIPTS_DIR` at the **main** worktree's `transcripts/`, and `wt run` loads it. This lets a worktree reuse the already-downloaded corpus for read-mostly work and builds. ### Copying a corpus that has a relocated channel `channels//data` may be an absolute symlink to another drive (see AGENTS.md). `rsync -a` copies a symlink **as a symlink**, which is usually right — both checkouts then read the same media through the same absolute path, and nothing is duplicated. Use `rsync -a --copy-links` only when the copy has to carry the media itself (a shard bound for a machine that will not have that drive). Note what `--copy-links` gives you: a real directory where the source had a link, so the copy's `config.dataDir` still names a target it is no longer using — `inspectChannelMedia` reports that as `inconsistent`, and clearing `dataDir` in the copy's `config.json` is what makes it in-place again. > ⚠️ **Caveat:** the shared LMDB index (`transcripts/index.mdb`) is not safe for concurrent > **writes**. Use shared mode for reading/building, not for running ingestion (downloads / > indexing) in two worktrees at the same time — concurrent writers can corrupt the index.