commit bf0355624493ca065fa5a483dd4ff2b71f7f2921
parent 47dc212014f5ed779c09adeeb7e62001135a3e15
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Fri, 9 Oct 2026 14:04:47 -0400
plans: release 19 Track A records — A1 to A5 as shipped, with the gates
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Diffstat:
| M | plans/release-19.md | | | 154 | +++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ |
1 file changed, 154 insertions(+), 0 deletions(-)
diff --git a/plans/release-19.md b/plans/release-19.md
@@ -116,6 +116,160 @@ now: C1, C3 ──► C2 ──► C4, C5; C2 regenerated after A's actions la
### Track A
+Branch `worktree-agent-a103bef2ee4055362` (worktree `.claude/worktrees/agent-a103bef2ee4055362`, ports editor 4101,
+test 4111, export 4110) off `4cffda3f`, one Opus implementer, A1–A5 in order, one commit each; `r19/integration`
+merged in before this record. Scratch files `a-*` in the job's `tmp`.
+
+| commit | slice | one line |
+|---|---|---|
+| `988b282c` | A1 | `pnpm ops` reads `WORKER_TOKEN` / `ARCHILYZER_EDITOR_URL` from `editor/.env(.local)`; `--wait` asks `GET /api/ops/job/<id>` |
+| `5f86b10a` | A2 | `GET /api/ops/jobs`, `POST /api/ops/job {verb, ids}`; `get job`, `get jobs`, `job <verb>`, `job wait` |
+| `e065bd2d` | A2b | 19 of `ops-api.spec`'s request/response tests become route unit tests |
+| `30111658` | A3 | the read side (`get settings|storage|sites|workers|auto-queue|scheduler|cleanup`); `/api/auto-queue/control` behind the token |
+| `5f442885` | A4 | archival writes (settings patch, the publish lane, platform holds, workers, one video, cleanup, relocate dryRun); `archilyzer storage report` |
+| `c261b64a` | A5 | clip windows on `clips:<platform>`; an in-flight window answers its job |
+| `e53f3169` | — | `r19/integration` merged in |
+| `cde0e759` | A2b | the test corpus helper's paths opt out of tracing (`next-build-trace.test.mjs`) |
+
+#### Slice A1, as shipped — `pnpm ops` finds its token
+
+- `scripts/archilyzer-ops.mjs` `loadEditorEnv()` fills `WORKER_TOKEN` and `ARCHILYZER_EDITOR_URL` when unset — a
+ variable already set, even to `""`, wins — from `editor/.env.local` then `editor/.env` of the checkout the SCRIPT is
+ in (`REPO_ROOT` from `import.meta.url`, never the cwd), then of the main worktree when it runs in a linked one
+ (`mainWorktreeOf`: `.git` file → `gitdir` → `commondir`, read off disk, no git process). Only those two keys are
+ taken; the files also hold deploy credentials. A 401/503 prints `tokenHint`: where the token came from, never its
+ value. `parseDotenv` reads `KEY=value`, `export`, quotes, comments.
+- **What `--wait` polled and why it was a rewrite.** After three failed log polls it asked `/api/jobs/active`, which
+ since one-core phase 3 slice 2 is a `next.config` rewrite onto `/api/view/activeJobs` — kept at its old path for the
+ pages and pinned widgets that poll it — and that view builds the whole live payload (channel stats, the disk gate,
+ every runner's status) to answer "is job X still there". It now asks `GET /api/ops/job/<id>` (behind the token: one
+ registry record or `.meta.json` sidecar). An `archived` log status (an id the registry forgot) resolves to the
+ sidecar's terminal status instead of exiting 1 for a job that ended `done`; a 404 `{ok:false}` is "gone". While a
+ job waits, `followJob` prints `queued — position N on <queue>` once per change.
+- `GET /api/ops/job/<id>[?tail=N]` (`editor/app/api/ops/_jobs.ts`): the `getJobEntry` row without `logPath`, `queue:
+ {key, position, queued, head}` while queued or running, the runner's `progress`; `tail` (≤ 500) is
+ `common/jobs/listJobs.ts` `readLogTail` — the last N lines read from the end (256 KiB window, a partial first line
+ dropped).
+
+#### Slice A2, as shipped — jobs over ops
+
+- `GET /api/ops/jobs` — `?active` (registry queued/running, queue order, each with its place), `?failed`, `?kind`,
+ `?slug`, `?limit` (50, ≤ 500); a filtered history list reads the newest 2000 jobs (`JOBS_SCAN`); `active` with
+ `failed` is a 400; unknown query keys are a 400.
+- `POST /api/ops/job {verb, ids}` — `cancel | drain | promote | force-release | retry` per id through
+ `editor/app/jobs/actions.ts` (the row buttons' own actions), each answered in `results`; one refused id makes the
+ answer `ok: false` (400, naming it) without stopping the others. `retry-failed` (no ids) is Retry all;
+ `retryAllFailedAction` now also returns `jobIds` and cancels the new jobs' unread streams. `jobIds` = the jobs a
+ retry started, which `--wait` follows.
+- CLI: `get job <id> [--tail [N]]`, `get jobs [--active|--failed] [--kind] [--slug] [--limit]`, `job <verb> <id…>`,
+ `job wait <id…>` (no request; `followAll`, the `--wait` loop factored out). A read's flag on the wrong noun or on an
+ action is refused, not ignored.
+- `ops-api.spec` gains "jobs over ops": two `/api/test/stuck-job` holders on one queue — the live list, one job's
+ place and head, `?tail`, a promote refused ("already at the front"), a cancel of one known and one unknown id,
+ `force-release` of the head. The success paths revalidate `/jobs`, which only a running server can.
+
+#### Slice A2b, as shipped — the ops-api spec's request/response tests are unit tests
+
+- 19 of `ops-api.spec`'s 33 tests (32 at base + A2's) exercised only a route's request and response; they moved,
+ assertion for assertion, to `editor/app/api/ops/_door.test.ts` (the 401, the 503, unknown keys, the traversing slug
+ on fourteen routes) and route tests for retry-bucket, tag-videos, persist-videos, build-site (+ build-deploy),
+ deploy-site (+ build-deploy), deploy-hub, deploy-homepage and publish; fetch-posts and capture-posts gain one each.
+ **ops-api.spec: 33 → 14 tests.** What stays needs the running server: writes that revalidate, jobs that run, the
+ publish queue's runs, the jobs verbs.
+- `editor/app/api/ops/_testCorpus.ts` (imported only by tests): copies the named e2e fixture into a temp dir, sets the
+ e2e test server's environment (fake yt-dlp, gallery-dl, ffmpeg/ffprobe, whisper, chough, parakeet, diarize, claude,
+ findmnt, udisksctl, wrangler, next; the fake Cloudflare token; `ARCHILYZER_BRANCH=main`; `E2E_LIVE_CHECK=skip`) and
+ `holdPublishQueue()`s as e2e did, so a refusal that regresses can only queue a fake.
+- The 503 no longer needs `/api/test/worker-token`: the unit process unsets its own `WORKER_TOKEN`. That test route now
+ has no e2e caller (left in place).
+
+#### Slice A3, as shipped — the read side; the lane control route needs the token
+
+- `GET /api/ops/settings[?key]` (getSettings, migrated and defaulted), `storage` (`buildStorage`), `sites`
+ (`listSites`: id, title, `siteUrl`, Pages project, audience, `isListedSite`, search, publish policy, channels),
+ `workers` (`buildWorkersPayload`), `auto-queue` (`buildAutoQueueStatusPayload`), `scheduler`
+ (`buildSchedulerStatusPayload`), `cleanup/<slug>` (new `loadCleanupRow` — `loadCleanupSummary`'s row for one channel,
+ id lists as counts). One door, `editor/app/api/ops/_read.ts` `readRoute` (token, unknown query keys refused) and
+ `redactSecrets` (any string under a key naming a token/secret/password/api key/credential → `"<redacted>"`; a remote
+ worker's `token` is the one that exists today). CLI `get settings [<key>]`, `get storage|sites|workers|auto-queue|
+ scheduler`, `get cleanup <slug>`.
+- **`/api/auto-queue/control` answers `opsAuth`.** Its one UI caller — `RunnerOperationView`'s Start/Drain/Stop on
+ /operations — calls `startAutoQueueAction` / `stopAutoQueueAction` / new `drainAutoQueueAction` instead: `opsAuth`
+ has no exemption for the UI (the UI never called an ops route; it uses server actions, which carry Next's origin
+ check), and a page holds no token. The six e2e specs that post to the route (auto-queue, auto-subs-replace,
+ disk-space, lane-runner, pacing, rate-limit) send the test token.
+
+#### Slice A4, as shipped — archival writes
+
+- `POST /api/ops/settings {patch}` → `saveSettings` (one-level merge). Refused before anything is written
+ (`settings/patch.ts`): an unknown top-level key; **a key another writer owns** — `channelPriority` (compiled with the
+ lanes' trees in one write), `autoQueue`, `workers`, `storage` — naming that writer; and a value the schema would not
+ keep as sent (it coerces, never throws), every leaf of the patch compared with the schema's reading of the merged
+ settings and each difference named ("minFreeDiskGB: sent -5, would be saved as <the clamp>"). Answers `{changed, value}` read
+ back, redacted.
+- `lane` takes `publish`: `enabled` through `savePublishSettingsAction` with the stored fields, `held` through the
+ pause gate, Start/Drain/Stop the publish page's actions; the ingest lanes' drain goes through
+ `drainAutoQueueAction`.
+- `clear-platform-hold {platform}` (`clearPlatformHoldAction`, answers its sentence); `POST /api/ops/workers {op:
+ enable|disable|drain, ids}` (the live pool, per id); `transcribe-one {slug, id, file?}` (`transcribeOneAction` with a
+ file, else the video page's `whisperVideoAction`, which fetches audio through the managed path when there is none;
+ the file name is checked as one path segment); `delete-file {slug, id, file}` (`deleteVideoFileAction` —
+ `removeMediaFile`, refused while the media drive is not answering); `do-not-clean {slug, id, keep?}` (set, not
+ toggled); `POST /api/ops/cleanup {slug, sweep: transcribed|extra-formats|wrong-format|failed-transcriptions}` (the
+ channel page's four actions, as jobs); `relocate {dryRun}` (`previewRelocationAction` per slug — which, as on the
+ panel, tiers a never-tiered channel in place first; a slug with no channel is the bulk path's "channel not found").
+- `archilyzer storage report [<slug>…] [--json]` (`common/bin/storage-report.ts`): per channel the tierable media,
+ text and clip bytes off its report (`totalMediaBytes` is `isTierable`'s share) and where its media is — corpus disk,
+ a location (`config.mediaDir` under its root, longest match), a custom path, or legacy — with totals per place; a
+ missing figure is unknown (`—` / `null`), never 0; social channels left out; no drive touched, no editor needed.
+
+#### Slice A5, as shipped — fetch queue hygiene
+
+- **Its own queue.** `common/lib/queueKeys.ts` `clipWindowQueueKey`: a window job runs on `clips:<platform>` (from the
+ platform queue it would have used, so `platform:vimeo.com` → `clips:vimeo.com`) — single (`fetchWindowAction`) and
+ batch (`fetchWindowsAction`). A queue runs one job at a time and a tier orders only the waiting ones, so priority
+ alone could not get a window past a running multi-hour persist; the queue key can. **Ruling recorded:** one window
+ may now run beside one download on the same platform; windows stay one at a time per platform with the batches'
+ `clip-window:<platform>` gap, and the platform's hold, cooldown and 429 backoff are shared with the downloads.
+- **Deduped.** `common/jobs/windowJobs.ts` `windowInFlight`: a queued or running `fetch-window` whose window contains
+ the asked one, or a `fetch-windows` batch with such an item — same channel, video and height cap — read off the job's
+ replay spec. `/api/media/fetch-window` answers that job (202, `existing: true`, its window and file); a batch lists
+ such windows in `inFlight` and queues nothing for them, and the ops route adds their jobs to `jobIds` so `--wait`
+ follows the windows asked for. Retry (`stream: true`) is not deduplicated.
+- **Not done: `fetch-via-editor.mjs`.** It is `umtool/report-to-video/fetch-via-editor.mjs` — under the directory no
+ track touches — so its 10-minute poll timeout and its wait output are unchanged. The editor half covers its retry
+ loop (a re-run after its timeout gets the job already fetching), and `pnpm ops job wait <id>` waits with no timeout,
+ printing the queue position.
+
+#### Track A gates (A1–A5, on the merged tip)
+
+- `pnpm -r --no-bail --workspace-concurrency=1 exec tsc --noEmit` — clean before every commit and on the merged tip
+ (the A2/A2b/A3 commits under the 12:35 throttle ran the editor package's tsc alone; A1, A4, A5 and the tip ran the
+ whole workspace).
+- **editor unit** 205/205 (159 at base: +5 A1, +4 A2, +19 A2b, +7 A3, +9 A4, +2 A5). **test:scripts** 693: 689
+ pass, 3 skip, 1 fail — `next-build-trace.test.mjs`'s cwd-path check on the new test helper, fixed by `cde0e759`
+ (the rerun's one failure is `queue-lock.test.mjs` "prints a banner naming the holder", run while another session's
+ e2e held the machine lock; it passed in the gate run). `archilyzer-ops.test.mjs` 46 → 59. **common** 3473: 3469
+ pass, 4 fail — `fetchPosts.test.ts` and `relocateDir.test.ts` pass alone (load); `autoRunner.test.ts`'s two
+ ("computeLeafPending takes its channel listing from `shared`", "a channel with a move marker projects no work…")
+ fail the same way on an extracted `4cffda3f` tree: not this branch's. +8 here (listJobs 1, windowJobs 3,
+ storage-report 4). **mcp** 292/292.
+- **`pnpm --filter editor exec next build`** (5 GB scope) — ok.
+- **e2e** (detached, queued, memory-gated; `export/public` linked to the composed fixture as `r18-integration`'s is —
+ the primary's no longer has `summaries/`, so the export server never answered and the first launch timed out
+ waiting for it): `ops-api jobs fetch-window auto-queue lane-runner disk-space pacing rate-limit auto-subs-replace` —
+ **65 passed, 5 failed, 19.0 min**, at load average ~26–32. ops-api (14/14), fetch-window (9/9), lane-runner,
+ pacing and auto-subs-replace passed. The five failures are 30-s test timeouts before any code this branch changed ran: auto-queue's "UI:
+ build a policy…" (the Start click its own comment calls a race against the poll — it lost, the button went disabled
+ under it), "Drain completes when … parked" (timed out polling for the manual whisper job, before the Drain click),
+ disk-space's "prevents a download…" (the channel page's Download click), jobs' "kicks off the index update" (the
+ index stage still running at 30 s), rate-limit's "an overdue hold…" (`ECONNRESET` from the dev server on a control
+ POST; the same POST with the token passed in the other specs). **Recheck of the five alone** (load ~22, 2.0 min):
+ **4 passed, 1 failed** — the UI Start and the UI Drain (the two buttons A3 moved to server actions), disk-space and
+ rate-limit pass; jobs' index test hit its 30-s test timeout again while its own page snapshot shows the stage's log
+ ending `[stage] update-index _index: Done` — the index stage, which this branch does not touch, finishing late
+ under load.
+
### Track B
### Track C