Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit bf0355624493ca065fa5a483dd4ff2b71f7f2921
parent 47dc212014f5ed779c09adeeb7e62001135a3e15
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Fri,  9 Oct 2026 14:04:47 -0400

plans: release 19 Track A records — A1 to A5 as shipped, with the gates

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
Mplans/release-19.md | 154+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
1 file changed, 154 insertions(+), 0 deletions(-)

diff --git a/plans/release-19.md b/plans/release-19.md @@ -116,6 +116,160 @@ now: C1, C3 ──► C2 ──► C4, C5; C2 regenerated after A's actions la ### Track A +Branch `worktree-agent-a103bef2ee4055362` (worktree `.claude/worktrees/agent-a103bef2ee4055362`, ports editor 4101, +test 4111, export 4110) off `4cffda3f`, one Opus implementer, A1–A5 in order, one commit each; `r19/integration` +merged in before this record. Scratch files `a-*` in the job's `tmp`. + +| commit | slice | one line | +|---|---|---| +| `988b282c` | A1 | `pnpm ops` reads `WORKER_TOKEN` / `ARCHILYZER_EDITOR_URL` from `editor/.env(.local)`; `--wait` asks `GET /api/ops/job/<id>` | +| `5f86b10a` | A2 | `GET /api/ops/jobs`, `POST /api/ops/job {verb, ids}`; `get job`, `get jobs`, `job <verb>`, `job wait` | +| `e065bd2d` | A2b | 19 of `ops-api.spec`'s request/response tests become route unit tests | +| `30111658` | A3 | the read side (`get settings|storage|sites|workers|auto-queue|scheduler|cleanup`); `/api/auto-queue/control` behind the token | +| `5f442885` | A4 | archival writes (settings patch, the publish lane, platform holds, workers, one video, cleanup, relocate dryRun); `archilyzer storage report` | +| `c261b64a` | A5 | clip windows on `clips:<platform>`; an in-flight window answers its job | +| `e53f3169` | — | `r19/integration` merged in | +| `cde0e759` | A2b | the test corpus helper's paths opt out of tracing (`next-build-trace.test.mjs`) | + +#### Slice A1, as shipped — `pnpm ops` finds its token + +- `scripts/archilyzer-ops.mjs` `loadEditorEnv()` fills `WORKER_TOKEN` and `ARCHILYZER_EDITOR_URL` when unset — a + variable already set, even to `""`, wins — from `editor/.env.local` then `editor/.env` of the checkout the SCRIPT is + in (`REPO_ROOT` from `import.meta.url`, never the cwd), then of the main worktree when it runs in a linked one + (`mainWorktreeOf`: `.git` file → `gitdir` → `commondir`, read off disk, no git process). Only those two keys are + taken; the files also hold deploy credentials. A 401/503 prints `tokenHint`: where the token came from, never its + value. `parseDotenv` reads `KEY=value`, `export`, quotes, comments. +- **What `--wait` polled and why it was a rewrite.** After three failed log polls it asked `/api/jobs/active`, which + since one-core phase 3 slice 2 is a `next.config` rewrite onto `/api/view/activeJobs` — kept at its old path for the + pages and pinned widgets that poll it — and that view builds the whole live payload (channel stats, the disk gate, + every runner's status) to answer "is job X still there". It now asks `GET /api/ops/job/<id>` (behind the token: one + registry record or `.meta.json` sidecar). An `archived` log status (an id the registry forgot) resolves to the + sidecar's terminal status instead of exiting 1 for a job that ended `done`; a 404 `{ok:false}` is "gone". While a + job waits, `followJob` prints `queued — position N on <queue>` once per change. +- `GET /api/ops/job/<id>[?tail=N]` (`editor/app/api/ops/_jobs.ts`): the `getJobEntry` row without `logPath`, `queue: + {key, position, queued, head}` while queued or running, the runner's `progress`; `tail` (≤ 500) is + `common/jobs/listJobs.ts` `readLogTail` — the last N lines read from the end (256 KiB window, a partial first line + dropped). + +#### Slice A2, as shipped — jobs over ops + +- `GET /api/ops/jobs` — `?active` (registry queued/running, queue order, each with its place), `?failed`, `?kind`, + `?slug`, `?limit` (50, ≤ 500); a filtered history list reads the newest 2000 jobs (`JOBS_SCAN`); `active` with + `failed` is a 400; unknown query keys are a 400. +- `POST /api/ops/job {verb, ids}` — `cancel | drain | promote | force-release | retry` per id through + `editor/app/jobs/actions.ts` (the row buttons' own actions), each answered in `results`; one refused id makes the + answer `ok: false` (400, naming it) without stopping the others. `retry-failed` (no ids) is Retry all; + `retryAllFailedAction` now also returns `jobIds` and cancels the new jobs' unread streams. `jobIds` = the jobs a + retry started, which `--wait` follows. +- CLI: `get job <id> [--tail [N]]`, `get jobs [--active|--failed] [--kind] [--slug] [--limit]`, `job <verb> <id…>`, + `job wait <id…>` (no request; `followAll`, the `--wait` loop factored out). A read's flag on the wrong noun or on an + action is refused, not ignored. +- `ops-api.spec` gains "jobs over ops": two `/api/test/stuck-job` holders on one queue — the live list, one job's + place and head, `?tail`, a promote refused ("already at the front"), a cancel of one known and one unknown id, + `force-release` of the head. The success paths revalidate `/jobs`, which only a running server can. + +#### Slice A2b, as shipped — the ops-api spec's request/response tests are unit tests + +- 19 of `ops-api.spec`'s 33 tests (32 at base + A2's) exercised only a route's request and response; they moved, + assertion for assertion, to `editor/app/api/ops/_door.test.ts` (the 401, the 503, unknown keys, the traversing slug + on fourteen routes) and route tests for retry-bucket, tag-videos, persist-videos, build-site (+ build-deploy), + deploy-site (+ build-deploy), deploy-hub, deploy-homepage and publish; fetch-posts and capture-posts gain one each. + **ops-api.spec: 33 → 14 tests.** What stays needs the running server: writes that revalidate, jobs that run, the + publish queue's runs, the jobs verbs. +- `editor/app/api/ops/_testCorpus.ts` (imported only by tests): copies the named e2e fixture into a temp dir, sets the + e2e test server's environment (fake yt-dlp, gallery-dl, ffmpeg/ffprobe, whisper, chough, parakeet, diarize, claude, + findmnt, udisksctl, wrangler, next; the fake Cloudflare token; `ARCHILYZER_BRANCH=main`; `E2E_LIVE_CHECK=skip`) and + `holdPublishQueue()`s as e2e did, so a refusal that regresses can only queue a fake. +- The 503 no longer needs `/api/test/worker-token`: the unit process unsets its own `WORKER_TOKEN`. That test route now + has no e2e caller (left in place). + +#### Slice A3, as shipped — the read side; the lane control route needs the token + +- `GET /api/ops/settings[?key]` (getSettings, migrated and defaulted), `storage` (`buildStorage`), `sites` + (`listSites`: id, title, `siteUrl`, Pages project, audience, `isListedSite`, search, publish policy, channels), + `workers` (`buildWorkersPayload`), `auto-queue` (`buildAutoQueueStatusPayload`), `scheduler` + (`buildSchedulerStatusPayload`), `cleanup/<slug>` (new `loadCleanupRow` — `loadCleanupSummary`'s row for one channel, + id lists as counts). One door, `editor/app/api/ops/_read.ts` `readRoute` (token, unknown query keys refused) and + `redactSecrets` (any string under a key naming a token/secret/password/api key/credential → `"<redacted>"`; a remote + worker's `token` is the one that exists today). CLI `get settings [<key>]`, `get storage|sites|workers|auto-queue| + scheduler`, `get cleanup <slug>`. +- **`/api/auto-queue/control` answers `opsAuth`.** Its one UI caller — `RunnerOperationView`'s Start/Drain/Stop on + /operations — calls `startAutoQueueAction` / `stopAutoQueueAction` / new `drainAutoQueueAction` instead: `opsAuth` + has no exemption for the UI (the UI never called an ops route; it uses server actions, which carry Next's origin + check), and a page holds no token. The six e2e specs that post to the route (auto-queue, auto-subs-replace, + disk-space, lane-runner, pacing, rate-limit) send the test token. + +#### Slice A4, as shipped — archival writes + +- `POST /api/ops/settings {patch}` → `saveSettings` (one-level merge). Refused before anything is written + (`settings/patch.ts`): an unknown top-level key; **a key another writer owns** — `channelPriority` (compiled with the + lanes' trees in one write), `autoQueue`, `workers`, `storage` — naming that writer; and a value the schema would not + keep as sent (it coerces, never throws), every leaf of the patch compared with the schema's reading of the merged + settings and each difference named ("minFreeDiskGB: sent -5, would be saved as <the clamp>"). Answers `{changed, value}` read + back, redacted. +- `lane` takes `publish`: `enabled` through `savePublishSettingsAction` with the stored fields, `held` through the + pause gate, Start/Drain/Stop the publish page's actions; the ingest lanes' drain goes through + `drainAutoQueueAction`. +- `clear-platform-hold {platform}` (`clearPlatformHoldAction`, answers its sentence); `POST /api/ops/workers {op: + enable|disable|drain, ids}` (the live pool, per id); `transcribe-one {slug, id, file?}` (`transcribeOneAction` with a + file, else the video page's `whisperVideoAction`, which fetches audio through the managed path when there is none; + the file name is checked as one path segment); `delete-file {slug, id, file}` (`deleteVideoFileAction` — + `removeMediaFile`, refused while the media drive is not answering); `do-not-clean {slug, id, keep?}` (set, not + toggled); `POST /api/ops/cleanup {slug, sweep: transcribed|extra-formats|wrong-format|failed-transcriptions}` (the + channel page's four actions, as jobs); `relocate {dryRun}` (`previewRelocationAction` per slug — which, as on the + panel, tiers a never-tiered channel in place first; a slug with no channel is the bulk path's "channel not found"). +- `archilyzer storage report [<slug>…] [--json]` (`common/bin/storage-report.ts`): per channel the tierable media, + text and clip bytes off its report (`totalMediaBytes` is `isTierable`'s share) and where its media is — corpus disk, + a location (`config.mediaDir` under its root, longest match), a custom path, or legacy — with totals per place; a + missing figure is unknown (`—` / `null`), never 0; social channels left out; no drive touched, no editor needed. + +#### Slice A5, as shipped — fetch queue hygiene + +- **Its own queue.** `common/lib/queueKeys.ts` `clipWindowQueueKey`: a window job runs on `clips:<platform>` (from the + platform queue it would have used, so `platform:vimeo.com` → `clips:vimeo.com`) — single (`fetchWindowAction`) and + batch (`fetchWindowsAction`). A queue runs one job at a time and a tier orders only the waiting ones, so priority + alone could not get a window past a running multi-hour persist; the queue key can. **Ruling recorded:** one window + may now run beside one download on the same platform; windows stay one at a time per platform with the batches' + `clip-window:<platform>` gap, and the platform's hold, cooldown and 429 backoff are shared with the downloads. +- **Deduped.** `common/jobs/windowJobs.ts` `windowInFlight`: a queued or running `fetch-window` whose window contains + the asked one, or a `fetch-windows` batch with such an item — same channel, video and height cap — read off the job's + replay spec. `/api/media/fetch-window` answers that job (202, `existing: true`, its window and file); a batch lists + such windows in `inFlight` and queues nothing for them, and the ops route adds their jobs to `jobIds` so `--wait` + follows the windows asked for. Retry (`stream: true`) is not deduplicated. +- **Not done: `fetch-via-editor.mjs`.** It is `umtool/report-to-video/fetch-via-editor.mjs` — under the directory no + track touches — so its 10-minute poll timeout and its wait output are unchanged. The editor half covers its retry + loop (a re-run after its timeout gets the job already fetching), and `pnpm ops job wait <id>` waits with no timeout, + printing the queue position. + +#### Track A gates (A1–A5, on the merged tip) + +- `pnpm -r --no-bail --workspace-concurrency=1 exec tsc --noEmit` — clean before every commit and on the merged tip + (the A2/A2b/A3 commits under the 12:35 throttle ran the editor package's tsc alone; A1, A4, A5 and the tip ran the + whole workspace). +- **editor unit** 205/205 (159 at base: +5 A1, +4 A2, +19 A2b, +7 A3, +9 A4, +2 A5). **test:scripts** 693: 689 + pass, 3 skip, 1 fail — `next-build-trace.test.mjs`'s cwd-path check on the new test helper, fixed by `cde0e759` + (the rerun's one failure is `queue-lock.test.mjs` "prints a banner naming the holder", run while another session's + e2e held the machine lock; it passed in the gate run). `archilyzer-ops.test.mjs` 46 → 59. **common** 3473: 3469 + pass, 4 fail — `fetchPosts.test.ts` and `relocateDir.test.ts` pass alone (load); `autoRunner.test.ts`'s two + ("computeLeafPending takes its channel listing from `shared`", "a channel with a move marker projects no work…") + fail the same way on an extracted `4cffda3f` tree: not this branch's. +8 here (listJobs 1, windowJobs 3, + storage-report 4). **mcp** 292/292. +- **`pnpm --filter editor exec next build`** (5 GB scope) — ok. +- **e2e** (detached, queued, memory-gated; `export/public` linked to the composed fixture as `r18-integration`'s is — + the primary's no longer has `summaries/`, so the export server never answered and the first launch timed out + waiting for it): `ops-api jobs fetch-window auto-queue lane-runner disk-space pacing rate-limit auto-subs-replace` — + **65 passed, 5 failed, 19.0 min**, at load average ~26–32. ops-api (14/14), fetch-window (9/9), lane-runner, + pacing and auto-subs-replace passed. The five failures are 30-s test timeouts before any code this branch changed ran: auto-queue's "UI: + build a policy…" (the Start click its own comment calls a race against the poll — it lost, the button went disabled + under it), "Drain completes when … parked" (timed out polling for the manual whisper job, before the Drain click), + disk-space's "prevents a download…" (the channel page's Download click), jobs' "kicks off the index update" (the + index stage still running at 30 s), rate-limit's "an overdue hold…" (`ECONNRESET` from the dev server on a control + POST; the same POST with the token passed in the other specs). **Recheck of the five alone** (load ~22, 2.0 min): + **4 passed, 1 failed** — the UI Start and the UI Drain (the two buttons A3 moved to server actions), disk-space and + rate-limit pass; jobs' index test hit its 30-s test timeout again while its own page snapshot shows the stage's log + ending `[stage] update-index _index: Done` — the index stage, which this branch does not touch, finishing late + under load. + ### Track B ### Track C