Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 85bec1b8348279434f26e41711572445e3aea7a8
parent b35c82948eb312c54997bf6650aaafc3de3a50e2
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sat, 26 Sep 2026 14:49:15 -0400

plans: release 10 slice M review fixes — S1 job id kept mid-poll, S2 the channel requirement, S3 15 s per request (+ N1 cause), S4 progress notifications and the 60 s clause; changelog bullet

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
Meditor/CHANGELOG.md | 2+-
Mplans/release-10.md | 85+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++--
2 files changed, 84 insertions(+), 3 deletions(-)

diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -6,7 +6,7 @@ - **Jobs a restart left queued are settled even when a drive hangs.** At boot the editor settles those jobs after it has checked where its storage locations are. A hung network mount could stall that check forever, and the jobs then stayed "queued" on `/jobs`. The settling now waits at most 60 seconds, logs `[boot] storage pass still running after 60 s …` and carries on. The storage check keeps running and logs when it ends. A job re-queued for a channel on the hung drive itself still waits for that drive, and any re-queued after it wait too. - **`/jobs` says why a job was cancelled at boot.** A job the boot settled shows its reason under its status on `/jobs` and as *Cancelled because* on its own page: for example "server restarted; the scheduler re-derives syncs" or "superseded by a newer queued job (…)". The reason used to be only in the job's log. - **The server log says how often a queued job skips its page refresh.** When a queued job finishes outside any request, the editor skips its page refresh and notes it in the log. The note used to appear once and never again. Now the first one after a quiet spell is logged at once, any more in the next 10 minutes are counted, and one line at the end gives the count, with a running total. -- **An agent working through the MCP server asks the editor for a clip instead of running yt-dlp.** The MCP server has a new tool, `fetch_clip`. Given a citation's channel, video id, start and end and a one-line reason, it asks the local editor for that window through `POST /api/media/fetch-window`: the same paced, cookie-aware job umtool uses, which records who asked and why beside the file. It answers with the file's path in the corpus (`channels/<slug>/data/<id>/clips/`). The window is the cited span with 3 seconds either side, at most 15 minutes. `full: true` asks for the whole recording instead, which lands in the saved-video store and needs a video the editor already knows. A Rumble citation's id (the embed id the archive publishes) is mapped to the id the editor names the video's folder by, through the archive record's link. The tool waits up to 90 seconds by default (at most 300) and otherwise returns the job's id, to wait on with `job`; the fetch carries on in the editor either way. The `/ask` and `/sweep` plans now tell the agent to use it and never to run yt-dlp itself. The MCP needs `ARCHILYZER_EDITOR_URL` and `WORKER_TOKEN` (the editor's own) in its environment, so re-register it with the two `--env` lines in the README; without them the tool says so and fetches nothing. The MCP server itself still writes nothing. The README's `yt-dlp --download-sections` command is now only the fallback for a machine with no editor. +- **An agent working through the MCP server asks the editor for a clip instead of running yt-dlp.** The MCP server has a new tool, `fetch_clip`. Given a citation's channel, video id, start and end and a one-line reason, it asks the local editor for that window through `POST /api/media/fetch-window`: the same paced, cookie-aware job umtool uses, which records who asked and why beside the file. It answers with the file's path in the corpus (`channels/<slug>/data/<id>/clips/`). The window is the cited span with 3 seconds either side, at most 15 minutes. `full: true` asks for the whole recording instead, which lands in the saved-video store and needs a video the editor already knows. A Rumble citation's id (the embed id the archive publishes) is mapped to the id the editor names the video's folder by, through the archive record's link. The editor must already archive the channel: pointed at a public site with a fresh editor, every clip gets a 404 `Channel "<slug>" not found`. The tool waits up to 90 seconds by default (at most 300) and otherwise returns the job's id, to wait on with `job`; the fetch carries on in the editor either way. While it waits it sends a progress notification per poll to a client that asks for progress. A client whose requests time out at 60 seconds (the MCP SDK's default) must raise that or pass `wait_seconds` of 50 or less. No request to the editor waits more than 15 seconds. If the editor stops answering mid-fetch, the answer gives the job's id and says not to ask again from scratch. The `/ask` and `/sweep` plans now tell the agent to use it and never to run yt-dlp itself. The MCP needs `ARCHILYZER_EDITOR_URL` and `WORKER_TOKEN` (the editor's own) in its environment, so re-register it with the two `--env` lines in the README; without them the tool says so and fetches nothing. The MCP server itself still writes nothing. The README's `yt-dlp --download-sections` command is now only the fallback for a machine with no editor. ## [0.9.0] - 2026-09-26 - **Every page now has a ground and an accent to choose, and the five theme families are gone.** The theme menu (the palette button beside the quick toggle, in the editor's sidebar and in the header of every published site, the hub and the homepage) has two groups. **Base** is System, Light, Sepia or Dark; Sepia is new, a warm paper ground for long reading. **Accent** is Signal, Brass, Vermilion, Violet, Sakura, Blue or Green, with the site's own tagged *default*; a site with a custom hex offers it first as *Site colour*. The quick toggle cycles System → Light → Sepia → Dark. A published site opens on the reader's system setting, in the accent its site form sets. The hub and the homepage open on Dark, in Signal, even with JavaScript off, and the editor follows the system, in Signal. Each accent has a value for each ground that reads at 4.5:1, and a custom hex is darkened or lightened per ground to match. A reader's accent is remembered only while it differs from the site's: picking the site's own again forgets it, so the reader follows the site if its accent changes later. Base, Archive, Selenized, Swiss and Archilyzer are gone. A choice made before this update carries over once: light stays light (Archive light becomes Sepia), dark stays dark and system stays system; the family itself is dropped. Headings are Archivo, text is IBM Plex Sans and figures are IBM Plex Mono everywhere, with one corner radius. Success, warning and other status text reads at 4.5:1 on its own tinted fill on every ground; on Light, success and warning are a shade deeper than before for it. Chart colours are fixed per ground and never follow the accent; the third is a violet, well clear of the red that marks a recording as gone. The phone's browser bar takes the page's ground, not the accent. Needs a rebuild and deploy of every site, the hub and the homepage. diff --git a/plans/release-10.md b/plans/release-10.md @@ -851,7 +851,7 @@ clause. `wait_seconds` and `full` are not backticked, and `NOT_TOOLS` is unchang before the first commit (`m-gate1.log`). The mcp package's `tsc --noEmit` is also clean on each of the three code commits checked out alone (`m-tsc-per-commit.log`). The docs commit has no code. - **mcp 259/259** (219 + 40: `fetchClip.test.ts` 30, `fetchClip.tool.test.ts` 9, - `instructions.test.ts` +1; `protocol.test.ts` still 2 tests, now with 15 names), 23 s. + `instructions.test.ts` +1; `protocol.test.ts` still 8 tests, its list now 15 names), 23 s. - **`test:scripts` 173 + 1 skip of 174**; **common 1,954/1,954** (`m-gate2.log`). Both are `main`'s counts: the prompt's 162 + 1 and 1,845 predate L1, L2 and S4. This slice changes neither (`git diff --stat 4ac32a2e a0326dfd -- common scripts umtool editor export homepage` is empty; @@ -875,7 +875,9 @@ clause. `wait_seconds` and `full` are not backticked, and `NOT_TOOLS` is unchang - **`mcp/README.md`'s own "Add to Claude Code" and `mcp.json` examples** do not carry the two env lines. The plan named only README.md's two blocks and AGENTS.md's; the tool row names both variables. -- **The plan's known limitations stand.** +- **The plan's known limitations stand**, plus one the review found (S2, now documented): + - The editor fetches only for a channel it already archives: a fresh editor behind an MCP + pointed at a public site answers 404 `Channel "<slug>" not found` on every clip. - A video absent from `source` is passed through as cited, with the note. - Full mode sends no `webpageUrl`, so a video the editor has never seen gets the editor's 404 (with the added sentence), where a window can still fetch it by URL. @@ -884,6 +886,85 @@ clause. `wait_seconds` and `full` are not backticked, and `NOT_TOOLS` is unchang is the only MCP tool that causes a write, and the editor does it". - **Registration is owed by the operator** (rollout step 2): this machine's `archilyzer` entry has `"env": {}`, so until it is re-registered `fetch_clip` answers "no editor configured". +- **Review nits not taken** (`m-review.md`): + - N2: the `queued` text still ends "the next ask finds it cached". It is the plan's exact + text, and whether to add "don't repeat the request" is the plan owner's call. For a window a + repeat is harmless (the fetch re-checks its cache when it runs). For `full: true` it queues + a second download. + - N3: a poll-phase disabled 503 gets the token sentence but not "the endpoint is off". Cosmetic. + - N4: the sweep's step says "(or a report needs one)", so extractors could fetch per finding. + Plan text, left for the operator. + - N5 (pre-existing): `AGENTS.md`'s "What it needs: **yt-dlp**" line for umtool's report-to-video + is stale, because umtool asks the editor by default. +- **A request whose body stalls after its headers** reads as `{}`. `readJson` swallows the + aborted body, so a POST reads as `The editor refused (HTTP 200): no reason given` and a poll + keeps polling until the wait ends. Rare; the timeout still bounds it. + +**Review fixes** (review SHIP AFTER FIXES, `m-review.md`: no must-fix, four should-fix, all fixed; +nit N1 taken). + +1. **S1: the job id survives an editor that stops answering mid-poll** (`868d4ab7`). + - **The failure.** A network error while polling answered "Could not reach the editor" with no + job id. The natural retry was the same call. With `full: true` that queues a second + whole-recording download, because `fetchFullSourceAction` checks the saved-video pointer + only when a request arrives. + - **The fix.** The poll-phase outcome carries the `jobId`, and the text is `Could not reach the + editor at <url> while polling job <jobId>: <message>. Is it running? Call fetch_clip again + with job: "<jobId>" — do not repeat the original request, the fetch may still be running.` +2. **S2: the docs say the editor must already archive the channel** (`d117e533`). The route answers + 404 `Channel "<slug>" not found` for a channel with no dir under the editor's `transcripts/`, so + pointing the MCP at a public site with a fresh editor gets a 404 on every clip. One sentence + each in README "Clips and video" (whose yt-dlp fallback now also covers an editor that does not + archive the channel), AGENTS.md's clips loop and the `mcp/README.md` row. The tool description + says it too. The `download-sections` grep still gives two lines, both fallback sentences + (`AGENTS.md:87`, `README.md:369`). +3. **S3 + N1: every request is bounded, and a failure says why** (`0777bfd1`). + - Each request, the POST and every poll, carries its own `AbortSignal.timeout(REQUEST_TIMEOUT_MS + = 15 s)`. Before, a POST that stats a hung mount, or an editor whose event loop had stalled, + held the call until undici's 300 s headers timeout. Now `wait_seconds` is the bound, give or + take one request (at most POST 15 s + the wait + one poll of 15 s). + - A timeout reads `no answer within 15 s (request timed out)`. A timed-out POST adds `It may + have queued the fetch anyway: check the editor's /jobs page before asking again.` + - Node's bare `fetch failed` carries its cause, e.g. `fetch failed (connect ECONNREFUSED + 127.0.0.1:3001)`. + - `requestTimeoutMs` is injectable in `FetchClipDeps` for the tests only. +4. **S4: the 60 s client timeout** (`8cf50cf6`). + - (a) The tool description and the `mcp/README.md` row say that a client with a 60 s default + request timeout must raise it or pass `wait_seconds` ≤ 50. The fetch continues on the editor + either way, and the next call finds it cached. + - (b) **Progress notifications: done, because it was cheap.** The `tools/call` handler now takes + `ctx`. When `ctx.mcpReq._meta.progressToken` is set, `progressNotifier` turns `fetchClip`'s new + `onPoll` hook into one `notifications/progress` per poll that finds the job still waiting + (`progress` = the poll count, `message` `editor job <id>: <status>, <n>s waited`, no `total`), + sent with `ctx.mcpReq.notify`. A client's `resetTimeoutOnProgress` then keeps the call alive. + A failing notify never fails the fetch, and with no token none is sent. + - Claude Code sets its own, longer tool timeout. **The rollout's in-memory proof script must + still pass `{ timeout: 330_000 }`** (or `onprogress` with `resetTimeoutOnProgress`) to + `callTool`. + +| sha | what | +|---|---| +| `868d4ab7` | S1: the poll-phase `unreachable` keeps `jobId`; the resume text; `fetchClip.test.ts` +1 (POST 202, poll throws, then the named resume polls once and posts nothing) | +| `d117e533` | S2: README, AGENTS.md, mcp/README: the editor must already archive the channel | +| `0777bfd1` | S3 + N1: `AbortSignal.timeout(15 s)` per request, `describeFetchError` (cause, timeout), the POST-timeout `/jobs` sentence; `fetchClip.test.ts` +6 (a signal per request, the cause, a fake `TimeoutError`, and over a real socket: a POST never answered, a poll never answered → the job id, a refused connection) | +| `8cf50cf6` | S4: `onPoll` → `notifications/progress`; the 60 s clause in the tool description and the README row; `fetchClip.test.ts` +1, `fetchClip.tool.test.ts` +1 (in-memory progress), `protocol.test.ts` +1 (the real process on the modern era with a stub editor) | +| _this_ | `plans:` these review fixes; the changelog bullet names the channel requirement, the progress and the 60 s clause | + +**Gates on `8cf50cf6`** (`m-gate3.log`): +- **tsc:** + - The whole workspace is clean. + - The mcp `tsc --noEmit` is clean on `868d4ab7` and `0777bfd1`, each checked out alone. + - `d117e533` is docs only. +- **mcp 269/269** (259 + 10), 17 s. +- common and `test:scripts` were not re-run: nothing outside `mcp/` and the three docs changed. +- **The new tests bite:** + - Without the timeout signal, the signal test fails, and the real never-answering POST hangs + past an 8 s test timeout. + - With the progress wiring disabled, the in-memory progress test fails. +- **The modern era carries progress.** The protocol test drives the real `src/index.ts` over stdio, + negotiated `modern`, with `ARCHILYZER_EDITOR_URL` and `WORKER_TOKEN` passed through the + transport's `env`. The stub editor sees three polls, each with `Bearer tok-proto`, and the client + gets progress `[1, 2]`. ## Rollout