commit 983b04a5a79a1a20c491e0c87f70221ede013ba3
parent 85bec1b8348279434f26e41711572445e3aea7a8
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Sat, 26 Sep 2026 14:54:42 -0400
mcp: fetch_clip's queued text never invites repeating a running request (re-read R2, R1)
The editor does not dedupe a running whole-recording job, so "the next
ask finds it cached" held only once the job had finished; a repeated
full: true while it ran queued a second download. The queued text now
says to resume with job and never repeat the original request while it
runs, and that the same request finds it cached once it has finished;
the 60 s clauses in the tool description and the mcp/README row say the
same. The plan's queued text (the contract) is amended with the reason,
and the record strikes N2. R1: the three real-socket tests carry
{ timeout: 10_000 }. mcp 269/269.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Diffstat:
6 files changed, 46 insertions(+), 16 deletions(-)
diff --git a/mcp/README.md b/mcp/README.md
@@ -21,7 +21,7 @@ clip window; the MCP itself still writes nothing.
| `get_transcript` | One video's full transcript as clean markdown (metadata + **linked** timestamped captions). |
| `get_post` / `get_thread` | One archived social post, or its whole thread. Posts have no timeline — cite them with no `@ mm:ss`. |
| `get_video_metadata` | Everything known about one video without the transcript body: metadata, plus **view/like counts, cue count and transcript coverage** (`stats/`), **other archived copies of the same recording** with an explicit timings-aligned verdict (`duplicates.json`), and **AI chapters/tags** where they exist (`digests/`). |
-| `fetch_clip` | The media behind a cited moment, **fetched by the local editor** (`POST /api/media/fetch-window`) through its paced, cookie-aware, provenanced job — never a yt-dlp run by hand. Needs `ARCHILYZER_EDITOR_URL` (default `http://localhost:3001`) and `WORKER_TOKEN` (the editor's own) in this server's env; without them it says so and fetches nothing. The editor must already archive the cited channel (a channel dir under its `transcripts/`), else it answers 404 `Channel "<slug>" not found`: an MCP pointed at a public site with a fresh editor gets that on every clip. A window is the cited span ± `pad` (default 3 s), at most 15 min, and lands at `channels/<slug>/data/<id>/clips/`; `full: true` fetches the whole recording into the saved-video store (needs a video the editor already knows). Waits up to `wait_seconds` (default 90, max 300), then returns the job id to resume with `job`; a client with a 60 s default request timeout must raise it or pass `wait_seconds` ≤ 50 — the fetch continues on the editor either way and the next call finds it cached. While it waits it sends one progress notification per poll to a client that asked for progress (a `progressToken`), which keeps a reset-on-progress timeout alive. A Rumble embed id is mapped to the editor's slug id through the record's `webpageUrl`, so pass the citing corpus as `source`; a video not in `source` is passed through as cited (known limitation). The file is a read-only corpus artifact. |
+| `fetch_clip` | The media behind a cited moment, **fetched by the local editor** (`POST /api/media/fetch-window`) through its paced, cookie-aware, provenanced job — never a yt-dlp run by hand. Needs `ARCHILYZER_EDITOR_URL` (default `http://localhost:3001`) and `WORKER_TOKEN` (the editor's own) in this server's env; without them it says so and fetches nothing. The editor must already archive the cited channel (a channel dir under its `transcripts/`), else it answers 404 `Channel "<slug>" not found`: an MCP pointed at a public site with a fresh editor gets that on every clip. A window is the cited span ± `pad` (default 3 s), at most 15 min, and lands at `channels/<slug>/data/<id>/clips/`; `full: true` fetches the whole recording into the saved-video store (needs a video the editor already knows). Waits up to `wait_seconds` (default 90, max 300), then returns the job id to resume with `job`; a client with a 60 s default request timeout must raise it or pass `wait_seconds` ≤ 50 — the fetch continues on the editor either way; resume it with `job`, and once it has finished the same request finds it cached. While it waits it sends one progress notification per poll to a client that asked for progress (a `progressToken`), which keeps a reset-on-progress timeout alive. A Rumble embed id is mapped to the editor's slug id through the record's `webpageUrl`, so pass the citing corpus as `source`; a video not in `source` is passed through as cited (known limitation). The file is a read-only corpus artifact. |
| `open_link` | Paste an archilyzer viewer **share link** to re-run that exact search here (query tree + every filter, at full fidelity) — plan, results and corpus handle in **one** call. `dry_run:true` for the plan alone. |
| `list_sources` | Show the **default** corpus and, with a hub, its member sites as ready-to-paste handles. |
| `resolve_source` | Turn a URL or site name into the canonical `source` handle and check it can be read. Changes nothing. |
diff --git a/mcp/src/fetchClip.test.ts b/mcp/src/fetchClip.test.ts
@@ -374,8 +374,9 @@ test("the wait runs out → queued with the job id, without real timers", async
assert.equal(
r.text,
'Still running on the editor (job j1, waited 3s). Call fetch_clip again with job: "j1" ' +
- "to keep waiting — nothing is lost, the fetch continues on the editor and the next " +
- "ask finds it cached.",
+ "to keep waiting — never repeat the original request while it runs (that would queue " +
+ "a second fetch). Nothing is lost: the fetch continues on the editor, and once it has " +
+ "finished the same request finds it cached.",
);
});
@@ -682,7 +683,7 @@ function realDeps(url: string, requestTimeoutMs: number): FetchClipDeps {
};
}
-test("a real editor that never answers the POST is given up on at the request timeout", async () => {
+test("a real editor that never answers the POST is given up on at the request timeout", { timeout: 10_000 }, async () => {
const ed = await silentEditor();
try {
const t0 = Date.now();
@@ -701,7 +702,7 @@ test("a real editor that never answers the POST is given up on at the request ti
}
});
-test("a real editor that stops answering a poll hands back the job", async () => {
+test("a real editor that stops answering a poll hands back the job", { timeout: 10_000 }, async () => {
const ed = await silentEditor({ cached: false, jobId: "j4", file: null, from: 0, to: 0 });
try {
const r = renderFetchClip(await fetchClip(fullRequest(), realDeps(ed.url, 150)), CTX);
@@ -716,7 +717,7 @@ test("a real editor that stops answering a poll hands back the job", async () =>
}
});
-test("a real refused connection names its cause", async () => {
+test("a real refused connection names its cause", { timeout: 10_000 }, async () => {
const ed = await silentEditor();
const url = ed.url;
await ed.close(); // nothing listens there now
diff --git a/mcp/src/fetchClip.ts b/mcp/src/fetchClip.ts
@@ -628,8 +628,10 @@ export function renderFetchClip(
text:
`Still ${outcome.status} on the editor (job ${outcome.jobId}, waited ` +
`${outcome.waited}s). Call fetch_clip again with job: ` +
- `"${outcome.jobId}" to keep waiting — nothing is lost, the fetch ` +
- `continues on the editor and the next ask finds it cached.`,
+ `"${outcome.jobId}" to keep waiting — never repeat the original ` +
+ `request while it runs (that would queue a second fetch). Nothing is ` +
+ `lost: the fetch continues on the editor, and once it has finished ` +
+ `the same request finds it cached.`,
isError: false,
};
case "cooldown":
diff --git a/mcp/src/server.ts b/mcp/src/server.ts
@@ -730,7 +730,8 @@ export const TOOLS: Tool[] = [
"waiting. It sends a progress notification per poll when the client asks " +
"for progress. A client with a 60 s default request timeout must raise " +
"it or pass wait_seconds ≤ 50 — the fetch continues on the editor either " +
- "way and the next call finds it cached. The editor must already archive " +
+ "way; resume it with job, and once it has finished the same request " +
+ "finds it cached. The editor must already archive " +
"the cited channel. Needs ARCHILYZER_EDITOR_URL and WORKER_TOKEN in this server's " +
"environment; without them it says so and fetches nothing.",
inputSchema: {
diff --git a/plans/mcp-fetch-clip.md b/plans/mcp-fetch-clip.md
@@ -106,8 +106,13 @@ it as <from>-<to>.json.`
- 202 → done: `Fetched <from>–<to> of <channel>/<canonical> (job <id>, <n>s waited).` + footer from
the poll's `file/from/to/bytes`.
- wait expired, still queued/running (NOT `isError`): `Still <status> on the editor (job <jobId>,
- waited <n>s). Call fetch_clip again with job: "<jobId>" to keep waiting — nothing is lost, the fetch
- continues on the editor and the next ask finds it cached.`
+ waited <n>s). Call fetch_clip again with job: "<jobId>" to keep waiting — never repeat the original
+ request while it runs (that would queue a second fetch). Nothing is lost: the fetch continues on
+ the editor, and once it has finished the same request finds it cached.` (Amended 2026-09-26 after
+ the review: the editor does not dedupe a running whole-recording job — `archiveSourceVideo` →
+ `runManagedFunction` has no running-job check — so a repeated `full: true` request while it runs
+ queues a second download, and "the next ask finds it cached" was true only once the job had
+ finished.)
- 409: `The editor is in a <platform> rate-limit cooldown — <ceil(cooldownMs/1000)>s remaining. Wait,
then call fetch_clip again. (<error>)`
- 401 / 503: `The editor refused the token (HTTP <s>): <error>. WORKER_TOKEN must equal the value the
diff --git a/plans/release-10.md b/plans/release-10.md
@@ -886,11 +886,9 @@ clause. `wait_seconds` and `full` are not backticked, and `NOT_TOOLS` is unchang
is the only MCP tool that causes a write, and the editor does it".
- **Registration is owed by the operator** (rollout step 2): this machine's `archilyzer` entry has
`"env": {}`, so until it is re-registered `fetch_clip` answers "no editor configured".
+- ~~N2: the `queued` text ends "the next ask finds it cached"~~: taken by the plan owner with the
+ re-read's R2, below.
- **Review nits not taken** (`m-review.md`):
- - N2: the `queued` text still ends "the next ask finds it cached". It is the plan's exact
- text, and whether to add "don't repeat the request" is the plan owner's call. For a window a
- repeat is harmless (the fetch re-checks its cache when it runs). For `full: true` it queues
- a second download.
- N3: a poll-phase disabled 503 gets the token sentence but not "the endpoint is off". Cosmetic.
- N4: the sweep's step says "(or a report needs one)", so extractors could fetch per finding.
Plan text, left for the operator.
@@ -931,7 +929,7 @@ nit N1 taken).
4. **S4: the 60 s client timeout** (`8cf50cf6`).
- (a) The tool description and the `mcp/README.md` row say that a client with a 60 s default
request timeout must raise it or pass `wait_seconds` ≤ 50. The fetch continues on the editor
- either way, and the next call finds it cached.
+ either way; resume it with `job` (wording corrected in the re-read, R2 below).
- (b) **Progress notifications: done, because it was cheap.** The `tools/call` handler now takes
`ctx`. When `ctx.mcpReq._meta.progressToken` is set, `progressNotifier` turns `fetchClip`'s new
`onPoll` hook into one `notifications/progress` per poll that finds the job still waiting
@@ -966,6 +964,29 @@ nit N1 taken).
transport's `env`. The stub editor sees three polls, each with `Bearer tok-proto`, and the client
gets progress `[1, 2]`.
+**Re-read of the fixes** (`m-review.md`: SHIP). The plan owner took R1 and R2 in one commit.
+- **R2: the queued text no longer invites a repeat.** The editor does not dedupe a running
+ whole-recording job (`archiveSourceVideo` → `runManagedFunction` has no running-job check), so
+ "the next ask finds it cached" held only once the job had finished. A repeated `full: true`
+ while it ran queued a second download.
+ - The queued text now reads `Still <status> on the editor (job <jobId>, waited <n>s). Call
+ fetch_clip again with job: "<jobId>" to keep waiting — never repeat the original request while
+ it runs (that would queue a second fetch). Nothing is lost: the fetch continues on the editor,
+ and once it has finished the same request finds it cached.`
+ - The two 60 s clauses (tool description, `mcp/README.md` row) now end: "the fetch continues on
+ the editor either way; resume it with job, and once it has finished the same request finds it
+ cached".
+ - The instructions step already said "call it again with the job it names" and is unchanged.
+ - The plan's queued text in `mcp-fetch-clip.md` is amended to match, with the reason.
+- **R1:** the three real-socket tests carry `{ timeout: 10_000 }`, so a timeout regression fails in
+ 10 s instead of hanging.
+
+| sha | what |
+|---|---|
+| _this_ | R2's queued text and the 60 s clauses; the test that asserts the queued text; R1's test timeouts; the plan's queued text; this paragraph and N2 struck above |
+
+**Gates** (`m-gate4.log`): tsc clean (whole workspace); **mcp 269/269**, 16 s.
+
## Rollout
Nothing is rolled out. The live :3001 editor still runs `0213f6c8` (the pre-brand build); the five