Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit bfb05bd869ef3b14ad314c4fd4858cb14d61c725
parent bbb4ed7e3f018f3e750f897df94dd0da8a8adb90
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Tue, 11 Aug 2026 00:51:11 -0400

Make the sweep prompt refuse a word-split request instead of running it

The rebuild moved Claude Code onto /sweep, but left the nine-argument MCP
prompt in place for form-based clients — and it stays visible in Claude
Code's slash list, still looking like the entry point and still being
shredded by the tokenizer. Routing it through the same parser was not
enough: a parser that defaults a bad value still runs a sweep for something
nobody asked for.

The prompt path is now strict where the tool path stays forgiving. A typed
argument holding prose is the tokenizer's fingerprint, so the prompt answers
with what it detected, the words that survived in the order they were typed
(the declared argument order is exactly the recipe for reassembling them),
and the /sweep line to use instead. A form client only trips it by genuinely
typing a bad value, which deserves the same error.

Its description now leads with the redirect, since that string is what a
Claude Code user actually reads next to the entry.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Diffstat:
Mexport/CHANGELOG.md | 2+-
Mmcp/README.md | 9+++++++++
Mmcp/src/promptRequest.test.ts | 62++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mmcp/src/promptRequest.ts | 63+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mmcp/src/server.ts | 65+++++++++++++++++++++++++++++++++++++++++++++++++++++++++--------
5 files changed, 192 insertions(+), 9 deletions(-)

diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md @@ -3,7 +3,7 @@ ## [0.8.5] - 2026-08-11 - **MCP: the server no longer remembers which corpus you're reading — because remembering it was silently getting it wrong.** `use_source` switched a mutable "active corpus" and persisted the choice to a state file so it survived reconnects. That was the bug. The server registered here runs `--local …/export/public`, but its state file held `{"activeSpec":{"kind":"remote","url":"https://hasanalyzer.pages.dev"}}` from some earlier session — so **every call since had been reading a different archive, and nothing in any result said so**. There is now no active source and nothing is persisted: **every read tool takes its own `source` handle**, and a call that omits it reads the server's startup corpus. The handle *is* the serialised spec in canonical form — `default`, `local:/dir`, `remote:https://site`, `hub:https://hub`, or `hub:https://hub#alpha,beta` for a subset — not an opaque token, so it survives a restart and a human reading one in a transcript knows exactly what was searched. Shorthands (a bare site or hub URL, probed to tell one from the other; a hub member's siteId or title) normalise to canonical and are echoed back. **Every result now ends with `(corpus: <handle>)`** — errors included, implemented once in the dispatch wrapper so a new tool cannot forget it; it is `corpus:` and not `source:` because `- source:` already means "this video's URL" in the output. `use_source` survives one release as an unadvertised alias that resolves a target and tells you the handle to pass; `reset_source` is gone. Leftover state files are inert and can be deleted. Caching moved with it: a source instance is built once per handle and `listChannels` is memoised per instance, which also fixes a pre-existing cost — a 20-id `get_transcripts` batch against a remote used to fetch `corpus.json` twenty times. See `mcp/src/sourceRegistry.ts` (new; `sourceController.ts` deleted), `mcp/src/{server,source,index}.ts`, `mcp/src/sourceRegistry.test.ts`. - **MCP: claiming you covered the corpus is now hard to do by accident.** A sweep returned `total 319; showing 1–200; has_more: yes`, was never paged, and reported **319 videos swept** having seen 200. The paging instruction was already in the sweep prompt and was ignored — so prose is not the enforcement mechanism. Worse, paging was also the *expensive* option: the engine materialises the entire match set and only then slices, so each page re-scanned the whole corpus. New **`enumerate_matches`** returns a query's complete id/title/channel/date worklist plus the batch count in **one scan** — a new tool rather than a flag, because a flag that silently changes the output shape is exactly what gets ignored. If a cap is hit, the **first line** reads `⚠ COVERAGE PARTIAL … this is a SAMPLE, not the full set`, never a quiet footnote. `search_transcripts` now prints `⚠ INCOMPLETE PAGE — N total, showing a–b. Do NOT report a count from this page.` **above** the hits (the old footer stays, so existing consumers keep working). Verified on the real 30-channel corpus: enumerate and search agree exactly at 57, 86 and at the 2000-video cap, where both correctly flag partial coverage. -- **MCP: `/sweep` and `/ask` stop shredding what you type.** Claude Code parses an MCP prompt's arguments as whitespace-splitting zipped against the declared argument names — not quote-aware, the last argument does *not* absorb the remainder, and tokens past the declared count are **dropped silently**. With nine declared arguments, a real request became `link="This"`, `channel="search"`, `group="deleted"`, `directive="videos."`, `batch_size="Why"` — rendered literally into `` ceil(N / Why) `` — and the entire actual question vanished without a warning. The entry points are now **tools** (`sweep_plan`, `ask_plan`) taking one free-text `request`, called by two thin `.claude/commands/` shims that pass `$ARGUMENTS`, the whole raw string, untokenised. A URL with `?a=b&c=d` and a full sentence of punctuation now arrive intact. Settings are still available as `key=value`, but against a **closed whitelist**: an unrecognised `x=y` **stays in the question** and warns (with a typo hint if it's one edit from a real key) instead of being eaten, and every value is validated — `batch_size` an integer 1–20, `parse_model` a single token, `report` a single `.md` path with no `..` — so a bad one becomes a default plus a `⚠` line, never arithmetic. The MCP `sweep` prompt still serves form-based clients (Claude Desktop, Cursor) through the same parser. The plans also **pre-resolve** what they can — the corpus handle, your channel/group tokens validated against the live corpus, the group roster when you gave no scope — turning a three-call preamble into none, and an unresolvable channel now **halts** the plan rather than quietly widening it. See `mcp/src/{promptRequest,instructions}.ts` (new) and their tests. +- **MCP: `/sweep` and `/ask` stop shredding what you type.** Claude Code parses an MCP prompt's arguments as whitespace-splitting zipped against the declared argument names — not quote-aware, the last argument does *not* absorb the remainder, and tokens past the declared count are **dropped silently**. With nine declared arguments, a real request became `link="This"`, `channel="search"`, `group="deleted"`, `directive="videos."`, `batch_size="Why"` — rendered literally into `` ceil(N / Why) `` — and the entire actual question vanished without a warning. The entry points are now **tools** (`sweep_plan`, `ask_plan`) taking one free-text `request`, called by two thin `.claude/commands/` shims that pass `$ARGUMENTS`, the whole raw string, untokenised. A URL with `?a=b&c=d` and a full sentence of punctuation now arrive intact. Settings are still available as `key=value`, but against a **closed whitelist**: an unrecognised `x=y` **stays in the question** and warns (with a typo hint if it's one edit from a real key) instead of being eaten, and every value is validated — `batch_size` an integer 1–20, `parse_model` a single token, `report` a single `.md` path with no `..` — so a bad one becomes a default plus a `⚠` line, never arithmetic. The MCP `sweep` prompt still serves form-based clients (Claude Desktop, Cursor) through the same parser — but since it also still appears in Claude Code's slash list, where it is unusable, it now **refuses** a word-split request instead of sweeping the wrong thing: it names the arguments that cannot be what they claim to be, replays the words that survived in the order they were typed, and hands back the `/sweep` line to use instead. The plans also **pre-resolve** what they can — the corpus handle, your channel/group tokens validated against the live corpus, the group roster when you gave no scope — turning a three-call preamble into none, and an unresolvable channel now **halts** the plan rather than quietly widening it. See `mcp/src/{promptRequest,instructions}.ts` (new) and their tests. - **MCP: three parameters that were declared and silently ignored now work.** `open_link`'s `overrides.query_scope:"posts"` was in the schema, dropped by the parser and excluded by the type, so asking to re-target a link at the social-post corpus quietly searched transcripts instead. `get_transcripts.content_types` was declared and never read, so a batch of post ids came back "not found" — ids now fall through to the post corpus. And a posts search over channels with no posts index reported `total 0; scanned 0 page(s) across 0 channel(s)`, indistinguishable from "searched everything, found nothing"; it now says **"no channel in scope ships a posts index — the post corpus is EMPTY here, not merely unmatched"**, with the posts pass counted separately from the video pass. - **MCP: `open_link` does the whole job in one call, and `get_transcripts` can carry several queries.** `open_link` was preview-then-`apply:true`, where apply *switched the global active source* — one extra round trip and the mutation this release exists to remove. It now decodes, resolves the origin to a handle, searches, and returns plan + results + handle together; `dry_run:true` gets the plan alone. `get_transcripts` gains **`queries`** (up to 8): the windows merge in one pass per video and each header reports a **per-query count**, so a term that matched nothing in that video is visible rather than absorbed — which kills the re-read-per-quote pattern. - **MCP: on the 2026-07-28 protocol.** The server now speaks MCP revision 2026-07-28 via `@modelcontextprotocol/server@2`'s `serveStdio`, which owns the era decision — `modern` (negotiated by `server/discover`) or `legacy` (the 2025 `initialize` handshake) — and serves both from one definition. Confirmed negotiating `modern` in practice, not just in principle. Cache hints are a construction-time policy (`tools/list`, `prompts/list` and `server/discover` are literal constants with no corpus data, so `public` for an hour; every read result stays uncached) and an invalid one throws at startup rather than on the wire. Because `InMemoryTransport` only ever exercises the 2025 era, a new `protocol.test.ts` spawns the real process over stdio and asserts both eras serve an identical, order-pinned tool list. See `mcp/src/protocol.test.ts`. diff --git a/mcp/README.md b/mcp/README.md @@ -181,6 +181,15 @@ intact. The MCP `sweep` prompt still exists for form-based clients (Claude Desktop, Cursor), where each argument gets its own field; it routes through the same parser and validators. +**It cannot be used as the shredding trap it used to be.** In Claude Code that +prompt still appears in the slash list as `/mcp__<server>__sweep`, so it now +**refuses** rather than sweeping the wrong thing: a typed argument holding prose +(`link` = `"search"`, `batch_size` = `"might"`) is the tokenizer's fingerprint, +and the prompt answers with what it detected, the words that survived in the +order they were typed, and the `/sweep` line to use instead. A form client only +trips it by genuinely typing a bad value — which deserves the error too. Note +the tool path stays forgiving (default + `⚠`); only the prompt path is strict. + **Settings.** Append `key=value` for any of `channels=`, `groups=`, `batch_size=`, `parse_model=`, `report=`, `directive=`, `source=`, `content_types=`, `regex=` (quote a multi-word value). The whitelist is closed: diff --git a/mcp/src/promptRequest.test.ts b/mcp/src/promptRequest.test.ts @@ -4,6 +4,7 @@ import { parsePromptRequest, requestFromArguments, renderWarnings, + validateSweepArguments, DEFAULT_BATCH_SIZE, DEFAULT_PARSE_MODEL, } from "./promptRequest"; @@ -263,3 +264,64 @@ test("warnings render as a leading block, and nothing when clean", () => { assert.ok(out.startsWith("⚠ a\n⚠ b")); assert.ok(out.endsWith("\n\n")); }); + +// ─── The prompt form refuses a shredded request ─── + +test("the exact Claude Code shredding of a real request is detected", () => { + // What `text.trim().split(/\s+/)` + zipObject(declaredArgs, tokens) does to + // "This search finds deleted videos. Why might have Rekieta privated these? + // Look for context around each one." — nine words kept, the rest dropped. + const shredded = { + query: "This", + link: "search", + channel: "finds", + channels: "deleted", + group: "videos.", + directive: "Why", + batch_size: "might", + parse_model: "have", + report_path: "Rekieta", + }; + const { problems, reassembled } = validateSweepArguments(shredded); + const named = problems.map((p) => p.arg).sort(); + // "have" in parse_model is not flagged — it has the shape of a model name. + // Three independent signals is already conclusive; the check only has to + // fire, not catch every slot. + assert.deepEqual(named, ["batch_size", "link", "report_path"]); + // The surviving words come back in the order they were typed, so the user + // can recognise their own sentence and see where it was cut. + assert.equal( + reassembled, + "This search finds deleted videos. Why might have Rekieta", + ); +}); + +test("a legitimate form-client invocation trips nothing", () => { + for (const args of [ + { query: "k cups" }, + { query: "k cups", group: "other" }, + { query: "k cups", channels: "chan-a,chan-b", batch_size: "12" }, + { + query: "deleted videos", + link: "https://site.example/?qt=abc&fav=deleted", + parse_model: "haiku", + report_path: "./out.md", + directive: "key claims & contradictions", + }, + ]) { + assert.deepEqual(validateSweepArguments(args).problems, [], JSON.stringify(args)); + } +}); + +test("a genuinely bad value is refused even from a form client", () => { + // The prompt path is strict where the tool path is forgiving: this is the + // value that used to render into `ceil(N / Why)`. + assert.deepEqual( + validateSweepArguments({ query: "k cups", batch_size: "Why" }).problems.map((p) => p.arg), + ["batch_size"], + ); + assert.deepEqual( + validateSweepArguments({ query: "x", report_path: "../secrets.md" }).problems.map((p) => p.arg), + ["report_path"], + ); +}); diff --git a/mcp/src/promptRequest.ts b/mcp/src/promptRequest.ts @@ -349,6 +349,69 @@ export function parsePromptRequest(input: string): PromptRequest { return req; } +// The `sweep` prompt's declared argument names, IN ORDER. Claude Code zips the +// user's whitespace-split words against exactly this list, so the order is also +// the recipe for reassembling what they typed when it has shredded a request. +export const SWEEP_PROMPT_ARGS = [ + "query", + "link", + "channel", + "channels", + "group", + "directive", + "batch_size", + "parse_model", + "report_path", +] as const; + +export type ArgumentProblem = { arg: string; value: string; why: string }; + +// Hard validation for the PROMPT form. The tool path is forgiving (a bad value +// becomes a default plus a ⚠, because failing a whole sweep over a typo is +// worse than running it at 8 and saying so). The prompt path must not be: a +// typed argument holding prose is the fingerprint of Claude Code's tokenizer +// having word-split the request, and continuing would run a sweep for +// something the user never asked for. +export function validateSweepArguments(args: Record<string, unknown>): { + problems: ArgumentProblem[]; + reassembled: string; +} { + const str = (k: string): string | undefined => { + const v = args[k]; + return typeof v === "string" && v.trim() !== "" ? v.trim() : undefined; + }; + const problems: ArgumentProblem[] = []; + const bad = (arg: string, value: string, why: string): void => { + problems.push({ arg, value, why }); + }; + + const link = str("link"); + if (link && !/^<?https?:\/\//i.test(link)) { + bad("link", link, "is not an http(s) URL"); + } + const batch = str("batch_size"); + if (batch) { + const n = Number(batch); + if (!Number.isInteger(n) || n < 1 || n > 20) { + bad("batch_size", batch, "is not an integer between 1 and 20"); + } + } + const report = str("report_path") ?? str("report"); + if (report && (!report.toLowerCase().endsWith(".md") || report.includes(".."))) { + bad("report_path", report, "is not a safe .md path"); + } + const model = str("parse_model"); + if (model && !/^[A-Za-z0-9._-]+$/.test(model)) { + bad("parse_model", model, "is not a model name"); + } + + // Whatever words did land, back in the order they were typed. + const reassembled = SWEEP_PROMPT_ARGS.map((a) => str(a) ?? "") + .filter((v) => v !== "") + .join(" "); + return { problems, reassembled }; +} + // The MCP prompt's declared arguments, routed through the SAME validators, so // a form-based client gets an error naming the offending argument rather than // a rendered `ceil(N / Why)`. diff --git a/mcp/src/server.ts b/mcp/src/server.ts @@ -35,6 +35,7 @@ import { import { parsePromptRequest, requestFromArguments, + validateSweepArguments, DEFAULT_REPORT_PATH, type PromptRequest, } from "./promptRequest"; @@ -2001,14 +2002,16 @@ const PROMPTS: Prompt[] = [ { name: "sweep", description: - "Sweep a query across the corpus (or a chosen group/channels): enumerate " + - "every matching video, batch-read the transcripts, and fold cited, " + - "cross-referenced findings into a markdown report — on plan usage, no " + - "API key. With no scope arg it lists the channel groups and asks you to " + - "pick a group/channels (or confirm 'all') before sweeping. This form is " + - "for clients that give each argument its own field (Claude Desktop, " + - "Cursor); in Claude Code use the `/sweep` command, which passes the " + - "whole line to the sweep_plan tool instead of word-splitting it.", + "IN CLAUDE CODE, USE /sweep INSTEAD — this form's arguments get " + + "word-split by the slash-command tokenizer and everything past the " + + "ninth word is dropped; it will refuse rather than sweep the wrong " + + "thing. This form is for clients that give each argument its own field " + + "(Claude Desktop, Cursor). Sweep a query across the corpus (or a chosen " + + "group/channels): enumerate every matching video, batch-read the " + + "transcripts, and fold cited, cross-referenced findings into a markdown " + + "report — on plan usage, no API key. With no scope arg it lists the " + + "channel groups and asks you to pick a group/channels (or confirm " + + "'all') before sweeping.", arguments: [ { name: "query", @@ -2080,6 +2083,52 @@ function buildSweepPrompt(args: Record<string, unknown>): { description: string; messages: { role: "user"; content: { type: "text"; text: string } }[]; } { + // Refuse a shredded request rather than sweeping for something nobody asked + // for. A typed argument holding prose ("link" = "search", "batch_size" = + // "Why") is the fingerprint of Claude Code's slash-command tokenizer, which + // whitespace-splits this prompt's arguments and drops the overflow. A + // form-based client, where each argument has its own field, only trips this + // by genuinely typing a bad value — which also deserves an error. + const { problems, reassembled } = validateSweepArguments(args); + if (problems.length > 0) { + const named = problems + .map((p) => ` - \`${p.arg}\` = "${p.value}" — ${p.why}`) + .join("\n"); + return { + description: "sweep: the request did not arrive intact", + messages: [ + { + role: "user" as const, + content: { + type: "text" as const, + text: + `Do NOT run a sweep. Tell me, briefly, that the request did not ` + + `arrive intact, and show me this:\n\n` + + `The \`sweep\` prompt received values that cannot be what they ` + + `claim to be:\n${named}\n\n` + + `In Claude Code this almost always means the slash-command ` + + `tokenizer word-split the line: it splits on whitespace, zips ` + + `the words onto this prompt's nine declared arguments in order, ` + + `and **silently drops everything past the ninth word**. It is ` + + `not quote-aware.\n\n` + + (reassembled + ? `The words that survived, in order, were:\n\n ${reassembled}\n\n` + + `Anything after them was discarded.\n\n` + : "") + + `**Use \`/sweep\` instead** — it passes the whole line through ` + + `untouched:\n\n` + + ` /sweep ${reassembled || "<link and/or what to sweep for, in your own words>"}\n\n` + + `(\`/ask\` is the same thing for a question answered in the ` + + `conversation rather than a report file. If \`/sweep\` is not in ` + + `the slash list, the commands live in \`.claude/commands/\` and ` + + `are picked up when a session starts — restart the session; ` + + `restarting the MCP server only reloads its tools.)`, + }, + }, + ], + }; + } + const req = requestFromArguments(args); if (!req.query && !req.link) { throw new Error("sweep requires a query argument (or a link)");