Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 40c8de2f769444cafcc13c638748dcdbf2155a4f
parent 1e447c717cf0ea8aca7d50881639451af5f49abe
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sun,  4 Oct 2026 16:28:36 -0400

mcp: fetch_clip takes an optional maxHeight and reports a file's height

maxHeight (144-2160) is checked before any HTTP and sent to the editor with a
window or a whole recording. The answer gives the file's height when the
editor knows it, and says when a cached file, or a whole recording, is
taller than the cap asked for.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
MREADME.md | 3++-
Mmcp/README.md | 2+-
Mmcp/src/fetchClip.test.ts | 79+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mmcp/src/fetchClip.tool.test.ts | 21++++++++++++++++++++-
Mmcp/src/fetchClip.ts | 85++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-------
Mmcp/src/server.ts | 20+++++++++++++++++++-
6 files changed, 199 insertions(+), 11 deletions(-)

diff --git a/README.md b/README.md @@ -362,7 +362,8 @@ seconds. So the loop is: window through its paced, cookie-aware, provenanced job; the file lands in the corpus beside the video (`channels/<slug>/data/<id>/clips/`). Seconds of media, not hours. `full: true` fetches the whole recording into the saved-video store instead - (it needs a video the editor already knows). + (it needs a video the editor already knows). `maxHeight` caps the source height; + at 720 or less a whole recording is saved as the 720p H.264 preset. 4. Optionally, render them into a finished video. You are never downloading a back catalogue to find a quote. You search text, then fetch diff --git a/mcp/README.md b/mcp/README.md @@ -21,7 +21,7 @@ clip window; the MCP itself still writes nothing. | `get_transcript` | One video's full transcript as clean markdown (metadata + **linked** timestamped captions). | | `get_post` / `get_thread` | One archived social post, or its whole thread. Posts have no timeline — cite them with no `@ mm:ss`. | | `get_video_metadata` | Everything known about one video without the transcript body: metadata, plus **view/like counts, cue count and transcript coverage** (`stats/`), **other archived copies of the same recording** with an explicit timings-aligned verdict (`duplicates.json`), and **AI chapters/tags** where they exist (`digests/`). | -| `fetch_clip` | The media behind a cited moment, **fetched by the local editor** (`POST /api/media/fetch-window`) through its paced, cookie-aware, provenanced job — never a yt-dlp run by hand. Needs `ARCHILYZER_EDITOR_URL` (default `http://localhost:3001`) and `WORKER_TOKEN` (the editor's own) in this server's env; without them it says so and fetches nothing. The editor must already archive the cited channel (a channel dir under its `transcripts/`), else it answers 404 `Channel "<slug>" not found`: an MCP pointed at a public site with a fresh editor gets that on every clip. A window is the cited span ± `pad` (default 3 s), at most 15 min, and lands at `channels/<slug>/data/<id>/clips/`; `full: true` fetches the whole recording into the saved-video store (needs a video the editor already knows). Waits up to `wait_seconds` (default 90, max 300), then returns the job id to resume with `job`; a client with a 60 s default request timeout must raise it or pass `wait_seconds` ≤ 50 — the fetch continues on the editor either way; resume it with `job`, and once it has finished the same request finds it cached. While it waits it sends one progress notification per poll to a client that asked for progress (a `progressToken`), which keeps a reset-on-progress timeout alive. A Rumble embed id is mapped to the editor's slug id through the record's `webpageUrl`, so pass the citing corpus as `source`; a video not in `source` is passed through as cited (known limitation). The file is a read-only corpus artifact. | +| `fetch_clip` | The media behind a cited moment, **fetched by the local editor** (`POST /api/media/fetch-window`) through its paced, cookie-aware, provenanced job — never a yt-dlp run by hand. Needs `ARCHILYZER_EDITOR_URL` (default `http://localhost:3001`) and `WORKER_TOKEN` (the editor's own) in this server's env; without them it says so and fetches nothing. The editor must already archive the cited channel (a channel dir under its `transcripts/`), else it answers 404 `Channel "<slug>" not found`: an MCP pointed at a public site with a fresh editor gets that on every clip. A window is the cited span ± `pad` (default 3 s), at most 15 min, and lands at `channels/<slug>/data/<id>/clips/`; `full: true` fetches the whole recording into the saved-video store (needs a video the editor already knows). `maxHeight` (144–2160) caps the source height: a window is fetched at or under it (default 720); a whole recording at 720 or less is saved as the editor's 720p H.264 preset and above 720 at the original quality (omitted, the channel's source-video quality applies). A file already on disk is returned as it is, never re-fetched for a different cap, and the answer gives its height and says when it is taller than asked. Waits up to `wait_seconds` (default 90, max 300), then returns the job id to resume with `job`; a client with a 60 s default request timeout must raise it or pass `wait_seconds` ≤ 50 — the fetch continues on the editor either way; resume it with `job`, and once it has finished the same request finds it cached. While it waits it sends one progress notification per poll to a client that asked for progress (a `progressToken`), which keeps a reset-on-progress timeout alive. A Rumble embed id is mapped to the editor's slug id through the record's `webpageUrl`, so pass the citing corpus as `source`; a video not in `source` is passed through as cited (known limitation). The file is a read-only corpus artifact. | | `open_link` | Paste an archilyzer viewer **share link** to re-run that exact search here (query tree + every filter, at full fidelity) — plan, results and corpus handle in **one** call. `dry_run:true` for the plan alone. | | `list_sources` | Show the **default** corpus and, with a hub, its member sites as ready-to-paste handles. | | `resolve_source` | Turn a URL or site name into the canonical `source` handle and check it can be read. Changes nothing. | diff --git a/mcp/src/fetchClip.test.ts b/mcp/src/fetchClip.test.ts @@ -744,3 +744,82 @@ test("onPoll is told after every poll that finds the job still waiting, and cann assert.deepEqual(seen, ["1 j1 queued 1s", "2 j1 running 2s"]); assert.equal(outcome.kind, "fetched"); }); + +// ─── maxHeight ─── + +test("maxHeight: a whole number from 144 to 2160 rides on the target; anything else is refused", () => { + const base = { channel: "chan", video: "vid", start: 10, end: 20, reason: "why" }; + for (const bad of [143, 2161, 720.5, "720", 0]) { + const v = validateFetchClipArgs({ ...base, maxHeight: bad }); + assert.equal(v.ok, false); + assert.equal( + v.ok ? "" : v.error, + `fetch_clip: maxHeight "${String(bad)}" must be a whole number of pixels from 144 to 2160`, + ); + } + const win = validateFetchClipArgs({ ...base, maxHeight: 480 }); + assert.ok(win.ok && "target" in win.request); + assert.equal(win.ok && "target" in win.request ? win.request.target.maxHeight : null, 480); + const full = validateFetchClipArgs({ ...base, full: true, maxHeight: 1080 }); + assert.ok(full.ok && "target" in full.request); + assert.deepEqual(full.ok && "target" in full.request ? full.request.target : null, { + kind: "full", + channel: "chan", + video: "vid", + maxHeight: 1080, + }); + // null is "not given", as JSON clients often send it. + const none = validateFetchClipArgs({ ...base, maxHeight: null }); + assert.ok(none.ok && "target" in none.request); + assert.equal(none.ok && "target" in none.request ? "maxHeight" in none.request.target : true, false); +}); + +test("maxHeight is sent in the POST for a window and for a whole recording", async () => { + for (const req of [windowRequest({ maxHeight: 480 }), fullRequest({ maxHeight: 720 })]) { + const { deps, calls } = editor( + seq({ status: 202, body: { cached: false, jobId: "j3", file: null, from: 0, to: 0 } }), + ); + await fetchClip({ ...req, waitSeconds: 0 } as FetchClipRequest, deps); + const body = calls[0].body as Record<string, unknown>; + assert.equal(body.maxHeight, "target" in req && req.target.kind === "full" ? 720 : 480); + } +}); + +test("a cached window's height is shown, and one taller than maxHeight is said to be", async () => { + const cached = (height: number) => + editor( + seq({ + status: 200, + body: { cached: true, file: CLIP_FILE, from: 7, to: 23, bytes: 10, provenance: null, height }, + }), + ).deps; + const fits = renderFetchClip(await fetchClip(windowRequest({ maxHeight: 720 }), cached(720)), CTX); + assert.match(fits.text, /\nheight: 720p\n/); + const tall = renderFetchClip(await fetchClip(windowRequest({ maxHeight: 480 }), cached(1080)), CTX); + assert.equal(tall.isError, false); + assert.match( + tall.text, + /\nheight: 1080p — taller than the maxHeight 480 asked for; this file was fetched earlier and is served as it is\n/, + ); + // No cap asked: the height is still given, with no comparison. + const plain = renderFetchClip(await fetchClip(windowRequest(), cached(1080)), CTX); + assert.match(plain.text, /\nheight: 1080p\n/); + // An editor that reports no height: no line at all. + const { deps } = editor( + seq({ status: 200, body: { cached: true, file: CLIP_FILE, from: 7, to: 23, bytes: 10, provenance: null } }), + ); + assert.ok(!/height:/.test(renderFetchClip(await fetchClip(windowRequest(), deps), CTX).text)); +}); + +test("full: a finished recording taller than maxHeight says why", async () => { + const saved = "/corpus/saved-videos/chan/vid/source.mp4"; + const { deps } = editor( + seq( + { status: 202, body: { cached: false, jobId: "j4", file: null, from: 0, to: 0 } }, + { status: 200, body: { status: "done", jobId: "j4", file: saved, bytes: 42, height: 1080 } }, + ), + ); + const r = renderFetchClip(await fetchClip(fullRequest({ maxHeight: 720 }), deps), CTX); + assert.equal(r.isError, false); + assert.match(r.text, /\nheight: 1080p — taller than the maxHeight 720 asked for; a whole recording is saved at 720p/); +}); diff --git a/mcp/src/fetchClip.tool.test.ts b/mcp/src/fetchClip.tool.test.ts @@ -265,7 +265,7 @@ test("the tool is advertised with job-only calls allowed", async () => { assert.ok(tool, "fetch_clip is listed"); assert.deepEqual(tool.inputSchema.required ?? [], []); const props = Object.keys(tool.inputSchema.properties ?? {}); - for (const p of ["source", "channel", "video", "start", "end", "pad", "full", "reason", "report", "wait_seconds", "job"]) { + for (const p of ["source", "channel", "video", "start", "end", "pad", "full", "maxHeight", "reason", "report", "wait_seconds", "job"]) { assert.ok(props.includes(p), `has ${p}`); } assert.match(tool.description ?? "", /NEVER run yt-dlp/); @@ -297,3 +297,22 @@ test("a client that asks for progress gets one notification per poll", async () { progress: 2, message: "editor job j6: running, 2s waited" }, ]); }); + +test("maxHeight goes to the editor as given, and a bad one is refused before any HTTP", async () => { + const { deps, calls } = fakeEditor(); + const client = await connect(deps); + await client.callTool({ + name: "fetch_clip", + arguments: { channel: RUMBLE, video: "vxe1ae", full: true, maxHeight: 720, reason: "summary" }, + }); + assert.equal((calls[0].body as Record<string, unknown>).maxHeight, 720); + + const bad = fakeEditor(); + const res = await (await connect(bad.deps)).callTool({ + name: "fetch_clip", + arguments: { channel: RUMBLE, video: "vxe1ae", start: 1, end: 5, maxHeight: 4320, reason: "why" }, + }); + assert.equal(isError(res), true); + assert.match(textOf(res), /^fetch_clip: maxHeight "4320" must be a whole number of pixels from 144 to 2160/); + assert.equal(bad.calls.length, 0); +}); diff --git a/mcp/src/fetchClip.ts b/mcp/src/fetchClip.ts @@ -1,4 +1,9 @@ -import { MAX_CLIP_WINDOW_SECONDS } from "yt-dlp-transcript-common/lib/clipWindow"; +import { + MAX_CLIP_WINDOW_SECONDS, + MAX_FETCH_MAX_HEIGHT, + MIN_FETCH_MAX_HEIGHT, + isFetchMaxHeight, +} from "yt-dlp-transcript-common/lib/clipWindow"; // ─── fetch_clip: ask the local Archilyzer editor for a clip's media ─── // @@ -142,6 +147,9 @@ export function planWindow(a: { // ─── Arguments ─── +// `maxHeight`, when the caller gave one: the source height to cap the fetch +// at. Absent, the editor's own default applies (720 for a window; the +// channel's, else the global, source-video quality for a whole recording). export type ClipTarget = | { kind: "window"; @@ -151,8 +159,9 @@ export type ClipTarget = from: number; to: number; pad: number; + maxHeight?: number; } - | { kind: "full"; channel: string; video: string }; + | { kind: "full"; channel: string; video: string; maxHeight?: number }; export type FetchClipRequest = | { job: string; waitSeconds: number } @@ -199,9 +208,24 @@ export function validateFetchClipArgs( return { ok: false, error: `fetch_clip: video "${video}" must match /^[\\w.-]+$/` }; } + // Checked here rather than passed through: it ends up inside a yt-dlp + // format selector on the editor, which refuses it too, but a refusal before + // any HTTP says what to fix in the tool's own words. + const rawHeight = args.maxHeight; + if (rawHeight !== undefined && rawHeight !== null && !isFetchMaxHeight(rawHeight)) { + return { + ok: false, + error: + `fetch_clip: maxHeight "${String(rawHeight)}" must be a whole number ` + + `of pixels from ${MIN_FETCH_MAX_HEIGHT} to ${MAX_FETCH_MAX_HEIGHT}`, + }; + } + const maxHeight = isFetchMaxHeight(rawHeight) ? rawHeight : undefined; + const capped = maxHeight !== undefined ? { maxHeight } : {}; + let target: ClipTarget; if (args.full === true) { - target = { kind: "full", channel, video }; + target = { kind: "full", channel, video, ...capped }; } else { const times: Record<"start" | "end", number> = { start: 0, end: 0 }; for (const key of ["start", "end"] as const) { @@ -234,7 +258,7 @@ export function validateFetchClipArgs( } const w = planWindow({ video, start: times.start, end: times.end, pad }); if ("error" in w) return { ok: false, error: w.error }; - target = { kind: "window", channel, video, from: w.from, to: w.to, pad }; + target = { kind: "window", channel, video, from: w.from, to: w.to, pad, ...capped }; } const reason = trimmed(args.reason); @@ -265,6 +289,10 @@ export type FetchClipOutcome = to: number; bytes: number; requestedBy?: string; + // How tall the file on disk is, when the editor knows; and the cap this + // call asked for, so the answer can say when a cached file is taller. + height?: number; + maxHeight?: number; } | { kind: "fetched"; @@ -277,6 +305,8 @@ export type FetchClipOutcome = to?: number; bytes?: number; waited: number; + height?: number; + maxHeight?: number; } | { kind: "queued"; jobId: string; status: string; waited: number } | { kind: "cooldown"; platform: string; cooldownMs: number; error: string } @@ -304,6 +334,18 @@ export type FetchClipOutcome = type Json = Record<string, unknown>; +// The two height fields of an outcome, each only when there is one — so an +// answer from an editor that reports no height reads exactly as it did. +function heightFields( + height: number | undefined, + maxHeight: number | undefined, +): { height?: number; maxHeight?: number } { + return { + ...(height !== undefined ? { height } : {}), + ...(maxHeight !== undefined ? { maxHeight } : {}), + }; +} + function isTimeout(e: unknown): boolean { return typeof e === "object" && e !== null && (e as { name?: unknown }).name === "TimeoutError"; } @@ -386,6 +428,7 @@ export async function fetchClip( jobId: string, mode: "window" | "full" | "unknown", pollFirst: boolean, + maxHeight?: number, ): Promise<FetchClipOutcome> { let status = "queued"; let skipSleep = pollFirst; @@ -433,6 +476,7 @@ export async function fetchClip( to, bytes: num(body.bytes), waited: waited(), + ...heightFields(num(body.height), maxHeight), }; } if (status === "failed" || status === "cancelled") { @@ -463,12 +507,14 @@ export async function fetchClip( // `full === true` before it validates from/to, and the saved-video path // resolves the URL itself, so a webpageUrl would describe a request the // editor does not have. + const capped = target.maxHeight !== undefined ? { maxHeight: target.maxHeight } : {}; const body = target.kind === "full" ? { channelSlug: target.channel, videoId: target.video, full: true, + ...capped, ...provenance, } : { @@ -478,6 +524,7 @@ export async function fetchClip( from: target.from, to: target.to, pad: target.pad, + ...capped, ...provenance, }; const res = await call(`${editor.url}/api/media/fetch-window`, { @@ -509,10 +556,11 @@ export async function fetchClip( to: num(answer.to) ?? reqTo, bytes: num(answer.bytes) ?? 0, requestedBy: prov && typeof prov === "object" ? str(prov.requestedBy) : undefined, + ...heightFields(num(answer.height), target.maxHeight), }; } if (res.status === 202 && str(answer.jobId)) { - return waitFor(str(answer.jobId)!, target.kind, false); + return waitFor(str(answer.jobId)!, target.kind, false, target.maxHeight); } if (res.status === 409) { return { @@ -539,7 +587,16 @@ const READ_ONLY_NOTE = function footer( mode: "window" | "full", - f: { file: string; from?: number; to?: number; bytes?: number; requestedBy?: string }, + f: { + file: string; + from?: number; + to?: number; + bytes?: number; + requestedBy?: string; + height?: number; + maxHeight?: number; + }, + cached = false, ): string { const lines = [`file: ${f.file}`]; if (mode === "window" && f.from !== undefined && f.to !== undefined) { @@ -547,6 +604,20 @@ function footer( `window: ${fmtSeconds(f.from)}–${fmtSeconds(f.to)} (${fmtSpan(f.to - f.from)})`, ); } + if (f.height !== undefined) { + // TALLER THAN ASKED is said, not hidden. A cached file was fetched for an + // earlier ask and is served as it is; a whole recording is saved at one of + // two qualities, not at the exact height asked for. + const over = + f.maxHeight !== undefined && f.height > f.maxHeight + ? ` — taller than the maxHeight ${f.maxHeight} asked for` + + (cached + ? "; this file was fetched earlier and is served as it is" + : "; a whole recording is saved at 720p (falling back to what " + + "the source has) or at its original quality, not at an exact height") + : ""; + lines.push(`height: ${f.height}p${over}`); + } if (f.bytes !== undefined) lines.push(`bytes: ${f.bytes}`); if (f.requestedBy) lines.push(`requested by ${f.requestedBy}`); if (mode === "window") { @@ -599,7 +670,7 @@ export function renderFetchClip( `${fmtSeconds(outcome.reqFrom)}–${fmtSeconds(outcome.reqTo)}: ` + `${fmtSeconds(outcome.from)}–${fmtSeconds(outcome.to)}.`; } - return { text: `${head}\n\n${footer(outcome.mode, outcome)}`, isError: false }; + return { text: `${head}\n\n${footer(outcome.mode, outcome, true)}`, isError: false }; } case "fetched": { if (!outcome.file) { diff --git a/mcp/src/server.ts b/mcp/src/server.ts @@ -724,7 +724,11 @@ export const TOOLS: Tool[] = [ "yourself instead. Give the citation's channel slug, video id, start " + "and end, and a one-line reason; the window is padded (pad, default 3 " + "s) and may be at most 15 min. full: true fetches the whole recording " + - "instead, into the editor's saved-video store. The answer names the " + + "instead, into the editor's saved-video store. maxHeight caps the " + + "source height (a window defaults to 720; a whole recording at or under " + + "720 is saved as 720p H.264, above it at the original quality); a " + + "cached file is served as it is, and the answer gives its height. The " + + "answer names the " + "file on disk: a read-only corpus artifact to play or copy, never to " + "move, edit or delete. The call waits up to wait_seconds; if the fetch " + "is still running it returns the job id — call again with job to keep " + @@ -773,6 +777,20 @@ export const TOOLS: Tool[] = [ "and needs a video the editor already knows (its metadata or " + "playlist entry). start, end and pad are ignored.", }, + maxHeight: { + type: "integer", + minimum: 144, + maximum: 2160, + description: + "Optional: the tallest source video to fetch, in pixels (144–2160). " + + "A window is fetched at or under it (default 720, H.264 preferred). " + + "With full: true, 720 or less saves the 720p H.264 preset (480p, " + + "then whatever the source has, when it has no 720p H.264) and more " + + "than 720 saves the original; omitted, the channel's own " + + "source-video quality applies. Nothing is re-fetched for it: a " + + "file already on disk is returned as it is, and the answer gives " + + "its height, so a taller one can be seen.", + }, reason: { type: "string", description: