commit ca1ccc9f52272deb312ec9adcae5ef3d430a4dec
parent cd0ffcddd55c037370572899081ef6f407a99249
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Sat, 26 Sep 2026 14:27:09 -0400
docs: fetch_clip is how clip media is fetched; yt-dlp --download-sections is only the no-editor fallback
README: both `claude mcp add` blocks carry the optional ARCHILYZER_EDITOR_URL
and WORKER_TOKEN; "Clips and video" step 3 is fetch_clip (window lands in
clips/, full: true in the saved-video store), the clips-raw claim is gone
from the editor path. mcp/README: the one exception to "writes nothing",
the tool row, and "Still read-only" says the editor writes. AGENTS.md: the
env lines and the clips loop.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Diffstat:
3 files changed, 45 insertions(+), 17 deletions(-)
diff --git a/AGENTS.md b/AGENTS.md
@@ -52,11 +52,17 @@ Register the MCP server against a public instance and the corpus is readable ove
```sh
claude mcp add archilyzer \
--env TRANSCRIPT_SITE_URL=https://jeralyzer.pages.dev \
+ --env ARCHILYZER_EDITOR_URL=http://localhost:3001 \
+ --env WORKER_TOKEN=… \
-- pnpm -C "$PWD" --filter yt-dlp-transcript-mcp exec tsx src/index.ts
```
+The two editor lines are optional: they let `fetch_clip` ask a local editor for clip
+media (`WORKER_TOKEN` is the editor's own, from `editor/.env`).
+
`TRANSCRIPT_HUB_URL` federates several sites; `TRANSCRIPT_LOCAL_DIR` reads a local
-build off disk. The server never writes to an archive.
+build off disk. The server never writes to an archive — `fetch_clip` asks the editor,
+and the editor writes.
**Register it as `archilyzer`.** `.claude/commands/{ask,sweep}.md` are tracked in git
and call `mcp__archilyzer__ask_plan` / `mcp__archilyzer__sweep_plan`. That tool name
@@ -74,9 +80,12 @@ written to a file.
## Clips and report-to-video
-The high-value loop for a repo with no corpus: point the MCP at a public instance, ask
-about a subject, then use `yt-dlp --download-sections` to pull **just the cited
-seconds** rather than whole videos. Searching text first is what makes fetching cheap.
+The high-value loop: point the MCP at a public instance, ask about a subject, then pull
+**just the cited seconds** rather than whole videos. Searching text first is what makes
+fetching cheap. Clip media for a cited moment goes through the `fetch_clip` MCP tool
+(→ the editor's `POST /api/media/fetch-window`, paced, cookie-aware and provenanced),
+never `yt-dlp --download-sections` by hand; that command is only the fallback for a
+machine with no editor.
`umtool/report-to-video/` renders a cited sweep report to an mp4. What it needs:
diff --git a/README.md b/README.md
@@ -68,10 +68,13 @@ Claude Code:
```bash
claude mcp add archilyzer \
--env TRANSCRIPT_SITE_URL=https://jeralyzer.pages.dev \
+ --env ARCHILYZER_EDITOR_URL=http://localhost:3001 \
+ --env WORKER_TOKEN=… \
-- pnpm -C /ABS/PATH/TO/this/repo --filter yt-dlp-transcript-mcp exec tsx src/index.ts
```
-Use `TRANSCRIPT_HUB_URL` instead to federate a whole hub of sites, or
+The two editor lines are optional: they let `fetch_clip` ask a local editor for clip
+media. Use `TRANSCRIPT_HUB_URL` instead to federate a whole hub of sites, or
`TRANSCRIPT_LOCAL_DIR` to read a local build off disk.
It gives a client full-text search with clickable second-level citations, complete
@@ -318,10 +321,15 @@ quietly sampling.
pnpm install
claude mcp add archilyzer \
--env TRANSCRIPT_SITE_URL=https://jeralyzer.pages.dev \
+ --env ARCHILYZER_EDITOR_URL=http://localhost:3001 \
+ --env WORKER_TOKEN=… \
-- pnpm -C "$PWD" --filter yt-dlp-transcript-mcp exec tsx src/index.ts
claude # then try: /ask what has he said about magic tournaments?
```
+The two editor lines are optional: they let `fetch_clip` ask a local editor for clip
+media (`WORKER_TOKEN` is the editor's own). Leave them out for research alone.
+
> **Register the server as `archilyzer`.** The shipped commands call
> `mcp__archilyzer__ask_plan` / `mcp__archilyzer__sweep_plan`, and that tool name
> embeds the server name **as you registered it**. Under any other name the commands
@@ -338,24 +346,31 @@ no GPU, nothing hosted.** Point `TRANSCRIPT_SITE_URL` at any instance, or use
### Clips and video: download the moment, not the movie
This is the part that makes an archive worth more than a search box. A transcript gives
-you the **exact second** something was said, and `yt-dlp --download-sections` can fetch
-just those seconds. So the loop is:
+you the **exact second** something was said, and the editor can fetch just those
+seconds. So the loop is:
1. Point the MCP server at an instance and `/ask` or `/sweep` about a subject.
2. Get back citations that resolve to precise moments in real recordings.
-3. Pull **just those clips** with yt-dlp — seconds of media, not hours.
+3. Ask for **just those clips** with the `fetch_clip` MCP tool — the editor fetches the
+ window through its paced, cookie-aware, provenanced job; the file lands in the
+ corpus beside the video (`channels/<slug>/data/<id>/clips/`). Seconds of media, not
+ hours. `full: true` fetches the whole recording into the saved-video store instead
+ (it needs a video the editor already knows).
4. Optionally, render them into a finished video.
You are never downloading a back catalogue to find a quote. You search text, then fetch
the few seconds you actually want. Steps 1–2 need no corpus and no media at all; step 3
-needs `yt-dlp`, and step 4 adds `ffmpeg`/`ffprobe` and **ImageMagick with Pango** for
-the chrome.
-
-The clips **are** kept on disk — they land under the report's own `out/` directory
-(`clips-raw/`, `segments/`, `cards/`, and the finished `<slug>.mp4`) and are reused on
-a rebuild. "No corpus" means you are not mirroring a channel's back catalogue, not that
-nothing is stored: your disk use scales with the clips you actually pull, which for a
-report is minutes of video rather than years of it.
+needs a local Archilyzer editor (the MCP registered with `ARCHILYZER_EDITOR_URL` and
+`WORKER_TOKEN`), and step 4 adds `ffmpeg`/`ffprobe` and **ImageMagick with Pango** for
+the chrome. With no editor (a public-only setup) the fallback is running yt-dlp
+yourself — `yt-dlp --download-sections` fetches just the cited seconds.
+
+A window fetched through the editor is kept in the corpus and reused by every later ask
+and render; the editor prunes clips by age. A render keeps its own `segments/`,
+`cards/` and the finished `<slug>.mp4` under the report's `out/`. "No corpus" means you
+are not mirroring a channel's back catalogue, not that nothing is stored: your disk use
+scales with the clips you actually pull, which for a report is minutes of video rather
+than years of it.
### Rendering a report to video
diff --git a/mcp/README.md b/mcp/README.md
@@ -7,6 +7,8 @@ Desktop, Cursor, and any other MCP client.
It is a **local tool you run yourself**. It does not change the archive: it only
reads the site's already-published static JSON shards (`corpus.json` +
`transcripts/<slug>/…`), either from disk or over HTTP. Nothing is hosted for you.
+The one exception is `fetch_clip`, which asks a local Archilyzer editor to fetch a
+clip window; the MCP itself still writes nothing.
## Tools
@@ -19,6 +21,7 @@ reads the site's already-published static JSON shards (`corpus.json` +
| `get_transcript` | One video's full transcript as clean markdown (metadata + **linked** timestamped captions). |
| `get_post` / `get_thread` | One archived social post, or its whole thread. Posts have no timeline — cite them with no `@ mm:ss`. |
| `get_video_metadata` | Everything known about one video without the transcript body: metadata, plus **view/like counts, cue count and transcript coverage** (`stats/`), **other archived copies of the same recording** with an explicit timings-aligned verdict (`duplicates.json`), and **AI chapters/tags** where they exist (`digests/`). |
+| `fetch_clip` | The media behind a cited moment, **fetched by the local editor** (`POST /api/media/fetch-window`) through its paced, cookie-aware, provenanced job — never a yt-dlp run by hand. Needs `ARCHILYZER_EDITOR_URL` (default `http://localhost:3001`) and `WORKER_TOKEN` (the editor's own) in this server's env; without them it says so and fetches nothing. A window is the cited span ± `pad` (default 3 s), at most 15 min, and lands at `channels/<slug>/data/<id>/clips/`; `full: true` fetches the whole recording into the saved-video store (needs a video the editor already knows). Waits up to `wait_seconds` (default 90, max 300), then returns the job id to resume with `job`. A Rumble embed id is mapped to the editor's slug id through the record's `webpageUrl`, so pass the citing corpus as `source`; a video not in `source` is passed through as cited (known limitation). The file is a read-only corpus artifact. |
| `open_link` | Paste an archilyzer viewer **share link** to re-run that exact search here (query tree + every filter, at full fidelity) — plan, results and corpus handle in **one** call. `dry_run:true` for the plan alone. |
| `list_sources` | Show the **default** corpus and, with a hub, its member sites as ready-to-paste handles. |
| `resolve_source` | Turn a URL or site name into the canonical `source` handle and check it can be read. Changes nothing. |
@@ -474,7 +477,8 @@ nothing to reset).
**Still read-only.** A `source` handle only changes *which* already-published
static shards are read — the same capability the startup flags already grant
-this locally-run tool. Nothing is ever written to any corpus.
+this locally-run tool. This server never writes to any corpus; `fetch_clip` asks
+the editor, and the editor writes.
## Protocol