Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit ca1ccc9f52272deb312ec9adcae5ef3d430a4dec
parent cd0ffcddd55c037370572899081ef6f407a99249
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sat, 26 Sep 2026 14:27:09 -0400

docs: fetch_clip is how clip media is fetched; yt-dlp --download-sections is only the no-editor fallback

README: both `claude mcp add` blocks carry the optional ARCHILYZER_EDITOR_URL
and WORKER_TOKEN; "Clips and video" step 3 is fetch_clip (window lands in
clips/, full: true in the saved-video store), the clips-raw claim is gone
from the editor path. mcp/README: the one exception to "writes nothing",
the tool row, and "Still read-only" says the editor writes. AGENTS.md: the
env lines and the clips loop.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
MAGENTS.md | 17+++++++++++++----
MREADME.md | 39+++++++++++++++++++++++++++------------
Mmcp/README.md | 6+++++-
3 files changed, 45 insertions(+), 17 deletions(-)

diff --git a/AGENTS.md b/AGENTS.md @@ -52,11 +52,17 @@ Register the MCP server against a public instance and the corpus is readable ove ```sh claude mcp add archilyzer \ --env TRANSCRIPT_SITE_URL=https://jeralyzer.pages.dev \ + --env ARCHILYZER_EDITOR_URL=http://localhost:3001 \ + --env WORKER_TOKEN=… \ -- pnpm -C "$PWD" --filter yt-dlp-transcript-mcp exec tsx src/index.ts ``` +The two editor lines are optional: they let `fetch_clip` ask a local editor for clip +media (`WORKER_TOKEN` is the editor's own, from `editor/.env`). + `TRANSCRIPT_HUB_URL` federates several sites; `TRANSCRIPT_LOCAL_DIR` reads a local -build off disk. The server never writes to an archive. +build off disk. The server never writes to an archive — `fetch_clip` asks the editor, +and the editor writes. **Register it as `archilyzer`.** `.claude/commands/{ask,sweep}.md` are tracked in git and call `mcp__archilyzer__ask_plan` / `mcp__archilyzer__sweep_plan`. That tool name @@ -74,9 +80,12 @@ written to a file. ## Clips and report-to-video -The high-value loop for a repo with no corpus: point the MCP at a public instance, ask -about a subject, then use `yt-dlp --download-sections` to pull **just the cited -seconds** rather than whole videos. Searching text first is what makes fetching cheap. +The high-value loop: point the MCP at a public instance, ask about a subject, then pull +**just the cited seconds** rather than whole videos. Searching text first is what makes +fetching cheap. Clip media for a cited moment goes through the `fetch_clip` MCP tool +(→ the editor's `POST /api/media/fetch-window`, paced, cookie-aware and provenanced), +never `yt-dlp --download-sections` by hand; that command is only the fallback for a +machine with no editor. `umtool/report-to-video/` renders a cited sweep report to an mp4. What it needs: diff --git a/README.md b/README.md @@ -68,10 +68,13 @@ Claude Code: ```bash claude mcp add archilyzer \ --env TRANSCRIPT_SITE_URL=https://jeralyzer.pages.dev \ + --env ARCHILYZER_EDITOR_URL=http://localhost:3001 \ + --env WORKER_TOKEN=… \ -- pnpm -C /ABS/PATH/TO/this/repo --filter yt-dlp-transcript-mcp exec tsx src/index.ts ``` -Use `TRANSCRIPT_HUB_URL` instead to federate a whole hub of sites, or +The two editor lines are optional: they let `fetch_clip` ask a local editor for clip +media. Use `TRANSCRIPT_HUB_URL` instead to federate a whole hub of sites, or `TRANSCRIPT_LOCAL_DIR` to read a local build off disk. It gives a client full-text search with clickable second-level citations, complete @@ -318,10 +321,15 @@ quietly sampling. pnpm install claude mcp add archilyzer \ --env TRANSCRIPT_SITE_URL=https://jeralyzer.pages.dev \ + --env ARCHILYZER_EDITOR_URL=http://localhost:3001 \ + --env WORKER_TOKEN=… \ -- pnpm -C "$PWD" --filter yt-dlp-transcript-mcp exec tsx src/index.ts claude # then try: /ask what has he said about magic tournaments? ``` +The two editor lines are optional: they let `fetch_clip` ask a local editor for clip +media (`WORKER_TOKEN` is the editor's own). Leave them out for research alone. + > **Register the server as `archilyzer`.** The shipped commands call > `mcp__archilyzer__ask_plan` / `mcp__archilyzer__sweep_plan`, and that tool name > embeds the server name **as you registered it**. Under any other name the commands @@ -338,24 +346,31 @@ no GPU, nothing hosted.** Point `TRANSCRIPT_SITE_URL` at any instance, or use ### Clips and video: download the moment, not the movie This is the part that makes an archive worth more than a search box. A transcript gives -you the **exact second** something was said, and `yt-dlp --download-sections` can fetch -just those seconds. So the loop is: +you the **exact second** something was said, and the editor can fetch just those +seconds. So the loop is: 1. Point the MCP server at an instance and `/ask` or `/sweep` about a subject. 2. Get back citations that resolve to precise moments in real recordings. -3. Pull **just those clips** with yt-dlp — seconds of media, not hours. +3. Ask for **just those clips** with the `fetch_clip` MCP tool — the editor fetches the + window through its paced, cookie-aware, provenanced job; the file lands in the + corpus beside the video (`channels/<slug>/data/<id>/clips/`). Seconds of media, not + hours. `full: true` fetches the whole recording into the saved-video store instead + (it needs a video the editor already knows). 4. Optionally, render them into a finished video. You are never downloading a back catalogue to find a quote. You search text, then fetch the few seconds you actually want. Steps 1–2 need no corpus and no media at all; step 3 -needs `yt-dlp`, and step 4 adds `ffmpeg`/`ffprobe` and **ImageMagick with Pango** for -the chrome. - -The clips **are** kept on disk — they land under the report's own `out/` directory -(`clips-raw/`, `segments/`, `cards/`, and the finished `<slug>.mp4`) and are reused on -a rebuild. "No corpus" means you are not mirroring a channel's back catalogue, not that -nothing is stored: your disk use scales with the clips you actually pull, which for a -report is minutes of video rather than years of it. +needs a local Archilyzer editor (the MCP registered with `ARCHILYZER_EDITOR_URL` and +`WORKER_TOKEN`), and step 4 adds `ffmpeg`/`ffprobe` and **ImageMagick with Pango** for +the chrome. With no editor (a public-only setup) the fallback is running yt-dlp +yourself — `yt-dlp --download-sections` fetches just the cited seconds. + +A window fetched through the editor is kept in the corpus and reused by every later ask +and render; the editor prunes clips by age. A render keeps its own `segments/`, +`cards/` and the finished `<slug>.mp4` under the report's `out/`. "No corpus" means you +are not mirroring a channel's back catalogue, not that nothing is stored: your disk use +scales with the clips you actually pull, which for a report is minutes of video rather +than years of it. ### Rendering a report to video diff --git a/mcp/README.md b/mcp/README.md @@ -7,6 +7,8 @@ Desktop, Cursor, and any other MCP client. It is a **local tool you run yourself**. It does not change the archive: it only reads the site's already-published static JSON shards (`corpus.json` + `transcripts/<slug>/…`), either from disk or over HTTP. Nothing is hosted for you. +The one exception is `fetch_clip`, which asks a local Archilyzer editor to fetch a +clip window; the MCP itself still writes nothing. ## Tools @@ -19,6 +21,7 @@ reads the site's already-published static JSON shards (`corpus.json` + | `get_transcript` | One video's full transcript as clean markdown (metadata + **linked** timestamped captions). | | `get_post` / `get_thread` | One archived social post, or its whole thread. Posts have no timeline — cite them with no `@ mm:ss`. | | `get_video_metadata` | Everything known about one video without the transcript body: metadata, plus **view/like counts, cue count and transcript coverage** (`stats/`), **other archived copies of the same recording** with an explicit timings-aligned verdict (`duplicates.json`), and **AI chapters/tags** where they exist (`digests/`). | +| `fetch_clip` | The media behind a cited moment, **fetched by the local editor** (`POST /api/media/fetch-window`) through its paced, cookie-aware, provenanced job — never a yt-dlp run by hand. Needs `ARCHILYZER_EDITOR_URL` (default `http://localhost:3001`) and `WORKER_TOKEN` (the editor's own) in this server's env; without them it says so and fetches nothing. A window is the cited span ± `pad` (default 3 s), at most 15 min, and lands at `channels/<slug>/data/<id>/clips/`; `full: true` fetches the whole recording into the saved-video store (needs a video the editor already knows). Waits up to `wait_seconds` (default 90, max 300), then returns the job id to resume with `job`. A Rumble embed id is mapped to the editor's slug id through the record's `webpageUrl`, so pass the citing corpus as `source`; a video not in `source` is passed through as cited (known limitation). The file is a read-only corpus artifact. | | `open_link` | Paste an archilyzer viewer **share link** to re-run that exact search here (query tree + every filter, at full fidelity) — plan, results and corpus handle in **one** call. `dry_run:true` for the plan alone. | | `list_sources` | Show the **default** corpus and, with a hub, its member sites as ready-to-paste handles. | | `resolve_source` | Turn a URL or site name into the canonical `source` handle and check it can be read. Changes nothing. | @@ -474,7 +477,8 @@ nothing to reset). **Still read-only.** A `source` handle only changes *which* already-published static shards are read — the same capability the startup flags already grant -this locally-run tool. Nothing is ever written to any corpus. +this locally-run tool. This server never writes to any corpus; `fetch_clip` asks +the editor, and the editor writes. ## Protocol