Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit ff900508518f15ce9284ceb80512b4bf639f8334
parent f0c61a10bc43afecc58d18f97b016ba0c0c78ddb
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sat, 26 Sep 2026 14:43:37 -0400

docs: fetch_clip needs an editor that already archives the cited channel (review S2)

The route answers 404 `Channel "<slug>" not found` for a channel with no
dir under the editor's transcripts/, so an MCP pointed at a public site
with a fresh editor gets a 404 on every clip. One sentence each in
README "Clips and video" (whose fallback now also covers an editor that
does not archive the channel), AGENTS.md's clips loop and the mcp/README
row.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
MAGENTS.md | 4+++-
MREADME.md | 7+++++--
Mmcp/README.md | 2+-
3 files changed, 9 insertions(+), 4 deletions(-)

diff --git a/AGENTS.md b/AGENTS.md @@ -85,7 +85,9 @@ The high-value loop: point the MCP at a public instance, ask about a subject, th fetching cheap. Clip media for a cited moment goes through the `fetch_clip` MCP tool (→ the editor's `POST /api/media/fetch-window`, paced, cookie-aware and provenanced), never `yt-dlp --download-sections` by hand; that command is only the fallback for a -machine with no editor. +machine with no editor. The editor fetches only for a channel it already archives (a +`transcripts/channels/<slug>/`), so an MCP pointed at a public site with a fresh editor +gets a 404 `Channel "<slug>" not found` on every clip. `umtool/report-to-video/` renders a cited sweep report to an mp4. What it needs: diff --git a/README.md b/README.md @@ -361,8 +361,11 @@ seconds. So the loop is: You are never downloading a back catalogue to find a quote. You search text, then fetch the few seconds you actually want. Steps 1–2 need no corpus and no media at all; step 3 needs a local Archilyzer editor (the MCP registered with `ARCHILYZER_EDITOR_URL` and -`WORKER_TOKEN`), and step 4 adds `ffmpeg`/`ffprobe` and **ImageMagick with Pango** for -the chrome. With no editor (a public-only setup) the fallback is running yt-dlp +`WORKER_TOKEN`) that already archives the cited channel: the editor fetches only for a +channel it has under its `transcripts/`, so an MCP pointed at a public site with a fresh +editor gets a 404 (`Channel "<slug>" not found`) on every clip. Step 4 adds +`ffmpeg`/`ffprobe` and **ImageMagick with Pango** for the chrome. With no editor (a +public-only setup), or none that archives the channel, the fallback is running yt-dlp yourself — `yt-dlp --download-sections` fetches just the cited seconds. A window fetched through the editor is kept in the corpus and reused by every later ask diff --git a/mcp/README.md b/mcp/README.md @@ -21,7 +21,7 @@ clip window; the MCP itself still writes nothing. | `get_transcript` | One video's full transcript as clean markdown (metadata + **linked** timestamped captions). | | `get_post` / `get_thread` | One archived social post, or its whole thread. Posts have no timeline — cite them with no `@ mm:ss`. | | `get_video_metadata` | Everything known about one video without the transcript body: metadata, plus **view/like counts, cue count and transcript coverage** (`stats/`), **other archived copies of the same recording** with an explicit timings-aligned verdict (`duplicates.json`), and **AI chapters/tags** where they exist (`digests/`). | -| `fetch_clip` | The media behind a cited moment, **fetched by the local editor** (`POST /api/media/fetch-window`) through its paced, cookie-aware, provenanced job — never a yt-dlp run by hand. Needs `ARCHILYZER_EDITOR_URL` (default `http://localhost:3001`) and `WORKER_TOKEN` (the editor's own) in this server's env; without them it says so and fetches nothing. A window is the cited span ± `pad` (default 3 s), at most 15 min, and lands at `channels/<slug>/data/<id>/clips/`; `full: true` fetches the whole recording into the saved-video store (needs a video the editor already knows). Waits up to `wait_seconds` (default 90, max 300), then returns the job id to resume with `job`. A Rumble embed id is mapped to the editor's slug id through the record's `webpageUrl`, so pass the citing corpus as `source`; a video not in `source` is passed through as cited (known limitation). The file is a read-only corpus artifact. | +| `fetch_clip` | The media behind a cited moment, **fetched by the local editor** (`POST /api/media/fetch-window`) through its paced, cookie-aware, provenanced job — never a yt-dlp run by hand. Needs `ARCHILYZER_EDITOR_URL` (default `http://localhost:3001`) and `WORKER_TOKEN` (the editor's own) in this server's env; without them it says so and fetches nothing. The editor must already archive the cited channel (a channel dir under its `transcripts/`), else it answers 404 `Channel "<slug>" not found`: an MCP pointed at a public site with a fresh editor gets that on every clip. A window is the cited span ± `pad` (default 3 s), at most 15 min, and lands at `channels/<slug>/data/<id>/clips/`; `full: true` fetches the whole recording into the saved-video store (needs a video the editor already knows). Waits up to `wait_seconds` (default 90, max 300), then returns the job id to resume with `job`. A Rumble embed id is mapped to the editor's slug id through the record's `webpageUrl`, so pass the citing corpus as `source`; a video not in `source` is passed through as cited (known limitation). The file is a read-only corpus artifact. | | `open_link` | Paste an archilyzer viewer **share link** to re-run that exact search here (query tree + every filter, at full fidelity) — plan, results and corpus handle in **one** call. `dry_run:true` for the plan alone. | | `list_sources` | Show the **default** corpus and, with a hub, its member sites as ready-to-paste handles. | | `resolve_source` | Turn a URL or site name into the canonical `source` handle and check it can be read. Changes nothing. |