Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 5397363fa1ff78a0c1bf031a768b6ed3b71e8b50
parent 087ddf9e6a7a3bd219f32ebc409be27b97359fc9
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sat, 26 Sep 2026 17:25:26 -0400

editor: the Source video card, the changelog bullet and AGENTS.md no longer say the transcript is "kept" — YouTube's subtitles are fetched again first, as on any re-download; a Whisper transcript is not touched, and no audio is extracted beside a transcript (review M1)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
MAGENTS.md | 8+++++---
Meditor/CHANGELOG.md | 2+-
Meditor/app/channels/[slug]/videos/[id]/components/cards/SourceVideoSection.tsx | 5+++--
3 files changed, 9 insertions(+), 6 deletions(-)

diff --git a/AGENTS.md b/AGENTS.md @@ -92,9 +92,11 @@ gets a 404 `Channel "<slug>" not found` on every clip. **`full: true` and "Persist source video" download the source even when a transcript or captions exist** (`forceMedia` on `downloadOneManaged`, release 10 slice N; "Persist kept now" too). On a youtube-handling channel that pass used to be skipped for any video with a -transcript — the job ended `done` with no file. With a transcript on disk the forced pass -moves the container into the saved-video store and does nothing else: no subtitles, no -`audio.<fmt>`, no transcription. **Every rewrite of `metadata.info.json` appends to +transcript or captions — the job ended `done` with no file. With a transcript on disk the +forced pass moves the container into the saved-video store and does nothing else: no +subtitles, no `audio.<fmt>`, no transcription. YouTube's subtitles are fetched again first, +as on any re-download; a Whisper transcript is not touched, and no audio is extracted beside +a transcript. **Every rewrite of `metadata.info.json` appends to `metadata.history.json`** beside it (`common/lib/metadataHistory.ts`): the old and new values of changed content keys (title, description, …), the counters that moved, and formats/thumbnails/caption URLs only by fingerprint; 200 entries, newest last. The video diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -7,7 +7,7 @@ - **`/jobs` says why a job was cancelled at boot.** A job the boot settled shows its reason under its status on `/jobs` and as *Cancelled because* on its own page: for example "server restarted; the scheduler re-derives syncs" or "superseded by a newer queued job (…)". The reason used to be only in the job's log. - **The server log says how often a queued job skips its page refresh.** When a queued job finishes outside any request, the editor skips its page refresh and notes it in the log. The note used to appear once and never again. Now the first one after a quiet spell is logged at once, any more in the next 10 minutes are counted, and one line at the end gives the count, with a running total. - **An agent working through the MCP server asks the editor for a clip instead of running yt-dlp.** The MCP server has a new tool, `fetch_clip`. Given a citation's channel, video id, start and end and a one-line reason, it asks the local editor for that window through `POST /api/media/fetch-window`: the same paced, cookie-aware job umtool uses, which records who asked and why beside the file. It answers with the file's path in the corpus (`channels/<slug>/data/<id>/clips/`). The window is the cited span with 3 seconds either side, at most 15 minutes. `full: true` asks for the whole recording instead, which lands in the saved-video store and needs a video the editor already knows. A Rumble citation's id (the embed id the archive publishes) is mapped to the id the editor names the video's folder by, through the archive record's link. The editor must already archive the channel: pointed at a public site with a fresh editor, every clip gets a 404 `Channel "<slug>" not found`. The tool waits up to 90 seconds by default (at most 300) and otherwise returns the job's id, to wait on with `job`; the fetch carries on in the editor either way. While it waits it sends a progress notification per poll to a client that asks for progress. A client whose requests time out at 60 seconds (the MCP SDK's default) must raise that or pass `wait_seconds` of 50 or less. No request to the editor waits more than 15 seconds. If the editor stops answering mid-fetch, the answer gives the job's id and says not to ask again from scratch. The `/ask` and `/sweep` plans now tell the agent to use it and never to run yt-dlp itself. The MCP needs `ARCHILYZER_EDITOR_URL` and `WORKER_TOKEN` (the editor's own) in its environment, so re-register it with the two `--env` lines in the README; without them the tool says so and fetches nothing. The MCP server itself still writes nothing. The README's `yt-dlp --download-sections` command is now only the fallback for a machine with no editor. -- **"Persist source video" and a whole-recording fetch download the video even when it already has a transcript.** On a channel that takes YouTube's subtitles, the button — and a whole-recording request from umtool or the MCP server's `fetch_clip` with `full: true`, which run the same job — fetched only the subtitles again when the video already had a transcript or captions, and finished with no file. It now downloads the source and moves it into the saved-video store. A transcript already on disk is kept: no audio file is extracted beside it and nothing is transcribed. **Persist kept now** on a channel's Cleanup stage does the same for every kept video, so on such a channel it now downloads each kept video's source. The video page's Source video card offers the button on these channels too; it used to say persistence was for transcribe-handling channels only. +- **"Persist source video" and a whole-recording fetch download the video even when it already has a transcript.** On a channel that takes YouTube's subtitles, the button — and a whole-recording request from umtool or the MCP server's `fetch_clip` with `full: true`, which run the same job — fetched only the subtitles again when the video already had a transcript or captions, and finished with no file. It now downloads the source and moves it into the saved-video store. YouTube's subtitles are fetched again first, as on any re-download; a Whisper transcript is not touched, and no audio is extracted beside a transcript. **Persist kept now** on a channel's Cleanup stage does the same for every kept video, so on such a channel it now downloads each kept video's source. The video page's Source video card offers the button on these channels too; it used to say persistence was for transcribe-handling channels only. - **Each video keeps a history of how its metadata changed at the source.** Every download that rewrites a video's `metadata.info.json` and changes anything in it adds one entry to `metadata.history.json` beside it: the old and new value of each field that changed (title, description, duration, availability, chapters and the rest), the view, like and comment counts that moved, and which of the fields that change on every fetch (format URLs, thumbnails, caption URLs) differed, compared by fingerprint only. A caption language appearing or disappearing counts as a change. The newest 200 entries are kept. The video page shows the history under the description: "Metadata rewritten N× · last … by …: <what changed>", with each entry's old → new values when opened. The history starts with the first rewrite after this update. ## [0.9.0] - 2026-09-26 diff --git a/editor/app/channels/[slug]/videos/[id]/components/cards/SourceVideoSection.tsx b/editor/app/channels/[slug]/videos/[id]/components/cards/SourceVideoSection.tsx @@ -53,8 +53,9 @@ export function SourceVideoSection({ ) : handling === "youtube" ? ( <p className="text-sm text-muted-foreground"> Download this video&apos;s full source container and move it into - the saved-video store. A transcript already on disk is kept, and - no audio is extracted beside it. + the saved-video store. YouTube&apos;s subtitles are fetched again + first, as on any re-download; a Whisper transcript is not touched, + and no audio is extracted beside a transcript. </p> ) : ( <p className="text-sm text-muted-foreground">