Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit dd53e9831b1d30ac760eb80c6173f5cca40bb471
parent e7231d42d93739f222087cbb88c32b133b24018c
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sun, 20 Sep 2026 19:38:49 -0400

docs: record why clips/ is invisible, and where a window's bytes come from

plans/FACTS.md gets the load-bearing rule (a window fetch writes no
metadata.info.json, because buildIndex keys a video's presence on that file)
and the re-verification behind it: every video-dir enumerator ignores clips/
because its predicates are ANCHORED, not because anything was added to exclude
it — with the file:line for each one. reconcileVideoDirs.mergeDir is named as
the one thing that moves a video dir's contents and the one place that needed a
change. Also why CHANNELS_DIR joins READ_ROOTS and not WRITE_ROOTS, and why
WORKER_TOKEN is shared.

umtool's docs say where a window now comes from, what the sidecar records, that
a window is not a download, and what UMTOOL_LOCAL_FETCH=1 is for.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Diffstat:
Mplans/FACTS.md | 67+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mumtool/docs/clip-bench.md | 27+++++++++++++++++++++++----
Mumtool/docs/report-video.md | 43+++++++++++++++++++++++++++++++++++++++++++
3 files changed, 133 insertions(+), 4 deletions(-)

diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -4148,3 +4148,70 @@ never pulls `reader-fs.ts` into a client chunk. has no `runner`, so the operator presses Run. The obvious next step is for the download lane to run it for a channel whose filter has unscanned listed videos — which is exactly the backlog the snapshot already carries. + +--- + +## Clip windows: media sourced for another tool (2026-09-20) + +umtool asks the editor for a clip window instead of running yt-dlp +(`POST /api/media/fetch-window`). Verified while building it: + +- **A window fetch MUST NOT write `metadata.info.json`, and that is the whole + design constraint.** `common/controller/buildIndex.ts:309-315` stats that file + per video dir and `continue`s when it throws — so the file's presence IS a + video's presence in the LMDB index. Thirty seconds of audio would otherwise + put an undownloaded video into the corpus. This is why + `common/ytdlp/fetchWindowManaged.ts` assembles its own argv and goes through + neither `outputArgsForUrl` (`ytdlp/runYtdlp.ts`) nor `downloadOneManaged`'s + prefetch. `editor/e2e/fetch-window.spec.ts` pins it, and the fake yt-dlp's + `--download-sections` branch deliberately writes only the window. +- **`channels/<slug>/data/<id>/clips/` is invisible to every video-dir + enumerator**, and it is invisible because the predicates are ANCHORED, not + because anything was added to exclude it. Re-verified, each one: + - `readVideoFiles` (`common/lib/videoStatus.ts:143-158`) filters a plain + `readdir` with the anchored regexes in `common/lib/mediaFiles.ts` + (`^audio\.(ext)$`, `^source-media\.(ext)$`, `^transcript\.[^.]+\.vtt$`) — + a directory named `clips` matches none. + - `controller/scanCorruptMedia.ts:112` filters the same listing by + EXTENSION (`MEDIA_EXT_SET.has(fileExt(e))`); `clips` has none. + - `controller/pruneSavedVideos.ts:51` and `controller/savedVideoInventory.ts:57` + `readdir` **dataDir** (the video dirs, one level above `clips/`) and are + pointer-driven from there. + - `controller/channelSnapshot.ts` decides undownloaded with + `videoHasAnyArtifact(readVideoFiles(dir))`, so a video dir holding only + `clips/` stays undownloaded. The e2e asserts the count does not move. + None needed a guard. A NEW reader of `data/<id>/` must still ask whether a + subdirectory is a file. +- **`reconcileVideoDirs.mergeDir` is the one thing that MOVES a video dir's + contents**, and it is written for files: a colliding entry is stashed as + `<name>.dup-<srcName>`. It now special-cases `clips` and merges the two + subdirectories, because stashing the directory would hide every window in it + from `listClipWindows` (which reads `clips/` and nothing else). A file present + in both is the same bytes — a window's name IS its span. +- **`umtool/lib/paths.mjs` READ_ROOTS now includes `CHANNELS_DIR`, and + WRITE_ROOTS deliberately does not.** `/api/report/raw` resolves the file it + serves through `resolveInRoots`, so without the read root a corpus window + 400s as "outside the roots". The corpus is production data; nothing in umtool + should be able to render over it. That two-list split is exactly what makes + this safe to widen. +- **`WORKER_TOKEN` is shared with `/api/worker/*` by design.** + `common/lib/workerToken.ts` validates against the env var, never settings, and + a missing token DISABLES the endpoint (503, not 401). The fetch endpoint + spends a source's patience, so it must be no easier to reach than the executor + protocol. It is NOT put in a umtool job step's `env`: `lib/jobs.ts:262` + (`jobView`) echoes a step's env back to the browser. +- **The clip format selector is umtool's, not `downloadFormatPreset`'s.** + `resolveDownloadFormatSelector` answers "what should we archive"; the window + needs "what can ffmpeg cut and re-encode cheaply". Unpinned, yt-dlp picks + VP9+Opus at these heights and `--force-keyframes-at-cuts` then re-encodes + through libvpx-vp9 (27 s to cut a 5 s clip, measured) and writes `.webm`. The + two copies must stay byte-identical or a window one side fetched is invisible + to the other. +- **`WIN_EPS = 0.02` and 2 dp naming are a cross-repo contract.** + `common/lib/clipWindow.ts` and `umtool/report-to-video/build-video.mjs` both + declare them. A request that rounds differently addresses a different file and + the cache misses forever. +- **`SavedVideoPointer.origin` is separate from `keepReason` on purpose.** + `keepReason` stays `override`/`pin` so `pruneSavedVideos` never evicts a + container somebody asked for; `origin` is who asked. `parseSavedVideoPointer` + is tolerant, so a legacy pointer parses unchanged (`clipWindow.test.ts`). diff --git a/umtool/docs/clip-bench.md b/umtool/docs/clip-bench.md @@ -59,10 +59,29 @@ for its video's files and picks by cite) still gets every file, widest first. Dragging past the cached window **clamps** and offers a button. A handle that silently starts a 12-second network fetch is a handle you stop trusting. -The fetch runs the pipeline's own `--fetch-only` path, so the file lands named the -way the build expects, with the same format pin and the same Rumble HLS retry — -and containing-window reuse then makes that generous fetch **be** the build's -cache rather than a second one. +The fetch **asks the editor** (`POST /api/media/fetch-window`), which downloads +the window through its managed path — the channel's cookie policy, its extra +args, the per-platform sleeps, the 429 cooldown and the backoff — and writes it +into the corpus at `channels/<slug>/data/<id>/clips/<from>-<to>.mp4` with a +sidecar saying who asked and why. Containing-window reuse then makes that +generous fetch **be** the build's cache rather than a second one, and because +the corpus is keyed by video it is the *next report's* cache too: `out/clips-raw` +is one project's directory, and a report that cites the same stream used to pay +for the same seconds again. + +The window is not a download. No `metadata.info.json` is written, so the video +stays exactly as undownloaded — and as absent from the archive index — as it +was. See [report-video.md](report-video.md#where-the-bytes-come-from-the-editor-not-yt-dlp). + +`UMTOOL_LOCAL_FETCH=1` runs the pipeline's own `--fetch-only` path instead, for a +machine with no editor to ask. Same file naming, same format pin, same Rumble +HLS retry, same NDJSON events. + +**"Fetch N unfetched clips via the editor"**, above the cut on the project page, +closes the whole `ready N of M` gap in one press — the clips that need judgement +with nothing cached holding them, one at a time, awaiting each before starting +the next. Stop abandons rather than kills: the in-flight fetch is already paid +for, so Stop means *start no more*. Past the source's own duration the handle stops for good. diff --git a/umtool/docs/report-video.md b/umtool/docs/report-video.md @@ -334,6 +334,49 @@ from the manifest alone, and why the clip bench edits stage 2 and shows you stag to the API. The fetched file's own start (`fetchStart`) is the only place a relative number appears, and it is always derived, never stored. +## Where the bytes come from: the editor, not yt-dlp + +A clip window is no longer downloaded by this pipeline. `POST /api/report/fetch` +asks the **editor** (`POST /api/media/fetch-window`), which fetches it through +its managed download path and writes it into the corpus: + +``` +channels/<slug>/data/<videoId>/clips/<from>-<to>.mp4 +channels/<slug>/data/<videoId>/clips/<from>-<to>.json # who asked, and why +``` + +Four things follow, and each is the reason: + +- **Politeness.** The editor owns the channel's cookie policy and its + `ytdlpExtraArgs`, the `--sleep-requests`, the per-platform 429 cooldown and the + backoff that opens it. Running yt-dlp from here had none of them, and a burst + is what trips a bot check. +- **Reuse across reports.** `out/clips-raw` is one project's directory. The + corpus is keyed by video, so a window one report paid for is a window the next + one citing the same stream gets for free. Containment is the predicate, so a + deliberately generous fetch covers its neighbours too. +- **Provenance.** The sidecar records `requestedBy`, the manifest and clip it + was wanted for, the reason in the author's own words, the pad, the size and + the argv. A job log is pruned; the window is not, and a directory of windows + nobody can account for is worse than no directory. +- **It is not a download.** The fetch deliberately writes no + `metadata.info.json`. The archive index keys a video's presence on that file, + so a window must not make an undownloaded video look downloaded — the video's + pipeline state and its absence from the index are unchanged. The editor's + video page says so, in those words, under *Fetched windows*. + +`UMTOOL_LOCAL_FETCH=1` restores the old path (`build-video.mjs --fetch-only`) +for a machine with no editor to ask. The step shape, the argv flags and the +NDJSON events are identical either way, so the job runner, the bench's progress +readout and the cache predicate cannot tell the two apart. + +The editor is named by `ARCHILYZER_EDITOR_URL` (default `http://localhost:3001`) +and asked with the shared `WORKER_TOKEN` — the same secret `/api/worker/*` takes, +because this endpoint spends a source's patience and must be no easier to reach. +Both come from the environment, never from a request and never from a job step's +`env` (which `jobView()` shows the browser). An unset token is an error that +names the variable. + ## Writing a window: the four rules umtool is a *second* writer of a file the CLI also writes. So: