commit d794fe8acc859f9c00d17d23e7f6ef4fa596f64a
parent 49932bd2ee3f581d06a012ead7fc043359098ce2
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Sun, 4 Oct 2026 16:36:06 -0400
records: report-to-video README on where a clip's source comes from and --no-network; an [Unreleased] bullet
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Diffstat:
2 files changed, 45 insertions(+), 2 deletions(-)
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -1,6 +1,7 @@
# Changelog
## [Unreleased]
+- **A report build cuts from media already on disk before it downloads anything, and `--no-network` makes sure it never does.** For each clip, `build-video.mjs` now looks, in order, in the project's own `out/clips-raw`, in the clip windows the editor fetched into the channel (`channels/<slug>/data/<id>/clips/`), and in a saved whole source video (through the saved-video store's pointer, or a `source-media` file still in the video's folder), and cuts from the first that holds the clip plus its fetch pad; only when none does is the window downloaded. A file that is a link to a drive that is not mounted counts as not there, and the next place is tried. The build prints one line per clip naming where its source came from (`raw-cache`, `corpus-window`, `saved-video`, or a network fetch). With `--no-network`, every clip's source is found before anything is rendered, and if any clip would need a download the build stops at once and lists each one (its position in the timeline, channel, video and the span it needs). umtool's clip bench reads the same three places, so a clip it shows as fetched is one the build cuts from without downloading.
- **umtool's report videos can show a highlighted sentence from a saved article.** `node umtool/report-to-video/shoot-page.mjs --page <saved page.html> --quote "<sentence>" --out <shot.png>` opens a web page saved to disk, finds the sentence in its text, highlights it and saves a PNG of the paragraph that holds it, ready to be a report manifest's `image` entry. `--batch <items.json> --out <dir>` does a list of `{ id, page, quote, context? }` at once and writes `<id>.png` for each plus a `results.json` recording each shot's crop, the matched text and the block it shot. The page is opened offline: nothing is fetched except files saved beside it, and its own scripts do not run unless `--js` is given. The sentence is found whether its quotes and apostrophes are curly or straight, across links and emphasis, and through non-breaking spaces, soft hyphens and line breaks in the page's source. A sentence that is not on the page is listed in `results.json` and on the terminal, and the run ends with an error rather than leaving it out. `--color` sets the highlight; `context` picks one occurrence of a sentence that appears more than once.
- **A clip or whole-recording fetch can name the tallest source video it wants.** The MCP's `fetch_clip` takes `maxHeight`, `fetch-via-editor.mjs` takes `--max-height`, and the editor's fetch endpoint takes `maxHeight`: a whole number of pixels from 144 to 2160; anything else is refused before anything is fetched. A window is fetched at or under that height (720 when none is given, as before). A whole recording asked for at 720 or less is saved as the **Video 720p** quality, and above 720 at the original quality; with no height it follows the channel's, else the global, source video quality, as before. umtool's whole-source fetch from the clip bench now asks at the report's `render.maxHeightSource`. A file already on disk is returned as it is and never fetched again for a different height; the answer now gives its height (a window's is read from the file, a whole recording's from what its persist recorded) and says when it is taller than the height asked for.
- **Persist a list of videos, across channels, to the saved-video store.** `pnpm ops persist-videos --json '{"items":[{"slug":"<channel>","id":"<video id>"}, …]}'` (or `--file list.json` for a long list) re-fetches the source container of each video that is not saved yet, the way **Persist source video** does on a video's page. `"format": "original"` or `"video_720"` picks the quality (default: each channel's own), and `"replace": "above-height"` also re-fetches a saved video whose recorded height is unknown or above that quality — the old file is removed only after the new one is saved, and the saved video keeps its retention class. Downloads run one at a time, one job per channel on that channel's download queue, with the usual gap between them (`"gapMs"` overrides it); `"minFreeMemMb"` holds each download until that much memory is free. A low disk or a rate limit stops the job, and running the same list again picks up where it left off: saved videos are skipped. `"dryRun": true` answers with what a run would do — saved, saved above the height, to fetch, no source URL, unknown — and starts nothing. Jobs show as **Persist videos**, can be drained, and can be retried from /jobs.
diff --git a/umtool/report-to-video/README.md b/umtool/report-to-video/README.md
@@ -108,12 +108,19 @@ Everything is cached by content, so iteration is cheap:
nothing at all. Only a window that escapes every cached file downloads again.
- **Preview one entry** — `--only <id>` builds a single segment and stops.
- **Work offline** — `--skip-fetch` fails loudly instead of downloading, so you
- can confirm you are working entirely from cache.
+ can confirm you are working entirely from cache. `--no-network` is the
+ stronger form: every clip's source is found on disk **before** anything is
+ rendered, and if even one clip would need a download the build refuses at
+ once, listing each such clip as `timeline[<i>] <id> <channel>/<video>
+ <from>–<to>` (the padded span it needs). It governs clip **media** only: cue
+ windows and metadata still come from wherever `--cue-source` says, so add
+ `--cue-source local` to keep those off the network too.
- **Fetch one clip, wide** — `--fetch-only <id> --pad 20` puts a generous window
in the cache without building anything. This is what the umtool clip bench runs,
and containing-window reuse is what makes that fetch double as the build's cache.
- **Force a refetch** — delete `out/clips-raw/`, or pass `--no-reuse` to require an
- exact-window file.
+ exact-window file (the corpus and saved-source tiers below are reuse too, and
+ are skipped).
- **Iterate on the rail** — `--rail-only` re-runs only the rail chain over a cached
`out/<slug>.prerail.mp4`; `--preview <start> <dur>` does the same over a window.
`--no-rail` builds the cut without one. See
@@ -135,6 +142,41 @@ cue and runs the search on to the *next* sentence. Left alone, every re-run grew
the same clip. Hence `EPS` in the end lookup and the 0.05 s deadband on applying a
change.
+### Where a clip's source comes from
+
+Before a clip is fetched, `sources.mjs` looks for its source on disk, in this
+order, and the first place holding the clip's **padded** span (its extent plus
+`fetchPad`, to the same 0.02 s tolerance) wins:
+
+1. **`raw-cache`** — this build's `out/clips-raw/<video>_<from>-<to>.mp4`: the
+ exact-window file first, else the tightest containing one, as above.
+2. **`corpus-window`** — the editor's window cache,
+ `channels/<slug>/data/<id>/clips/<from>-<to>.mp4` (`.mkv`/`.webm` from older
+ fetches; the `.json` sidecar and in-flight `.part.mp4` are ignored), tightest
+ first. This is where umtool's bench and the MCP's `fetch_clip` put windows.
+3. **`saved-video`** — the whole recording: the saved-video store's container
+ through `data/<id>/saved-video.json`, else a `data/<id>/source-media.<ext>`
+ not yet moved there. It counts as the window `[0, duration]`, the duration
+ (and picture height) read by one `ffprobe` — run only when tiers 1 and 2
+ missed.
+
+Only a miss in all three fetches. The channels tree is the one the clip bench
+reads for the project: `provenance.channelsDir`, else a `.shadow-channels/`
+beside the manifest, else `CHANNELS_DIR` / the checkout's `transcripts/channels`.
+A file reached through a link into `media/` (perhaps on another drive) is
+followed; a dangling one — a drive not mounted — is "not here", and the next
+candidate or tier answers. The build logs one line per clip naming the kind it
+used (`--progress ndjson`: the `fetch` event's `source`, which is `"network"`
+for a download, plus `local` and `window` for a hit).
+
+A corpus window or a whole source is wider than the fetch would have been, so
+silence detection measures just the padded span of it — the same seconds a
+fetch would have produced — rather than decoding a fifteen-minute window or a
+three-hour container. A raw-cache window is measured whole, as before, so no
+cached cut moves. umtool's bench lists tiers 2 and 3 through the same module
+(`clipWindowDirs`), with a container's span from its cue doc rather than an
+ffprobe, so what the bench calls fetched is what the build cuts from.
+
## Driven from umtool
The three CLIs are the source of truth and stay usable on their own; umtool drives