commit c88be0c27d824ef9b63511f31b5ba8215660e291
parent 740c745cff5bd9fd1af7f6d13763f2051b3f2fe0
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Sun, 4 Oct 2026 17:24:43 -0400
report-to-video: shoot-page's mark crop in the README and the changelog bullet
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Diffstat:
2 files changed, 14 insertions(+), 3 deletions(-)
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -3,7 +3,7 @@
## [Unreleased]
- **A report video can play a clip that has only sound, and a clip can be a file beside the manifest.** When a clip's source has no picture, `build-video.mjs` plays it under a poster: a card with the clip's channel, title and date, the size of the picture area, with the sound's waveform moving along its foot (`render.audioPoster.waveform: false` keeps it still). The segment matches every other one in size, frame rate and sound, and the header, footer and on-screen deck are drawn over it as over footage. A video's saved sound (`audio.mp3` and the like in its folder) is now a source the build can cut from, after every saved picture: before any download when the clip has no picture to fetch (`"audioOnly": true` on the clip, `"preferLocalAudio": true` in `render`, or a podcast or feed record), and otherwise only when nothing can be downloaded (`--no-network`, `--skip-fetch`, or a record with no page); `--no-network` lists such clips as playing from audio only instead of refusing them. A clip may also give `"src"` (a video or audio file) and `"cues"` (its transcript, either a `transcript.cues.json` or a `parakeet-stitch` transcript), both relative to the manifest, instead of a channel and video: it plays the whole file unless `start`/`end` cut inside it, `resolve-windows.mjs` widens it with those cues, it gets no QR unless it has a `citeUrl`, and a path that leaves the manifest's folder (or an absolute one, without `"allowAbsoluteSrc": true` in `render`), a missing file or an unreadable transcript stops the build before anything runs, naming the clip.
- **A report build cuts from media already on disk before it downloads anything, and `--no-network` makes sure it never does.** For each clip, `build-video.mjs` now looks, in order, in the project's own `out/clips-raw`, in the clip windows the editor fetched into the channel (`channels/<slug>/data/<id>/clips/`), and in a saved whole source video (through the saved-video store's pointer, or a `source-media` file still in the video's folder), and cuts from the first that holds the clip plus its fetch pad; only when none does is the window downloaded. A file that is a link to a drive that is not mounted counts as not there, and the next place is tried. The build prints one line per clip naming where its source came from (`raw-cache`, `corpus-window`, `saved-video`, or a network fetch). With `--no-network`, every clip's source is found before anything is rendered, and if any clip would need a download the build stops at once and lists each one (its position in the timeline, channel, video and the span it needs). umtool's clip bench reads the same three places, so a clip it shows as fetched is one the build cuts from without downloading.
-- **umtool's report videos can show a highlighted sentence from a saved article.** `node umtool/report-to-video/shoot-page.mjs --page <saved page.html> --quote "<sentence>" --out <shot.png>` opens a web page saved to disk, finds the sentence in its text, highlights it and saves a PNG of the paragraph that holds it, ready to be a report manifest's `image` entry. `--batch <items.json> --out <dir>` does a list of `{ id, page, quote, context? }` at once and writes `<id>.png` for each plus a `results.json` recording each shot's crop, the matched text and the block it shot. The page is opened offline: nothing is fetched except files saved beside it, and its own scripts do not run unless `--js` is given. The sentence is found whether its quotes and apostrophes are curly or straight, across links and emphasis, and through non-breaking spaces, soft hyphens and line breaks in the page's source. A sentence that is not on the page is listed in `results.json` and on the terminal, and the run ends with an error rather than leaving it out. `--color` sets the highlight; `context` picks one occurrence of a sentence that appears more than once.
+- **umtool's report videos can show a highlighted sentence from a saved article.** `node umtool/report-to-video/shoot-page.mjs --page <saved page.html> --quote "<sentence>" --out <shot.png>` opens a web page saved to disk, finds the sentence in its text, highlights it and saves a PNG of the paragraph that holds it, ready to be a report manifest's `image` entry. `--batch <items.json> --out <dir>` does a list of `{ id, page, quote, context? }` at once and writes `<id>.png` for each plus a `results.json` recording each shot's crop, the matched text and the block it shot. The page is opened offline: nothing is fetched except files saved beside it, and its own scripts do not run unless `--js` is given. The sentence is found whether its quotes and apostrophes are curly or straight, across links and emphasis, and through non-breaking spaces, soft hyphens and line breaks in the page's source. A sentence that is not on the page is listed in `results.json` and on the terminal, and the run ends with an error rather than leaving it out. `--color` sets the highlight; `context` picks one occurrence of a sentence that appears more than once. On a page where a whole post is one block of paragraphs separated by line breaks, `--crop mark` (or an item's `"crop": "mark"`) shoots only the sentence's own lines and one whole line above and below (`--context-lines` sets how many) instead of the whole post; `results.json` records which crop each shot used.
- **A clip or whole-recording fetch can name the tallest source video it wants.** The MCP's `fetch_clip` takes `maxHeight`, `fetch-via-editor.mjs` takes `--max-height`, and the editor's fetch endpoint takes `maxHeight`: a whole number of pixels from 144 to 2160; anything else is refused before anything is fetched. A window is fetched at or under that height (720 when none is given, as before). A whole recording asked for at 720 or less is saved as the **Video 720p** quality, and above 720 at the original quality; with no height it follows the channel's, else the global, source video quality, as before. umtool's whole-source fetch from the clip bench now asks at the report's `render.maxHeightSource`. A file already on disk is returned as it is and never fetched again for a different height; the answer now gives its height (a window's is read from the file, a whole recording's from what its persist recorded) and says when it is taller than the height asked for.
- **Persist a list of videos, across channels, to the saved-video store.** `pnpm ops persist-videos --json '{"items":[{"slug":"<channel>","id":"<video id>"}, …]}'` (or `--file list.json` for a long list) re-fetches the source container of each video that is not saved yet, the way **Persist source video** does on a video's page. `"format": "original"` or `"video_720"` picks the quality (default: each channel's own), and `"replace": "above-height"` also re-fetches a saved video whose recorded height is unknown or above that quality — the old file is removed only after the new one is saved, and the saved video keeps its retention class. Downloads run one at a time, one job per channel on that channel's download queue, with the usual gap between them (`"gapMs"` overrides it); `"minFreeMemMb"` holds each download until that much memory is free. A low disk or a rate limit stops the job, and running the same list again picks up where it left off: saved videos are skipped. `"dryRun": true` answers with what a run would do — saved, saved above the height, to fetch, no source URL, unknown — and starts nothing. Jobs show as **Persist videos**, can be drained, and can be retried from /jobs.
- **Capture specific X posts: a screenshot of each, and its attached media.** `pnpm ops capture-posts --json '{"slug":"<channel>","ids":["<post id>", …]}'` shoots each post as X shows it, through the connected X profile, and downloads its pictures and videos with gallery-dl, into the channel's `posts-media/<post id>/` beside a `capture.json` that records when, from which URLs, and each file's size and SHA-256. Every id must already be in the channel's posts archive; one that is not is refused by name and nothing runs. `"shots": false` or `"media": false` skips that half, and posts already captured are skipped unless `"force": true`. The job runs on the X queue with a post fetch, so the two never run at once, and waits a random 4–10 seconds before each request to X, as fetches do. A deleted post, or one behind its account's wall (protected, suspended, gone), is recorded as such in the channel's deleted-post record; a post behind a sensitive-media warning is opened and shot. If X asks to log in, or answers "Something went wrong", the job stops at that post and leaves the rest for a later run. Captures are never published: the export does not read them.
diff --git a/umtool/report-to-video/README.md b/umtool/report-to-video/README.md
@@ -662,7 +662,8 @@ node umtool/report-to-video/shoot-page.mjs --batch stills.json --out stills/
```jsonc
// stills.json — `page` is relative to this file
[{ "id": "a01", "page": "saved/article.html", "quote": "…",
- "context": "…" }] // optional: a longer stretch around the quote, to pick one occurrence
+ "context": "…", // optional: a longer stretch around the quote, to pick one occurrence
+ "crop": "mark" }] // optional: this item's --crop
```
- **The page is loaded offline.** Every request is refused except a `file://`
@@ -689,13 +690,23 @@ node umtool/report-to-video/shoot-page.mjs --batch stills.json --out stills/
to that height around the quote, and says `trimmed`. The highlight
(`--color`, `#ffe14d`) never reflows the text: one `<mark>` per text node it
covers, padded vertically only.
+- **`--crop mark` shoots the quote's lines, not the block** (`--crop block` is
+ the default; an item's `crop` overrides either). On a forum-style page the
+ whole post is ONE block of `<br>`-separated paragraphs, so the block crop is a
+ slab of embeds around a two-line quote. The mark crop is the quote's own lines
+ plus `--context-lines` (1) whole lines above and below, stepped in the
+ computed `line-height` of the text the quote sits in from the lines the
+ highlight is on, so a context line is never cut mid-glyph; it spans the
+ block's content box, never leaves it, and adds `--padding`. That padding would
+ show the next line's glyphs cut, so above and below the lines it is painted
+ in the block's own background. `--max-height` does not apply.
- **A miss is never skipped.** It is in `results.json` with its `reason`
(`quote not found`, `context not found`, `quote not found inside its context`,
`page not found`, or the load error), it is summarised on stderr, and the run
exits 1 after shooting every other item.
Each result carries `crop` (the CSS-pixel rectangle, in document coordinates),
-`pixels` (the PNG's size), `matched` (the page's own text of the match, curly
+`cropMode` (`block` or `mark`), `pixels` (the PNG's size), `matched` (the page's own text of the match, curly
quotes and all), `block` (a CSS selector for the block shot) and `blocked`. The
matcher is pure and is what `shoot-page.test.mjs` tests; the one test that
launches Chromium runs only with `SHOOT_PAGE_BROWSER=1`.