Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 7e814d3bce49b921958265d2695c4b1126d27f92
parent 151ca982e565ef49dd2d4fb4123a449136dde3b3
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sun,  4 Oct 2026 17:25:09 -0400

Merge shoot-page-crop (shoot-page --crop mark: the quote's lines plus whole context lines, not the whole block)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
Meditor/CHANGELOG.md | 2+-
Mumtool/report-to-video/README.md | 15+++++++++++++--
Mumtool/report-to-video/shoot-page.mjs | 157++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-------------
Mumtool/report-to-video/shoot-page.test.mjs | 64+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-
4 files changed, 209 insertions(+), 29 deletions(-)

diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -3,7 +3,7 @@ ## [Unreleased] - **A report video can play a clip that has only sound, and a clip can be a file beside the manifest.** When a clip's source has no picture, `build-video.mjs` plays it under a poster: a card with the clip's channel, title and date, the size of the picture area, with the sound's waveform moving along its foot (`render.audioPoster.waveform: false` keeps it still). The segment matches every other one in size, frame rate and sound, and the header, footer and on-screen deck are drawn over it as over footage. A video's saved sound (`audio.mp3` and the like in its folder) is now a source the build can cut from, after every saved picture: before any download when the clip has no picture to fetch (`"audioOnly": true` on the clip, `"preferLocalAudio": true` in `render`, or a podcast or feed record), and otherwise only when nothing can be downloaded (`--no-network`, `--skip-fetch`, or a record with no page); `--no-network` lists such clips as playing from audio only instead of refusing them. A clip may also give `"src"` (a video or audio file) and `"cues"` (its transcript, either a `transcript.cues.json` or a `parakeet-stitch` transcript), both relative to the manifest, instead of a channel and video: it plays the whole file unless `start`/`end` cut inside it, `resolve-windows.mjs` widens it with those cues, it gets no QR unless it has a `citeUrl`, and a path that leaves the manifest's folder (or an absolute one, without `"allowAbsoluteSrc": true` in `render`), a missing file or an unreadable transcript stops the build before anything runs, naming the clip. - **A report build cuts from media already on disk before it downloads anything, and `--no-network` makes sure it never does.** For each clip, `build-video.mjs` now looks, in order, in the project's own `out/clips-raw`, in the clip windows the editor fetched into the channel (`channels/<slug>/data/<id>/clips/`), and in a saved whole source video (through the saved-video store's pointer, or a `source-media` file still in the video's folder), and cuts from the first that holds the clip plus its fetch pad; only when none does is the window downloaded. A file that is a link to a drive that is not mounted counts as not there, and the next place is tried. The build prints one line per clip naming where its source came from (`raw-cache`, `corpus-window`, `saved-video`, or a network fetch). With `--no-network`, every clip's source is found before anything is rendered, and if any clip would need a download the build stops at once and lists each one (its position in the timeline, channel, video and the span it needs). umtool's clip bench reads the same three places, so a clip it shows as fetched is one the build cuts from without downloading. -- **umtool's report videos can show a highlighted sentence from a saved article.** `node umtool/report-to-video/shoot-page.mjs --page <saved page.html> --quote "<sentence>" --out <shot.png>` opens a web page saved to disk, finds the sentence in its text, highlights it and saves a PNG of the paragraph that holds it, ready to be a report manifest's `image` entry. `--batch <items.json> --out <dir>` does a list of `{ id, page, quote, context? }` at once and writes `<id>.png` for each plus a `results.json` recording each shot's crop, the matched text and the block it shot. The page is opened offline: nothing is fetched except files saved beside it, and its own scripts do not run unless `--js` is given. The sentence is found whether its quotes and apostrophes are curly or straight, across links and emphasis, and through non-breaking spaces, soft hyphens and line breaks in the page's source. A sentence that is not on the page is listed in `results.json` and on the terminal, and the run ends with an error rather than leaving it out. `--color` sets the highlight; `context` picks one occurrence of a sentence that appears more than once. +- **umtool's report videos can show a highlighted sentence from a saved article.** `node umtool/report-to-video/shoot-page.mjs --page <saved page.html> --quote "<sentence>" --out <shot.png>` opens a web page saved to disk, finds the sentence in its text, highlights it and saves a PNG of the paragraph that holds it, ready to be a report manifest's `image` entry. `--batch <items.json> --out <dir>` does a list of `{ id, page, quote, context? }` at once and writes `<id>.png` for each plus a `results.json` recording each shot's crop, the matched text and the block it shot. The page is opened offline: nothing is fetched except files saved beside it, and its own scripts do not run unless `--js` is given. The sentence is found whether its quotes and apostrophes are curly or straight, across links and emphasis, and through non-breaking spaces, soft hyphens and line breaks in the page's source. A sentence that is not on the page is listed in `results.json` and on the terminal, and the run ends with an error rather than leaving it out. `--color` sets the highlight; `context` picks one occurrence of a sentence that appears more than once. On a page where a whole post is one block of paragraphs separated by line breaks, `--crop mark` (or an item's `"crop": "mark"`) shoots only the sentence's own lines and one whole line above and below (`--context-lines` sets how many) instead of the whole post; `results.json` records which crop each shot used. - **A clip or whole-recording fetch can name the tallest source video it wants.** The MCP's `fetch_clip` takes `maxHeight`, `fetch-via-editor.mjs` takes `--max-height`, and the editor's fetch endpoint takes `maxHeight`: a whole number of pixels from 144 to 2160; anything else is refused before anything is fetched. A window is fetched at or under that height (720 when none is given, as before). A whole recording asked for at 720 or less is saved as the **Video 720p** quality, and above 720 at the original quality; with no height it follows the channel's, else the global, source video quality, as before. umtool's whole-source fetch from the clip bench now asks at the report's `render.maxHeightSource`. A file already on disk is returned as it is and never fetched again for a different height; the answer now gives its height (a window's is read from the file, a whole recording's from what its persist recorded) and says when it is taller than the height asked for. - **Persist a list of videos, across channels, to the saved-video store.** `pnpm ops persist-videos --json '{"items":[{"slug":"<channel>","id":"<video id>"}, …]}'` (or `--file list.json` for a long list) re-fetches the source container of each video that is not saved yet, the way **Persist source video** does on a video's page. `"format": "original"` or `"video_720"` picks the quality (default: each channel's own), and `"replace": "above-height"` also re-fetches a saved video whose recorded height is unknown or above that quality — the old file is removed only after the new one is saved, and the saved video keeps its retention class. Downloads run one at a time, one job per channel on that channel's download queue, with the usual gap between them (`"gapMs"` overrides it); `"minFreeMemMb"` holds each download until that much memory is free. A low disk or a rate limit stops the job, and running the same list again picks up where it left off: saved videos are skipped. `"dryRun": true` answers with what a run would do — saved, saved above the height, to fetch, no source URL, unknown — and starts nothing. Jobs show as **Persist videos**, can be drained, and can be retried from /jobs. - **Capture specific X posts: a screenshot of each, and its attached media.** `pnpm ops capture-posts --json '{"slug":"<channel>","ids":["<post id>", …]}'` shoots each post as X shows it, through the connected X profile, and downloads its pictures and videos with gallery-dl, into the channel's `posts-media/<post id>/` beside a `capture.json` that records when, from which URLs, and each file's size and SHA-256. Every id must already be in the channel's posts archive; one that is not is refused by name and nothing runs. `"shots": false` or `"media": false` skips that half, and posts already captured are skipped unless `"force": true`. The job runs on the X queue with a post fetch, so the two never run at once, and waits a random 4–10 seconds before each request to X, as fetches do. A deleted post, or one behind its account's wall (protected, suspended, gone), is recorded as such in the channel's deleted-post record; a post behind a sensitive-media warning is opened and shot. If X asks to log in, or answers "Something went wrong", the job stops at that post and leaves the rest for a later run. Captures are never published: the export does not read them. diff --git a/umtool/report-to-video/README.md b/umtool/report-to-video/README.md @@ -662,7 +662,8 @@ node umtool/report-to-video/shoot-page.mjs --batch stills.json --out stills/ ```jsonc // stills.json — `page` is relative to this file [{ "id": "a01", "page": "saved/article.html", "quote": "…", - "context": "…" }] // optional: a longer stretch around the quote, to pick one occurrence + "context": "…", // optional: a longer stretch around the quote, to pick one occurrence + "crop": "mark" }] // optional: this item's --crop ``` - **The page is loaded offline.** Every request is refused except a `file://` @@ -689,13 +690,23 @@ node umtool/report-to-video/shoot-page.mjs --batch stills.json --out stills/ to that height around the quote, and says `trimmed`. The highlight (`--color`, `#ffe14d`) never reflows the text: one `<mark>` per text node it covers, padded vertically only. +- **`--crop mark` shoots the quote's lines, not the block** (`--crop block` is + the default; an item's `crop` overrides either). On a forum-style page the + whole post is ONE block of `<br>`-separated paragraphs, so the block crop is a + slab of embeds around a two-line quote. The mark crop is the quote's own lines + plus `--context-lines` (1) whole lines above and below, stepped in the + computed `line-height` of the text the quote sits in from the lines the + highlight is on, so a context line is never cut mid-glyph; it spans the + block's content box, never leaves it, and adds `--padding`. That padding would + show the next line's glyphs cut, so above and below the lines it is painted + in the block's own background. `--max-height` does not apply. - **A miss is never skipped.** It is in `results.json` with its `reason` (`quote not found`, `context not found`, `quote not found inside its context`, `page not found`, or the load error), it is summarised on stderr, and the run exits 1 after shooting every other item. Each result carries `crop` (the CSS-pixel rectangle, in document coordinates), -`pixels` (the PNG's size), `matched` (the page's own text of the match, curly +`cropMode` (`block` or `mark`), `pixels` (the PNG's size), `matched` (the page's own text of the match, curly quotes and all), `block` (a CSS selector for the block shot) and `blocked`. The matcher is pure and is what `shoot-page.test.mjs` tests; the one test that launches Chromium runs only with `SHOOT_PAGE_BROWSER=1`. diff --git a/umtool/report-to-video/shoot-page.mjs b/umtool/report-to-video/shoot-page.mjs @@ -26,7 +26,7 @@ // node umtool/report-to-video/shoot-page.mjs --page <file.html> --quote "<text>" --out <shot.png> // node umtool/report-to-video/shoot-page.mjs --batch <items.json> --out <dir> // -// Batch items are `[{ id, page, quote, context? }]` (or `{ items: [...] }`), +// Batch items are `[{ id, page, quote, context?, crop? }]` (or `{ items: [...] }`), // `page` relative to the items file. Each writes `<dir>/<id>.png`, and the run // writes `<dir>/results.json`. `context` is a longer stretch of text around the // quote, for a quote that occurs more than once. @@ -41,6 +41,9 @@ // quote (default: 1200) // --js Run the page's own scripts // --no-isolate Leave the neighbouring content visible in the padding +// --crop block|mark Shoot the whole block (default), or only the quote's +// lines and their context; an item's `crop` overrides it +// --context-lines <n> (mark crop) whole lines kept above and below (default: 1) // --timeout <ms> Page load timeout (default: 20000) import { existsSync, statSync } from "node:fs"; @@ -57,9 +60,13 @@ export const DEFAULTS = Object.freeze({ maxHeight: 1200, js: false, isolate: true, + crop: "block", + contextLines: 1, timeout: 20_000, }); +export const CROP_MODES = Object.freeze(["block", "mark"]); + // --------------------------------------------------------------------------- // The matcher. Pure, and SELF-CONTAINED: these two functions are also sent into // the page as source text (see browserShoot), so neither may reference anything @@ -143,13 +150,47 @@ export function matchInSegments(segments, quote, context) { return { ok: true, start, end: { segment: last.segment, offset: last.offset + 1 }, occurrences }; } +// The `mark` crop: the quote's own lines and `contextLines` lines either side, +// across the block's content box, plus `padding` — for a block that is a whole +// forum post of `<br>`-separated paragraphs, where the block crop is a slab. +// Pure and self-contained like the matcher; it runs in the page too. +// +// `rects` are the marks' line fragments, `content` the block's content box, +// both in document coordinates. A fragment's box is its glyphs' content area, +// centred in its line box, so the quote's first line box is centred on the +// first fragment and its last on the last; from there the crop moves in whole +// `lineHeight`s, which is what keeps a context line from being cut mid-glyph. +// It never leaves the content box, so a quote on the block's first line has no +// context above it rather than the bottom of whatever sits above the block. +// +// Returns `{ crop, band }`: `band` is the lines' own top and bottom. Between it +// and the crop's edges is padding, which on a page of text is the next line's +// glyphs, cut — so the page side masks it. +export function markCrop({ rects, lineHeight, content, contextLines, padding, docW, docH }) { + let first = rects[0]; + let last = rects[0]; + for (const r of rects) { + if (r.top < first.top) first = r; + if (r.bottom > last.bottom) last = r; + } + const lh = lineHeight; + const n = contextLines; + const top = Math.max(content.top, (first.top + first.bottom) / 2 - lh / 2 - n * lh); + const bottom = Math.min(content.bottom, (last.top + last.bottom) / 2 + lh / 2 + n * lh); + const x0 = Math.max(0, Math.floor(content.left - padding)); + const y0 = Math.max(0, Math.floor(top - padding)); + const x1 = Math.min(docW, Math.ceil(content.right + padding)); + const y1 = Math.min(docH, Math.ceil(bottom + padding)); + return { crop: { x: x0, y: y0, width: x1 - x0, height: y1 - y0 }, band: { top, bottom } }; +} + // --------------------------------------------------------------------------- // The page side. Runs INSIDE the browser (serialised by `pageExpression`): walk // the visible text, match, wrap the match in marks, measure what to shoot. // --------------------------------------------------------------------------- function browserShoot(args) { - const { quote, context, color, padding, maxHeight, isolate } = args; + const { quote, context, color, padding, maxHeight, isolate, cropMode, contextLines } = args; const doc = document; const SKIP = new Set(["SCRIPT", "STYLE", "NOSCRIPT", "TEMPLATE", "TITLE", "HEAD", "SVG", "MATH", "IFRAME", "OBJECT", "SELECT", "TEXTAREA"]); const blockCache = new Map(); @@ -264,33 +305,82 @@ function browserShoot(args) { const docW = Math.max(de.scrollWidth, doc.body?.scrollWidth ?? 0); const docH = Math.max(de.scrollHeight, doc.body?.scrollHeight ?? 0); const br = block.getBoundingClientRect(); + // The marks' line fragments. Their vertical padding is symmetric, so it + // moves no fragment's centre, which is all the mark crop reads. + const rects = []; + for (const mk of marks) { + for (const r of mk.getClientRects()) rects.push({ top: r.top + sy, bottom: r.bottom + sy }); + } let top = br.top + sy; let bottom = br.bottom + sy; let trimmed = false; - if (bottom - top > maxHeight) { - let mt = Infinity; - let mb = -Infinity; - for (const mk of marks) { - for (const r of mk.getClientRects()) { - mt = Math.min(mt, r.top + sy); - mb = Math.max(mb, r.bottom + sy); + let crop; + if (cropMode === "mark") { + const bs = getComputedStyle(block); + const px = (v) => parseFloat(v) || 0; + const content = { + left: br.left + sx + px(bs.borderLeftWidth) + px(bs.paddingLeft), + right: br.right + sx - px(bs.borderRightWidth) - px(bs.paddingRight), + top: br.top + sy + px(bs.borderTopWidth) + px(bs.paddingTop), + bottom: br.bottom + sy - px(bs.borderBottomWidth) - px(bs.paddingBottom), + }; + // The line-height of the text the quote sits in. `normal` has no number in + // computed style; 1.2 × the font size is what browsers use for most fonts. + const ps = getComputedStyle(marks[0].parentElement); + const lineHeight = parseFloat(ps.lineHeight) || parseFloat(ps.fontSize) * 1.2; + const mc = markCrop({ rects, lineHeight, content, contextLines, padding, docW, docH }); + crop = mc.crop; + // Mask the padding above and below the lines with the block's own ground, + // across its padding box: past that the ground is the page's, and the + // isolation already blanks whatever else stands there. The masks hang off + // <html>, so they are positioned in document coordinates and the + // isolation rule (`body *`) does not hide them. + let ground = "#ffffff"; + for (let el = block; el; el = el.parentElement) { + const bg = getComputedStyle(el).backgroundColor; + if (bg && bg !== "transparent" && !/^rgba\(.*,\s*0\)$/.test(bg)) { + ground = bg; + break; } } - if (mb - mt >= maxHeight) { - top = mt; - bottom = mb; - } else { - const mid = (mt + mb) / 2; - const t = Math.min(Math.max(top, mid - maxHeight / 2), bottom - maxHeight); - top = t; - bottom = t + maxHeight; + const padLeft = Math.max(crop.x, br.left + sx + px(bs.borderLeftWidth)); + const padRight = Math.min(crop.x + crop.width, br.right + sx - px(bs.borderRightWidth)); + const mask = (y0, y1) => { + if (y1 <= y0 || padRight <= padLeft) return; + const d = doc.createElement("div"); + d.setAttribute("data-shoot-page-mask", ""); + d.style.cssText = + `position:absolute;left:${padLeft}px;top:${y0}px;width:${padRight - padLeft}px;height:${y1 - y0}px;` + + `background:${ground};z-index:2147483647;pointer-events:none;margin:0;padding:0;border:0;`; + de.appendChild(d); + }; + mask(crop.y, mc.band.top); + mask(mc.band.bottom, crop.y + crop.height); + } else { + if (bottom - top > maxHeight) { + let mt = Infinity; + let mb = -Infinity; + for (const r of rects) { + mt = Math.min(mt, r.top); + mb = Math.max(mb, r.bottom); + } + if (mb - mt >= maxHeight) { + top = mt; + bottom = mb; + } else { + const mid = (mt + mb) / 2; + const t = Math.min(Math.max(top, mid - maxHeight / 2), bottom - maxHeight); + top = t; + bottom = t + maxHeight; + } + trimmed = true; } - trimmed = true; + const x0 = Math.max(0, Math.floor(br.left + sx - padding)); + const y0 = Math.max(0, Math.floor(top - padding)); + const x1 = Math.min(docW, Math.ceil(br.right + sx + padding)); + const y1 = Math.min(docH, Math.ceil(bottom + padding)); + crop = { x: x0, y: y0, width: x1 - x0, height: y1 - y0 }; } - const x0 = Math.max(0, Math.floor(br.left + sx - padding)); - const y0 = Math.max(0, Math.floor(top - padding)); - const x1 = Math.min(docW, Math.ceil(br.right + sx + padding)); - const y1 = Math.min(docH, Math.ceil(bottom + padding)); const cssPath = (el) => { const parts = []; @@ -313,7 +403,8 @@ function browserShoot(args) { ok: true, matched, block: cssPath(block), - crop: { x: x0, y: y0, width: x1 - x0, height: y1 - y0 }, + crop, + cropMode: cropMode === "mark" ? "mark" : "block", trimmed, occurrences: m.occurrences, marks: marks.length, @@ -323,7 +414,7 @@ function browserShoot(args) { // The expression handed to page.evaluate: the matcher and the walker as source, // then a call. A string rather than a function so the helpers travel with it. export function pageExpression(args) { - return `(() => {\n${normaliseWithMap}\n${matchInSegments}\n${browserShoot}\nreturn browserShoot(${JSON.stringify(args)});\n})()`; + return `(() => {\n${normaliseWithMap}\n${matchInSegments}\n${markCrop}\n${browserShoot}\nreturn browserShoot(${JSON.stringify(args)});\n})()`; } // --------------------------------------------------------------------------- @@ -409,6 +500,8 @@ async function shootOne(ctx, item, o) { padding: o.padding, maxHeight: o.maxHeight, isolate: o.isolate, + cropMode: item.crop ?? o.crop, + contextLines: o.contextLines, }), ); if (!r.ok) return { ...base, ok: false, reason: r.reason, occurrences: r.occurrences, blocked }; @@ -421,6 +514,7 @@ async function shootOne(ctx, item, o) { matched: r.matched, block: r.block, crop: r.crop, + cropMode: r.cropMode, scale: o.scale, pixels: { width: Math.round(r.crop.width * o.scale), height: Math.round(r.crop.height * o.scale) }, trimmed: r.trimmed, @@ -456,6 +550,7 @@ export function validateItems(data) { if (typeof it.page !== "string" || !it.page) problems.push(`${where}: page is required`); if (typeof it.quote !== "string" || !it.quote.trim()) problems.push(`${where}: quote is required`); if (it.context != null && typeof it.context !== "string") problems.push(`${where}: context must be a string`); + if (it.crop != null && !CROP_MODES.includes(it.crop)) problems.push(`${where}: crop must be "block" or "mark"`); }); if (problems.length) throw new Error(problems.join("\n")); return items; @@ -489,6 +584,18 @@ export function parseArgs(argv) { case "--timeout": opts.timeout = num(a, next()); break; case "--js": opts.js = true; break; case "--no-isolate": opts.isolate = false; break; + case "--crop": { + const v = next(); + if (!CROP_MODES.includes(v)) throw new Error(`--crop must be "block" or "mark", not "${v}"`); + opts.crop = v; + break; + } + case "--context-lines": { + const v = num(a, next()); + if (!Number.isInteger(v)) throw new Error("--context-lines needs a whole number"); + opts.contextLines = v; + break; + } default: throw new Error(`unknown argument: ${a}`); } } @@ -504,7 +611,7 @@ const USAGE = "usage: shoot-page.mjs --page <file.html> --quote <text> [--context <text>] --out <shot.png>\n" + " shoot-page.mjs --batch <items.json> --out <dir>\n" + " [--color <css>] [--padding <px>] [--width <px>] [--scale <n>] [--max-height <px>] [--js] [--no-isolate]\n" + - " [--timeout <ms>]"; + " [--crop block|mark] [--context-lines <n>] [--timeout <ms>]"; async function main() { let args; diff --git a/umtool/report-to-video/shoot-page.test.mjs b/umtool/report-to-video/shoot-page.test.mjs @@ -14,7 +14,7 @@ import test from "node:test"; import { fileURLToPath } from "node:url"; import { - allowedRequest, matchInSegments, normaliseWithMap, pageExpression, parseArgs, shootPages, validateItems, + allowedRequest, markCrop, matchInSegments, normaliseWithMap, pageExpression, parseArgs, shootPages, validateItems, } from "./shoot-page.mjs"; // The raw text a match covers, read back off the segments. @@ -159,6 +159,47 @@ test("arguments: two modes, never both", () => { assert.throws(() => parseArgs(["--page", "p.html", "--quote", "q"]), /--out is required/); assert.throws(() => parseArgs(["--padding", "-1"]), /non-negative/); assert.throws(() => parseArgs(["--bogus"]), /unknown argument/); + const m = parseArgs(["--batch", "i.json", "--out", "d", "--crop", "mark", "--context-lines", "2"]); + assert.deepEqual(m.opts, { crop: "mark", contextLines: 2 }); + assert.throws(() => parseArgs(["--crop", "page"]), /block" or "mark/); + assert.throws(() => parseArgs(["--context-lines", "1.5"]), /whole number/); +}); + +test("an item's crop is block or mark", () => { + assert.equal(validateItems([{ id: "a", page: "p.html", quote: "q", crop: "mark" }]).length, 1); + assert.throws(() => validateItems([{ id: "a", page: "p.html", quote: "q", crop: "lines" }]), /crop must be/); +}); + +// A forum post: one block, 27 px lines (18px/1.5), its content box 674 px wide +// from x 52 and 1400 px tall from y 100. A fragment's box is its glyphs (21 px +// here, plus the mark's padding), centred in its line. +const POST = { lineHeight: 27, content: { left: 52, right: 726, top: 100, bottom: 1500 }, padding: 10, docW: 1280, docH: 3000 }; +const lineRect = (k) => ({ top: 100 + 27 * k + 3, bottom: 100 + 27 * k + 24 }); + +test("the mark crop is the quote's lines and whole context lines, not the block", () => { + // A quote over lines 10 and 11, one line of context either side: four lines. + const { crop: c, band } = markCrop({ ...POST, rects: [lineRect(10), lineRect(11)], contextLines: 1 }); + assert.deepEqual(c, { x: 42, y: 100 + 27 * 9 - 10, width: 674 + 20, height: 4 * 27 + 20 }); + // The band is the four lines alone; the padding outside it is masked. + assert.deepEqual(band, { top: 100 + 27 * 9, bottom: 100 + 27 * 13 }); + // Its edges are line boundaries: (y + padding - content.top) is whole lines. + assert.equal((c.y + 10 - 100) % 27, 0); + // No context: exactly the quote's lines. + assert.equal(markCrop({ ...POST, rects: [lineRect(10)], contextLines: 0 }).crop.height, 27 + 20); + // Fragments in any order, several on a line (a link mid-quote), the same crop. + const mixed = [lineRect(11), { top: lineRect(10).top, bottom: lineRect(10).bottom }, lineRect(10)]; + assert.deepEqual(markCrop({ ...POST, rects: mixed, contextLines: 1 }).crop, c); +}); + +test("the mark crop stays inside the block's content box", () => { + // On the first line there is no line above to keep, and nothing above the + // block is taken in its place. + const top = markCrop({ ...POST, rects: [lineRect(0)], contextLines: 2 }).crop; + assert.equal(top.y, 100 - 10); + assert.equal(top.height, 3 * 27 + 20); + const lastLine = (1500 - 100) / 27 - 1; // not whole: the box ends mid-line + const bottom = markCrop({ ...POST, rects: [lineRect(Math.floor(lastLine))], contextLines: 3 }).crop; + assert.equal(bottom.y + bottom.height, 1500 + 10); }); // --------------------------------------------------------------------------- @@ -178,6 +219,7 @@ const FIXTURE = `<!doctype html> <div style="display:none">a hidden sentence that must not match</div> <img src="https://example.invalid/tracker.png" alt=""> <img src="../outside.png" alt=""> +<div id="post" style="padding: 12px; border: 1px solid #999">${Array.from({ length: 24 }, (_, i) => `Paragraph ${i} says something short.`).join("<br><br>")}<div style="height: 300px; background: #000"></div>The post goes on after the embed.</div> <div id="long">${Array.from({ length: 80 }, (_, i) => `Line ${i} of the long block.`).join("<br>")}</div> </body></html>`; @@ -206,6 +248,10 @@ test("a real page: shots, marks, a miss, and nothing fetched from outside", { { id: "nopage", page: "saved/missing.html", quote: "anything" }, // Across a <br>, in a block far taller than --max-height. { id: "long", page: "saved/article.html", quote: "Line 40 of the long block. Line 41" }, + // A forum post: one block of <br>-separated paragraphs. The same quote, + // shot whole and (an item's own crop) as its line with one either side. + { id: "post-block", page: "saved/article.html", quote: "Paragraph 12 says something short" }, + { id: "post-mark", page: "saved/article.html", quote: "Paragraph 12 says something short", crop: "mark" }, ].map((it) => ({ ...it, out: path.join(out, `${it.id}.png`) })); const results = await shootPages(items, { baseDir: dir, padding: 10 }); @@ -244,6 +290,22 @@ test("a real page: shots, marks, a miss, and nothing fetched from outside", { // (Within a pixel: the crop floors its top edge and ceils its bottom.) assert.ok(Math.abs(by.long.crop.height - (1200 + 2 * 10)) <= 1, String(by.long.crop.height)); + // The post: the block crop is the whole post; the mark crop is three lines + // (the quote's, and the empty line from the <br><br> either side) across + // its content box (700 less 2 x 12 padding and 2 x 1 border). + for (const id of ["post-block", "post-mark"]) { + assert.equal(by[id].ok, true, id); + assert.equal(by[id].block, "#post", id); + pngSize(await readFile(by[id].png)); + } + assert.equal(by["post-block"].cropMode, "block"); + assert.equal(by["post-mark"].cropMode, "mark"); + assert.ok(by["post-block"].crop.height > 1000, String(by["post-block"].crop.height)); + const pm = by["post-mark"].crop; + assert.ok(Math.abs(pm.height - (3 * 27 + 2 * 10)) <= 1, String(pm.height)); + assert.ok(Math.abs(pm.width - (674 + 2 * 10)) <= 1, String(pm.width)); + assert.equal(by["post-mark"].trimmed, false); + // The CLI: a batch with a miss writes what it shot, lists the miss in the // results, says so on stderr, and exits 1. await writeFile(path.join(dir, "items.json"), JSON.stringify([