Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit ee8b98a934c0bc42f73dbf0b8362b1882fd67c41
parent 1bf5096d17eb06612adce750ed4b85c7b15f8359
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sun,  4 Oct 2026 20:13:50 -0400

editor: capture-posts takes articles (default true) through the route, action, replay and pnpm ops; a route test; the [Unreleased] bullet

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
Meditor/CHANGELOG.md | 1+
Aeditor/app/api/ops/capture-posts/route.test.ts | 73+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/app/api/ops/capture-posts/route.ts | 15++++++++++-----
Meditor/app/channels/[slug]/socialActions.ts | 10+++++++---
Meditor/app/jobs/jobReplayRegistry.ts | 1+
Mscripts/archilyzer-ops.mjs | 5++++-
Mscripts/archilyzer-ops.test.mjs | 1+
7 files changed, 97 insertions(+), 9 deletions(-)

diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,6 +1,7 @@ # Changelog ## [Unreleased] +- **Capturing an X post that is an Article also saves the article.** An X Article (a long-form post) is archived as nothing but its link, and gallery-dl cannot read its body. When `pnpm ops capture-posts` meets a post whose archived text, or whose card on the page, links to an article, it now opens the article in the same X profile, after the same 4–10 second pause, and saves beside the post's capture: `article.json` (the title, author, date and every heading, paragraph, quote, list item, image, link and embedded post in reading order, an embedded post by its URL), `article.md` (the same as readable text), `article.png` (the whole article as shown, cut off at 16,000 pixels tall and marked `trimmed` when longer), `article.html` (the article as X served it, so it can be read again without going back to X) and the article's pictures as `article-img-1.jpg`, `article-img-2.png`, … at full size, fetched through the same browser session. `capture.json` records the article's state, title, block count and every file's size and SHA-256. `"articles": false` leaves articles alone; `"shots": false, "media": false, "articles": true` reads only the articles. An article already captured, deleted or unavailable is not opened again unless `"force": true`; one that failed is tried on the next run. If X asks to log in on the article page, the job stops there as it does for a post. The post viewer's capture panel in the editor shows the article's title with a link to `article.md`. - **A report video can play a clip that has only sound, and a clip can be a file beside the manifest.** When a clip's source has no picture, `build-video.mjs` plays it under a poster: a card with the clip's channel, title and date, the size of the picture area, with the sound's waveform moving along its foot (`render.audioPoster.waveform: false` keeps it still). The segment matches every other one in size, frame rate and sound, and the header, footer and on-screen deck are drawn over it as over footage. A video's saved sound (`audio.mp3` and the like in its folder) is now a source the build can cut from, after every saved picture: before any download when the clip has no picture to fetch (`"audioOnly": true` on the clip, `"preferLocalAudio": true` in `render`, a podcast or feed record, or a record with no page). A clip that should have a picture is not quietly played from its sound: with `--no-network` or `--skip-fetch`, one whose picture is missing still stops the build, listed as needing a download with a note that its sound is on disk, so `--no-network` still proves every picture is there. Add `--audio-fallback` to play such clips from their sound under the poster instead; they are listed apart (and logged as `audio-fallback`), and clips that have no picture to fetch are listed as playing from audio only rather than refused. A clip may also give `"src"` (a video or audio file) and `"cues"` (its transcript, either a `transcript.cues.json` or a `parakeet-stitch` transcript), both relative to the manifest, instead of a channel and video: it plays the whole file unless `start`/`end` cut inside it, `resolve-windows.mjs` widens it with those cues, it gets no QR unless it has a `citeUrl`, and a path that leaves the manifest's folder (or an absolute one, without `"allowAbsoluteSrc": true` in `render`), a missing file or an unreadable transcript stops the build before anything runs, naming the clip. - **A report build cuts from media already on disk before it downloads anything, and `--no-network` makes sure it never does.** For each clip, `build-video.mjs` now looks, in order, in the project's own `out/clips-raw`, in the clip windows the editor fetched into the channel (`channels/<slug>/data/<id>/clips/`), and in a saved whole source video (through the saved-video store's pointer, or a `source-media` file still in the video's folder), and cuts from the first that holds the clip plus its fetch pad; only when none does is the window downloaded. A file that is a link to a drive that is not mounted counts as not there, and the next place is tried. The build prints one line per clip naming where its source came from (`raw-cache`, `corpus-window`, `saved-video`, or a network fetch). With `--no-network`, every clip's source is found before anything is rendered, and if any clip would need a download the build stops at once and lists each one (its position in the timeline, channel, video and the span it needs). umtool's clip bench reads the same three places, so a clip it shows as fetched is one the build cuts from without downloading. - **umtool's report videos can show a highlighted sentence from a saved article.** `node umtool/report-to-video/shoot-page.mjs --page <saved page.html> --quote "<sentence>" --out <shot.png>` opens a web page saved to disk, finds the sentence in its text, highlights it and saves a PNG of the paragraph that holds it, ready to be a report manifest's `image` entry. `--batch <items.json> --out <dir>` does a list of `{ id, page, quote, context? }` at once and writes `<id>.png` for each plus a `results.json` recording each shot's crop, the matched text and the block it shot. The page is opened offline: nothing is fetched except files saved beside it, and its own scripts do not run unless `--js` is given. The sentence is found whether its quotes and apostrophes are curly or straight, across links and emphasis, and through non-breaking spaces, soft hyphens and line breaks in the page's source. A sentence that is not on the page is listed in `results.json` and on the terminal, and the run ends with an error rather than leaving it out. `--color` sets the highlight; `context` picks one occurrence of a sentence that appears more than once. On a page where a whole post is one block of paragraphs separated by line breaks, `--crop mark` (or an item's `"crop": "mark"`) shoots only the sentence's own lines and one whole line above and below (`--context-lines` sets how many) instead of the whole post; `results.json` records which crop each shot used. diff --git a/editor/app/api/ops/capture-posts/route.test.ts b/editor/app/api/ops/capture-posts/route.test.ts @@ -0,0 +1,73 @@ +import test from "node:test"; +import assert from "node:assert/strict"; +import { mkdir, mkdtemp, readdir, rm, writeFile } from "node:fs/promises"; +import os from "node:os"; +import path from "node:path"; + +// Run with: +// pnpm -C editor exec tsx --test "app/api/ops/capture-posts/route.test.ts" +// +// The body's shape and the refusals that come before any job: every case here +// is answered from the disk alone, so a temp corpus with one X channel is the +// whole world — no job is queued and nothing reaches X. + +const ROOT = await mkdtemp(path.join(os.tmpdir(), "capture-posts-route-")); +const SLUG = "demo-x"; +const CHANNEL = path.join(ROOT, "channels", SLUG); +// Set before the route (and getPaths, which caches) is first imported. +process.env.WORKER_TOKEN = "test-token"; +process.env.TRANSCRIPTS_DIR = ROOT; +process.env.SETTINGS_FILE = path.join(ROOT, "settings.json"); +await mkdir(CHANNEL, { recursive: true }); +await writeFile( + path.join(CHANNEL, "config.json"), + JSON.stringify({ + handling: "transcribe", + sourceKind: "social", + platform: "twitter", + postFetcher: "x-gallery-dl", + socialHandle: "example_user", + name: "Example (X)", + url: "https://x.com/example_user", + }), +); +await writeFile(path.join(CHANNEL, "posts-archive"), "twitter 111\n"); +const { POST } = await import("./route"); +test.after(() => rm(ROOT, { recursive: true, force: true })); + +async function post(body: Record<string, unknown>): Promise<{ status: number; error: string }> { + const res = await POST( + new Request("http://localhost/api/ops/capture-posts", { + method: "POST", + headers: { + authorization: "Bearer test-token", + "content-type": "application/json", + }, + body: JSON.stringify(body), + }), + ); + return { status: res.status, error: ((await res.json()) as { error?: string }).error ?? "" }; +} + +test("articles must be a boolean; the route names it among its keys", async () => { + const bad = await post({ slug: SLUG, ids: ["111"], articles: "yes" }); + assert.equal(bad.status, 400); + assert.match(bad.error, /"articles" must be a boolean/); + const unknown = await post({ slug: SLUG, ids: ["111"], article: true }); + assert.equal(unknown.status, 400); + assert.match(unknown.error, /unknown key\(s\): article — this route accepts .*articles/); +}); + +test("both halves off is nothing to do unless the articles are asked for by name", async () => { + for (const articles of [undefined, false]) { + const res = await post({ slug: SLUG, ids: ["111"], shots: false, media: false, articles }); + assert.equal(res.status, 400, String(articles)); + assert.match(res.error, /both the screenshot and the media are turned off/); + } + // Asked for: past that check, to the next refusal — an id not archived. + const stray = await post({ slug: SLUG, ids: ["999"], shots: false, media: false, articles: true }); + assert.equal(stray.status, 400); + assert.equal(stray.error, "1 id(s) not in demo-x's posts archive: 999"); + // No job was ever written. + assert.deepEqual(await readdir(path.join(ROOT, ".jobs")).catch(() => []), []); +}); diff --git a/editor/app/api/ops/capture-posts/route.ts b/editor/app/api/ops/capture-posts/route.ts @@ -10,20 +10,24 @@ import { export const dynamic = "force-dynamic"; -// POST { slug, ids, shots?, media?, force?, queueKey? } -> { ok: true, jobId } +// POST { slug, ids, shots?, media?, force?, articles?, queueKey? } -> { ok: true, jobId } // // Capture specific archived posts of a social channel: a screenshot of each // (`shots`, default true) and its attached media (`media`, default true), into -// the channel's posts-media/<id>/. Posts already captured are skipped unless -// `force`. The job runs on the platform's queue, as a post fetch does. +// the channel's posts-media/<id>/. A post that links to an X Article (its +// archived text, or its card) also gets the article — read into article.json, +// article.md and article.png, with its images — unless `articles` is false +// (default true). Posts already captured are skipped unless `force`. The job +// runs on the platform's queue, as a post fetch does. // -// Every refusal is the action's own sentence: both halves off, a channel that +// Every refusal is the action's own sentence: both halves off (an +// articles-only run names `articles: true`), a channel that // is not a social one, a fetcher that cannot capture, an id not in the // channel's posts archive. export async function POST(request: Request) { return ops( request, - ["slug", "ids", "shots", "media", "force", "queueKey"], + ["slug", "ids", "shots", "media", "force", "articles", "queueKey"], async (body) => { const slug = reqSlug(body, "slug"); return jobResponse( @@ -34,6 +38,7 @@ export async function POST(request: Request) { optBool(body, "shots"), optBool(body, "media"), optBool(body, "force"), + optBool(body, "articles"), ), ); }, diff --git a/editor/app/channels/[slug]/socialActions.ts b/editor/app/channels/[slug]/socialActions.ts @@ -34,6 +34,7 @@ import { capturePosts, capturePostsProblem, NOTHING_TO_CAPTURE, + nothingToCapture, strayCaptureIds, strayIdsRefusal, } from "yt-dlp-transcript-common/controller/capturePosts"; @@ -211,7 +212,8 @@ export async function fetchPostsAction( // two never run against the same source at once. Refused HERE, before a job // exists: nothing asked for, no ids, a channel that is not social, a fetcher // that cannot capture, or an id that is not in the channel's posts archive -// (named). Posts already captured are skipped unless `force`. +// (named). Posts already captured are skipped unless `force`. A post that +// links to an X Article gets the article too unless `articles` is false. export async function capturePostsAction( slug: string, ids: string[], @@ -219,8 +221,9 @@ export async function capturePostsAction( shots?: boolean, media?: boolean, force?: boolean, + articles?: boolean, ): Promise<StreamActionResult> { - if (shots === false && media === false) return { ok: false, error: NOTHING_TO_CAPTURE }; + if (nothingToCapture({ shots, media, articles })) return { ok: false, error: NOTHING_TO_CAPTURE }; const wanted = [...new Set(ids)]; if (wanted.length === 0) return { ok: false, error: "No post ids to capture." }; const paths = getPaths(); @@ -249,7 +252,7 @@ export async function capturePostsAction( spec: { kind: "capture-posts", slug, - params: { queueKey, ids: wanted, shots, media, force }, + params: { queueKey, ids: wanted, shots, media, force, articles }, }, fn: async (onLog, signal, _progress, ctx) => { const result = await capturePosts({ @@ -260,6 +263,7 @@ export async function capturePostsAction( shots, media, force, + articles, onLog, signal, drain: ctx.drainSignal, diff --git a/editor/app/jobs/jobReplayRegistry.ts b/editor/app/jobs/jobReplayRegistry.ts @@ -273,6 +273,7 @@ export const JOB_REPLAY_HANDLERS: Record<string, ReplayHandler> = { bool(p.shots), bool(p.media), bool(p.force), + bool(p.articles), ); }, "download-missing-subs": (spec) => { diff --git a/scripts/archilyzer-ops.mjs b/scripts/archilyzer-ops.mjs @@ -337,7 +337,10 @@ export function usage() { ' media through gallery-dl, into the channel\'s posts-media/<id>/:', ' {"slug", "ids": [...]}. Every id must be in the channel\'s posts archive.', ' "shots": false or "media": false skips that half; posts already captured', - ' are skipped unless "force": true. Paced like a post fetch, on its queue.', + ' are skipped unless "force": true. A post that links to an X Article also', + ' gets the article (article.json, .md, .png and its images) unless', + ' "articles": false; both halves off with "articles": true reads only the', + ' articles. Paced like a post fetch, on its queue.', "", 'persist-videos saves specific videos, across channels, to the saved-video', ' store: {"items": [{"slug", "id"}, ...]}. "format": "original" |', diff --git a/scripts/archilyzer-ops.test.mjs b/scripts/archilyzer-ops.test.mjs @@ -404,6 +404,7 @@ test("capture-posts is a POST to its route, named in the usage", () => { assert.deepEqual(p.body, { slug: "example-x", ids: ["123"], media: false }); assert.match(usage(), /Actions:.*fetch-posts, capture-posts/); assert.match(usage(), /Every id must be in the channel's posts archive/); + assert.match(usage(), /unless\s+"articles": false/); }); test("persist-videos is a POST to its route, named in the usage", () => {