Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 8180a19bff3a3744cb58acf054ead19c33749fda
parent 25f7d342f297ee59448a5e55520f282258677e82
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sun,  4 Oct 2026 16:15:51 -0400

editor: capture-posts — server action, ops route, replay, pnpm ops

capturePostsAction refuses, before a job exists, what the controller
refuses (neither half, no ids, not social, a fetcher that cannot capture,
ids not in the posts archive) and runs the job on the channel's platform
queue, carrying ids, shots, media and force in its spec so it replays.
POST /api/ops/capture-posts { slug, ids, shots?, media?, force?, queueKey? }
is the adapter over it; pnpm ops capture-posts posts to it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
MRUNNING_IN_DOCKER.md | 1+
Meditor/CHANGELOG.md | 1+
Aeditor/app/api/ops/capture-posts/route.ts | 41+++++++++++++++++++++++++++++++++++++++++
Meditor/app/channels/[slug]/socialActions.ts | 73+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/app/jobs/jobReplayRegistry.ts | 14++++++++++++++
Meditor/e2e/ops-api.spec.ts | 63+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mscripts/archilyzer-ops.mjs | 9+++++++++
Mscripts/archilyzer-ops.test.mjs | 15+++++++++++++++
8 files changed, 217 insertions(+), 0 deletions(-)

diff --git a/RUNNING_IN_DOCKER.md b/RUNNING_IN_DOCKER.md @@ -255,6 +255,7 @@ pnpm ops lane --json '{"lane":"download","held":true}' pnpm ops refresh-report --json '{"all":true}' pnpm ops keep-videos --json '{"slug":"paramount-tactical","match":"TheQuartering","dryRun":true}' pnpm ops fetch-posts --json '{"slug":"example-x","older":true}' --wait +pnpm ops capture-posts --json '{"slug":"example-x","ids":["1234567890"]}' --wait pnpm ops get channel the-quartering pnpm ops list # every action name ``` diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,6 +1,7 @@ # Changelog ## [Unreleased] +- **Capture specific X posts: a screenshot of each, and its attached media.** `pnpm ops capture-posts --json '{"slug":"<channel>","ids":["<post id>", …]}'` shoots each post as X shows it, through the connected X profile, and downloads its pictures and videos with gallery-dl, into the channel's `posts-media/<post id>/` beside a `capture.json` that records when, from which URLs, and each file's size and SHA-256. Every id must already be in the channel's posts archive; one that is not is refused by name and nothing runs. `"shots": false` or `"media": false` skips that half, and posts already captured are skipped unless `"force": true`. The job runs on the X queue with a post fetch, so the two never run at once, and waits a random 4–10 seconds before each request to X, as fetches do. A deleted post, or one behind its account's wall (protected, suspended, gone), is recorded as such in the channel's deleted-post record; a post behind a sensitive-media warning is opened and shot. If X asks to log in, or answers "Something went wrong", the job stops at that post and leaves the rest for a later run. Captures are never published: the export does not read them. - **X posts are fetched more slowly, with random gaps.** Every read of X now waits a random 4 to 10 seconds before each request to X, where it used to page as fast as X answered, and always waits out a rate limit rather than pushing through. When fetching older posts, the pause between one three-month window and the next is a random 45 to 120 seconds instead of a fixed 15. A deep walk of an account's history takes longer; a routine fetch of new posts takes a few seconds more. - **The MCP's search tools take `date_from` and `date_to` as `2024-10-26` as well as `20241026`, and refuse a date they cannot read.** `search_transcripts` and `enumerate_matches` used to accept only `YYYYMMDD`: any other spelling was dropped with a footer warning and the search ran with no date bound, so a whole-corpus count could be read as the bounded one. Dashed, slashed and dotted dates and ISO timestamps are now normalised, and anything else is an error and nothing is searched. - **An X channel can fetch posts older than its timeline reaches.** X's timeline only pages back so far, so a fetch could end, and call the history done, well short of an account's first post. The new **Fetch older posts** button on an X channel's page (or `pnpm ops fetch-posts --json '{"slug":"<channel>","older":true}'`) walks back from the oldest archived post through X search, three months at a time, and saves posts the same way a normal fetch does; posts already archived are skipped. It needs a login, as search does: without one it stops at once and the channel shows **Needs credentials**. A run saves its place as it goes and stops after three hours; the next run continues from there. The walk ends at the account's creation date, after a year of windows with no posts, or at a date you give as `"floor": "YYYY-MM-DD"`, and the page's **Older posts** line then says it is complete; running it again says so and fetches nothing. A normal **Fetch posts** is unaffected and still fetches new posts from the top. Bluesky channels have no such button: their fetch already reads the whole history. diff --git a/editor/app/api/ops/capture-posts/route.ts b/editor/app/api/ops/capture-posts/route.ts @@ -0,0 +1,41 @@ +import { capturePostsAction } from "../../../channels/[slug]/socialActions"; +import { + jobResponse, + ops, + optBool, + optString, + reqSlug, + reqStringArray, +} from "../_lib"; + +export const dynamic = "force-dynamic"; + +// POST { slug, ids, shots?, media?, force?, queueKey? } -> { ok: true, jobId } +// +// Capture specific archived posts of a social channel: a screenshot of each +// (`shots`, default true) and its attached media (`media`, default true), into +// the channel's posts-media/<id>/. Posts already captured are skipped unless +// `force`. The job runs on the platform's queue, as a post fetch does. +// +// Every refusal is the action's own sentence: both halves off, a channel that +// is not a social one, a fetcher that cannot capture, an id not in the +// channel's posts archive. +export async function POST(request: Request) { + return ops( + request, + ["slug", "ids", "shots", "media", "force", "queueKey"], + async (body) => { + const slug = reqSlug(body, "slug"); + return jobResponse( + await capturePostsAction( + slug, + reqStringArray(body, "ids"), + optString(body, "queueKey"), + optBool(body, "shots"), + optBool(body, "media"), + optBool(body, "force"), + ), + ); + }, + ); +} diff --git a/editor/app/channels/[slug]/socialActions.ts b/editor/app/channels/[slug]/socialActions.ts @@ -6,6 +6,7 @@ // which applies to a post fetch. A social channel has exactly two stages // (Fetch → Index), and this file owns the first. +import path from "node:path"; import { revalidatePath } from "next/cache"; import { safeRevalidate } from "../../lib/safeRevalidate"; import { getPaths } from "yt-dlp-transcript-common/lib/paths"; @@ -30,6 +31,14 @@ import { type CheckPostAvailabilityMode, } from "yt-dlp-transcript-common/controller/checkPostAvailability"; import { + capturePosts, + capturePostsProblem, + NOTHING_TO_CAPTURE, + strayCaptureIds, + strayIdsRefusal, +} from "yt-dlp-transcript-common/controller/capturePosts"; +import { readSeenPostIds } from "yt-dlp-transcript-common/lib/posts-server"; +import { getSocialFetcher, listSocialFetchers, resolveSocialFetcher, @@ -196,3 +205,67 @@ export async function fetchPostsAction( }, }); } + +// A screenshot and the attached media of specific archived posts, into the +// channel's `posts-media/<id>/`. On the platform queue, as a fetch is, so the +// two never run against the same source at once. Refused HERE, before a job +// exists: nothing asked for, no ids, a channel that is not social, a fetcher +// that cannot capture, or an id that is not in the channel's posts archive +// (named). Posts already captured are skipped unless `force`. +export async function capturePostsAction( + slug: string, + ids: string[], + queueKey?: string, + shots?: boolean, + media?: boolean, + force?: boolean, +): Promise<StreamActionResult> { + if (shots === false && media === false) return { ok: false, error: NOTHING_TO_CAPTURE }; + const wanted = [...new Set(ids)]; + if (wanted.length === 0) return { ok: false, error: "No post ids to capture." }; + const paths = getPaths(); + const config = await readChannelConfig(paths, slug); + if (!config) return { ok: false, error: `No such channel: ${slug}` }; + if (!isSocialChannel(config)) { + return { ok: false, error: `${slug} is not a social channel.` }; + } + await registerBuiltinSocialFetchers(); + const problem = capturePostsProblem( + resolveSocialFetcher(config.postFetcher, config.url), + ); + if (problem) return { ok: false, error: problem }; + const stray = strayCaptureIds( + wanted, + await readSeenPostIds(path.join(paths.channelsDir, slug)), + ); + if (stray) return { ok: false, error: strayIdsRefusal(slug, stray) }; + + const key = resolveQueueKey(downloadQueueKey(config), queueKey); + return runManagedFunction({ + kind: "capture-posts", + queueKey: key, + paths, + channelSlug: slug, + spec: { + kind: "capture-posts", + slug, + params: { queueKey, ids: wanted, shots, media, force }, + }, + fn: async (onLog, signal, _progress, ctx) => { + const result = await capturePosts({ + paths, + slug, + settings: getSettings(), + ids: wanted, + shots, + media, + force, + onLog, + signal, + drain: ctx.drainSignal, + }); + safeRevalidate([`/channels/${slug}`]); + if (!result.ok) throw new Error(result.error ?? "Post capture failed"); + }, + }); +} diff --git a/editor/app/jobs/jobReplayRegistry.ts b/editor/app/jobs/jobReplayRegistry.ts @@ -40,6 +40,7 @@ import { import { backfillChannelAction } from "../channels/[slug]/backfillActions"; import { persistKeptAction } from "../channels/[slug]/persistActions"; import { + capturePostsAction, checkPostAvailabilityAction, fetchPostsAction, } from "../channels/[slug]/socialActions"; @@ -255,6 +256,19 @@ export const JOB_REPLAY_HANDLERS: Record<string, ReplayHandler> = { str(p.floor), ); }, + // The ids are the spec's own (a capture is OF specific posts, unlike a + // bucket); the re-run skips whatever the first run already captured. + "capture-posts": (spec) => { + const { p, queueKey } = params(spec); + return capturePostsAction( + spec.slug, + strings(p.ids) ?? [], + queueKey, + bool(p.shots), + bool(p.media), + bool(p.force), + ); + }, "download-missing-subs": (spec) => { const { p, queueKey } = params(spec); return downloadMissingSubsAction(spec.slug, queueKey, bool(p.abortOnError)); diff --git a/editor/e2e/ops-api.spec.ts b/editor/e2e/ops-api.spec.ts @@ -193,6 +193,7 @@ test("a traversing slug is refused at the door, on every route that takes one", ["relocate", { slugs: ["../../escape"], root: "/tmp/ops-api-never" }], ["relocate-back", { slugs: ["../../escape"] }], ["fetch-posts", { slug: "../../escape", older: true }], + ["capture-posts", { slug: "../../escape", ids: ["1"] }], ]; for (const [action, data] of cases) { const { status, body } = await ops(request, action, data); @@ -335,6 +336,68 @@ test("fetch-posts refuses a channel that is not social, full with older, and an expect(await listJobIds()).toEqual(before); }); +test("capture-posts refuses what it cannot capture, and an id not in the archive — before any job", async ({ + request, +}) => { + await resetData("title-filter-channel"); + await settings(); + // Nothing below reaches X: every case is refused before a job exists. + await writeChannelConfig("example-bsky", { + handling: "transcribe", + sourceKind: "social", + platform: "bluesky", + postFetcher: "bluesky-atproto", + socialHandle: "example.bsky.social", + name: "Example (Bluesky)", + url: "https://bsky.app/profile/example.bsky.social", + }); + await writeChannelConfig("example-x", { + handling: "transcribe", + sourceKind: "social", + platform: "twitter", + postFetcher: "x-gallery-dl", + socialHandle: "example_user", + name: "Example (X)", + url: "https://x.com/example_user", + }); + const before = await listJobIds(); + + const video = await ops(request, "capture-posts", { slug: "test-filter", ids: ["1"] }); + expect(video.status).toBe(400); + expect(video.body.error).toBe("test-filter is not a social channel."); + + const bsky = await ops(request, "capture-posts", { slug: "example-bsky", ids: ["1"] }); + expect(bsky.status).toBe(400); + expect(bsky.body.error).toMatch(/cannot capture posts/); + + const neither = await ops(request, "capture-posts", { + slug: "example-x", + ids: ["1"], + shots: false, + media: false, + }); + expect(neither.status).toBe(400); + expect(neither.body.error).toMatch(/both the screenshot and the media are turned off/); + + // The channel's posts archive is empty: every id is a stray, named. + const stray = await ops(request, "capture-posts", { slug: "example-x", ids: ["111", "222"] }); + expect(stray.status).toBe(400); + expect(stray.body.error).toBe("2 id(s) not in example-x's posts archive: 111, 222"); + + // The body's shape. + const noIds = await ops(request, "capture-posts", { slug: "example-x" }); + expect(noIds.status).toBe(400); + expect(noIds.body.error).toMatch(/"ids" is required/); + const badFlag = await ops(request, "capture-posts", { slug: "example-x", ids: ["1"], shots: "yes" }); + expect(badFlag.status).toBe(400); + expect(badFlag.body.error).toMatch(/"shots" must be a boolean/); + const unknown = await ops(request, "capture-posts", { slug: "example-x", ids: ["1"], limit: 5 }); + expect(unknown.status).toBe(400); + expect(unknown.body.error).toMatch(/unknown key\(s\): limit/); + + expect(await listJobIds()).toEqual(before); +}); + test("channel-config round-trips a download filter and refuses a bad regex", async ({ page, request, diff --git a/scripts/archilyzer-ops.mjs b/scripts/archilyzer-ops.mjs @@ -103,6 +103,8 @@ const ACTIONS = [ // A social channel's post fetch: new posts, the full re-walk ("full"), or // the walk back below the oldest archived post ("older"). "fetch-posts", + // A screenshot and the attached media of specific archived posts. + "capture-posts", "build-index", "build-deploy", "build-site", @@ -326,6 +328,13 @@ export function usage() { ' the next run, down to "floor": "YYYY-MM-DD" when given. "limit": N caps', ' the posts one run reads. "full" and "older" together are refused.', "", + 'capture-posts captures archived posts of a social channel (X): a', + ' screenshot of each through the connected X profile, and its attached', + ' media through gallery-dl, into the channel\'s posts-media/<id>/:', + ' {"slug", "ids": [...]}. Every id must be in the channel\'s posts archive.', + ' "shots": false or "media": false skips that half; posts already captured', + ' are skipped unless "force": true. Paced like a post fetch, on its queue.', + "", '"preview": "<branch>" on deploy-site or build-deploy makes it a Cloudflare', " Pages PREVIEW instead of production: the same bundle goes to a branch", " alias, https://<branch>.<project>.pages.dev, and the live site is left", diff --git a/scripts/archilyzer-ops.test.mjs b/scripts/archilyzer-ops.test.mjs @@ -390,3 +390,18 @@ test("fetch-posts is a POST to its route, named in the usage", () => { assert.match(usage(), /Actions:.*transcribe-bucket, fetch-posts/); assert.match(usage(), /"older": true walks back from the oldest/); }); + +// A post capture: a POST to its route, the body passed through untouched — the +// route and the action judge it (ids in the archive, both halves off). +test("capture-posts is a POST to its route, named in the usage", () => { + const p = parseArgs([ + "capture-posts", + "--json", + '{"slug":"example-x","ids":["123"],"media":false}', + ]); + assert.equal(p.method, "POST"); + assert.equal(p.path, "/api/ops/capture-posts"); + assert.deepEqual(p.body, { slug: "example-x", ids: ["123"], media: false }); + assert.match(usage(), /Actions:.*fetch-posts, capture-posts/); + assert.match(usage(), /Every id must be in the channel's posts archive/); +});