Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 4bec96293d43351cd7e9354c62ba5ec38eb587bb
parent ea1e02ae3fafc5f997645efc31c739179abb87fc
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sun,  4 Oct 2026 22:09:13 -0400

editor: fetch-posts takes force (route, action, replay, pnpm ops usage) and passes the job's drain; an empty-account older walk is refused before any job; a route test; the [Unreleased] bullet

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
Meditor/CHANGELOG.md | 1+
Aeditor/app/api/ops/fetch-posts/route.test.ts | 75+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/app/api/ops/fetch-posts/route.ts | 9++++++---
Meditor/app/channels/[slug]/socialActions.ts | 29+++++++++++++++++++++++++----
Meditor/app/jobs/jobReplayRegistry.ts | 1+
Mscripts/archilyzer-ops.mjs | 5++++-
Mscripts/archilyzer-ops.test.mjs | 1+
7 files changed, 113 insertions(+), 8 deletions(-)

diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,6 +1,7 @@ # Changelog ## [Unreleased] +- **X post fetches stop on Drain, and an account with no posts is not searched.** Draining a `fetch-posts` job used to do nothing until gallery-dl finished its whole run. Now the timeline fetch stops at the next page boundary (at once when gallery-dl is between pages or waiting out a rate limit), the older-posts walk stops its current window's search at once and never starts the 45–120 second pause between windows, and both keep their resume point: the job ends done, not failed, and the log says "Drained; the next run resumes …". An older-posts walk is refused when nothing is archived and the last timeline fetch finished having read no posts, since it would only repeat empty searches; `"force": true` (`--force` on `archilyzer posts fetch`) walks anyway. A walk with nothing archived that finds nothing ends after two empty three-month windows instead of four, and records why; a walk that has posts keeps the year-of-empty-windows rule. Capture-posts already stopped between posts on Drain. - **Capturing an X post that is an Article also saves the article.** An X Article (a long-form post) is archived as nothing but its link, and gallery-dl cannot read its body. When `pnpm ops capture-posts` meets a post whose archived text, or whose card on the page, links to an article, it now opens the article in the same X profile, after the same 4–10 second pause, and saves beside the post's capture: `article.json` (the title, author, date and every heading, paragraph, quote, list item, image, link and embedded post in reading order, an embedded post by its URL), `article.md` (the same as readable text), `article.png` (the whole article as shown, cut off at 16,000 pixels tall and marked `trimmed` when longer), `article.html` (the article as X served it, so it can be read again without going back to X) and the article's pictures as `article-img-1.jpg`, `article-img-2.png`, … at full size, fetched through the same browser session. `capture.json` records the article's state, title, block count and every file's size and SHA-256. `"articles": false` leaves articles alone; `"shots": false, "media": false, "articles": true` reads only the articles. An article already captured, deleted or unavailable is not opened again unless `"force": true`; one that failed is tried on the next run. If X asks to log in on the article page, the job stops there as it does for a post. The post viewer's capture panel in the editor shows the article's title with a link to `article.md`. - **A report video can play a clip that has only sound, and a clip can be a file beside the manifest.** When a clip's source has no picture, `build-video.mjs` plays it under a poster: a card with the clip's channel, title and date, the size of the picture area, with the sound's waveform moving along its foot (`render.audioPoster.waveform: false` keeps it still). The segment matches every other one in size, frame rate and sound, and the header, footer and on-screen deck are drawn over it as over footage. A video's saved sound (`audio.mp3` and the like in its folder) is now a source the build can cut from, after every saved picture: before any download when the clip has no picture to fetch (`"audioOnly": true` on the clip, `"preferLocalAudio": true` in `render`, a podcast or feed record, or a record with no page). A clip that should have a picture is not quietly played from its sound: with `--no-network` or `--skip-fetch`, one whose picture is missing still stops the build, listed as needing a download with a note that its sound is on disk, so `--no-network` still proves every picture is there. Add `--audio-fallback` to play such clips from their sound under the poster instead; they are listed apart (and logged as `audio-fallback`), and clips that have no picture to fetch are listed as playing from audio only rather than refused. A clip may also give `"src"` (a video or audio file) and `"cues"` (its transcript, either a `transcript.cues.json` or a `parakeet-stitch` transcript), both relative to the manifest, instead of a channel and video: it plays the whole file unless `start`/`end` cut inside it, `resolve-windows.mjs` widens it with those cues, it gets no QR unless it has a `citeUrl`, and a path that leaves the manifest's folder (or an absolute one, without `"allowAbsoluteSrc": true` in `render`), a missing file or an unreadable transcript stops the build before anything runs, naming the clip. - **A report build cuts from media already on disk before it downloads anything, and `--no-network` makes sure it never does.** For each clip, `build-video.mjs` now looks, in order, in the project's own `out/clips-raw`, in the clip windows the editor fetched into the channel (`channels/<slug>/data/<id>/clips/`), and in a saved whole source video (through the saved-video store's pointer, or a `source-media` file still in the video's folder), and cuts from the first that holds the clip plus its fetch pad; only when none does is the window downloaded. A file that is a link to a drive that is not mounted counts as not there, and the next place is tried. The build prints one line per clip naming where its source came from (`raw-cache`, `corpus-window`, `saved-video`, or a network fetch). With `--no-network`, every clip's source is found before anything is rendered, and if any clip would need a download the build stops at once and lists each one (its position in the timeline, channel, video and the span it needs). umtool's clip bench reads the same three places, so a clip it shows as fetched is one the build cuts from without downloading. diff --git a/editor/app/api/ops/fetch-posts/route.test.ts b/editor/app/api/ops/fetch-posts/route.test.ts @@ -0,0 +1,75 @@ +import test from "node:test"; +import assert from "node:assert/strict"; +import { mkdir, mkdtemp, readdir, rm, writeFile } from "node:fs/promises"; +import os from "node:os"; +import path from "node:path"; + +// Run with: +// pnpm -C editor exec tsx --test "app/api/ops/fetch-posts/route.test.ts" +// +// The body's shape and the refusals that come before any job: every case here +// is answered from the disk alone, so a temp corpus with one X channel is the +// whole world — no job is queued and nothing reaches X. + +const ROOT = await mkdtemp(path.join(os.tmpdir(), "fetch-posts-route-")); +const SLUG = "demo-x"; +const CHANNEL = path.join(ROOT, "channels", SLUG); +// Set before the route (and getPaths, which caches) is first imported. +process.env.WORKER_TOKEN = "test-token"; +process.env.TRANSCRIPTS_DIR = ROOT; +process.env.SETTINGS_FILE = path.join(ROOT, "settings.json"); +await mkdir(CHANNEL, { recursive: true }); +await writeFile( + path.join(CHANNEL, "config.json"), + JSON.stringify({ + handling: "transcribe", + sourceKind: "social", + platform: "twitter", + postFetcher: "x-gallery-dl", + socialHandle: "example_user", + name: "Example (X)", + url: "https://x.com/example_user", + }), +); +// An account that shows no posts: nothing archived, and the last timeline +// fetch finished having read none. +await writeFile( + path.join(CHANNEL, "posts-state.json"), + JSON.stringify({ lastFetchedAt: "2026-01-01T00:00:00.000Z", lastFetchedCount: 0 }), +); +const { POST } = await import("./route"); +const { EMPTY_ACCOUNT_OLDER_REFUSAL } = await import( + "yt-dlp-transcript-common/controller/fetchPosts" +); +test.after(() => rm(ROOT, { recursive: true, force: true })); + +async function post(body: Record<string, unknown>): Promise<{ status: number; error: string }> { + const res = await POST( + new Request("http://localhost/api/ops/fetch-posts", { + method: "POST", + headers: { + authorization: "Bearer test-token", + "content-type": "application/json", + }, + body: JSON.stringify(body), + }), + ); + return { status: res.status, error: ((await res.json()) as { error?: string }).error ?? "" }; +} + +test("force must be a boolean, and applies only to an older walk", async () => { + const bad = await post({ slug: SLUG, older: true, force: "yes" }); + assert.equal(bad.status, 400); + assert.match(bad.error, /"force" must be a boolean/); + const alone = await post({ slug: SLUG, force: true }); + assert.equal(alone.status, 400); + assert.equal(alone.error, '"force" applies only to an older-posts fetch.'); +}); + +test("an older walk over an account that shows no posts is refused before any job", async () => { + const res = await post({ slug: SLUG, older: true }); + assert.equal(res.status, 400); + assert.equal(res.error, EMPTY_ACCOUNT_OLDER_REFUSAL); + // No job was ever written. + assert.deepEqual(await readdir(path.join(ROOT, ".jobs")).catch(() => []), []); +}); diff --git a/editor/app/api/ops/fetch-posts/route.ts b/editor/app/api/ops/fetch-posts/route.ts @@ -10,7 +10,7 @@ import { export const dynamic = "force-dynamic"; -// POST { slug, full?, older?, floor?, limit?, queueKey? } -> { ok: true, jobId } +// POST { slug, full?, older?, floor?, force?, limit?, queueKey? } -> { ok: true, jobId } // // A social channel's "Fetch posts" button, over HTTP — and its "Re-fetch full // history" (`full`) and "Fetch older posts" (`older`, with an optional `floor` @@ -19,11 +19,13 @@ export const dynamic = "force-dynamic"; // // Every refusal is the action's own sentence: a channel that is not a social // one, `full` and `older` together, a fetcher with no older walk, a floor -// without `older` or that is not a date. +// without `older` or that is not a date, `force` without `older`, and an older +// walk over an account that shows no posts — nothing archived, and the last +// timeline fetch read none — unless `force` is true. export async function POST(request: Request) { return ops( request, - ["slug", "full", "older", "floor", "limit", "queueKey"], + ["slug", "full", "older", "floor", "force", "limit", "queueKey"], async (body) => { const slug = reqSlug(body, "slug"); return jobResponse( @@ -34,6 +36,7 @@ export async function POST(request: Request) { optPositiveInt(body, "limit"), optBool(body, "older"), optString(body, "floor"), + optBool(body, "force"), ), ); }, diff --git a/editor/app/channels/[slug]/socialActions.ts b/editor/app/channels/[slug]/socialActions.ts @@ -21,6 +21,7 @@ import { } from "yt-dlp-transcript-common/controller/channels"; import { isSocialChannel } from "yt-dlp-transcript-common/lib/channelConfig"; import { + emptyAccountOlderProblem, fetchPosts, FULL_AND_OLDER_REFUSAL, olderPostsProblem, @@ -38,7 +39,10 @@ import { strayCaptureIds, strayIdsRefusal, } from "yt-dlp-transcript-common/controller/capturePosts"; -import { readSeenPostIds } from "yt-dlp-transcript-common/lib/posts-server"; +import { + readPostFetchState, + readSeenPostIds, +} from "yt-dlp-transcript-common/lib/posts-server"; import { getSocialFetcher, listSocialFetchers, @@ -141,7 +145,8 @@ export async function checkPostAvailabilityAction( // (the fetcher's `fetchOlder` — X: search windows) instead of fetching new // posts; `floor` (YYYY-MM-DD) is the date that walk stops at. Both are refused // HERE, before a job exists, when they cannot run: with `full`, on a fetcher -// that has no older walk, or with a floor that is not a date. +// that has no older walk, with a floor that is not a date, or on an account +// that shows no posts (emptyAccountOlderProblem) unless `force`. export async function fetchPostsAction( slug: string, queueKey?: string, @@ -149,11 +154,15 @@ export async function fetchPostsAction( limit?: number, older?: boolean, floor?: string, + force?: boolean, ): Promise<StreamActionResult> { if (full && older) return { ok: false, error: FULL_AND_OLDER_REFUSAL }; if (floor !== undefined && !older) { return { ok: false, error: "A floor date applies only to an older-posts fetch." }; } + if (force && !older) { + return { ok: false, error: "\"force\" applies only to an older-posts fetch." }; + } if (floor !== undefined && !isUtcDay(floor)) { return { ok: false, error: `"${floor}" is not a date (YYYY-MM-DD).` }; } @@ -169,6 +178,14 @@ export async function fetchPostsAction( resolveSocialFetcher(config.postFetcher, config.url), ); if (problem) return { ok: false, error: problem }; + if (!force) { + const channelRoot = path.join(paths.channelsDir, slug); + const empty = emptyAccountOlderProblem( + (await readSeenPostIds(channelRoot)).size, + await readPostFetchState(channelRoot), + ); + if (empty) return { ok: false, error: empty }; + } } // queueKeyForUrl() already routes x.com / bsky.app to platform:x.com / @@ -184,9 +201,9 @@ export async function fetchPostsAction( spec: { kind: "fetch-posts", slug, - params: { queueKey, full, limit, older, floor }, + params: { queueKey, full, limit, older, floor, force }, }, - fn: async (onLog, signal) => { + fn: async (onLog, signal, _progress, ctx) => { const result = await fetchPosts({ paths, slug, @@ -194,9 +211,13 @@ export async function fetchPostsAction( full, older, floor, + force, limit, onLog, signal, + // A drained fetch keeps its resume point and returns ok: the job + // ends done, not failed. + drain: ctx.drainSignal, }); safeRevalidate([`/channels/${slug}`]); // Surface a failed fetch as a failed JOB (the managed wrapper turns a diff --git a/editor/app/jobs/jobReplayRegistry.ts b/editor/app/jobs/jobReplayRegistry.ts @@ -260,6 +260,7 @@ export const JOB_REPLAY_HANDLERS: Record<string, ReplayHandler> = { num(p.limit), bool(p.older), str(p.floor), + bool(p.force), ); }, // The ids are the spec's own (a capture is OF specific posts, unlike a diff --git a/scripts/archilyzer-ops.mjs b/scripts/archilyzer-ops.mjs @@ -330,7 +330,10 @@ export function usage() { ' re-walks the whole timeline; "older": true walks back from the oldest', " archived post through search (X; needs a login), saving its place for", ' the next run, down to "floor": "YYYY-MM-DD" when given. "limit": N caps', - ' the posts one run reads. "full" and "older" together are refused.', + ' the posts one run reads. "full" and "older" together are refused. An', + ' older walk over an account that shows no posts (nothing archived, and', + ' the last timeline fetch read none) is refused unless "force": true. A', + " drained fetch stops at its next resume point and the next run resumes.", "", 'capture-posts captures archived posts of a social channel (X): a', ' screenshot of each through the connected X profile, and its attached', diff --git a/scripts/archilyzer-ops.test.mjs b/scripts/archilyzer-ops.test.mjs @@ -389,6 +389,7 @@ test("fetch-posts is a POST to its route, named in the usage", () => { assert.equal(p.wait, true); assert.match(usage(), /Actions:.*transcribe-bucket, fetch-posts/); assert.match(usage(), /"older": true walks back from the oldest/); + assert.match(usage(), /read none\) is refused unless "force": true/); }); // A post capture: a POST to its route, the body passed through untouched — the