Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 21982be9852a50527e9ac65d3a9f4ef0ef40aa2e
parent dd55e7f7c97c335a28f9802b20744058e2ac88cd
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Thu,  6 Aug 2026 18:52:38 -0400

Add an account-free X fetcher via Nitter, and a scraping-method picker

x-nitter: read an X timeline through a Nitter instance, in a real browser.

Why a third X path: the guest-token GraphQL route is throttled hard (gallery-dl
blocks for minutes and returned only a few stale posts for the account this was
built for), and the authenticated route needs an account kept alive. A Nitter
instance does the talking to X, so there is no account to juggle and no X quota
to burn.

Why a browser rather than plain HTTP: nitter.poast.org answers non-browser
clients with a hard 503 — curl, and gallery-dl's own Nitter extractor, both get
it — while a real browser gets the fully rendered page. So the HTTP status is
NOT the success signal here; presence of `.timeline-item` is. Public instances
are short-lived, so the list is ordered, tried in turn, and overridable with
NITTER_INSTANCES.

Verified live against the real channel: 60 posts over 4 cursor-paginated pages,
no credentials, ids intact, permalinks pointing at x.com rather than at the
instance (which will likely be dead by the time anyone clicks).

Two defects found while building it, both the same shape as the earlier 64-bit
id bug — a test that passed while the real path was broken:
  - The stats scraper matched the wrapping `div.icon-container` before the
    inner `span.icon-comment`, so every post archived with zero engagement.
    The unit test fed synthetic icon names and never exercised the selector.
  - Playwright could not be resolved at all from `common` (it installs only
    under editor/ and export/ in this pnpm workspace), and @playwright/test is
    CommonJS so a resolved-path import yields `{default:{chromium}}`. Both are
    now handled in one shared playwrightRuntime module, which the session
    broker and the x-playwright fallback also use instead of three separate
    fragile copies.

Also: the scraping method is now switchable from the channel page. It was fixed
at creation time and only editable by hand, which is wrong for a choice that
decides how much history you can actually reach — the picker lists only the
fetchers for that channel's platform, so a Bluesky channel (one method) shows
none.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Diffstat:
Mcommon/controller/fetchPosts.ts | 1+
Acommon/social/playwrightRuntime.ts | 113+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/social/xNitterFetcher.test.ts | 132+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/social/xNitterFetcher.ts | 384+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/social/xPlaywrightFetcher.ts | 31+------------------------------
Mcommon/social/xSessionBroker.ts | 31+------------------------------
Meditor/app/channels/[slug]/components/SocialChannelPanel.tsx | 57+++++++++++++++++++++++++++++++++++++++++++++++++++++++--
Meditor/app/channels/[slug]/page.tsx | 4++++
Meditor/app/channels/[slug]/socialActions.ts | 58+++++++++++++++++++++++++++++++++++++++++++++++++++++++++-
Meditor/app/channels/actions.ts | 1+
Aeditor/e2e/social-fetcher-picker.spec.ts | 61+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
11 files changed, 810 insertions(+), 63 deletions(-)

diff --git a/common/controller/fetchPosts.ts b/common/controller/fetchPosts.ts @@ -37,6 +37,7 @@ import "../social/xGalleryDlFetcher"; // The fallback is registered LAST and never claims a URL by detection, so it is // only ever used when a channel opts into it via postFetcher: "x-playwright". import "../social/xPlaywrightFetcher"; +import "../social/xNitterFetcher"; export type FetchPostsOptions = { paths: Paths; diff --git a/common/social/playwrightRuntime.ts b/common/social/playwrightRuntime.ts @@ -0,0 +1,113 @@ +// Locating Playwright at runtime, for the modules that need a real browser +// (the Nitter fetcher, the X fallback fetcher and the X session broker). +// +// Two constraints pull against each other: +// - `common` must NOT depend on Playwright. It is imported by the Docker +// build image, which installs no browsers, and by the client bundle path. +// So there can be no static import and no entry in its package.json. +// - The code still has to FIND Playwright when it runs, and in this pnpm +// workspace it is installed only under `editor/` and `export/` — not +// hoisted to the root and not visible from `common/`. A bare dynamic +// import resolves relative to THIS file and therefore fails, which is +// exactly how the Nitter fetcher first failed from the CLI. +// +// So: try the bare specifier first (works when the importing app owns it, e.g. +// an editor server action), then resolve from the running process's directory, +// then from the workspace packages that actually install it. + +import { createRequire } from "node:module"; +import path from "node:path"; +import { pathToFileURL } from "node:url"; +import { getPaths } from "../lib/paths"; + +// The slice of Playwright's chromium API these fetchers use. Structural, so no +// Playwright types are needed at compile time. +export type ChromiumLike = { + launch: (opts: Record<string, unknown>) => Promise<BrowserLike>; + launchPersistentContext: ( + dir: string, + opts: Record<string, unknown>, + ) => Promise<BrowserContextLike>; +}; + +export type BrowserLike = { + newContext: (opts: Record<string, unknown>) => Promise<BrowserContextLike>; + close: () => Promise<void>; +}; + +export type PageLike = { + goto: (url: string, opts?: unknown) => Promise<unknown>; + waitForSelector: (sel: string, opts?: unknown) => Promise<unknown>; + waitForTimeout: (ms: number) => Promise<void>; + evaluate: (fn: string) => Promise<unknown>; + on: (event: string, cb: (arg: never) => void) => void; +}; + +export type BrowserContextLike = { + newPage: () => Promise<PageLike>; + pages: () => PageLike[]; + cookies: () => Promise<unknown[]>; + close: () => Promise<void>; + on: (event: string, cb: () => void) => void; +}; + +// Assembled at runtime so TypeScript never tries to resolve it. +const SPECIFIER = ["@playwright", "test"].join("/"); + +function candidateDirs(): string[] { + const dirs = [process.cwd()]; + try { + const root = getPaths().monorepoRoot; + // The workspace packages that actually declare @playwright/test. + dirs.push(path.join(root, "editor"), path.join(root, "export"), root); + } catch { + /* getPaths can throw outside the repo — the cwd attempt still stands */ + } + return dirs; +} + +// @playwright/test is CommonJS, so `await import()` of a resolved file path +// yields `{ default: { chromium, ... } }` while a bare specifier may yield the +// namespace directly. Accept either rather than silently handing back an +// object whose `chromium` is undefined. +function unwrap(mod: unknown): { chromium: ChromiumLike } | null { + const m = mod as { chromium?: ChromiumLike; default?: { chromium?: ChromiumLike } }; + const chromium = m?.chromium ?? m?.default?.chromium; + return chromium ? { chromium } : null; +} + +export async function importPlaywright(): Promise<{ chromium: ChromiumLike }> { + const attempts: string[] = []; + + // 1. Plain specifier — succeeds when the importing app has it in scope. + try { + const hit = unwrap(await import(/* webpackIgnore: true */ SPECIFIER)); + if (hit) return hit; + attempts.push("bare: no chromium export"); + } catch (err) { + attempts.push(`bare: ${(err as Error).message.slice(0, 60)}`); + } + + // 2. Resolve from the running process / workspace packages. + for (const dir of candidateDirs()) { + try { + // createRequire needs a file path, not a directory. + const req = createRequire(path.join(dir, "__resolve__.cjs")); + const resolved = req.resolve(SPECIFIER); + const hit = unwrap( + await import(/* webpackIgnore: true */ pathToFileURL(resolved).href), + ); + if (hit) return hit; + attempts.push(`${dir}: no chromium export`); + } catch (err) { + attempts.push(`${dir}: ${(err as Error).message.slice(0, 40)}`); + } + } + + throw new Error( + "Playwright is not available on this host. It ships with the editor and " + + "export packages; it is intentionally absent from `common` and from the " + + "minimal Docker build image, so browser-backed fetching is an " + + `editor-host concern. Tried — ${attempts.join(" | ")}`, + ); +} diff --git a/common/social/xNitterFetcher.test.ts b/common/social/xNitterFetcher.test.ts @@ -0,0 +1,132 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + cursorFromShowMore, + nitterHandleFromHref, + nitterIdFromHref, + normalizeNitterItem, + parseNitterDate, + parseNitterStat, + xNitterFetcher, + type NitterRawItem, +} from "./xNitterFetcher"; +import { parsePost } from "../lib/posts"; + +// Every literal below was captured from a real nitter.poast.org render of +// /TheQuartering — no network in tests. + +test("parses nitter's UTC timestamp format", () => { + assert.equal(parseNitterDate("Aug 6, 2026 · 8:51 PM UTC"), "2026-08-06T20:51:00.000Z"); + assert.equal(parseNitterDate("Aug 5, 2026 · 1:13 AM UTC"), "2026-08-05T01:13:00.000Z"); + // Midnight/noon are the classic 12-hour conversion traps. + assert.equal(parseNitterDate("Jan 1, 2024 · 12:00 AM UTC"), "2024-01-01T00:00:00.000Z"); + assert.equal(parseNitterDate("Jan 1, 2024 · 12:30 PM UTC"), "2024-01-01T12:30:00.000Z"); + // A drifted format must be rejected, not mis-dated. + assert.equal(parseNitterDate("6 August 2026 20:51"), null); + assert.equal(parseNitterDate(""), null); +}); + +test("extracts a 64-bit id from the tweet link, as a string", () => { + const href = "/TheQuartering/status/2085468926012702808#m"; + assert.equal(nitterIdFromHref(href), "2085468926012702808"); + assert.equal(nitterHandleFromHref(href), "TheQuartering"); + // Never through a Number — X ids exceed 2^53. + assert.ok(!Number.isSafeInteger(Number("2085468926012702808"))); + assert.equal(nitterIdFromHref("/TheQuartering"), null); +}); + +test("parses stat counts, including thousands separators", () => { + assert.equal(parseNitterStat("7,602"), 7602); + assert.equal(parseNitterStat("47"), 47); + assert.equal(parseNitterStat(""), undefined); + assert.equal(parseNitterStat(null), undefined); + assert.equal(parseNitterStat("—"), undefined); +}); + +test("reads the pagination cursor off the Load-more link", () => { + assert.equal( + cursorFromShowMore("?cursor=DAAHCgABHPEsST3__-sLAAIAAAATMjA4NDgx"), + "DAAHCgABHPEsST3__-sLAAIAAAATMjA4NDgx", + ); + assert.equal(cursorFromShowMore(null), null); + assert.equal(cursorFromShowMore("/TheQuartering"), null); +}); + +const REAL_ITEM: NitterRawItem = { + href: "/TheQuartering/status/2085468926012702808#m", + dateTitle: "Aug 6, 2026 · 8:51 PM UTC", + text: "Coffee Brand Coffee has just shipped our first shipment to Walmart. seethe losers.", + fullname: "TheQuartering", + username: "@TheQuartering", + isRetweet: false, + isReply: false, + replyingTo: [], + quoteHref: null, + stats: [ + { icon: "icon-comment", value: "47" }, + { icon: "icon-retweet", value: "6" }, + { icon: "icon-heart", value: "194" }, + { icon: "icon-quote", value: "7,602" }, + ], + mediaCount: 0, +}; + +test("normalizes a real nitter item into a valid Post", () => { + const post = normalizeNitterItem(REAL_ITEM, "thequartering-X")!; + assert.equal(post.id, "2085468926012702808"); + assert.equal(post.slug, "thequartering-X/2085468926012702808"); + assert.equal(post.platform, "twitter"); + assert.equal(post.author, "TheQuartering"); + assert.equal(post.authorName, "TheQuartering"); + assert.equal(post.createdAt, "2026-08-06T20:51:00.000Z"); + assert.equal(post.uploadDate, "20260806"); + assert.match(post.text, /^Coffee Brand Coffee/); + assert.deepEqual(post.engagement, { + replies: 47, + reposts: 6, + likes: 194, + quotes: 7602, + }); + // Survives the on-disk round trip. + assert.deepEqual(parsePost(JSON.parse(JSON.stringify(post))), post); +}); + +test("links to x.com, never to the nitter instance", () => { + // The instance is a transport detail and is likely dead by the time anyone + // clicks the citation. + const post = normalizeNitterItem(REAL_ITEM, "c")!; + assert.equal(post.url, "https://x.com/TheQuartering/status/2085468926012702808"); + assert.ok(!post.url.includes("nitter")); +}); + +test("carries retweet / reply / quote structure", () => { + const rt = normalizeNitterItem({ ...REAL_ITEM, isRetweet: true }, "c")!; + assert.equal(rt.isRepost, true); + + const reply = normalizeNitterItem( + { ...REAL_ITEM, isReply: true, replyingTo: ["someoneelse"] }, + "c", + )!; + assert.equal(reply.isReply, true); + assert.equal(reply.replyTo?.author, "someoneelse"); + + const quote = normalizeNitterItem( + { ...REAL_ITEM, quoteHref: "/OtherUser/status/1788248211448004625#m" }, + "c", + )!; + assert.equal(quote.quoted?.id, "1788248211448004625"); + assert.equal(quote.quoted?.author, "OtherUser"); +}); + +test("rejects an item with no id or an unparseable date", () => { + assert.equal(normalizeNitterItem({ ...REAL_ITEM, href: "/TheQuartering" }, "c"), null); + assert.equal(normalizeNitterItem({ ...REAL_ITEM, dateTitle: "yesterday" }, "c"), null); +}); + +test("claims nitter URLs only, so gallery-dl stays the default X path", () => { + assert.equal(xNitterFetcher.detect("https://nitter.poast.org/TheQuartering"), true); + assert.equal(xNitterFetcher.detect("https://nitter.net/nasa"), true); + // An x.com channel must opt in explicitly via postFetcher: "x-nitter". + assert.equal(xNitterFetcher.detect("https://x.com/TheQuartering"), false); + assert.equal(xNitterFetcher.platform, "twitter"); +}); diff --git a/common/social/xNitterFetcher.ts b/common/social/xNitterFetcher.ts @@ -0,0 +1,384 @@ +// X/Twitter via a Nitter instance, driven through a real browser. +// +// Why this exists as a THIRD X path: the guest-token GraphQL route +// (x-gallery-dl without cookies) is throttled hard and returned only a handful +// of stale posts for some accounts, while an authenticated route needs an +// account to keep alive. A Nitter instance does the talking to X for us, so +// there is no account to juggle and no X quota to burn — the archive only ever +// talks to the instance. +// +// Why a BROWSER rather than plain HTTP: nitter.poast.org answers non-browser +// clients with a hard `503` (verified — curl and gallery-dl both get it, and +// gallery-dl's own Nitter extractor therefore fails), but a real browser gets +// the fully rendered page. The browser IS the mechanism that works, so this +// fetcher owns one. +// +// Because of that, the HTTP status is NOT authoritative: poast serves 503 +// alongside a perfectly good timeline. Presence of `.timeline-item` is the +// real success signal. +// +// Public Nitter instances are famously short-lived, so the instance list is +// configurable and tried in order. + +import { + postPermalink, + uploadDateFromCreatedAt, + type Post, + type PostRef, +} from "../lib/posts"; +import { importPlaywright } from "./playwrightRuntime"; +import { + registerSocialFetcher, + type PostFetchInput, + type PostFetchResult, + type SocialFetcher, + type SocialFetcherProbe, +} from "./fetchers"; + +// Ordered by preference. Override with NITTER_INSTANCES (comma-separated) — +// instances die often enough that hardcoding one is a liability. +export function nitterInstances(): string[] { + const configured = process.env.NITTER_INSTANCES?.trim(); + if (configured) { + return configured + .split(",") + .map((s) => s.trim().replace(/\/+$/, "")) + .filter(Boolean); + } + return [ + "https://nitter.poast.org", + "https://nitter.net", + "https://nitter.space", + ]; +} + +// Bound how far back one run walks, so a huge account can't spin forever. +const MAX_PAGES = 40; + +// --------------------------------------------------------------------------- +// Pure parsing helpers (no DOM, no network — unit-tested against real values) +// --------------------------------------------------------------------------- + +const MONTHS: Record<string, number> = { + Jan: 1, Feb: 2, Mar: 3, Apr: 4, May: 5, Jun: 6, + Jul: 7, Aug: 8, Sep: 9, Oct: 10, Nov: 11, Dec: 12, +}; + +// Nitter renders its timestamps as `Aug 6, 2026 · 8:51 PM UTC` in the +// `.tweet-date a[title]` attribute. Always UTC, so the conversion is exact +// rather than locale-dependent. Returns an ISO-8601 string, or null when the +// format drifts (in which case the item is skipped rather than mis-dated). +export function parseNitterDate(title: string): string | null { + const m = + /^([A-Z][a-z]{2})\s+(\d{1,2}),\s+(\d{4})\s*·\s*(\d{1,2}):(\d{2})\s*(AM|PM)\s*UTC$/.exec( + title.trim(), + ); + if (!m) return null; + const [, mon, day, year, hourRaw, minute, meridiem] = m; + const month = MONTHS[mon]; + if (!month) return null; + let hour = Number(hourRaw) % 12; + if (meridiem === "PM") hour += 12; + const iso = `${year}-${String(month).padStart(2, "0")}-${String(Number(day)).padStart(2, "0")}T${String(hour).padStart(2, "0")}:${minute}:00.000Z`; + return Number.isFinite(Date.parse(iso)) ? iso : null; +} + +// `/TheQuartering/status/2085468926012702808#m` -> the id, AS A STRING. +// X ids exceed 2^53, so they must never round-trip through a number. +export function nitterIdFromHref(href: string): string | null { + const m = /\/status\/(\d+)/.exec(href); + return m ? m[1] : null; +} + +// `/TheQuartering/status/123#m` -> `TheQuartering` +export function nitterHandleFromHref(href: string): string | null { + const m = /^\/([^/]+)\/status\/\d+/.exec(href); + return m ? m[1] : null; +} + +// Nitter prints counts with thousands separators ("7,602") and an empty string +// when a stat is zero/absent. +export function parseNitterStat(raw: string | null | undefined): number | undefined { + if (!raw) return undefined; + const cleaned = raw.replace(/[,\s]/g, ""); + if (!/^\d+$/.test(cleaned)) return undefined; + return Number(cleaned); +} + +// The `?cursor=…` of the "Load more" link, which is how Nitter paginates. +export function cursorFromShowMore(href: string | null | undefined): string | null { + if (!href) return null; + const m = /[?&]cursor=([^&]+)/.exec(href); + return m ? decodeURIComponent(m[1]) : null; +} + +// One `.timeline-item`, already flattened out of the DOM by the page-side +// scraper below. Kept as a plain shape so normalization is pure and testable. +export type NitterRawItem = { + href: string | null; + dateTitle: string | null; + text: string | null; + fullname: string | null; + username: string | null; // "@handle" + isRetweet: boolean; + isReply: boolean; + replyingTo: string[]; // handles, without "@" + quoteHref: string | null; + stats: { icon: string; value: string }[]; + mediaCount: number; +}; + +export function normalizeNitterItem( + item: NitterRawItem, + channelSlug: string, +): Post | null { + const href = item.href ?? ""; + const id = nitterIdFromHref(href); + if (!id) return null; + const createdAt = item.dateTitle ? parseNitterDate(item.dateTitle) : null; + if (!createdAt) return null; + + const author = + (item.username ?? "").replace(/^@/, "").trim() || + nitterHandleFromHref(href) || + ""; + + const stat = (icon: string) => + parseNitterStat(item.stats.find((s) => s.icon.includes(icon))?.value); + + const post: Post = { + id, + slug: `${channelSlug}/${id}`, + channelSlug, + author, + createdAt, + uploadDate: uploadDateFromCreatedAt(createdAt), + text: (item.text ?? "").trim(), + // Always link to x.com, never to the Nitter instance: the instance is a + // transport detail and is likely to be dead by the time anyone clicks. + url: postPermalink("twitter", author, id), + platform: "twitter", + isReply: item.isReply, + isRepost: item.isRetweet, + links: [], + threadId: id, + }; + + if (item.fullname) post.authorName = item.fullname; + if (item.mediaCount > 0) post.mediaCount = item.mediaCount; + + const parent = item.replyingTo[0]; + if (parent) { + // Nitter names the handle being replied to but not the parent's id, so the + // ref carries the author only — enough to render "replying to @x". + const ref: PostRef = { platform: "twitter", id: "", author: parent }; + if (ref.author) post.replyTo = ref; + } + const quotedId = item.quoteHref ? nitterIdFromHref(item.quoteHref) : null; + if (quotedId) { + const quotedAuthor = item.quoteHref + ? nitterHandleFromHref(item.quoteHref) + : null; + post.quoted = { + platform: "twitter", + id: quotedId, + ...(quotedAuthor + ? { author: quotedAuthor, url: postPermalink("twitter", quotedAuthor, quotedId) } + : {}), + }; + } + + const engagement: NonNullable<Post["engagement"]> = {}; + const replies = stat("comment"); + const reposts = stat("retweet"); + const likes = stat("heart"); + const quotes = stat("quote"); + if (likes !== undefined) engagement.likes = likes; + if (reposts !== undefined) engagement.reposts = reposts; + if (replies !== undefined) engagement.replies = replies; + if (quotes !== undefined) engagement.quotes = quotes; + if (Object.keys(engagement).length > 0) post.engagement = engagement; + + return post; +} + +// The page-side scraper, as a string so it can be handed to page.evaluate() +// without pulling DOM types into this module. Returns NitterRawItem[] plus the +// next cursor. +const SCRAPE = `(() => { + const items = [...document.querySelectorAll('.timeline-item')].map((el) => ({ + href: el.querySelector('a.tweet-link')?.getAttribute('href') ?? null, + dateTitle: el.querySelector('.tweet-date a')?.getAttribute('title') ?? null, + text: el.querySelector('.tweet-content')?.textContent ?? null, + fullname: el.querySelector('a.fullname')?.getAttribute('title') ?? null, + username: el.querySelector('a.username')?.getAttribute('title') ?? null, + isRetweet: !!el.querySelector('.retweet-header'), + isReply: !!el.querySelector('.replying-to'), + replyingTo: [...el.querySelectorAll('.replying-to a')].map((a) => (a.textContent || '').replace(/^@/, '').trim()).filter(Boolean), + quoteHref: el.querySelector('.quote a.quote-link')?.getAttribute('href') + ?? el.querySelector('.quote .tweet-link')?.getAttribute('href') ?? null, + stats: [...el.querySelectorAll('.tweet-stat')].map((s) => ({ + // The icon is the inner SPAN (span.icon-comment); the wrapping + // div.icon-container also matches [class*="icon-"] and would shadow it, + // which silently produced zero engagement stats. + icon: (s.querySelector('span[class*="icon-"]')?.className) || '', + value: (s.textContent || '').trim(), + })), + mediaCount: el.querySelectorAll('.attachments .attachment').length, + })); + const more = [...document.querySelectorAll('.show-more a')].pop(); + return { items, next: more ? more.getAttribute('href') : null }; +})()`; + +export const xNitterFetcher: SocialFetcher = { + id: "x-nitter", + label: "X / Twitter (Nitter, no account)", + platform: "twitter", + fields: { limit: true }, + + // Claims Nitter URLs only. An x.com channel opts in explicitly via + // `postFetcher: "x-nitter"`, so URL detection keeps preferring gallery-dl. + detect(url: string): boolean { + try { + return /(^|\.)nitter\./i.test(new URL(url).hostname); + } catch { + return false; + } + }, + + async probe(url: string): Promise<SocialFetcherProbe> { + const handle = handleOf(url); + if (!handle) { + return { ok: false, error: "Could not read an X handle from that URL" }; + } + // No network: a probe would cost a full browser launch for a handle we can + // already read off the URL. + return { ok: true, name: handle, handle, url: `https://x.com/${handle}` }; + }, + + async fetch(input: PostFetchInput): Promise<PostFetchResult> { + const { accountUrl, handle: rawHandle, channelSlug, since, seenIds, limit, signal, onLog } = + input; + const handle = rawHandle || handleOf(accountUrl); + if (!handle) throw new Error("No X handle to fetch"); + + const { chromium } = await importPlaywright(); + const browser = await chromium.launch({ headless: true }); + const posts: Post[] = []; + let complete = false; + let lastError = ""; + + try { + const ctx = await browser.newContext({ + userAgent: + "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36", + }); + const page = await ctx.newPage(); + + for (const instance of nitterInstances()) { + if (signal.aborted) break; + posts.length = 0; + let cursor: string | null = input.cursor ?? null; + let pages = 0; + let ok = false; + + try { + for (; pages < MAX_PAGES; pages++) { + if (signal.aborted) break; + const url = `${instance}/${handle}${cursor ? `?cursor=${encodeURIComponent(cursor)}` : ""}`; + await page.goto(url, { waitUntil: "domcontentloaded", timeout: 45_000 }); + // The HTTP status is NOT the success signal — poast serves 503 + // alongside a good timeline. Wait for real content instead. + await page + .waitForSelector(".timeline-item", { timeout: 20_000 }) + .catch(() => {}); + + const { items, next } = (await page.evaluate(SCRAPE)) as { + items: NitterRawItem[]; + next: string | null; + }; + if (items.length === 0) { + if (pages === 0) throw new Error("no timeline items rendered"); + complete = true; + break; + } + ok = true; + + let stop = false; + for (const raw of items) { + const post = normalizeNitterItem(raw, channelSlug); + if (!post) continue; + if (seenIds.has(post.id)) { + stop = true; + continue; + } + if (since && post.createdAt <= since) { + stop = true; + continue; + } + posts.push(post); + } + onLog?.( + `Nitter page ${pages + 1} (${instanceHost(instance)}): ${items.length} items, ${posts.length} new so far`, + ); + + if (stop) { + complete = true; + break; + } + if (limit && posts.length >= limit) { + posts.length = limit; + break; + } + cursor = cursorFromShowMore(next); + if (!cursor) { + complete = true; + break; + } + } + } catch (err) { + lastError = (err as Error).message; + onLog?.( + `Nitter instance ${instanceHost(instance)} failed: ${lastError} — trying the next one.`, + ); + continue; + } + + if (ok) { + return { posts, complete, ...(cursor && !complete ? { cursor } : {}) }; + } + } + } finally { + await browser.close().catch(() => {}); + } + + throw new Error( + `No Nitter instance returned a timeline for @${handle}` + + (lastError ? ` (last error: ${lastError})` : "") + + ". Public instances are short-lived — set NITTER_INSTANCES to a working one.", + ); + }, +}; + +function instanceHost(instance: string): string { + try { + return new URL(instance).hostname; + } catch { + return instance; + } +} + +function handleOf(url: string): string | null { + const trimmed = (url ?? "").trim(); + if (!trimmed) return null; + if (!/^https?:\/\//i.test(trimmed)) return trimmed.replace(/^@/, "") || null; + try { + const segments = new URL(trimmed).pathname.split("/").filter(Boolean); + return segments[0] ? decodeURIComponent(segments[0]).replace(/^@/, "") : null; + } catch { + return null; + } +} + + +registerSocialFetcher(xNitterFetcher); diff --git a/common/social/xPlaywrightFetcher.ts b/common/social/xPlaywrightFetcher.ts @@ -33,6 +33,7 @@ import type { Post } from "../lib/posts"; import { normalizeXTweets, type XTweetRaw } from "./xNormalize"; import { xProfileDir } from "./xSessionBroker"; import { getPaths } from "../lib/paths"; +import { importPlaywright } from "./playwrightRuntime"; import { registerSocialFetcher, type PostFetchInput, @@ -217,35 +218,5 @@ export const xPlaywrightFetcher: SocialFetcher = { }, }; -type PlaywrightPage = { - goto: (url: string, opts?: unknown) => Promise<unknown>; - evaluate: (fn: string) => Promise<unknown>; - waitForTimeout: (ms: number) => Promise<void>; - on: (event: string, cb: (res: never) => void) => void; -}; - -async function importPlaywright(): Promise<{ - chromium: { - launchPersistentContext: ( - dir: string, - opts: Record<string, unknown>, - ) => Promise<{ - pages: () => PlaywrightPage[]; - newPage: () => Promise<PlaywrightPage>; - close: () => Promise<void>; - }>; - }; -}> { - try { - // Assembled at runtime so TypeScript does not resolve it — `common` does - // not depend on Playwright (see xSessionBroker.ts for the same rationale). - const specifier = ["@playwright", "test"].join("/"); - return (await import(/* webpackIgnore: true */ specifier)) as never; - } catch { - throw new Error( - "Playwright is not available on this host — the X fallback fetcher needs it.", - ); - } -} registerSocialFetcher(xPlaywrightFetcher); diff --git a/common/social/xSessionBroker.ts b/common/social/xSessionBroker.ts @@ -22,6 +22,7 @@ import path from "node:path"; import { mkdir, readFile, rename, writeFile, rm, stat } from "node:fs/promises"; import type { Paths } from "../lib/paths"; +import { importPlaywright } from "./playwrightRuntime"; // Where the persistent browser profile lives. One profile per instance: a // single X identity is all the archive needs. @@ -228,33 +229,3 @@ export async function clearXSession(paths: Paths): Promise<void> { await rm(path.dirname(xCookieFile(paths)), { recursive: true, force: true }); } -// Dynamic import so a host without Playwright can still import this module. -async function importPlaywright(): Promise<{ - chromium: { - launchPersistentContext: ( - dir: string, - opts: Record<string, unknown>, - ) => Promise<{ - pages: () => { goto: (u: string, o?: unknown) => Promise<unknown> }[]; - newPage: () => Promise<{ goto: (u: string, o?: unknown) => Promise<unknown> }>; - cookies: () => Promise<unknown[]>; - close: () => Promise<void>; - on: (event: string, cb: () => void) => void; - }>; - }; -}> { - try { - // The specifier is assembled at runtime so TypeScript does not try to - // resolve it: `common` deliberately does NOT depend on Playwright (the - // editor and export packages do, and they are what run this code). A - // static import would break `common`'s typecheck and its Docker install. - const specifier = ["@playwright", "test"].join("/"); - return (await import(/* webpackIgnore: true */ specifier)) as never; - } catch { - throw new Error( - "Playwright is not available on this host — the X session broker needs " + - "it (it ships with the editor; it is intentionally absent from the " + - "minimal Docker build image).", - ); - } -} diff --git a/editor/app/channels/[slug]/components/SocialChannelPanel.tsx b/editor/app/channels/[slug]/components/SocialChannelPanel.tsx @@ -6,8 +6,8 @@ // Its only failure surfaces are "the fetch failed" and "we need credentials", // which is exactly what the state sidecar records. -import { useTransition } from "react"; -import { fetchPostsAction } from "../socialActions"; +import { useState, useTransition } from "react"; +import { fetchPostsAction, setPostFetcherAction } from "../socialActions"; export type SocialChannelState = { postCount: number; @@ -19,6 +19,9 @@ export type SocialChannelState = { fetcherLabel: string; handle: string; accountUrl?: string; + // Which fetcher currently drives this channel, and every fetcher that could. + fetcherId: string; + availableFetchers: { id: string; label: string }[]; }; export function SocialChannelPanel({ @@ -29,6 +32,22 @@ export function SocialChannelPanel({ state: SocialChannelState; }) { const [pending, startTransition] = useTransition(); + const [fetcherId, setFetcherId] = useState(state.fetcherId); + const [switchNote, setSwitchNote] = useState<string | null>(null); + const [switchErr, setSwitchErr] = useState<string | null>(null); + + const switchFetcher = (next: string) => + startTransition(async () => { + setSwitchNote(null); + setSwitchErr(null); + const res = await setPostFetcherAction(slug, next); + if (res.ok) { + setFetcherId(next); + setSwitchNote("Fetcher updated — the next fetch uses it."); + } else { + setSwitchErr(res.error); + } + }); const run = (full: boolean) => startTransition(async () => { @@ -72,6 +91,40 @@ export function SocialChannelPanel({ /> </dl> + {state.availableFetchers.length > 1 && ( + <label className="mt-4 flex flex-col gap-1 text-sm"> + <span className="font-medium">Scraping method</span> + <select + value={fetcherId} + onChange={(e) => switchFetcher(e.target.value)} + disabled={pending} + aria-label="post fetcher" + className="max-w-sm rounded border border-border bg-card px-2 py-1 text-sm" + > + {state.availableFetchers.map((f) => ( + <option key={f.id} value={f.id}> + {f.label} + </option> + ))} + </select> + <span className="text-xs text-muted-foreground"> + These reach different amounts of history. For X, the Nitter path + needs no account at all; gallery-dl can go deeper but is rate- + limited without credentials. + </span> + </label> + )} + {switchNote && ( + <p role="status" className="mt-2 text-xs text-success"> + {switchNote} + </p> + )} + {switchErr && ( + <p role="alert" className="mt-2 text-xs text-destructive"> + {switchErr} + </p> + )} + <div className="mt-4 flex flex-wrap gap-2"> <button type="button" diff --git a/editor/app/channels/[slug]/page.tsx b/editor/app/channels/[slug]/page.tsx @@ -15,7 +15,9 @@ import { resolveSocialFetcher } from "yt-dlp-transcript-common/social/fetchers"; import "yt-dlp-transcript-common/social/blueskyFetcher"; import "yt-dlp-transcript-common/social/xGalleryDlFetcher"; import "yt-dlp-transcript-common/social/xPlaywrightFetcher"; +import "yt-dlp-transcript-common/social/xNitterFetcher"; import { SocialChannelPanel } from "./components/SocialChannelPanel"; +import { listPostFetchersFor } from "./socialActions"; import { NoReportYet } from "./components/NoReportYet"; import type { ChannelSnapshot } from "yt-dlp-transcript-common/controller/channelSnapshot"; import { @@ -201,6 +203,8 @@ export default async function ChannelDetailPage({ lastError: fetchState?.lastError, needsCookies: fetchState?.needsCookies, fetcherLabel: fetcher?.label ?? "no fetcher configured", + fetcherId: fetcher?.id ?? "", + availableFetchers: await listPostFetchersFor(config.platform ?? ""), handle: config.socialHandle ?? slug, accountUrl: config.url, }} diff --git a/editor/app/channels/[slug]/socialActions.ts b/editor/app/channels/[slug]/socialActions.ts @@ -13,14 +13,70 @@ import { downloadQueueKey, resolveQueueKey, } from "yt-dlp-transcript-common/lib/queueKeys"; -import { readChannelConfig } from "yt-dlp-transcript-common/controller/channels"; +import { + readChannelConfig, + writeChannelConfig, +} from "yt-dlp-transcript-common/controller/channels"; import { isSocialChannel } from "yt-dlp-transcript-common/lib/channelConfig"; import { fetchPosts } from "yt-dlp-transcript-common/controller/fetchPosts"; import { + getSocialFetcher, + listSocialFetchers, +} from "yt-dlp-transcript-common/social/fetchers"; +import { runManagedFunction, type StreamActionResult, } from "yt-dlp-transcript-common/jobs/streamCommand"; +// Switch which SocialFetcher drives this channel. The fetchers differ in what +// they can actually reach — for X, gallery-dl needs credentials for depth, +// while the Nitter path needs none — so this is an operational choice that +// belongs on the channel page rather than only at creation time. +export async function setPostFetcherAction( + slug: string, + fetcherId: string, +): Promise<{ ok: true } | { ok: false; error: string }> { + const paths = getPaths(); + const config = await readChannelConfig(paths, slug); + if (!config) return { ok: false, error: `No such channel: ${slug}` }; + if (!isSocialChannel(config)) { + return { ok: false, error: `${slug} is not a social channel.` }; + } + await registerBuiltinSocialFetchers(); + const wanted = fetcherId.trim(); + const fetcher = getSocialFetcher(wanted); + if (!fetcher) return { ok: false, error: `Unknown fetcher: ${wanted}` }; + if (fetcher.platform !== config.platform) { + return { + ok: false, + error: `${fetcher.label} handles ${fetcher.platform}, not ${config.platform}.`, + }; + } + await writeChannelConfig(paths, slug, { ...config, postFetcher: wanted }); + revalidatePath(`/channels/${slug}`); + return { ok: true }; +} + +// The fetchers that can drive this channel's platform, as client-safe +// descriptors for the picker. +export async function listPostFetchersFor( + platform: string, +): Promise<{ id: string; label: string }[]> { + await registerBuiltinSocialFetchers(); + return listSocialFetchers() + .filter((f) => f.platform === platform) + .map((f) => ({ id: f.id, label: f.label })); +} + +// Deferred side-effect imports so the heavy fetchers stay out of any module +// graph a client component pulls in. +async function registerBuiltinSocialFetchers(): Promise<void> { + await import("yt-dlp-transcript-common/social/blueskyFetcher"); + await import("yt-dlp-transcript-common/social/xGalleryDlFetcher"); + await import("yt-dlp-transcript-common/social/xPlaywrightFetcher"); + await import("yt-dlp-transcript-common/social/xNitterFetcher"); +} + export async function fetchPostsAction( slug: string, queueKey?: string, diff --git a/editor/app/channels/actions.ts b/editor/app/channels/actions.ts @@ -51,6 +51,7 @@ async function registerBuiltinSocialFetchers(): Promise<void> { await import("yt-dlp-transcript-common/social/blueskyFetcher"); await import("yt-dlp-transcript-common/social/xGalleryDlFetcher"); await import("yt-dlp-transcript-common/social/xPlaywrightFetcher"); + await import("yt-dlp-transcript-common/social/xNitterFetcher"); } export type ProbeChannelResult = diff --git a/editor/e2e/social-fetcher-picker.spec.ts b/editor/e2e/social-fetcher-picker.spec.ts @@ -0,0 +1,61 @@ +// The scraping method is an operational choice, not a create-time one: the X +// paths reach very different amounts of history, so it must be switchable from +// the channel page. +import { test, expect } from "@playwright/test"; +import { mkdir, writeFile } from "node:fs/promises"; +import { resetData, readJson, resolvePath } from "./helpers"; + +type Cfg = { postFetcher?: string; platform?: string }; + +async function makeSocialChannel(slug: string, platform: string, fetcher: string) { + const dir = resolvePath(`test-transcripts/channels/${slug}`); + await mkdir(dir, { recursive: true }); + await writeFile( + `${dir}/config.json`, + JSON.stringify({ + handling: "transcribe", + sourceKind: "social", + platform, + postFetcher: fetcher, + socialHandle: "someone", + name: `@someone on ${platform}`, + url: platform === "twitter" ? "https://x.com/someone" : "https://bsky.app/profile/someone", + }, null, 2), + ); +} + +test("an X channel offers every X scraping method and switching persists", async ({ page }) => { + await resetData("empty"); + await makeSocialChannel("x-chan", "twitter", "x-gallery-dl"); + await page.goto("/channels/x-chan"); + + const picker = page.getByLabel("post fetcher"); + await expect(picker).toBeVisible(); + await expect(picker).toHaveValue("x-gallery-dl"); + + // All three X paths are offered — and no Bluesky one. + const ids = await picker.locator("option").evaluateAll((os) => + os.map((o) => (o as HTMLOptionElement).value), + ); + expect(ids.sort()).toEqual(["x-gallery-dl", "x-nitter", "x-playwright"]); + + await picker.selectOption("x-nitter"); + await expect(page.getByText("Fetcher updated")).toBeVisible(); + + // The choice is persisted to the channel config, not just local state. + await expect + .poll(async () => (await readJson<Cfg>("test-transcripts/channels/x-chan/config.json")).postFetcher, + { timeout: 10_000 }) + .toBe("x-nitter"); + + await page.reload(); + await expect(page.getByLabel("post fetcher")).toHaveValue("x-nitter"); +}); + +test("a Bluesky channel shows no picker — it has only one method", async ({ page }) => { + await resetData("empty"); + await makeSocialChannel("bsky-chan", "bluesky", "bluesky-atproto"); + await page.goto("/channels/bsky-chan"); + await expect(page.locator("[data-social-panel]")).toBeVisible(); + await expect(page.getByLabel("post fetcher")).toHaveCount(0); +});