Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit bb3a2179c9115bcb99056644e6ca06637e5128ac
parent 595df2710ea55ee8591cf5fa0af2506538f5bd9b
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sat, 25 Jul 2026 22:35:33 -0400

Social posts (X/Twitter + Bluesky) as a parallel corpus to transcripts

Archive social posts from the same third-party commentators as a parallel
dataset to video transcripts: one corpus, one set of searches, both feeding
the AI integrations together.

Posts are modelled on the live-chat layer, not the video layer — their own
month-sharded storage, their own page tree, their own LayerScope — and they
compose into the SAME boolean query tree, so `(transcripts:"x" OR posts:"x")`
is a single unified search rather than a second search surface.

Model + ingest
- common/lib/posts.ts (client-safe) + posts-server.ts: month-sharded
  channels/<slug>/posts/YYYY-MM.jsonl plus a posts-archive of seen ids, so a
  re-run is a no-op. Mirrors how yt-dlp's download-archive drives sync().
- common/social/fetchers.ts: a pluggable SocialFetcher registry modelled on
  transcriptionApps.ts. Every X retrieval path rots, so the seam matters more
  than any single implementation.
  - bluesky-atproto: public AT Protocol, no auth and no binary.
  - x-gallery-dl: primary X path, a light headless subprocess reusing
    cookiePolicy.ts.
  - x-playwright: fallback, immune to the GraphQL query-id rotations that
    periodically break gallery-dl. Shares the Post normalizer with gallery-dl,
    so it is a transport swap rather than a rewrite.
  - xSessionBroker.ts: a persistent logged-in browser profile that re-exports
    fresh cookies, removing gallery-dl's cookies-expire-in-days problem.
- One fetch-posts job kind (drainable, bookmarkable); queueKeyForUrl already
  routes x.com / bsky.app to their own serial queues.

Index + export
- SCHEMA_VERSION 11 -> 12: a posts LMDB sub-DB keyed [createdAt, channelSlug,
  id]. ISO-8601 sorts chronologically, fixing the intra-day ordering the
  videos' YYYYMMDD key has. uploadDate is derived and stored so every existing
  date filter keeps working untouched.
- Shared /posts/<slug>/ page tree with the same byte-cap + incremental skip as
  transcripts; per-site posts manifest; compose-site reconcile; CORS; service
  workers; .gitignore. corpus.json -> spec 2 with a postScheme.

Search
- LayerScope += "posts", whitelisted in qt= deserialization so a shared link
  round-trips a posts leaf. Post and video slugs are disjoint namespaces,
  partitioned per-leaf by the eval engine: without that, an AND against the
  global scope collapses to nothing and video leaves waste a fetch per post.
- Posts are a THIRD media kind beside videos and livestreams — folding them
  into the video toggle would silently drop the whole corpus.
- Post result cards drop what doesn't apply (no seek, no livestream/age
  badges, no VOD expiry); PostModal is a sibling of TranscriptModal, not a
  generalization of it.

AI
- /ask: a posts leaf per keyword sharing the term key with t#/m# so ranking
  counts a keyword once; post excerpts render without a clock and cite as a
  bare [n]; fetch_context on a post returns its thread.
- MCP: postsManifest/postsPage on the single ShardSource boundary (so local,
  remote and hub gain it at once), a posts search scope, content_types
  defaulting to both, get_post / get_thread, and post-aware sweep citations.

Editor
- sourceKind is a separate axis from handling — deliberately NOT a third
  handling value, so existing binary `handling === "transcribe" ? …` branches
  can never misroute a social channel. Social channels get a two-stage
  Fetch -> Index rail instead of the six video stages, and the form hides
  every video-only control.
- The scheduler's eligibility rules were already source-agnostic, but its
  DISPATCH was not: runTick would have run a yt-dlp video sync against a
  Bluesky profile URL. Now routed by source kind.

Verification
- Bluesky proven live end-to-end: 829 real posts across 38 month shards;
  capped run -> cursor resume -> full backfill -> re-run writes 0 duplicates.
- 352 unit tests (common 192, MCP 99, export 61); 137/137 export e2e including
  6 new posts-search tests; 5 new editor social-channel e2e. All four packages
  typecheck.
- X is NEVER contacted by the test suite: gallery-dl is driven through a
  fake-gallery-dl.mjs fixture and the Playwright fallback through recorded
  GraphQL payloads, so runs stay deterministic and no test can risk an account.

Not done: the live X spike. gallery-dl is not installed and a real run needs a
throwaway X account with cookies, so neither X fetcher has touched x.com yet.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Diffstat:
M.gitignore | 1+
Mcommon/bin/compose-site.ts | 47+++++++++++++++++++++++++++++++++++++++++++++--
Acommon/bin/fetch-posts.ts | 50++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/components/PlayerProvider.tsx | 19+++++++++++++++++--
Acommon/components/PostModal.tsx | 210+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/components/QueryLeafView.tsx | 12+++++++++++-
Mcommon/components/SearchDataContext.tsx | 74++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++--
Mcommon/components/SearchResults.tsx | 68+++++++++++++++++++++++++++++++++++++++++++++++++++-----------------
Mcommon/components/SearchSessionContext.tsx | 147++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++---
Mcommon/components/WorkspaceSearchBar.tsx | 21+++++++++++++++++++++
Mcommon/components/exportFilterStorage.ts | 3+++
Acommon/components/postsCache.ts | 166+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/components/searchPipeline.ts | 140+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/components/urlState.ts | 15+++++++++++----
Mcommon/controller/buildIndex.ts | 274++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-
Acommon/controller/fetchPosts.ts | 192+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/jobs/jobKinds.ts | 34++++++++++++++++++++++++++++++++++
Mcommon/lib/channelConfig.ts | 39+++++++++++++++++++++++++++++++++++++++
Mcommon/lib/corpus.ts | 76++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++------------
Mcommon/lib/paths.ts | 12++++++++++++
Mcommon/lib/platform.ts | 35++++++++++++++++++++++++++++++++++-
Acommon/lib/posts-server.test.ts | 148+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/posts-server.ts | 254+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/posts.ts | 274+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/lib/searchEval.ts | 19+++++++++++++++++++
Mcommon/lib/searchQuery.ts | 11++++++++++-
Mcommon/lib/site.ts | 4++++
Acommon/social/__fixtures__/bluesky-author-feed.json | 814+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/social/blueskyFetcher.test.ts | 99+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/social/blueskyFetcher.ts | 486+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/social/fetchers.ts | 203+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/social/xGalleryDlFetcher.ts | 310+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/social/xNormalize.test.ts | 389+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/social/xNormalize.ts | 235+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/social/xPlaywrightFetcher.ts | 251+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/social/xSessionBroker.ts | 260+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/CHANGELOG.md | 3+++
Aeditor/app/channels/[slug]/components/SocialChannelPanel.tsx | 148+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/app/channels/[slug]/page.tsx | 66++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aeditor/app/channels/[slug]/socialActions.ts | 65+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/app/channels/actions.ts | 57+++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/app/channels/components/ChannelForm.tsx | 98++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-
Meditor/app/channels/components/parseChannelForm.ts | 34++++++++++++++++++++++++++++++++++
Meditor/app/jobs/jobReplayRegistry.ts | 31+++++++++++++++++++++++++++++++
Meditor/app/scheduler/runTick.ts | 11++++++++++-
Aeditor/app/settings/components/XSessionSection.tsx | 129+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/app/settings/page.tsx | 12+++++++++++-
Aeditor/app/settings/xSessionActions.ts | 63+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aeditor/e2e/fixtures/bin/fake-gallery-dl.mjs | 88+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aeditor/e2e/social-channel.spec.ts | 185+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aeditor/e2e/x-session.spec.ts | 16++++++++++++++++
Meditor/package.json | 4++--
Mexport/CHANGELOG.md | 5+++++
Mexport/app/(workspace)/SiteWorkspace.tsx | 3+++
Mexport/app/ask/citations.tsx | 12+++++++++++-
Mexport/app/ask/useAskChat.ts | 45+++++++++++++++++++++++++++++++++++++++++++--
Mexport/app/components/hub/HubHome.tsx | 3+++
Mexport/app/duplicates/page.tsx | 3+++
Mexport/app/lib/askConversation.ts | 11+++++++----
Mexport/app/lib/askRetrieval.ts | 72++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++----
Mexport/app/lib/searchAgent.ts | 33+++++++++++++++++++++++++++++++++
Mexport/e2e/fixtures/data.ts | 86+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mexport/e2e/helpers.ts | 12++++++++++++
Aexport/e2e/posts-search.spec.ts | 187+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mexport/service-worker/site-sw.js | 2+-
Mexport/service-worker/sw-hub.js | 2+-
Mmcp/src/search.test.ts | 175+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mmcp/src/search.ts | 236+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mmcp/src/server.ts | 199+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++----
Mmcp/src/source.ts | 79+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mmcp/src/sourceController.test.ts | 6++++++
71 files changed, 7497 insertions(+), 76 deletions(-)

diff --git a/.gitignore b/.gitignore @@ -54,6 +54,7 @@ yarn-error.log* # checkout's dirs (see WORKTREES.md), and a directory-only pattern wouldn't match # them, leaving generated data showing up as untracked in every worktree. /export/public/subs +/export/public/posts /export/public/summaries /export/public/transcripts /export/public/stats diff --git a/common/bin/compose-site.ts b/common/bin/compose-site.ts @@ -26,6 +26,7 @@ import { type DuplicateReport, } from "../lib/duplicates"; import type { Manifest, SubsManifest } from "../lib/manifest"; +import type { PostsManifest } from "../lib/posts"; import { buildSiteDescriptor, type PublicSiteDescriptor } from "../lib/siteDescriptor"; import { effectiveSiteAliases } from "../lib/aliasesStore"; import { @@ -64,6 +65,8 @@ const CORS_HEADERS = `# Generated by compose-site.ts — do not edit by hand. Access-Control-Allow-Origin: * /transcripts/* Access-Control-Allow-Origin: * +/posts/* + Access-Control-Allow-Origin: * /stats/* Access-Control-Allow-Origin: * /archives/* @@ -130,7 +133,21 @@ async function emitAiFiles(paths: ReturnType<typeof getPaths>): Promise<void> { } } - const corpus = buildSiteCorpus(descriptor, { hasArchives }); + // Per-channel post counts come from the site posts manifest just composed + // above; absent (no social channels) leaves corpus.json shaped as before. + const postCounts: Record<string, number> = {}; + try { + const raw = await readFile( + path.join(paths.exportPostsDir, "manifest.json"), + "utf8", + ); + const pm = JSON.parse(raw) as PostsManifest; + for (const ch of pm.channels ?? []) postCounts[ch.slug] = ch.postCount; + } catch { + /* no posts manifest for this site */ + } + + const corpus = buildSiteCorpus(descriptor, { hasArchives, postCounts }); await writeFile( path.join(paths.exportPublicDir, "corpus.json"), JSON.stringify(corpus), @@ -469,6 +486,9 @@ async function replaceDir(src: string, dest: string): Promise<void> { type ComposeCache = { transcripts: Record<string, string>; subs: Record<string, string>; + // Optional for backwards compat: a cache written before the posts corpus + // existed simply has no entry, so every social channel composes once. + posts?: Record<string, string>; summaries?: string; stats?: string; duplicates?: string; @@ -491,12 +511,13 @@ async function readComposeCache(p: string): Promise<ComposeCache> { return { transcripts: parsed?.transcripts ?? {}, subs: parsed?.subs ?? {}, + posts: parsed?.posts ?? {}, summaries: parsed?.summaries, stats: parsed?.stats, duplicates: parsed?.duplicates, }; } catch { - return { transcripts: {}, subs: {} }; + return { transcripts: {}, subs: {}, posts: {} }; } } @@ -659,6 +680,17 @@ async function main(): Promise<void> { cache.subs, console.log, ); + // The social-post corpus: same shared-tree shape, same incremental reconcile. + // Only social member channels have a source dir; reconcileChannelTree treats a + // missing one as "nothing to copy", so passing every member slug is correct. + cache.posts = await reconcileChannelTree( + "posts", + paths.exportSharedPostsDir, + paths.exportPostsDir, + memberSlugs, + cache.posts ?? {}, + console.log, + ); // Subs also carries a per-site manifest.json (tiny — copied every build). const subsManifestSrc = path.join( paths.exportSitesIndexDir, @@ -669,6 +701,17 @@ async function main(): Promise<void> { if (await exists(subsManifestSrc)) { await cp(subsManifestSrc, path.join(paths.exportSubsDir, "manifest.json")); } + // Same for the per-site posts manifest (which channels carry posts). + const postsManifestSrc = path.join( + paths.exportSitesIndexDir, + siteId, + "posts", + "manifest.json", + ); + if (await exists(postsManifestSrc)) { + await mkdir(paths.exportPostsDir, { recursive: true }); + await cp(postsManifestSrc, path.join(paths.exportPostsDir, "manifest.json")); + } // --- charts dashboard --- const templatesSrc = path.join( diff --git a/common/bin/fetch-posts.ts b/common/bin/fetch-posts.ts @@ -0,0 +1,50 @@ +#!/usr/bin/env tsx +// Fetch social posts for one channel into its on-disk posts corpus. +// +// pnpm --filter yt-dlp-transcript-common exec tsx bin/fetch-posts.ts \ +// --slug <channel-slug> [--full] [--limit N] +// +// --full re-walks the account's whole history instead of stopping at the +// stored watermark. New posts are still deduped against the posts-archive, so +// it repairs a gap rather than creating duplicates. + +import { getPaths } from "../lib/paths"; +import { getSettings } from "../lib/settings"; +import { fetchPosts } from "../controller/fetchPosts"; +import { parseFlags } from "./_parseFlags"; + +const flags = parseFlags(process.argv.slice(2)); +const slug = flags.slug; +if (!slug) { + console.error("Usage: fetch-posts.ts --slug <channel-slug> [--full] [--limit N]"); + process.exit(2); +} + +const limitRaw = flags.limit ? Number(flags.limit) : undefined; +const limit = + typeof limitRaw === "number" && Number.isFinite(limitRaw) && limitRaw > 0 + ? Math.floor(limitRaw) + : undefined; + +fetchPosts({ + paths: getPaths(), + slug, + settings: getSettings(), + full: flags.full === "true", + limit, + onLog: (line) => console.log(line), +}) + .then((result) => { + if (!result.ok) { + console.error(`Failed: ${result.error}`); + process.exit(1); + } + console.log( + `Done: ${result.written} new, ${result.skipped} skipped, ` + + `${result.complete ? "caught up" : "more history remains"}.`, + ); + }) + .catch((err) => { + console.error(err); + process.exit(1); + }); diff --git a/common/components/PlayerProvider.tsx b/common/components/PlayerProvider.tsx @@ -309,7 +309,10 @@ export function PlayerProvider({ const modalMode: ModalMode = urlVm; const displayMode: DisplayMode = displayState?.slug === activeSlug ? displayState.mode : "modal"; - const modalOpen = activeSlug !== null && displayMode === "modal"; + // A post slug opens PostModal instead (a sibling component); keep the + // transcript modal closed so the two never stack. + const modalOpen = + activeSlug !== null && displayMode === "modal" && modalMode !== "post"; const detailMatches = detail?.slug === activeSlug; const data: TranscriptData | null = useMemo(() => { if (!detailMatches || !detail) return null; @@ -351,7 +354,12 @@ export function PlayerProvider({ writeUrlParams({ v: slug, t, - vm: opts?.mode === "chat" ? "chat" : "transcript", + vm: + opts?.mode === "chat" + ? "chat" + : opts?.mode === "post" + ? "post" + : "transcript", }); }, [activeSlug], @@ -529,6 +537,12 @@ export function PlayerProvider({ chatInFlightForRef.current = null; setKickError(false); if (!urlSlug) return; + // Posts are not transcripts — never run the video fetch for one. + // A post slug always arrives together with a slug change, so this closure's + // urlVm is fresh without adding it to the deps — and it MUST stay out of + // them, or a plain transcript<->chat mode toggle would re-run this effect + // and clobber the chat-missing snap-back notice. + if (urlVm === "post") return; readyForSlugRef.current = null; pendingSeekRef.current = urlTime ?? 0; let cancelled = false; @@ -563,6 +577,7 @@ export function PlayerProvider({ }; // urlTime intentionally omitted — t changes shouldn't refetch // eslint-disable-next-line react-hooks/exhaustive-deps + // eslint-disable-next-line react-hooks/exhaustive-deps }, [urlSlug]); // Lazy-fetch live chat only when the modal enters chat mode for the diff --git a/common/components/PostModal.tsx b/common/components/PostModal.tsx @@ -0,0 +1,210 @@ +"use client"; + +// The viewer for one archived social post. +// +// A SIBLING of TranscriptModal, deliberately not a generalization of it: +// TranscriptModal is a transcript-follows-playhead component (16:9 reserved +// layout, playhead auto-scroll, per-row seek buttons, clip markers) and none of +// that applies to a post. This renders the post, its thread context and its +// outbound links, and nothing else. + +import { useEffect, useMemo } from "react"; +import { useQuery } from "@tanstack/react-query"; +import { XIcon } from "lucide-react"; +import { Button } from "./ui/button"; +import { useUrlParams, writeUrlParams } from "./urlState"; +import { fetchPost, fetchThread } from "./postsCache"; +import type { Post } from "../lib/posts"; + +function formatWhen(iso: string): string { + const ms = Date.parse(iso); + if (!Number.isFinite(ms)) return iso; + const d = new Date(ms); + return d.toLocaleString(undefined, { + year: "numeric", + month: "short", + day: "numeric", + hour: "2-digit", + minute: "2-digit", + }); +} + +function platformLabel(platform: Post["platform"]): string { + return platform === "twitter" ? "X" : "Bluesky"; +} + +export default function PostModal() { + const { v: slug, vm } = useUrlParams(); + const open = vm === "post" && !!slug; + + const post = useQuery({ + queryKey: ["post", slug], + queryFn: () => fetchPost(slug!), + enabled: open, + }); + const thread = useQuery({ + queryKey: ["post-thread", slug], + queryFn: () => fetchThread(slug!), + enabled: open, + }); + + const close = () => writeUrlParams({ v: null, t: null }); + + // Escape closes, matching the transcript modal's behavior. + useEffect(() => { + if (!open) return; + const onKey = (e: KeyboardEvent) => { + if (e.key === "Escape") close(); + }; + window.addEventListener("keydown", onKey); + return () => window.removeEventListener("keydown", onKey); + // eslint-disable-next-line react-hooks/exhaustive-deps + }, [open]); + + if (!open) return null; + + return ( + <div + role="dialog" + aria-modal="true" + aria-label="Post" + data-post-modal="" + className="fixed inset-0 z-50 flex items-start justify-center overflow-y-auto bg-black/60 p-4 sm:p-8" + onClick={(e) => { + if (e.target === e.currentTarget) close(); + }} + > + <div className="w-full max-w-2xl rounded-lg border border-border bg-card shadow-lg"> + <div className="flex items-center gap-2 border-b border-border px-4 py-2"> + <span className="text-xs uppercase tracking-wide text-muted-foreground"> + {post.data ? platformLabel(post.data.platform) : "Post"} + </span> + <Button + type="button" + variant="ghost" + aria-label="Close" + onClick={close} + className="ml-auto h-auto px-2 py-1" + > + <XIcon className="size-4" /> + </Button> + </div> + + {post.isPending && ( + <p className="px-4 py-6 text-sm text-muted-foreground">Loading…</p> + )} + {post.isError && ( + <p className="px-4 py-6 text-sm text-destructive"> + Could not load this post. + </p> + )} + + {post.data && ( + <div className="flex flex-col gap-4 px-4 py-4"> + <PostBody post={post.data} primary /> + + {/* Thread context: the archived parent + replies around this post. + The post-corpus analogue of a transcript's surrounding cues. */} + {thread.data && thread.data.length > 1 && ( + <section className="flex flex-col gap-2"> + <h3 className="text-xs uppercase tracking-wide text-muted-foreground"> + Thread ({thread.data.length} posts) + </h3> + <ol className="flex flex-col gap-2 border-l border-border pl-3"> + {thread.data.map((p) => ( + <li key={p.id}> + <PostBody post={p} primary={p.slug === slug} compact /> + </li> + ))} + </ol> + </section> + )} + </div> + )} + </div> + </div> + ); +} + +function PostBody({ + post, + primary = false, + compact = false, +}: { + post: Post; + primary?: boolean; + compact?: boolean; +}) { + const links = useMemo(() => post.links ?? [], [post.links]); + return ( + <article + data-post-id={post.id} + className={ + "rounded-md " + + (primary && compact ? "bg-warning-soft " : "") + + (compact ? "px-2 py-1.5" : "") + } + > + <header className="flex flex-wrap items-baseline gap-x-2 text-xs text-muted-foreground"> + <span className="font-medium text-foreground"> + {post.authorName || post.author} + </span> + {post.authorName && <span>@{post.author}</span>} + <span>·</span> + <time dateTime={post.createdAt}>{formatWhen(post.createdAt)}</time> + {post.isRepost && <Badge>repost</Badge>} + {post.isReply && <Badge>reply</Badge>} + </header> + + <p className={"whitespace-pre-wrap " + (compact ? "text-sm" : "text-base mt-2")}> + {post.text} + </p> + + {links.length > 0 && ( + <ul className="mt-2 flex flex-col gap-1"> + {links.map((href) => ( + <li key={href}> + <a + href={href} + target="_blank" + rel="noopener noreferrer" + className="text-xs text-primary underline break-all" + > + {href} + </a> + </li> + ))} + </ul> + )} + + <footer className="mt-2 flex flex-wrap items-center gap-x-3 text-xs text-muted-foreground"> + <a + href={post.url} + target="_blank" + rel="noopener noreferrer" + className="underline" + > + Open original + </a> + {post.engagement?.likes != null && ( + <span>{post.engagement.likes} likes</span> + )} + {post.engagement?.reposts != null && ( + <span>{post.engagement.reposts} reposts</span> + )} + {post.engagement?.replies != null && ( + <span>{post.engagement.replies} replies</span> + )} + {post.mediaCount ? <span>{post.mediaCount} media (not archived)</span> : null} + </footer> + </article> + ); +} + +function Badge({ children }: { children: React.ReactNode }) { + return ( + <span className="rounded bg-muted px-1.5 py-0.5 text-[10px] uppercase tracking-wide"> + {children} + </span> + ); +} diff --git a/common/components/QueryLeafView.tsx b/common/components/QueryLeafView.tsx @@ -37,6 +37,7 @@ type Props = { const SCOPE_LABELS: Record<LayerScope, string> = { transcripts: "Transcripts", chat: "Live chat", + posts: "Posts", metadata: "Title / channel", description: "Description", tags: "Tags", @@ -184,7 +185,16 @@ export default function QueryLeafView({ data-testid={`leaf-scope-${leaf.id}`} className="rounded border border-border bg-card text-foreground px-1.5 py-1 text-xs font-mono" > - {(["transcripts", "chat", "metadata", "description", "tags"] as const).map((s) => ( + {( + [ + "transcripts", + "chat", + "posts", + "metadata", + "description", + "tags", + ] as const + ).map((s) => ( <option key={s} value={s}> {SCOPE_LABELS[s]} </option> diff --git a/common/components/SearchDataContext.tsx b/common/components/SearchDataContext.tsx @@ -17,11 +17,13 @@ import { import { useQueries } from "@tanstack/react-query"; import { useSummaries, type SummariesState } from "./summariesCache"; import { useSubsManifest } from "./subsCache"; +import { usePostsManifest } from "./postsCache"; import { useSearchAliases, fetchAliases } from "./aliasesCache"; import { mergeAliases, type SearchAlias } from "../lib/searchAliases"; import { idBaseUrl, makeId } from "./originId"; import type { DisplaySummary } from "../lib/transcripts"; import type { Manifest, SubsManifest } from "../lib/manifest"; +import type { PostsManifest } from "../lib/posts"; import { pageFileName } from "../lib/manifest"; import { DEFAULT_GROUP_FALLBACK_ID, @@ -49,6 +51,10 @@ export type SearchDataValue = { summariesState: SummariesState; // Cross-channel subs manifest (merged across origins in hub mode), or null. subsManifest: SubsManifest | null; + // Cross-channel social-posts manifest (which channels carry posts, and how + // many). Merged across origins in hub mode. Null / empty channel list when + // this site ships no posts corpus. + postsManifest: PostsManifest | null; // The channel selection model, resolved mode-appropriately. channels: ChannelOption[]; groups: ChannelGroup[]; @@ -84,6 +90,7 @@ export function useSearchData(): SearchDataValue { export function SingleSiteDataProvider({ children }: { children: ReactNode }) { const summariesState = useSummaries(""); const subsManifest = useSubsManifest("").data ?? null; + const postsManifest = usePostsManifest("").data ?? null; const aliases = useSearchAliases(""); const manifest = summariesState.manifest; @@ -125,6 +132,7 @@ export function SingleSiteDataProvider({ children }: { children: ReactNode }) { () => ({ summariesState, subsManifest, + postsManifest, channels, groups, defaultGroupId, @@ -132,7 +140,15 @@ export function SingleSiteDataProvider({ children }: { children: ReactNode }) { accentOf: NO_ACCENT, aliases, }), - [summariesState, subsManifest, channels, groups, defaultGroupId, aliases], + [ + summariesState, + subsManifest, + postsManifest, + channels, + groups, + defaultGroupId, + aliases, + ], ); return ( @@ -316,6 +332,50 @@ export function MultiSiteDataProvider({ // eslint-disable-next-line react-hooks/exhaustive-deps }, [sites, subsSettled]); + // 5b. Merge posts manifests, same origin-qualification as subs. A member site + // with no posts corpus 404s; treat that as an empty contribution so one + // video-only origin can't blank the hub's posts scope. + const postsQueries = useQueries({ + queries: sites.map((s) => ({ + queryKey: ["posts-manifest", s.origin], + queryFn: () => + fetchJson<PostsManifest>( + `${idBaseUrl(s.origin)}/posts/manifest.json`, + ).catch( + (): PostsManifest => ({ + version: 0, + channels: [], + totalCount: 0, + generatedAt: "", + }), + ), + })), + }); + const postsSettled = postsQueries.every((q) => q.isSuccess || q.isError); + const postsManifest = useMemo<PostsManifest | null>(() => { + const loaded = sites + .map((s, i) => ({ origin: s.origin, data: postsQueries[i]?.data })) + .filter((e): e is { origin: string; data: PostsManifest } => !!e.data); + if (loaded.length === 0) return null; + const channelsOut: PostsManifest["channels"] = []; + let totalCount = 0; + let generatedAt = ""; + for (const { origin, data } of loaded) { + for (const c of data.channels) { + channelsOut.push({ ...c, slug: makeId(origin, c.slug) }); + } + totalCount += data.totalCount ?? 0; + if (data.generatedAt > generatedAt) generatedAt = data.generatedAt; + } + return { + version: loaded[0].data.version, + channels: channelsOut, + totalCount, + generatedAt, + }; + // eslint-disable-next-line react-hooks/exhaustive-deps + }, [sites, postsSettled]); + // Alias dictionaries per federated origin, merged into one list (later origins // shadow earlier ones by id). Missing files resolve to [] — additive only. const aliasQueries = useQueries({ @@ -385,6 +445,7 @@ export function MultiSiteDataProvider({ () => ({ summariesState, subsManifest, + postsManifest, channels, groups, defaultGroupId, @@ -394,7 +455,16 @@ export function MultiSiteDataProvider({ accentOf, aliases, }), - [summariesState, subsManifest, channels, groups, defaultGroupId, accentOf, aliases], + [ + summariesState, + subsManifest, + postsManifest, + channels, + groups, + defaultGroupId, + accentOf, + aliases, + ], ); return ( diff --git a/common/components/SearchResults.tsx b/common/components/SearchResults.tsx @@ -425,7 +425,7 @@ const ResultCard = memo(function ResultCard({ > <Checkbox checked={selected} - aria-label={`Select "${group.title}" for AI`} + aria-label={`Select "${group.post ? group.post.text.slice(0, 60) : group.title}" for AI`} onCheckedChange={() => onToggleSelect(group.slug)} /> </label> @@ -436,17 +436,29 @@ const ResultCard = memo(function ResultCard({ onClick={() => openWithMode(group.slug)} className="min-w-0 flex-1 h-auto justify-start text-left items-baseline gap-2 px-2 py-2 rounded-none bg-transparent font-normal" > - <span className="font-medium truncate flex-1 min-w-0"> - {group.title} - </span> - {group.isLivestream && <LivestreamBadge />} - {group.ageRestricted && <AgeRestrictedBadge />} - {(() => { - const e = vodExpiry(group.platform, group.uploadDate); - return e?.likelyExpired ? ( - <VodExpiredBadge tooltip={e.tooltip} /> - ) : null; - })()} + {/* A post has no title, no livestream/age state and no VOD expiry — + its body IS the headline, so the card leads with the text and a + platform badge instead of the video decorations. */} + {group.post ? ( + <> + <span className="truncate flex-1 min-w-0">{group.post.text}</span> + <PostBadge platform={group.post.platform} /> + </> + ) : ( + <> + <span className="font-medium truncate flex-1 min-w-0"> + {group.title} + </span> + {group.isLivestream && <LivestreamBadge />} + {group.ageRestricted && <AgeRestrictedBadge />} + {(() => { + const e = vodExpiry(group.platform, group.uploadDate); + return e?.likelyExpired ? ( + <VodExpiredBadge tooltip={e.tooltip} /> + ) : null; + })()} + </> + )} <span className="text-xs text-muted-foreground shrink-0"> {group.channel && `${group.channel} · `} {group.date} @@ -457,7 +469,9 @@ const ResultCard = memo(function ResultCard({ <button type="button" onClick={() => onAsk(group.slug)} - title="Ask the AI about this video" + title={ + group.post ? "Ask the AI about this post" : "Ask the AI about this video" + } className="flex shrink-0 items-center gap-1 border-l border-border px-3 text-xs text-muted-foreground transition-colors hover:bg-accent hover:text-accent-foreground" > <MessageSquareIcon className="size-3.5" /> Ask @@ -480,7 +494,9 @@ const ResultCard = memo(function ResultCard({ ? "Title / channel" : leafInfo.scope === "chat" ? "Live chat" - : "Transcripts"} + : leafInfo.scope === "posts" + ? "Posts" + : "Transcripts"} </span> <span className="font-mono text-xs text-muted-foreground truncate"> {leafInfo.query} @@ -495,6 +511,8 @@ const ResultCard = memo(function ResultCard({ key={i} slug={group.slug} hit={h} + // Posts have no timeline, so no seek affordance. + noSeek={!!group.post} query={leafInfo.query} useRegex={leafInfo.useRegex} isActive={ @@ -524,6 +542,7 @@ const HitRow = memo(function HitRow({ useRegex, isActive, onOpen, + noSeek = false, }: { slug: string; hit: LayerHit; @@ -531,6 +550,9 @@ const HitRow = memo(function HitRow({ useRegex: boolean; isActive: boolean; onOpen: (slug: string, hit?: LayerHit) => void; + // Post hits have no timeline position, so the timestamp gutter would only + // ever read "0:00". Suppress it rather than render a meaningless seek. + noSeek?: boolean; }) { return ( <li> @@ -542,9 +564,11 @@ const HitRow = memo(function HitRow({ isActive ? "bg-warning-soft ring-1 ring-inset ring-warning/60" : "" }`} > - <span className="text-xs font-mono text-muted-foreground shrink-0 w-16"> - {hit.scope === "metadata" ? "—" : formatSeconds(hit.start)} - </span> + {!noSeek && ( + <span className="text-xs font-mono text-muted-foreground shrink-0 w-16"> + {hit.scope === "metadata" ? "—" : formatSeconds(hit.start)} + </span> + )} {hit.track && hit.track !== "live_chat" && ( <TrackBadge track={hit.track} /> )} @@ -557,6 +581,16 @@ const HitRow = memo(function HitRow({ }); HitRow.displayName = "HitRow"; +// Which social platform a post came from. Reuses the TrackBadge shape so post +// rows sit visually alongside chat rows rather than inventing a new idiom. +function PostBadge({ platform }: { platform: string }) { + return ( + <span className="shrink-0 text-[10px] uppercase tracking-wide font-medium px-1.5 py-0.5 rounded bg-muted text-muted-foreground self-center"> + {platform === "twitter" ? "X" : "Bluesky"} + </span> + ); +} + function TrackBadge({ track }: { track: string }) { const label = track === "live_chat" ? "live chat" : track; return ( diff --git a/common/components/SearchSessionContext.tsx b/common/components/SearchSessionContext.tsx @@ -24,6 +24,7 @@ import { } from "react"; import { usePlayer } from "./PlayerProvider"; import { useChannelSubsManifests } from "./subsCache"; +import { peekPost, useChannelPostsManifests } from "./postsCache"; import { useSearchData } from "./SearchDataContext"; import type { LayerHit } from "./searchPipeline"; import { @@ -69,6 +70,7 @@ import { type ChartShape, } from "../lib/chartShare"; import type { DisplaySummary, Platform } from "../lib/transcripts"; +import type { Post } from "../lib/posts"; import { makeId, splitId } from "./originId"; import { sortGroups, type ChannelGroup } from "../lib/channelGroups"; import { buildSearchHandoff, type SearchHandoff } from "../lib/aiHandoff"; @@ -138,6 +140,10 @@ export type ResultGroup = { // Provenance accent of the source origin (hub mode only); undefined // single-site, so no marker renders. accent?: string; + // Set on rows from the social-post corpus. A post is not a video: it has no + // timeline to seek, no livestream/age state and no VOD expiry, so the result + // card renders a distinct variant rather than a degraded video card. + post?: Post; }; // A leaf's original terms so result rows can look up the query for <mark> @@ -330,6 +336,10 @@ function useSearchSessionState() { >(() => new Set()); const [committedNov, setCommittedNov] = useState(false); const [committedNol, setCommittedNol] = useState(false); + // Third media kind: social posts. Videos and livestreams are a binary split + // (isLivestream ? nol : nov); a post is neither, so it needs its own toggle + // or the existing two would silently drop the whole posts corpus. + const [committedNop, setCommittedNop] = useState(false); const [committedNaa, setCommittedNaa] = useState(false); const [committedNar, setCommittedNar] = useState(false); const [committedNav, setCommittedNav] = useState(false); @@ -360,6 +370,7 @@ function useSearchSessionState() { Set<string> >(() => new Set()); const [draftNov, setDraftNov] = useState(false); + const [draftNop, setDraftNop] = useState(false); const [draftNol, setDraftNol] = useState(false); const [draftNaa, setDraftNaa] = useState(false); const [draftNar, setDraftNar] = useState(false); @@ -422,6 +433,7 @@ function useSearchSessionState() { const { summariesState, subsManifest, + postsManifest, channels: channelModel, groups: manifestGroups, channelKeyOf, @@ -484,6 +496,67 @@ function useSearchSessionState() { subsRefs, channelSubsQueries.map((q) => (q.data ? 1 : 0)).join(""), ]); + // ─── Posts scope ────────────────────────────────────────────────────────── + // Mirrors the chat-scope block above: only load the per-channel posts + // manifests when a posts leaf is actually present, then flatten them into the + // set of post slugs the tree may match. + const draftHasPostsLeaf = useMemo( + () => anyLeafHasScope(draftRoot, "posts"), + [draftRoot], + ); + const committedHasPostsLeaf = useMemo( + () => anyLeafHasScope(committedRoot, "posts"), + [committedRoot], + ); + const needsPostsManifests = draftHasPostsLeaf || committedHasPostsLeaf; + const postsRefs = useMemo( + () => + postsManifest + ? postsManifest.channels.map((c) => { + const { origin, slug } = splitId(c.slug); + return { origin, channelSlug: slug, name: c.name }; + }) + : [], + [postsManifest], + ); + const channelPostsQueries = useChannelPostsManifests( + needsPostsManifests ? postsRefs : [], + ); + const postScopeSlugs = useMemo<Set<string>>(() => { + const set = new Set<string>(); + if (!needsPostsManifests) return set; + for (let i = 0; i < channelPostsQueries.length; i++) { + const data = channelPostsQueries[i]?.data; + const ref = postsRefs[i]; + if (!data || !ref) continue; + // Channel selection applies to posts exactly as it does to videos; the + // date filter can't be applied here (a manifest carries no dates) and is + // enforced on the result rows instead, where the post body is loaded. + if ( + committedChannels.has( + channelKeyOf({ channel: ref.name, channelSlug: makeId(ref.origin, ref.channelSlug) }), + ) + ) { + continue; + } + for (const id of Object.keys(data.slugToPage)) { + set.add(makeId(ref.origin, `${ref.channelSlug}/${id}`)); + } + } + return set; + // eslint-disable-next-line react-hooks/exhaustive-deps + }, [ + needsPostsManifests, + postsRefs, + committedChannelsKey, + channelKeyOf, + channelPostsQueries.map((q) => (q.data ? 1 : 0)).join(""), + ]); + const postsManifestReady = + !needsPostsManifests || + postsRefs.length === 0 || + channelPostsQueries.every((q) => q.data || q.isError); + const subsManifestReady = !needsChatManifests || (subsManifest !== null && @@ -638,15 +711,20 @@ function useSearchSessionState() { committedNu, ]); - const filterKey = `${committedChannelsKey}|${committedNov ? 1 : 0}|${committedNol ? 1 : 0}|${committedNaa ? 1 : 0}|${committedNar ? 1 : 0}|${committedNav ? 1 : 0}|${committedNd ? 1 : 0}|${committedNu ? 1 : 0}|${committedDateFrom}|${committedDateTo}`; + const filterKey = `${committedChannelsKey}|${committedNov ? 1 : 0}|${committedNop ? 1 : 0}|${committedNol ? 1 : 0}|${committedNaa ? 1 : 0}|${committedNar ? 1 : 0}|${committedNav ? 1 : 0}|${committedNd ? 1 : 0}|${committedNu ? 1 : 0}|${committedDateFrom}|${committedDateTo}`; + // The scope universe is videos AND posts. The two namespaces are disjoint; + // searchEval partitions them per-leaf (see EvalCtx.postScopeSlugs) so a posts + // leaf never fetches a video and vice versa, while OR across the two still + // unions into one result set. const globalScopeSlugs = useMemo<string[]>(() => { if (!transcripts) return []; const out: string[] = []; for (const t of transcripts) if (passesFilter(t)) out.push(t.slug); + if (!committedNop) for (const slug of postScopeSlugs) out.push(slug); return out; // eslint-disable-next-line react-hooks/exhaustive-deps - }, [transcripts, filterKey]); + }, [transcripts, filterKey, postScopeSlugs, committedNop]); const hasActiveQuery = useMemo( () => isNodeActive(committedRoot), @@ -668,12 +746,14 @@ function useSearchSessionState() { } if (!transcripts) return; if (needsChatManifests && !subsManifestReady) return; + if (needsPostsManifests && !postsManifestReady) return; const controller = runQueryTree({ root: committedRoot, globalScope: globalScopeSlugs, summaries: transcripts, chatScopeSlugs: needsChatManifests ? chatScopeSlugs : null, + postScopeSlugs: needsPostsManifests ? postScopeSlugs : null, initialHitLimit: hitLimit, concurrency: fetchConcurrency, flushIntervalMs, @@ -695,6 +775,8 @@ function useSearchSessionState() { filterKey, needsChatManifests, subsManifestReady, + needsPostsManifests, + postsManifestReady, fetchConcurrency, flushIntervalMs, ]); @@ -736,8 +818,10 @@ function useSearchSessionState() { const matched = treeProgress.slugs; const hitsBySlug = treeProgress.hits; const out: ResultGroup[] = []; + const videoSlugs = new Set<string>(); for (const t of transcripts) { if (!matched.has(t.slug)) continue; + videoSlugs.add(t.slug); out.push({ slug: t.slug, title: t.title, @@ -751,8 +835,46 @@ function useSearchSessionState() { accent: accentOf(splitId(t.slug).origin), }); } + // Matched POST slugs. They are not in `transcripts` (a disjoint namespace), + // so they're appended here from the posts page cache, which the pipeline + // has necessarily already warmed for anything it matched. The date filter + // is applied here rather than at scope-build time because a posts manifest + // carries no dates. + for (const slug of matched) { + if (videoSlugs.has(slug)) continue; + const post = peekPost(slug); + if (!post) continue; + if (committedDateFrom && post.uploadDate < committedDateFrom) continue; + if (committedDateTo && post.uploadDate > committedDateTo) continue; + out.push({ + slug, + // A post has no title; the card renders its body, and the author + // stands in for the channel. + title: "", + channel: post.authorName || post.author, + date: post.createdAt, + isLivestream: false, + ageRestricted: false, + platform: post.platform, + uploadDate: post.uploadDate, + hits: hitsBySlug.get(slug) ?? [], + accent: accentOf(splitId(slug).origin), + post, + }); + } + // One newest-first ordering across BOTH corpora, so a unified search reads + // as one result set rather than videos-then-posts. + out.sort((a, b) => (a.uploadDate === b.uploadDate ? 0 : a.uploadDate < b.uploadDate ? 1 : -1)); return out; - }, [hasActiveQuery, transcripts, treeProgress, passesFilter, accentOf]); + }, [ + hasActiveQuery, + transcripts, + treeProgress, + passesFilter, + accentOf, + committedDateFrom, + committedDateTo, + ]); const leafStates = treeProgress?.leafStates ?? new Map<string, LeafState>(); const groupStates = @@ -792,6 +914,7 @@ function useSearchSessionState() { const filtersDirty = !sameSet(draftExcludedChannels, committedChannels) || draftNov !== committedNov || + draftNop !== committedNop || draftNol !== committedNol || draftNaa !== committedNaa || draftNar !== committedNar || @@ -844,6 +967,7 @@ function useSearchSessionState() { ), }; if (draftNov) snap.nov = true; + if (draftNop) snap.nop = true; if (draftNol) snap.nol = true; if (draftNaa) snap.naa = true; if (draftNar) snap.nar = true; @@ -859,6 +983,7 @@ function useSearchSessionState() { channelOptions, defaultSelectedChannels, draftNov, + draftNop, draftNol, draftNaa, draftNar, @@ -873,6 +998,7 @@ function useSearchSessionState() { const promoteDraftsToCommitted = useCallback(() => { setCommittedExcludedChannels(new Set(draftExcludedChannels)); setCommittedNov(draftNov); + setCommittedNop(draftNop); setCommittedNol(draftNol); setCommittedNaa(draftNaa); setCommittedNar(draftNar); @@ -885,6 +1011,7 @@ function useSearchSessionState() { }, [ draftExcludedChannels, draftNov, + draftNop, draftNol, draftNaa, draftNar, @@ -1015,6 +1142,7 @@ function useSearchSessionState() { setDraftExcludedChannels(initialExcluded); setDraftNov(initialNov); + setDraftNop(false); setDraftNol(initialNol); setDraftNaa(initialNaa); setDraftNar(initialNar); @@ -1156,6 +1284,7 @@ function useSearchSessionState() { ); setDraftExcludedChannels(excluded); setDraftNov(snapshot?.nov === true); + setDraftNop(snapshot?.nop === true); setDraftNol(snapshot?.nol === true); setDraftNaa(snapshot?.naa === true); setDraftNar(snapshot?.nar === true); @@ -1446,7 +1575,15 @@ function useSearchSessionState() { const openWithMode = useCallback( (slug: string, hit?: LayerHit) => { - const modalMode = hit?.scope === "chat" ? "chat" : "transcript"; + // A post opens the PostModal, not the transcript reader — there is no + // video to load and no playhead to seek. peekPost covers the case where + // the row was opened from its header (no hit to read the scope from). + const isPost = hit?.scope === "posts" || !!peekPost(slug); + const modalMode = isPost + ? "post" + : hit?.scope === "chat" + ? "chat" + : "transcript"; openTranscript(slug, hit?.start, { mode: modalMode }); }, [openTranscript], @@ -1708,6 +1845,8 @@ function useSearchSessionState() { setSelected, draftNov, setDraftNov, + draftNop, + setDraftNop, draftNol, setDraftNol, draftNaa, diff --git a/common/components/WorkspaceSearchBar.tsx b/common/components/WorkspaceSearchBar.tsx @@ -20,6 +20,7 @@ import { DEFAULT_FETCH_CONCURRENCY, DEFAULT_FLUSH_INTERVAL_MS, } from "./SearchSessionContext"; +import { useSearchData } from "./SearchDataContext"; export default function WorkspaceSearchBar() { const { @@ -65,6 +66,8 @@ export default function WorkspaceSearchBar() { setSelected, channelLabelByKey, draftNov, + draftNop, + setDraftNop, setDraftNov, draftNol, setDraftNol, @@ -94,6 +97,11 @@ export default function WorkspaceSearchBar() { setHitLimit, } = useSearchSession(); + // Only offer the Posts type toggle on a site that actually ships a posts + // corpus, so a pure-video deployment's filter row is unchanged. + const { postsManifest } = useSearchData(); + const hasPostsCorpus = (postsManifest?.channels.length ?? 0) > 0; + return ( <div className="flex flex-col gap-6"> <ProfilesRow @@ -432,6 +440,19 @@ export default function WorkspaceSearchBar() { /> Livestreams </label> + {/* The social-post corpus is a third media kind, not a video + sub-type — without its own toggle the video/livestream pair + would silently drop every post. Only offered when the site + actually ships posts. */} + {hasPostsCorpus && ( + <label className="flex items-center gap-1.5 select-none"> + <Checkbox + checked={!draftNop} + onCheckedChange={(value) => setDraftNop(value !== true)} + /> + Posts + </label> + )} </div> <div className="flex flex-wrap items-center gap-x-3 gap-y-1"> <span className="text-xs uppercase tracking-wide text-muted-foreground"> diff --git a/common/components/exportFilterStorage.ts b/common/components/exportFilterStorage.ts @@ -21,6 +21,9 @@ const VERSION = 1; export type FilterSnapshot = { channels: { included: string[]; excluded: string[] }; nov?: boolean; + // Exclude the social-post corpus — the third media kind beside videos and + // livestreams (see SearchSessionContext.passesFilter). + nop?: boolean; nol?: boolean; naa?: boolean; nar?: boolean; diff --git a/common/components/postsCache.ts b/common/components/postsCache.ts @@ -0,0 +1,166 @@ +"use client"; + +// Client-side cache for the social-post corpus, mirroring subsCache.ts: a +// per-channel manifest hook plus lazily-fetched pages, keyed by OriginId so a +// federating hub can hold several origins' posts at once without collisions. + +import { useQueries, useQuery } from "@tanstack/react-query"; +import { + postsPageFileName, + type ChannelPostsManifest, + type Post, + type PostsManifest, +} from "../lib/posts"; +import { makeId, splitId, idBaseUrl } from "./originId"; + +async function fetchJson<T>(url: string): Promise<T> { + const r = await fetch(url); + if (!r.ok) throw new Error(`Failed to fetch ${url}: ${r.status}`); + return (await r.json()) as T; +} + +const resolved = new Map<string, Post>(); +const inFlight = new Map<string, Promise<Post>>(); +const channelManifests = new Map<string, Promise<ChannelPostsManifest>>(); +const pagePromises = new Map<string, Promise<Post[]>>(); + +export type PostsManifestRef = { channelSlug: string; origin?: string }; + +// `id` is an OriginId: bare `<channelSlug>/<postId>` same-origin, +// origin-prefixed cross-origin. +export function fetchPost(id: string): Promise<Post> { + const hit = resolved.get(id); + if (hit) return Promise.resolve(hit); + const flying = inFlight.get(id); + if (flying) return flying; + const p = load(id).then((post) => { + resolved.set(id, post); + inFlight.delete(id); + return post; + }); + p.catch(() => inFlight.delete(id)); + inFlight.set(id, p); + return p; +} + +// Synchronously readable cache hit — used by result cards that already pulled +// the page during search and shouldn't re-suspend to render. +export function peekPost(id: string): Post | undefined { + return resolved.get(id); +} + +export function fetchChannelPostsManifest( + channelSlug: string, + origin = "", +): Promise<ChannelPostsManifest> { + const key = makeId(origin, channelSlug); + let p = channelManifests.get(key); + if (!p) { + p = fetchJson<ChannelPostsManifest>( + `${idBaseUrl(origin)}/posts/${channelSlug}/manifest.json`, + ); + p.catch(() => channelManifests.delete(key)); + channelManifests.set(key, p); + } + return p; +} + +export function fetchPostsPage( + channelSlug: string, + pageIndex: number, + origin = "", +): Promise<Post[]> { + const key = `${makeId(origin, channelSlug)}:${pageIndex}`; + let p = pagePromises.get(key); + if (!p) { + p = fetchJson<Post[]>( + `${idBaseUrl(origin)}/posts/${channelSlug}/${postsPageFileName(pageIndex)}`, + ).then((page) => { + // Warm the by-id cache so a later fetchPost() for any post on this page + // is synchronous. + for (const entry of page) { + resolved.set(makeId(origin, entry.slug), entry); + } + return page; + }); + p.catch(() => pagePromises.delete(key)); + pagePromises.set(key, p); + } + return p; +} + +async function load(id: string): Promise<Post> { + const { origin, slug } = splitId(id); + const slashIdx = slug.indexOf("/"); + if (slashIdx < 0) throw new Error(`Malformed post slug: ${slug}`); + const channelSlug = slug.slice(0, slashIdx); + const postId = slug.slice(slashIdx + 1); + const manifest = await fetchChannelPostsManifest(channelSlug, origin); + const pageIndex = manifest.slugToPage[postId]; + if (pageIndex === undefined) throw new Error(`Unknown post slug: ${slug}`); + const page = await fetchPostsPage(channelSlug, pageIndex, origin); + const found = page.find((entry) => entry.slug === slug); + if (!found) throw new Error(`Post ${slug} missing from page ${pageIndex}`); + return found; +} + +// The thread a post belongs to (root + every archived reply), oldest first. +// This is the post-corpus analogue of a transcript's ±45s cue window: the +// natural "more context around this hit" unit. +export async function fetchThread(id: string): Promise<Post[]> { + const { origin, slug } = splitId(id); + const slashIdx = slug.indexOf("/"); + if (slashIdx < 0) throw new Error(`Malformed post slug: ${slug}`); + const channelSlug = slug.slice(0, slashIdx); + const post = await fetchPost(id); + const threadId = post.threadId || post.id; + + // A thread can straddle pages, so scan every page of the channel. Pages are + // cached, and a channel's page count is small (byte-capped shards). + const manifest = await fetchChannelPostsManifest(channelSlug, origin); + const thread: Post[] = []; + for (let i = 0; i < manifest.pageCount; i++) { + for (const entry of await fetchPostsPage(channelSlug, i, origin)) { + if ((entry.threadId || entry.id) === threadId) thread.push(entry); + } + } + thread.sort((a, b) => + a.createdAt === b.createdAt + ? a.id.localeCompare(b.id) + : a.createdAt.localeCompare(b.createdAt), + ); + return thread.length > 0 ? thread : [post]; +} + +export function usePostsManifest(origin = "") { + return useQuery<PostsManifest>({ + queryKey: ["posts-manifest", origin], + queryFn: () => + fetchJson<PostsManifest>(`${idBaseUrl(origin)}/posts/manifest.json`) + // A site with no social channels ships no posts manifest; treat that as + // an empty corpus rather than an error, so the UI degrades quietly. + .catch( + (): PostsManifest => ({ + version: 0, + channels: [], + totalCount: 0, + generatedAt: "", + }), + ), + }); +} + +export function useChannelPostsManifests( + refs: Array<string | PostsManifestRef>, +) { + return useQueries({ + queries: refs.map((ref) => { + const { channelSlug, origin = "" } = + typeof ref === "string" ? { channelSlug: ref, origin: "" } : ref; + return { + queryKey: ["channel-posts-manifest", origin, channelSlug], + queryFn: () => fetchChannelPostsManifest(channelSlug, origin), + }; + }), + }); +} diff --git a/common/components/searchPipeline.ts b/common/components/searchPipeline.ts @@ -1,5 +1,6 @@ import { fetchTranscript } from "./transcriptCache"; import { fetchSubs } from "./subsCache"; +import { fetchPost } from "./postsCache"; import type { DisplaySummary } from "../lib/transcripts"; import type { LayerScope, LeafNode } from "../lib/searchQuery"; @@ -287,6 +288,132 @@ function findFirstMatchInRange( return null; } +// Streaming search over the social-post corpus. Structurally the transcripts +// pipeline with a different fetch + match: the unit of iteration is a POST +// slug (`<channelSlug>/<postId>`), and fetchPost() resolves through the page +// cache, so the first post on a page warms every other post on it. +// +// Post bodies are short but unbounded, and unlike the cue path there is no +// natural per-line unit to clip to — so hits are truncated to a ±80-char +// window (the same one the description/tags scopes use) rather than shipping +// the whole body into a result card. +export function createPostsSearchPipeline( + config: Omit<PipelineConfig, "matchField">, +): PipelineController { + const { + slugs, + query, + useRegex, + regex, + initialHitLimit, + concurrency, + flushIntervalMs, + emit, + } = config; + + let cancelled = false; + let idx = 0; + let completed = 0; + let totalSoFar = 0; + let hitLimit = initialHitLimit; + let activeWorkers = 0; + let done = false; + const localHits: Record<string, Hit[]> = {}; + let flushTimer: number | null = null; + + const pushUpdate = (overrides: Partial<PipelineUpdate> = {}) => { + emit({ + hitsBySlug: { ...localHits }, + totalHits: totalSoFar, + processed: completed, + totalToProcess: slugs.length, + capped: totalSoFar >= hitLimit && idx < slugs.length, + done, + ...overrides, + }); + }; + + const scheduleFlush = () => { + if (flushTimer !== null || cancelled) return; + flushTimer = window.setTimeout(() => { + flushTimer = null; + if (cancelled) return; + pushUpdate(); + }, flushIntervalMs); + }; + + const finalize = () => { + if (done || cancelled) return; + done = true; + if (flushTimer !== null) { + window.clearTimeout(flushTimer); + flushTimer = null; + } + pushUpdate(); + }; + + const worker = async () => { + activeWorkers++; + try { + while (!cancelled) { + if (totalSoFar >= hitLimit) return; + if (idx >= slugs.length) return; + const my = idx++; + const slug = slugs[my]; + try { + const post = await fetchPost(slug); + if (cancelled) return; + if (totalSoFar < hitLimit) { + const hits = findHitsInText(post.text, query, useRegex, regex); + if (hits.length > 0) { + localHits[slug] = hits; + totalSoFar += hits.length; + } + } + } catch { + // ignore per-post failures + } + completed++; + scheduleFlush(); + } + } finally { + activeWorkers--; + if (activeWorkers === 0 && !cancelled) finalize(); + } + }; + + const ensureWorkers = () => { + if (cancelled || done) return; + if (totalSoFar >= hitLimit) return; + if (idx >= slugs.length) return; + const needed = Math.min(concurrency - activeWorkers, slugs.length - idx); + for (let i = 0; i < needed; i++) worker(); + }; + + pushUpdate(); + ensureWorkers(); + + return { + cancel() { + cancelled = true; + if (flushTimer !== null) { + window.clearTimeout(flushTimer); + flushTimer = null; + } + }, + setHitLimit(limit: number) { + if (cancelled) return; + if (limit <= hitLimit) return; + hitLimit = limit; + if (done) { + done = false; + pushUpdate(); + } + ensureWorkers(); + }, + }; +} + // Helper to build slug list from summaries + filter predicate. export function filterSlugs( summaries: DisplaySummary[], @@ -645,6 +772,19 @@ export function runLeafPipeline(opts: { : "cues", emit: (u) => onProgress(adaptTranscriptUpdate(u)), }); + } else if (leaf.scope === "posts") { + controller = createPostsSearchPipeline({ + slugs: scopeSlugs, + query: trimmed, + useRegex: leaf.useRegex, + regex, + initialHitLimit, + concurrency, + flushIntervalMs, + // A post has no timeline, so every hit is `start: 0` — the same + // convention description/tags/metadata hits already use. + emit: (u) => onProgress(adaptTranscriptUpdate(u)), + }); } else if (leaf.scope === "chat") { controller = createSubsSearchPipeline({ slugs: scopeSlugs, diff --git a/common/components/urlState.ts b/common/components/urlState.ts @@ -2,12 +2,15 @@ import { useMemo, useSyncExternalStore } from "react"; -export type SearchMode = "transcripts" | "subs"; +// Which corpus the legacy single-input search targets. "posts" is the social +// corpus (common/lib/posts.ts); the composite query tree can mix all of them +// freely, this only decides what a bare `?q=` URL means. +export type SearchMode = "transcripts" | "subs" | "posts"; // Per-video modal content mode. Independent from the search page's `mode` so // the modal can be toggled without disturbing search state. Absence on the // URL means "transcript" — only `"chat"` is persisted. -export type ModalMode = "transcript" | "chat"; +export type ModalMode = "transcript" | "chat" | "post"; export type UrlParams = { q: string; @@ -55,9 +58,11 @@ function parse(search: string): UrlParams { const tRaw = p.get("t"); const t = tRaw !== null && tRaw !== "" ? Number(tRaw) : null; const modeRaw = p.get("m"); - const mode: SearchMode = modeRaw === "subs" ? "subs" : "transcripts"; + const mode: SearchMode = + modeRaw === "subs" ? "subs" : modeRaw === "posts" ? "posts" : "transcripts"; const vmRaw = p.get("vm"); - const vm: ModalMode = vmRaw === "chat" ? "chat" : "transcript"; + const vm: ModalMode = + vmRaw === "chat" ? "chat" : vmRaw === "post" ? "post" : "transcript"; return { q: p.get("q") ?? "", re: p.get("re") === "1", @@ -120,6 +125,7 @@ export function writeUrlParams(patch: Patch) { } if (patch.vm !== undefined) { if (patch.vm === "chat") params.set("vm", "chat"); + else if (patch.vm === "post") params.set("vm", "post"); else params.delete("vm"); } if (patch.ch !== undefined) { @@ -156,6 +162,7 @@ export function writeUrlParams(patch: Patch) { } if (patch.mode !== undefined) { if (patch.mode === "subs") params.set("m", "subs"); + else if (patch.mode === "posts") params.set("m", "posts"); else params.delete("m"); } if (patch.tracks !== undefined) { diff --git a/common/controller/buildIndex.ts b/common/controller/buildIndex.ts @@ -65,7 +65,12 @@ import { subsPageFileName, } from "../lib/manifest"; import { getSettings } from "../lib/settings"; -import { listSites, siteSummariesDir, siteSubsDir } from "../lib/site"; +import { + listSites, + sitePostsDir, + siteSummariesDir, + siteSubsDir, +} from "../lib/site"; import { parseChannelConfig, type ChannelConfig, @@ -82,13 +87,44 @@ import { } from "../lib/videoStatus"; import { loadAvailability } from "../lib/availability-server"; import { AVAILABILITY_FILENAME } from "../lib/availability"; +import { + POSTS_MANIFEST_VERSION, + SITE_POSTS_MANIFEST_VERSION, + comparePostsNewestFirst, + postsPageFileName, + type ChannelPostsManifest, + type Post, + type PostPlatform, + type PostsChannelEntry, + type PostsManifest, +} from "../lib/posts"; +import { + channelPostsDir, + listPostShards, + readPostShard, +} from "../lib/posts-server"; +import { isSocialChannel } from "../lib/channelConfig"; // v10: multi-site build. Shared per-channel transcript/subs pages are written // once; per-site summaries + subs manifests are filtered selections. Bumped to // force a clean rebuild into the new shared/ + sites/ staging layout. // v11: transcript pages now carry `tags` (for the "tags" search scope) — force // a re-extract so existing pages re-emit with the field. -const SCHEMA_VERSION = 11; +// v12: the social-post corpus. A `posts` sub-DB keyed [createdAt, channelSlug, +// id] (ISO-8601 sorts correctly, unlike the [uploadDate, …] tuple videos use) +// plus a shared /posts/<slug>/ page tree. +const SCHEMA_VERSION = 12; + +// Per-channel post stats, persisted so per-site aggregates survive a no-op +// rebuild that doesn't re-encode the post pages. Mirrors ChannelSubsStat. +type ChannelPostsStat = { + name: string; + postCount: number; + platform: string; + // Signature of the channel's on-disk posts dir, so an unchanged channel + // skips re-encoding. Same idea as the mtime records the video scan keeps. + signature: string; +}; // Per-channel subtitle stats, collected while writing the shared subs pages and // persisted to LMDB so per-site subs manifests can be assembled on a no-op @@ -103,6 +139,8 @@ type ChannelSubsStat = { type IndexKey = [string, string, string]; type ChannelKey = [string, string, string]; type PageHashKey = [string, number]; +// [createdAt, channelSlug, postId] — see the `posts` sub-DB comment below. +type PostIndexKey = [string, string, string]; type PathKey = [string, string]; type MtimeRecord = { @@ -189,6 +227,10 @@ async function scanSource( continue; } channels.set(ch.name, cfg); + // A social channel carries posts, not videos: it has no data/ dir to walk, + // and routing it through the video scan would only ever produce noise. Its + // posts tree is built from the JSONL shards further down. + if (isSocialChannel(cfg)) continue; const dataDir = path.join(channelDir, "data"); let videoEntries: Dirent[]; try { @@ -310,15 +352,17 @@ export async function buildIndex({ // in the per-site loop below. const transcriptsOutDir = paths.exportSharedTranscriptsDir; const subsOutDir = paths.exportSharedSubsDir; + const postsOutDir = paths.exportSharedPostsDir; await mkdir(path.dirname(dbPath), { recursive: true }); await mkdir(transcriptsOutDir, { recursive: true }); await mkdir(subsOutDir, { recursive: true }); + await mkdir(postsOutDir, { recursive: true }); await mkdir(paths.exportSitesIndexDir, { recursive: true }); const root = open({ path: dbPath, - maxDbs: 14, + maxDbs: 17, compression: true, }); const sums = root.openDB<TranscriptSummary, IndexKey>({ @@ -355,6 +399,21 @@ export async function buildIndex({ name: "channelStats", encoding: "msgpack", }); + // The social-post corpus. Keyed [createdAt, channelSlug, id]: createdAt is + // ISO-8601 so a lexicographic key sort IS a chronological sort, which fixes + // the intra-day ordering problem the videos' [uploadDate, …] tuple has. + const posts = root.openDB<Post, PostIndexKey>({ + name: "posts", + encoding: "msgpack", + }); + const postPageHashes = root.openDB<PageHashRecord, PageHashKey>({ + name: "postPageHashes", + encoding: "msgpack", + }); + const channelPostsStatsDb = root.openDB<ChannelPostsStat, string>({ + name: "channelPostsStats", + encoding: "msgpack", + }); const meta = root.openDB<unknown, string>({ name: "meta", encoding: "msgpack", @@ -374,6 +433,9 @@ export async function buildIndex({ await pageHashes.clearAsync(); await subPageHashes.clearAsync(); await channelStatsDb.clearAsync(); + await posts.clearAsync(); + await postPageHashes.clearAsync(); + await channelPostsStatsDb.clearAsync(); await meta.put("schema", SCHEMA_VERSION); } @@ -1001,6 +1063,183 @@ export async function buildIndex({ } // --------------------------------------------------------------------------- + // The social-post corpus: a parallel per-channel page tree beside transcripts + // and subs. Social channels have no `data/` dir and so never appear in the + // video scan's mtime bookkeeping; incrementality here keys off a signature of + // each channel's on-disk posts shards instead. + // --------------------------------------------------------------------------- + const channelPostsStats = new Map<string, ChannelPostsStat>(); + let postsPagesWritten = 0; + let postsPagesSkipped = 0; + let postsTotalCount = 0; + + const socialSlugs = Array.from(channelConfigs.entries()) + .filter(([, cfg]) => isSocialChannel(cfg)) + .map(([slug]) => slug) + .sort(); + + for (const channelSlug of socialSlugs) { + const cfg = channelConfigs.get(channelSlug)!; + const channelRoot = path.join(channelsDir, channelSlug); + const postsChannelDir = path.join(postsOutDir, channelSlug); + + // Signature over the channel's month shards (name + size + mtime). Cheap + // to compute, and it changes exactly when a fetch appended something. + const shards = await listPostShards(channelRoot); + const sigParts: string[] = []; + for (const shard of shards) { + try { + const st = await stat( + path.join(channelPostsDir(channelRoot), `${shard}.jsonl`), + ); + sigParts.push(`${shard}:${st.size}:${st.mtimeMs}`); + } catch { + /* shard vanished mid-scan — treat as absent */ + } + } + const signature = sigParts.join("|"); + const prevStat = channelPostsStatsDb.get(channelSlug); + const manifestPresent = await readFile( + path.join(postsChannelDir, "manifest.json"), + "utf8", + ) + .then(() => true) + .catch(() => false); + + if ( + !schemaBumped && + prevStat && + prevStat.signature === signature && + manifestPresent + ) { + channelPostsStats.set(channelSlug, prevStat); + postsTotalCount += prevStat.postCount; + continue; + } + + // Load this channel's posts and refresh its slice of the posts DB. Clearing + // by range first means a deleted/rewritten shard can't leave orphans. + for (const { key } of posts.getRange({})) { + const pk = key as PostIndexKey; + if (pk[1] === channelSlug) posts.remove(pk); + } + const channelPosts: Post[] = []; + for (const shard of shards) { + channelPosts.push(...(await readPostShard(channelRoot, shard))); + } + // Dedupe by id (an interrupted archive write can duplicate across shards) + // then order newest-first, the order the page tree and the UI present. + const byId = new Map<string, Post>(); + for (const post of channelPosts) byId.set(post.id, post); + const ordered = [...byId.values()].sort(comparePostsNewestFirst); + for (const post of ordered) { + posts.put([post.createdAt, channelSlug, post.id], post); + } + + const postWriter = createPageWriter({ + outDir: postsChannelDir, + fileName: postsPageFileName, + maxPageBytes: maxTranscriptPageBytes, + ensureDir: true, + getPrevHash: (idx) => postPageHashes.get([channelSlug, idx])?.hash, + setHash: (idx, record) => { + postPageHashes.put([channelSlug, idx], record); + }, + onLog: log, + }); + for (const post of ordered) { + await postWriter.push(JSON.stringify(post), post.id); + } + const { + pageCount: postPageCount, + pagesWritten: chPostsWritten, + pagesSkipped: chPostsSkipped, + slugToPage: postSlugToPage, + } = await postWriter.finish(); + postsPagesWritten += chPostsWritten; + postsPagesSkipped += chPostsSkipped; + + if (ordered.length === 0) { + await rm(postsChannelDir, { recursive: true, force: true }); + for (const { key } of postPageHashes.getRange({ + start: [channelSlug], + end: [channelSlug, Number.MAX_SAFE_INTEGER], + })) { + postPageHashes.remove(key as PageHashKey); + } + channelPostsStatsDb.remove(channelSlug); + continue; + } + + const postKeep = new Set<string>(["manifest.json"]); + for (let i = 0; i < postPageCount; i++) postKeep.add(postsPageFileName(i)); + for (const name of await readdir(postsChannelDir).catch(() => [] as string[])) { + if (postKeep.has(name)) continue; + await rm(path.join(postsChannelDir, name), { force: true }); + } + for (const { key } of postPageHashes.getRange({ + start: [channelSlug, postPageCount], + end: [channelSlug, Number.MAX_SAFE_INTEGER], + })) { + postPageHashes.remove(key as PageHashKey); + } + + const postsManifest: ChannelPostsManifest = { + version: POSTS_MANIFEST_VERSION, + channelSlug, + pageCount: postPageCount, + maxPageBytes: maxTranscriptPageBytes, + generatedAt, + slugToPage: postSlugToPage, + }; + await writeJsonAtomic( + path.join(postsChannelDir, "manifest.json"), + postsManifest, + ); + + const statRecord: ChannelPostsStat = { + name: cfg.name ?? channelSlug, + postCount: ordered.length, + platform: cfg.platform ?? "", + signature, + }; + channelPostsStatsDb.put(channelSlug, statRecord); + channelPostsStats.set(channelSlug, statRecord); + postsTotalCount += ordered.length; + } + + // Load stats for channels skipped above (unchanged) that we didn't touch, and + // drop page trees for channels that are no longer social / no longer exist. + for (const { key, value } of channelPostsStatsDb.getRange()) { + const slug = key as string; + if (!channelPostsStats.has(slug)) { + if (!socialSlugs.includes(slug)) { + await rm(path.join(postsOutDir, slug), { recursive: true, force: true }); + channelPostsStatsDb.remove(slug); + continue; + } + channelPostsStats.set(slug, value as ChannelPostsStat); + } + } + for (const e of await readdir(postsOutDir, { withFileTypes: true }).catch( + () => [] as Dirent[], + )) { + if (e.isDirectory() && !channelPostsStats.has(e.name)) { + await rm(path.join(postsOutDir, e.name), { recursive: true, force: true }); + } + } + await posts.flushed; + await postPageHashes.flushed; + await channelPostsStatsDb.flushed; + + if (socialSlugs.length > 0) { + log( + `Post pages: ${postsPagesWritten} written, ${postsPagesSkipped} unchanged ` + + `across ${channelPostsStats.size} social channel(s) (${postsTotalCount} posts).`, + ); + } + + // --------------------------------------------------------------------------- // Per-site aggregates: a filtered summaries index (pages + manifest) and a // site-level subs manifest, one bundle per configured site. The heavy // per-channel page trees above are shared; here we only select + regroup. @@ -1037,8 +1276,10 @@ export async function buildIndex({ const summariesOut = siteSummariesDir(paths, site.siteId); const subsOut = siteSubsDir(paths, site.siteId); + const postsOut = sitePostsDir(paths, site.siteId); const summariesManifestPath = path.join(summariesOut, "manifest.json"); const subsManifestPath = path.join(subsOut, "manifest.json"); + const postsManifestPath = path.join(postsOut, "manifest.json"); const fingerprint = JSON.stringify({ gen: generation, @@ -1179,6 +1420,33 @@ export async function buildIndex({ }; await writeJsonAtomic(subsManifestPath, siteSubsManifest); + // --- site-level posts manifest (filtered to member channels) --- + // Only social members contribute; a site with none still gets a manifest + // with an empty channel list, so the client's fetch is unconditional. + const postsEntries: PostsChannelEntry[] = []; + let postsTotalForSite = 0; + for (const slug of memberSlugs) { + const statRec = channelPostsStats.get(slug); + if (!statRec || statRec.postCount === 0) continue; + postsEntries.push({ + name: statRec.name, + slug, + postCount: statRec.postCount, + platform: (statRec.platform || "bluesky") as PostPlatform, + groupId: slugGroup.get(slug) ?? site.defaultGroupId, + }); + postsTotalForSite += statRec.postCount; + } + const sitePostsManifest: PostsManifest = { + version: SITE_POSTS_MANIFEST_VERSION, + channels: postsEntries.sort((a, b) => a.name.localeCompare(b.name)), + totalCount: postsTotalForSite, + generatedAt: new Date().toISOString(), + siteId: site.siteId, + }; + await mkdir(postsOut, { recursive: true }); + await writeJsonAtomic(postsManifestPath, sitePostsManifest); + await meta.put(fpKey, fingerprint); sitesBuilt++; aggregateSummaryPages += pageIndex; diff --git a/common/controller/fetchPosts.ts b/common/controller/fetchPosts.ts @@ -0,0 +1,192 @@ +// Drive a SocialFetcher for one social channel and land the result on disk. +// +// This is the posts-corpus analogue of sync(): read the channel's already-seen +// ids, ask the fetcher for anything newer, append to the month-sharded JSONL +// and the posts-archive, and record a small state sidecar for the channel's +// snapshot buckets. Incremental by construction — a re-run over overlapping +// pages writes nothing new. + +import path from "node:path"; +import { readChannelConfig, writeChannelConfig } from "./channels"; +import type { Paths } from "../lib/paths"; +import { isSocialChannel } from "../lib/channelConfig"; +import { + alwaysCookies, + resolveCookiePolicy, + type CookiePolicyInputs, +} from "../lib/cookiePolicy"; +import { + latestPostCreatedAt, + readPostFetchState, + readSeenPostIds, + writePostFetchState, + writePosts, + type PostFetchState, +} from "../lib/posts-server"; +import { + handleFromAccountUrl, + resolveSocialFetcher, +} from "../social/fetchers"; +// Registering the built-in fetchers is a side effect of importing them. Keep +// this list here (rather than inside fetchers.ts) so the registry module stays +// free of imports from the heavier fetchers. +import "../social/blueskyFetcher"; +// Registration order encodes preference: gallery-dl is the primary X path, so +// it is registered before the Playwright fallback and wins URL detection. +import "../social/xGalleryDlFetcher"; +// The fallback is registered LAST and never claims a URL by detection, so it is +// only ever used when a channel opts into it via postFetcher: "x-playwright". +import "../social/xPlaywrightFetcher"; + +export type FetchPostsOptions = { + paths: Paths; + slug: string; + // Global settings, for the cookie policy. Structural so settings.ts need not + // be imported here. + settings: CookiePolicyInputs; + // Ignore the stored watermark and re-walk the account's full history. New + // posts are still deduped against the archive, so this is a safe repair + // operation rather than a duplicate-maker. + full?: boolean; + limit?: number; + onLog?: (line: string) => void; + signal?: AbortSignal; +}; + +export type FetchPostsResult = { + ok: boolean; + written: number; + skipped: number; + complete: boolean; + error?: string; + needsCookies?: boolean; +}; + +export async function fetchPosts( + opts: FetchPostsOptions, +): Promise<FetchPostsResult> { + const { paths, slug, settings, onLog, signal } = opts; + const log = (line: string) => onLog?.(line); + const channelRoot = path.join(paths.channelsDir, slug); + + const config = await readChannelConfig(paths, slug); + if (!config) { + return { ok: false, written: 0, skipped: 0, complete: false, error: `No such channel: ${slug}` }; + } + if (!isSocialChannel(config)) { + return { + ok: false, + written: 0, + skipped: 0, + complete: false, + error: `Channel ${slug} is not a social source (sourceKind=${config.sourceKind ?? "video"})`, + }; + } + + const accountUrl = config.url ?? ""; + const fetcher = resolveSocialFetcher(config.postFetcher, accountUrl); + if (!fetcher) { + return { + ok: false, + written: 0, + skipped: 0, + complete: false, + error: `No post fetcher for ${accountUrl || slug} (configured: ${config.postFetcher ?? "auto-detect"})`, + }; + } + + const handle = + config.socialHandle ?? handleFromAccountUrl(accountUrl) ?? ""; + if (!handle) { + return { + ok: false, + written: 0, + skipped: 0, + complete: false, + error: `Could not determine an account handle for ${slug}`, + }; + } + + const seenIds = await readSeenPostIds(channelRoot); + const priorState = await readPostFetchState(channelRoot); + // A stored cursor means the previous run stopped early (hit its limit, or was + // cancelled). Resume the backfill from there rather than applying the + // watermark, which would otherwise leave that history permanently missing. + // --full discards the cursor and re-walks from the top. + const resumeCursor = opts.full ? undefined : priorState?.cursor; + const since = + opts.full || resumeCursor + ? undefined + : ((await latestPostCreatedAt(channelRoot)) ?? undefined); + const cookies = alwaysCookies(resolveCookiePolicy(settings, config)); + + log( + `Fetching posts for ${slug} via ${fetcher.label} (@${handle})` + + (resumeCursor + ? " — resuming an unfinished backfill" + : since + ? ` since ${since}` + : " — full history") + + `; ${seenIds.size} already archived.`, + ); + + // A managed job's cancel/drain arrives as an AbortSignal; the fetcher + // interface requires one, so synthesize a never-aborting signal when the + // caller has none (a direct CLI invocation). + const controller = new AbortController(); + const effectiveSignal = signal ?? controller.signal; + + const state: PostFetchState = { + lastFetchedAt: new Date().toISOString(), + }; + + let result; + try { + result = await fetcher.fetch({ + accountUrl, + handle, + channelSlug: slug, + since, + cursor: resumeCursor, + seenIds, + cookies, + limit: opts.limit, + signal: effectiveSignal, + onLog: log, + }); + } catch (err) { + const message = (err as Error).message; + log(`[error] ${message}`); + state.lastError = message; + await writePostFetchState(channelRoot, state); + return { ok: false, written: 0, skipped: 0, complete: false, error: message }; + } + + const written = await writePosts(channelRoot, result.posts); + log( + `Wrote ${written.written} new post(s), skipped ${written.skipped} already archived` + + (written.shards.length > 0 ? ` (shards: ${written.shards.join(", ")})` : "") + + (result.complete ? "." : " — more history remains."), + ); + + state.lastFetchedCount = written.written; + if (result.cursor && !result.complete) state.cursor = result.cursor; + if (result.needsCookies) state.needsCookies = true; + await writePostFetchState(channelRoot, state); + + // Reuse the existing lastSyncedAt field so the scheduler + // (common/jobs/syncScheduler.ts) paces social channels with zero changes — + // it keys only off url / excludeFromSync / syncIntervalMinutes / lastSyncedAt. + await writeChannelConfig(paths, slug, { + ...config, + lastSyncedAt: new Date().toISOString(), + }); + + return { + ok: true, + written: written.written, + skipped: written.skipped, + complete: result.complete, + needsCookies: result.needsCookies, + }; +} diff --git a/common/jobs/jobKinds.ts b/common/jobs/jobKinds.ts @@ -77,6 +77,28 @@ const JOB_KINDS: Record<string, JobKindMeta> = { bookmarkable: true, queueKeyStrategy: "custom", }, + // Replace-auto-captions lane, transcribe half: whisper over videos whose only + // transcript is a YouTube ASR VTT (the downloadedAutoSubsOnly bucket). Same + // batch machinery as whisper-bucket-downloaded-no-transcript — drainable and + // bookmarkable, re-deriving the bucket's current members on replay. + "whisper-bucket-auto-subs": { + kind: "whisper-bucket-auto-subs", + label: "Replace auto-captions", + drainable: true, + bookmarkable: true, + queueKeyStrategy: "custom", + }, + // Delete the superseded English ASR VTTs kept as backups next to a finished + // whisper transcript. Manual only — never auto-queued — and the single + // irreversible step in the lane, so it is deliberately NOT drainable (it is a + // fast file sweep) but IS bookmarkable. + "purge-superseded-auto-subs": { + kind: "purge-superseded-auto-subs", + label: "Purge superseded auto-captions", + drainable: false, + bookmarkable: true, + queueKeyStrategy: "custom", + }, "redownload-incomplete-bucket": { kind: "redownload-incomplete-bucket", label: "Re-download truncated transcripts", @@ -161,6 +183,18 @@ const JOB_KINDS: Record<string, JobKindMeta> = { bookmarkable: false, queueKeyStrategy: "custom", }, + // Social-post ingest for a `sourceKind: "social"` channel. Drainable (the + // fetcher stops paging on the drain signal and keeps what it already has) and + // bookmarkable. queueKeyForUrl() routes x.com / bsky.app to + // `platform:x.com` / `platform:bsky.app`, so per-platform serialization and + // the existing 429 backoff come free. + "fetch-posts": { + kind: "fetch-posts", + label: "Fetch posts", + drainable: true, + bookmarkable: true, + queueKeyStrategy: "platform", + }, sync: { kind: "sync", label: "Sync", diff --git a/common/lib/channelConfig.ts b/common/lib/channelConfig.ts @@ -7,6 +7,26 @@ import { isCookieMode, type CookieMode } from "./cookiePolicy"; export type ChannelHandling = "youtube" | "transcribe"; +// What KIND of source this channel is. Deliberately a separate axis from +// `handling` rather than a third handling value: every existing +// `handling === "transcribe" ? … : …` branch stays binary and so can never +// misroute a social channel down a video path. +// "video" (default) — yt-dlp + whisper, the original pipeline. +// "social" — a social account fetched into the parallel posts corpus +// (see common/lib/posts.ts). Skipped by the video scan. +export type ChannelSourceKind = "video" | "social"; + +export const SOURCE_KIND_VALUES: ReadonlyArray<ChannelSourceKind> = [ + "video", + "social", +]; + +export function isSocialChannel( + config: Pick<ChannelConfig, "sourceKind"> | null | undefined, +): boolean { + return config?.sourceKind === "social"; +} + export type AudioFormat = "m4a" | "mp3" | "opus"; // Who owns audio extraction for a transcribe-handling download: @@ -38,6 +58,16 @@ export type AudioCheckConfig = { export type ChannelConfig = { handling: ChannelHandling; + // Omitted = "video" (every channel that predates the posts corpus). + sourceKind?: ChannelSourceKind; + // Social channels only: which SocialFetcher drives ingest (see + // common/social/fetchers.ts), e.g. "bluesky-atproto" / "x-gallery-dl". + // Omitted = resolve by URL detection. + postFetcher?: string; + // Social channels only: the bare account handle, without "@" or URL wrapper. + // Derived from `url` at creation time but stored so a later URL-format change + // upstream can't silently re-point ingest at a different account. + socialHandle?: string; platform?: Platform; name?: string; url?: string; @@ -167,6 +197,15 @@ export function parseChannelConfig(raw: unknown): ChannelConfig | null { const r = raw as Record<string, unknown>; if (r.handling !== "youtube" && r.handling !== "transcribe") return null; const config: ChannelConfig = { handling: r.handling }; + if (r.sourceKind === "video" || r.sourceKind === "social") { + config.sourceKind = r.sourceKind; + } + if (typeof r.postFetcher === "string" && r.postFetcher.trim()) { + config.postFetcher = r.postFetcher.trim(); + } + if (typeof r.socialHandle === "string" && r.socialHandle.trim()) { + config.socialHandle = r.socialHandle.trim().replace(/^@/, ""); + } if ( typeof r.platform === "string" && PLATFORM_VALUES.includes(r.platform as Platform) diff --git a/common/lib/corpus.ts b/common/lib/corpus.ts @@ -11,7 +11,9 @@ import type { PublicSiteDescriptor } from "./siteDescriptor"; // Pure module: builders take already-loaded data and return plain objects / // strings. All file I/O lives in compose-site.ts / compose-hub.ts. -export const CORPUS_SPEC_VERSION = 1; +// v2: the corpus now also describes the parallel social-post layer +// (postScheme + per-channel posts manifest pointers). +export const CORPUS_SPEC_VERSION = 2; // How to resolve a single transcript from the paginated shards, described once // and embedded in every corpus.json so any HTTP client can navigate without @@ -35,6 +37,32 @@ const SHARD_SCHEME = { pageNumberFormat: "zero-padded to 4 digits, e.g. page 0 -> page-0000.json", } as const; +// The social-post corpus: a parallel dataset to video transcripts, served under +// the same paginated-shard scheme. A post has no timeline, so it carries a +// `createdAt` instant instead of cue timings and its permalink needs no +// timestamp fragment. +const POST_SCHEME = { + description: + "Social posts (X/Twitter, Bluesky) are archived as a PARALLEL corpus to " + + "video transcripts and are served as paginated JSON shards under the same " + + "scheme: (1) GET the channel's posts manifest; (2) look up the post id in " + + "its `slugToPage` map to get a page number N; (3) GET page-<NNNN>.json and " + + "take the record whose `id` matches. Only channels whose source is a social " + + "account have a posts manifest.", + postsManifest: + "<channel.manifests.posts> -> { pageCount, slugToPage: { <postId>: <pageNumber> } }", + postPage: + "/posts/<slug>/page-<NNNN>.json -> array of { id, slug, channelSlug, author, " + + "authorName, createdAt, uploadDate, text, url, platform, threadId, replyTo, " + + "quoted, repostOf, isReply, isRepost, links, mediaCount, engagement }", + ordering: + "newest first, by `createdAt` (ISO-8601, ms precision where the source provides it)", + dateFilter: + "`uploadDate` (YYYYMMDD, derived from createdAt) is carried on every post so " + + "the same date filters work across transcripts and posts", + permalink: "each post carries its own canonical `url`; no timestamp fragment applies", +} as const; + export type CorpusChannel = { slug: string; name: string; @@ -43,7 +71,11 @@ export type CorpusChannel = { manifests: { transcripts: string; subs: string; + // Present only for social channels (the posts corpus). + posts?: string; }; + // Present only for social channels. + postCount?: number; }; export type SiteCorpus = { @@ -60,6 +92,8 @@ export type SiteCorpus = { totals: { channels: number; videos: number }; channels: CorpusChannel[]; shardScheme: typeof SHARD_SCHEME; + // Present when this site includes at least one social channel. + postScheme?: typeof POST_SCHEME; // Present when this build ships bulk-download archives (whole-channel zips). bulkArchives?: { manifest: string; note: string }; // Pointer to the human page and BYO-key chat. @@ -105,19 +139,32 @@ function join(base: string | undefined, p: string): string { // plus whether this build emitted bulk archives. export function buildSiteCorpus( descriptor: PublicSiteDescriptor, - opts: { hasArchives: boolean }, + opts: { + hasArchives: boolean; + // slug -> archived post count, for the social channels in this site. Absent + // / empty means the site has no posts corpus and postScheme is omitted. + postCounts?: Record<string, number>; + }, ): SiteCorpus { const base = descriptor.siteUrl; - const channels: CorpusChannel[] = descriptor.channels.map((c) => ({ - slug: c.slug, - name: c.name, - videoCount: c.count, - ...(c.groupId ? { groupId: c.groupId } : {}), - manifests: { - transcripts: join(base, `/transcripts/${c.slug}/manifest.json`), - subs: join(base, `/subs/${c.slug}/manifest.json`), - }, - })); + const postCounts = opts.postCounts ?? {}; + const channels: CorpusChannel[] = descriptor.channels.map((c) => { + const postCount = postCounts[c.slug]; + return { + slug: c.slug, + name: c.name, + videoCount: c.count, + ...(c.groupId ? { groupId: c.groupId } : {}), + ...(postCount ? { postCount } : {}), + manifests: { + transcripts: join(base, `/transcripts/${c.slug}/manifest.json`), + subs: join(base, `/subs/${c.slug}/manifest.json`), + ...(postCount + ? { posts: join(base, `/posts/${c.slug}/manifest.json`) } + : {}), + }, + }; + }); const videos = channels.reduce((n, c) => n + c.videoCount, 0); const corpus: SiteCorpus = { @@ -136,6 +183,11 @@ export function buildSiteCorpus( shardScheme: SHARD_SCHEME, useWithAi: join(base, "/use-with-ai"), }; + // Only advertise the post scheme when this site actually ships posts, so a + // pure-video site's corpus.json is unchanged apart from the spec bump. + if (Object.values(postCounts).some((n) => n > 0)) { + corpus.postScheme = POST_SCHEME; + } if (opts.hasArchives) { corpus.bulkArchives = { manifest: join(base, "/archives/manifest.json"), diff --git a/common/lib/paths.ts b/common/lib/paths.ts @@ -58,6 +58,9 @@ export type Paths = { exportSummariesDir: string; exportTranscriptsDir: string; exportSubsDir: string; + // Served per-channel social-post page tree (/posts/<slug>/{manifest,page-NNNN}.json). + // The parallel corpus to transcripts/subs — see common/lib/posts.ts. + exportPostsDir: string; exportStatsDir: string; // Staging area (NOT served) where the shared index + per-site aggregates are // built before composition. Shared per-channel transcript/subs pages are @@ -67,6 +70,7 @@ export type Paths = { exportSharedDir: string; exportSharedTranscriptsDir: string; exportSharedSubsDir: string; + exportSharedPostsDir: string; exportSitesIndexDir: string; // Per-site extracted static output (`out/`) from an isolated (Docker) build, // keyed exportBuildsDir/<siteId>. Sibling of .export-index. The wrangler deploy @@ -93,6 +97,11 @@ export type Paths = { // (Phase 4 of the video-persistence feature). See // common/controller/backupSavedVideos.ts. rsyncBin: string; + // gallery-dl, the primary X/Twitter post fetcher (see + // common/social/xGalleryDlFetcher.ts). A light headless subprocess — the same + // shape the codebase already manages for yt-dlp and whisper. NOT bundled; + // install it separately and point GALLERY_DL_BIN at it if it isn't on PATH. + galleryDlBin: string; // parakeet (overlapping-segment stitching) app. parakeetBin is the standalone // wrapper script invoked as the app binary; parakeetCliBin is the underlying // parakeet-cli it drives; parakeetModel is the default .gguf model. @@ -142,11 +151,13 @@ export function getPaths(): Paths { exportSummariesDir: path.join(exportPublicDir, "summaries"), exportTranscriptsDir: path.join(exportPublicDir, "transcripts"), exportSubsDir: path.join(exportPublicDir, "subs"), + exportPostsDir: path.join(exportPublicDir, "posts"), exportStatsDir: path.join(exportPublicDir, "stats"), exportIndexDir, exportSharedDir, exportSharedTranscriptsDir: path.join(exportSharedDir, "transcripts"), exportSharedSubsDir: path.join(exportSharedDir, "subs"), + exportSharedPostsDir: path.join(exportSharedDir, "posts"), exportSitesIndexDir: path.join(exportIndexDir, "sites"), exportBuildsDir: process.env.EXPORT_BUILDS_DIR ?? @@ -173,6 +184,7 @@ export function getPaths(): Paths { ffmpegBin: process.env.FFMPEG_BIN ?? "ffmpeg", ffprobeBin: process.env.FFPROBE_BIN ?? "ffprobe", rsyncBin: process.env.RSYNC_BIN ?? "rsync", + galleryDlBin: process.env.GALLERY_DL_BIN ?? "gallery-dl", parakeetBin: process.env.PARAKEET_STITCH_BIN ?? path.join(monorepoRoot, "scripts", "parakeet-stitch.mjs"), diff --git a/common/lib/platform.ts b/common/lib/platform.ts @@ -1,4 +1,15 @@ -export type Platform = "youtube" | "rumble" | "odysee" | "twitch" | "kick"; +// Video platforms plus the social-post platforms (see common/lib/posts.ts). +// Widening this allowlist is purely additive: it cannot invalidate existing +// cached entries, so no transcriptStore.ts DB_VERSION bump is needed. Posts get +// their own IDB store rather than sharing the transcript cache. +export type Platform = + | "youtube" + | "rumble" + | "odysee" + | "twitch" + | "kick" + | "twitter" + | "bluesky"; export const PLATFORM_VALUES: ReadonlyArray<Platform> = [ "youtube", @@ -6,8 +17,22 @@ export const PLATFORM_VALUES: ReadonlyArray<Platform> = [ "odysee", "twitch", "kick", + "twitter", + "bluesky", ]; +// The subset that carries social posts rather than videos. A channel on one of +// these is a `sourceKind: "social"` channel — the video scan skips it and the +// posts pipeline picks it up. +export const SOCIAL_PLATFORM_VALUES: ReadonlyArray<Platform> = [ + "twitter", + "bluesky", +]; + +export function isSocialPlatform(platform: Platform | null | undefined): boolean { + return platform === "twitter" || platform === "bluesky"; +} + export function detectPlatform( url: string | undefined | null, ): Platform | null { @@ -19,6 +44,9 @@ export function detectPlatform( if (host.endsWith("odysee.com")) return "odysee"; if (host.endsWith("twitch.tv")) return "twitch"; if (host.endsWith("kick.com")) return "kick"; + if (host === "x.com" || host.endsWith(".x.com")) return "twitter"; + if (host.endsWith("twitter.com")) return "twitter"; + if (host === "bsky.app" || host.endsWith(".bsky.app")) return "bluesky"; } catch { /* fall through */ } @@ -36,6 +64,11 @@ export function defaultWebpageUrl(platform: Platform, id: string): string { if (platform === "twitch") return `https://www.twitch.tv/videos/${id}`; // Kick's canonical id is the VOD UUID; /video/<uuid> resolves to the VOD. if (platform === "kick") return `https://kick.com/video/${id}`; + // Social posts: /i/status/<id> resolves without knowing the handle. Bluesky + // has no handle-free permalink, so this is only a last-resort fallback — + // every archived post carries its own canonical `url` (see postPermalink). + if (platform === "twitter") return `https://x.com/i/status/${id}`; + if (platform === "bluesky") return `https://bsky.app/profile/${id}`; return `https://www.youtube.com/watch?v=${id}`; } diff --git a/common/lib/posts-server.test.ts b/common/lib/posts-server.test.ts @@ -0,0 +1,148 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { mkdtemp, readFile, rm } from "node:fs/promises"; +import { tmpdir } from "node:os"; +import path from "node:path"; +import { + channelPostsArchivePath, + countPosts, + latestPostCreatedAt, + listPostShards, + readAllPosts, + readPostShard, + readSeenPostIds, + writePosts, +} from "./posts-server"; +import { uploadDateFromCreatedAt, type Post } from "./posts"; + +function makePost(id: string, createdAt: string): Post { + return { + id, + slug: `chan/${id}`, + channelSlug: "chan", + author: "someone.bsky.social", + createdAt, + uploadDate: uploadDateFromCreatedAt(createdAt), + text: `post ${id}`, + url: `https://bsky.app/profile/someone.bsky.social/post/${id}`, + platform: "bluesky", + isReply: false, + isRepost: false, + links: [], + }; +} + +async function withTempChannel( + fn: (root: string) => Promise<void>, +): Promise<void> { + const root = await mkdtemp(path.join(tmpdir(), "posts-store-")); + try { + await fn(root); + } finally { + await rm(root, { recursive: true, force: true }); + } +} + +test("writes posts into month shards and records their ids", async () => { + await withTempChannel(async (root) => { + const result = await writePosts(root, [ + makePost("a1", "2026-01-15T10:00:00.000Z"), + makePost("a2", "2026-01-20T10:00:00.000Z"), + makePost("b1", "2026-02-01T10:00:00.000Z"), + ]); + assert.equal(result.written, 3); + assert.equal(result.skipped, 0); + assert.deepEqual(result.shards, ["2026-01", "2026-02"]); + + assert.deepEqual(await listPostShards(root), ["2026-01", "2026-02"]); + assert.equal((await readPostShard(root, "2026-01")).length, 2); + assert.equal((await readPostShard(root, "2026-02")).length, 1); + + const archive = await readFile(channelPostsArchivePath(root), "utf8"); + assert.match(archive, /^bluesky a1$/m); + assert.match(archive, /^bluesky b1$/m); + assert.deepEqual([...(await readSeenPostIds(root))].sort(), ["a1", "a2", "b1"]); + }); +}); + +test("a re-run over overlapping pages writes no duplicates", async () => { + await withTempChannel(async (root) => { + await writePosts(root, [ + makePost("a1", "2026-01-15T10:00:00.000Z"), + makePost("a2", "2026-01-20T10:00:00.000Z"), + ]); + // Second run re-delivers a1/a2 and adds one new post — the incremental + // guarantee the whole fetch loop rests on. + const second = await writePosts(root, [ + makePost("a1", "2026-01-15T10:00:00.000Z"), + makePost("a2", "2026-01-20T10:00:00.000Z"), + makePost("a3", "2026-01-25T10:00:00.000Z"), + ]); + assert.equal(second.written, 1); + assert.equal(second.skipped, 2); + + const all = await readAllPosts(root); + assert.deepEqual(all.map((p) => p.id), ["a3", "a2", "a1"]); + assert.equal(await countPosts(root), 3); + }); +}); + +test("dedupes within a single batch", async () => { + await withTempChannel(async (root) => { + const result = await writePosts(root, [ + makePost("x", "2026-03-01T10:00:00.000Z"), + makePost("x", "2026-03-01T10:00:00.000Z"), + ]); + assert.equal(result.written, 1); + assert.equal(result.skipped, 1); + assert.equal(await countPosts(root), 1); + }); +}); + +test("readAllPosts returns newest first across shards", async () => { + await withTempChannel(async (root) => { + await writePosts(root, [ + makePost("old", "2025-11-02T00:00:00.000Z"), + makePost("new", "2026-04-09T00:00:00.000Z"), + makePost("mid", "2026-01-01T00:00:00.000Z"), + ]); + const all = await readAllPosts(root); + assert.deepEqual(all.map((p) => p.id), ["new", "mid", "old"]); + }); +}); + +test("latestPostCreatedAt is the incremental watermark", async () => { + await withTempChannel(async (root) => { + assert.equal(await latestPostCreatedAt(root), null); + await writePosts(root, [ + makePost("old", "2025-11-02T00:00:00.000Z"), + makePost("new", "2026-04-09T12:34:56.789Z"), + ]); + assert.equal(await latestPostCreatedAt(root), "2026-04-09T12:34:56.789Z"); + }); +}); + +test("an empty channel reads as empty rather than throwing", async () => { + await withTempChannel(async (root) => { + assert.deepEqual(await listPostShards(root), []); + assert.deepEqual(await readAllPosts(root), []); + assert.equal(await countPosts(root), 0); + assert.deepEqual([...(await readSeenPostIds(root))], []); + }); +}); + +test("a truncated JSONL line does not poison the rest of its shard", async () => { + await withTempChannel(async (root) => { + await writePosts(root, [makePost("good1", "2026-05-01T00:00:00.000Z")]); + // Simulate an append interrupted mid-line, then a later clean append. + const { appendFile } = await import("node:fs/promises"); + await appendFile( + path.join(root, "posts", "2026-05.jsonl"), + '{"id":"trunc","channelSl\n', + "utf8", + ); + await writePosts(root, [makePost("good2", "2026-05-02T00:00:00.000Z")]); + const posts = await readPostShard(root, "2026-05"); + assert.deepEqual(posts.map((p) => p.id).sort(), ["good1", "good2"]); + }); +}); diff --git a/common/lib/posts-server.ts b/common/lib/posts-server.ts @@ -0,0 +1,254 @@ +// Server-side on-disk store for the social-post corpus. +// +// Layout, per channel: +// transcripts/channels/<slug>/posts/YYYY-MM.jsonl — one post per line +// transcripts/channels/<slug>/posts-archive — "<platform> <id>" lines +// +// Month-sharded JSONL rather than a directory per post: an account with 50k +// posts would otherwise create 50k directories. The archive file mirrors how +// yt-dlp's `--download-archive` drives incremental sync today (sync() in +// common/ytdlp/runYtdlp.ts stops at the first already-archived entry; the post +// fetchers stop the same way), and uses the same "<extractor> <id>" line format +// so common/lib/archive.ts parses it unchanged. + +import path from "node:path"; +import { readdir, mkdir, readFile, rename, writeFile } from "node:fs/promises"; +import { appendFile } from "node:fs/promises"; +import { readArchive } from "./archive"; +import { + comparePostsNewestFirst, + monthShardFromCreatedAt, + parsePost, + type Post, + type PostPlatform, +} from "./posts"; + +export const POSTS_DIRNAME = "posts"; +export const POSTS_ARCHIVE_FILENAME = "posts-archive"; + +export function channelPostsDir(channelRoot: string): string { + return path.join(channelRoot, POSTS_DIRNAME); +} + +export function channelPostsArchivePath(channelRoot: string): string { + return path.join(channelRoot, POSTS_ARCHIVE_FILENAME); +} + +function shardPath(channelRoot: string, shard: string): string { + return path.join(channelPostsDir(channelRoot), `${shard}.jsonl`); +} + +// The set of post ids already archived for this channel, used by every fetcher +// to stop paging once it reaches previously-synced content. +export async function readSeenPostIds( + channelRoot: string, +): Promise<Set<string>> { + const archive = await readArchive(channelPostsArchivePath(channelRoot)); + return archive.ids; +} + +// Append ids to the archive. Idempotent at the caller's discretion — writePosts +// filters against the seen set before calling, so duplicate lines only appear +// if two fetches race, and readArchive dedupes on read anyway. +async function appendArchiveIds( + channelRoot: string, + platform: PostPlatform, + ids: ReadonlyArray<string>, +): Promise<void> { + if (ids.length === 0) return; + const lines = ids.map((id) => `${platform} ${id}\n`).join(""); + await appendFile(channelPostsArchivePath(channelRoot), lines, "utf8"); +} + +export type WritePostsResult = { + written: number; + skipped: number; // already present in the archive + shards: string[]; // YYYY-MM shards touched +}; + +// Append posts to their month shards and record their ids in the archive. +// Posts whose id is already archived are skipped, so a re-run over overlapping +// pages is a no-op rather than a source of duplicates. +export async function writePosts( + channelRoot: string, + posts: ReadonlyArray<Post>, +): Promise<WritePostsResult> { + if (posts.length === 0) return { written: 0, skipped: 0, shards: [] }; + await mkdir(channelPostsDir(channelRoot), { recursive: true }); + const seen = await readSeenPostIds(channelRoot); + + // Group by shard so each file is opened once, and dedupe within the batch + // itself (a fetcher may legitimately return the same post twice across a + // cursor boundary). + const byShard = new Map<string, Post[]>(); + const accepted: Post[] = []; + let skipped = 0; + for (const post of posts) { + if (seen.has(post.id)) { + skipped++; + continue; + } + seen.add(post.id); + accepted.push(post); + const shard = monthShardFromCreatedAt(post.createdAt); + const bucket = byShard.get(shard); + if (bucket) bucket.push(post); + else byShard.set(shard, [post]); + } + if (accepted.length === 0) { + return { written: 0, skipped, shards: [] }; + } + + for (const [shard, shardPosts] of byShard) { + const body = shardPosts.map((p) => JSON.stringify(p)).join("\n") + "\n"; + await appendFile(shardPath(channelRoot, shard), body, "utf8"); + } + + // Archive AFTER the shard writes: a crash between the two re-fetches the + // posts (harmless — the next writePosts dedupes them), whereas archiving + // first would lose them permanently. + const byPlatform = new Map<PostPlatform, string[]>(); + for (const post of accepted) { + const bucket = byPlatform.get(post.platform); + if (bucket) bucket.push(post.id); + else byPlatform.set(post.platform, [post.id]); + } + for (const [platform, ids] of byPlatform) { + await appendArchiveIds(channelRoot, platform, ids); + } + + return { + written: accepted.length, + skipped, + shards: [...byShard.keys()].sort(), + }; +} + +// List the channel's month shards, oldest first. +export async function listPostShards( + channelRoot: string, +): Promise<string[]> { + let entries: string[]; + try { + entries = await readdir(channelPostsDir(channelRoot)); + } catch { + return []; + } + return entries + .filter((e) => e.endsWith(".jsonl")) + .map((e) => e.slice(0, -".jsonl".length)) + .sort(); +} + +// Read one month shard. Malformed lines are skipped rather than failing the +// whole shard — an append that was interrupted mid-line must not make every +// earlier post in that month unreadable. +export async function readPostShard( + channelRoot: string, + shard: string, +): Promise<Post[]> { + let text: string; + try { + text = await readFile(shardPath(channelRoot, shard), "utf8"); + } catch { + return []; + } + const posts: Post[] = []; + for (const line of text.split("\n")) { + const trimmed = line.trim(); + if (!trimmed) continue; + let raw: unknown; + try { + raw = JSON.parse(trimmed); + } catch { + continue; + } + const post = parsePost(raw); + if (post) posts.push(post); + } + return posts; +} + +// Every archived post for a channel, newest first. Deduped by id: an +// interrupted archive write can leave the same post in two shards. +export async function readAllPosts(channelRoot: string): Promise<Post[]> { + const shards = await listPostShards(channelRoot); + const byId = new Map<string, Post>(); + for (const shard of shards) { + for (const post of await readPostShard(channelRoot, shard)) { + byId.set(post.id, post); + } + } + return [...byId.values()].sort(comparePostsNewestFirst); +} + +// Cheap count for the dashboard/snapshot without materializing every post. +export async function countPosts(channelRoot: string): Promise<number> { + const archive = await readArchive(channelPostsArchivePath(channelRoot)); + return archive.ids.size; +} + +// The newest archived post's createdAt, used as the incremental-fetch +// watermark. Only the newest shard is read. +export async function latestPostCreatedAt( + channelRoot: string, +): Promise<string | null> { + const shards = await listPostShards(channelRoot); + for (let i = shards.length - 1; i >= 0; i--) { + const posts = await readPostShard(channelRoot, shards[i]); + if (posts.length === 0) continue; + let newest = posts[0].createdAt; + for (const post of posts) { + if (post.createdAt > newest) newest = post.createdAt; + } + return newest; + } + return null; +} + +// --------------------------------------------------------------------------- +// Fetch state sidecar +// --------------------------------------------------------------------------- + +// Small per-channel record of the last fetch attempt. Drives the social +// channel's snapshot buckets (fetchFailed / needsCookies) without re-deriving +// them from job logs. +export type PostFetchState = { + lastFetchedAt?: string; + lastError?: string; + // Set when the fetcher stopped because credentials are missing/expired — + // the post-corpus analogue of the video pipeline's needs_auth outcome. + needsCookies?: boolean; + lastFetchedCount?: number; + cursor?: string; +}; + +export const POST_FETCH_STATE_FILENAME = "posts-state.json"; + +export function postFetchStatePath(channelRoot: string): string { + return path.join(channelRoot, POST_FETCH_STATE_FILENAME); +} + +export async function readPostFetchState( + channelRoot: string, +): Promise<PostFetchState | null> { + try { + const raw = await readFile(postFetchStatePath(channelRoot), "utf8"); + const parsed = JSON.parse(raw) as unknown; + if (!parsed || typeof parsed !== "object") return null; + return parsed as PostFetchState; + } catch { + return null; + } +} + +export async function writePostFetchState( + channelRoot: string, + state: PostFetchState, +): Promise<void> { + const file = postFetchStatePath(channelRoot); + await mkdir(path.dirname(file), { recursive: true }); + const tmp = `${file}.tmp-${process.pid}`; + await writeFile(tmp, JSON.stringify(state, null, 2) + "\n"); + await rename(tmp, file); +} diff --git a/common/lib/posts.ts b/common/lib/posts.ts @@ -0,0 +1,274 @@ +// The social-post corpus: a parallel content layer beside video transcripts. +// +// Posts are live-chat-shaped, not video-shaped. Like the `subs` layer they get +// their own page tree (/posts/<channelSlug>/{manifest,page-NNNN}.json), their +// own LayerScope ("posts") and their own manifest version — and they compose +// into the SAME boolean query tree as transcripts, so one search covers both. +// +// Client-safe: no node imports (same convention as cookiePolicy.ts / +// availability.ts). The server-side store lives in posts-server.ts. + +export type PostPlatform = "twitter" | "bluesky"; + +export const POST_PLATFORM_VALUES: ReadonlyArray<PostPlatform> = [ + "twitter", + "bluesky", +]; + +export function isPostPlatform(v: unknown): v is PostPlatform { + return v === "twitter" || v === "bluesky"; +} + +// A pointer to another post, which may or may not itself be archived. Used for +// reply/quote/repost edges so thread structure survives even when the +// referenced post is outside the archived account. +export type PostRef = { + platform: PostPlatform; + id: string; + url?: string; + author?: string; +}; + +export type PostEngagement = { + likes?: number; + reposts?: number; + replies?: number; + quotes?: number; +}; + +export type Post = { + // Native post id: tweet id, or the atproto record key (rkey). + id: string; + // `${channelSlug}/${id}` — the same identity convention videos use, so a post + // can be addressed by one opaque string across search/AI/viewer layers. + slug: string; + channelSlug: string; + author: string; // handle, e.g. "example.bsky.social" + authorName?: string; // display name + // ISO-8601, ms precision where the source provides it. THE sort key: it sorts + // lexicographically and so drops straight into an LMDB key tuple, and it + // fixes the intra-day ordering problem a YYYYMMDD key has. + createdAt: string; + // YYYYMMDD derived from createdAt. Retained (not derived at read time) so + // every existing date filter — the fdf/fdt share params, the lexicographic + // bounds in SearchSessionContext.passesFilter and mcp/src/search.ts + // passesFilters — keeps working against posts with zero changes. + uploadDate: string; + text: string; + url: string; // canonical permalink + platform: PostPlatform; + lang?: string; + threadId?: string; // root post id, for grouping + replyTo?: PostRef; + quoted?: PostRef; + repostOf?: PostRef; + isReply: boolean; + isRepost: boolean; + links: string[]; // expanded outbound urls + mediaCount?: number; // counted, not archived (v1) + engagement?: PostEngagement; +}; + +export type PostPage = Post[]; + +export type ChannelPostsManifest = { + version: number; + channelSlug: string; + pageCount: number; + maxPageBytes: number; + generatedAt: string; + // postId -> page index. Mirrors ChannelTranscriptsManifest.slugToPage so a + // single post can be fetched without scanning every page. + slugToPage: Record<string, number>; +}; + +export const POSTS_MANIFEST_VERSION = 1; + +// Site-level index of which channels carry posts — the posts analogue of +// SubsManifest. Lets the export client render the corpus toggle and the channel +// pickers without probing every channel's per-channel manifest. +export type PostsChannelEntry = { + name: string; + slug: string; + postCount: number; + platform: PostPlatform; + groupId?: string; +}; + +export type PostsManifest = { + version: number; + channels: PostsChannelEntry[]; + totalCount: number; + generatedAt: string; + siteId?: string; +}; + +export const SITE_POSTS_MANIFEST_VERSION = 1; + +export function postsPageFileName(index: number): string { + return `page-${String(index).padStart(4, "0")}.json`; +} + +// --------------------------------------------------------------------------- +// Derivations +// --------------------------------------------------------------------------- + +// YYYYMMDD in UTC from an ISO-8601 timestamp. Returns "" for an unparseable +// input so a malformed post degrades to "no date" rather than poisoning the +// lexicographic date filters with garbage. +export function uploadDateFromCreatedAt(createdAt: string): string { + const ms = Date.parse(createdAt); + if (!Number.isFinite(ms)) return ""; + const d = new Date(ms); + const y = d.getUTCFullYear(); + const m = d.getUTCMonth() + 1; + const day = d.getUTCDate(); + return `${String(y).padStart(4, "0")}${String(m).padStart(2, "0")}${String(day).padStart(2, "0")}`; +} + +// YYYY-MM shard key from an ISO-8601 timestamp — the on-disk JSONL month shard. +// Falls back to "unknown" so a post with a bad timestamp is still stored. +export function monthShardFromCreatedAt(createdAt: string): string { + const ms = Date.parse(createdAt); + if (!Number.isFinite(ms)) return "unknown"; + const d = new Date(ms); + return `${String(d.getUTCFullYear()).padStart(4, "0")}-${String( + d.getUTCMonth() + 1, + ).padStart(2, "0")}`; +} + +export function postSlug(channelSlug: string, id: string): string { + return `${channelSlug}/${id}`; +} + +// Split a post slug back into its parts. The id may itself contain "/" for no +// current platform, but splitting on the FIRST separator keeps that safe. +export function parsePostSlug( + slug: string, +): { channelSlug: string; id: string } | null { + const idx = slug.indexOf("/"); + if (idx <= 0 || idx === slug.length - 1) return null; + return { channelSlug: slug.slice(0, idx), id: slug.slice(idx + 1) }; +} + +// Canonical permalink for a post. Bluesky needs the handle (its URLs are +// /profile/<handle>/post/<rkey>); X only needs the numeric id but includes the +// handle for readability. +export function postPermalink( + platform: PostPlatform, + author: string, + id: string, +): string { + if (platform === "bluesky") { + return `https://bsky.app/profile/${author}/post/${id}`; + } + return `https://x.com/${author || "i"}/status/${id}`; +} + +// Runtime validation for a post read back off disk / off the wire. Mirrors +// parseChannelConfig's strict-allowlist stance: unknown keys are dropped, and a +// record missing any required field is rejected outright rather than repaired. +export function parsePost(raw: unknown): Post | null { + if (!raw || typeof raw !== "object") return null; + const r = raw as Record<string, unknown>; + if (typeof r.id !== "string" || !r.id) return null; + if (typeof r.channelSlug !== "string" || !r.channelSlug) return null; + if (typeof r.text !== "string") return null; + if (typeof r.createdAt !== "string" || !r.createdAt) return null; + if (!isPostPlatform(r.platform)) return null; + + const author = typeof r.author === "string" ? r.author : ""; + const post: Post = { + id: r.id, + slug: + typeof r.slug === "string" && r.slug + ? r.slug + : postSlug(r.channelSlug, r.id), + channelSlug: r.channelSlug, + author, + createdAt: r.createdAt, + uploadDate: + typeof r.uploadDate === "string" && r.uploadDate + ? r.uploadDate + : uploadDateFromCreatedAt(r.createdAt), + text: r.text, + url: + typeof r.url === "string" && r.url + ? r.url + : postPermalink(r.platform, author, r.id), + platform: r.platform, + isReply: r.isReply === true, + isRepost: r.isRepost === true, + links: + Array.isArray(r.links) && r.links.every((l) => typeof l === "string") + ? (r.links as string[]) + : [], + }; + if (typeof r.authorName === "string") post.authorName = r.authorName; + if (typeof r.lang === "string") post.lang = r.lang; + if (typeof r.threadId === "string") post.threadId = r.threadId; + const replyTo = parsePostRef(r.replyTo); + if (replyTo) post.replyTo = replyTo; + const quoted = parsePostRef(r.quoted); + if (quoted) post.quoted = quoted; + const repostOf = parsePostRef(r.repostOf); + if (repostOf) post.repostOf = repostOf; + if (typeof r.mediaCount === "number" && Number.isFinite(r.mediaCount)) { + post.mediaCount = Math.max(0, Math.floor(r.mediaCount)); + } + const engagement = parseEngagement(r.engagement); + if (engagement) post.engagement = engagement; + return post; +} + +function parsePostRef(raw: unknown): PostRef | null { + if (!raw || typeof raw !== "object") return null; + const r = raw as Record<string, unknown>; + if (typeof r.id !== "string" || !r.id) return null; + if (!isPostPlatform(r.platform)) return null; + const ref: PostRef = { platform: r.platform, id: r.id }; + if (typeof r.url === "string") ref.url = r.url; + if (typeof r.author === "string") ref.author = r.author; + return ref; +} + +function parseEngagement(raw: unknown): PostEngagement | null { + if (!raw || typeof raw !== "object") return null; + const r = raw as Record<string, unknown>; + const out: PostEngagement = {}; + let any = false; + for (const key of ["likes", "reposts", "replies", "quotes"] as const) { + const v = r[key]; + if (typeof v === "number" && Number.isFinite(v)) { + out[key] = Math.max(0, Math.floor(v)); + any = true; + } + } + return any ? out : null; +} + +// Newest-first ordering, the order both the page tree and the UI present. +// createdAt is ISO-8601 so a plain string compare is a chronological compare; +// the id tiebreak keeps the sort stable for same-instant posts. +export function comparePostsNewestFirst(a: Post, b: Post): number { + if (a.createdAt !== b.createdAt) return a.createdAt < b.createdAt ? 1 : -1; + return a.id < b.id ? 1 : a.id > b.id ? -1 : 0; +} + +// Group a flat post list into threads keyed by threadId (falling back to the +// post's own id for a root/standalone post). Used by the viewer's thread +// context and the MCP get_thread tool. +export function groupIntoThreads(posts: ReadonlyArray<Post>): Map<string, Post[]> { + const threads = new Map<string, Post[]>(); + for (const post of posts) { + const key = post.threadId || post.id; + const bucket = threads.get(key); + if (bucket) bucket.push(post); + else threads.set(key, [post]); + } + for (const bucket of threads.values()) { + // Threads read oldest-first — the opposite of the feed ordering. + bucket.sort((a, b) => -comparePostsNewestFirst(a, b)); + } + return threads; +} diff --git a/common/lib/searchEval.ts b/common/lib/searchEval.ts @@ -81,6 +81,13 @@ type EvalCtx = { // their input scope by this set before running, so we don't waste fetches // on videos that don't have a chat track at all. chatScopeSlugs: ReadonlySet<string> | null; + // Every post slug in the corpus. Posts live in a DISJOINT slug namespace from + // videos (`<channelSlug>/<postId>` vs a video id), so the two must be kept + // apart: a posts leaf narrows to this set, and every video-shaped leaf + // subtracts it. Without that, an AND against the global scope would + // intersect the two namespaces to nothing, and video leaves would waste a + // fetch per post. Null = no posts corpus on this site. + postScopeSlugs: ReadonlySet<string> | null; // Per-leaf live result state. Reused for cache-hit instant emission and // for incremental re-evaluation on each leaf progress. leafResults: Map<string, { slugs: Set<string>; hits: Map<string, LayerHit[]> }>; @@ -118,6 +125,7 @@ export function runQueryTree(opts: { globalScope: string[]; summaries: DisplaySummary[]; chatScopeSlugs?: ReadonlySet<string> | null; + postScopeSlugs?: ReadonlySet<string> | null; initialHitLimit: number; concurrency: number; flushIntervalMs: number; @@ -128,6 +136,7 @@ export function runQueryTree(opts: { globalScope, summaries, chatScopeSlugs = null, + postScopeSlugs = null, initialHitLimit, concurrency, flushIntervalMs, @@ -145,6 +154,7 @@ export function runQueryTree(opts: { flushIntervalMs, metadataIndex, chatScopeSlugs, + postScopeSlugs, leafResults: new Map(), leafStates: new Map(), controllers: new Map(), @@ -288,6 +298,15 @@ async function runLeaf( if (leaf.scope === "chat" && ctx.chatScopeSlugs) { effectiveScope = intersect(parentScope, ctx.chatScopeSlugs); } + // Keep the two corpora out of each other's way (see ctx.postScopeSlugs). + if (ctx.postScopeSlugs) { + effectiveScope = + leaf.scope === "posts" + ? intersect(effectiveScope, ctx.postScopeSlugs) + : diff(effectiveScope, ctx.postScopeSlugs); + } else if (leaf.scope === "posts") { + effectiveScope = new Set(); + } // Fully-network leaves (transcripts / chat). Try the layer cache first. const scopeArr = Array.from(effectiveScope); diff --git a/common/lib/searchQuery.ts b/common/lib/searchQuery.ts @@ -14,6 +14,10 @@ import type { SearchMode } from "../components/urlState"; export type LayerScope = | "transcripts" | "chat" + // The social-post corpus (common/lib/posts.ts). A parallel content layer, so + // it composes into the SAME boolean tree as transcripts — which is what makes + // `(transcripts:"foo" OR posts:"foo")` a single unified search. + | "posts" | "metadata" | "description" | "tags"; @@ -187,9 +191,13 @@ function deserialize(raw: unknown): QueryNode | null { const r = raw as Record<string, unknown>; if (r.k === "l") { const scope = r.s; + // Explicit whitelist: an unrecognized scope drops the leaf rather than + // silently evaluating it as something else. New scopes MUST be added here + // or a `qt=` URL carrying them loses the leaf. if ( scope !== "transcripts" && scope !== "chat" && + scope !== "posts" && scope !== "metadata" && scope !== "description" && scope !== "tags" @@ -254,7 +262,8 @@ export function rootFromLegacy( children: [ newLeaf({ query: q, - scope: mode === "subs" ? "chat" : "transcripts", + scope: + mode === "subs" ? "chat" : mode === "posts" ? "posts" : "transcripts", useRegex, contributeHits: true, }), diff --git a/common/lib/site.ts b/common/lib/site.ts @@ -146,6 +146,10 @@ export function siteSubsDir(paths: Paths, siteId: string): string { return path.join(siteIndexDir(paths, siteId), "subs"); } +export function sitePostsDir(paths: Paths, siteId: string): string { + return path.join(siteIndexDir(paths, siteId), "posts"); +} + export function siteStatsDir(paths: Paths, siteId: string): string { return path.join(siteIndexDir(paths, siteId), "stats"); } diff --git a/common/social/__fixtures__/bluesky-author-feed.json b/common/social/__fixtures__/bluesky-author-feed.json @@ -0,0 +1,813 @@ +{ + "feed": [ + { + "post": { + "uri": "at://did:plc:6kos45lixtga3pdwuncvh32x/app.bsky.feed.post/3mqc36slinc2m", + "cid": "bafyreifp32ys6w6ayfu3f2ayw4hdpsyaa7kyjz6gk3kfeblfwr2kbnao7e", + "author": { + "did": "did:plc:6kos45lixtga3pdwuncvh32x", + "handle": "paretooptimizer.bsky.social", + "displayName": "Pareto Optimizer", + "avatar": "https://cdn.bsky.app/img/avatar/plain/did:plc:6kos45lixtga3pdwuncvh32x/bafkreifvhi2e3intywaulvunc5wnva4fdrwhu7qjhwlnak77jbuzqiz7zq", + "associated": { + "chat": { + "allowIncoming": "none", + "allowGroupInvites": "none" + }, + "activitySubscription": { + "allowSubscriptions": "followers" + } + }, + "labels": [], + "createdAt": "2023-06-29T16:47:43.607Z" + }, + "record": { + "$type": "app.bsky.feed.post", + "createdAt": "2026-07-10T11:46:11.858Z", + "embed": { + "$type": "app.bsky.embed.record", + "record": { + "cid": "bafyreigvw7va73oqr3vkv74dch3uw7xubmyxddxf7f2houk5wrs6csrs64", + "uri": "at://did:plc:yfzqc42lkdurwcpslq7yp3ud/app.bsky.feed.post/3mqbs2zzdi22y" + } + }, + "langs": [ + "en" + ], + "text": "I know this guy has a Bluesky account." + }, + "embed": { + "record": { + "uri": "at://did:plc:yfzqc42lkdurwcpslq7yp3ud/app.bsky.feed.post/3mqbs2zzdi22y", + "cid": "bafyreigvw7va73oqr3vkv74dch3uw7xubmyxddxf7f2houk5wrs6csrs64", + "author": { + "did": "did:plc:yfzqc42lkdurwcpslq7yp3ud", + "handle": "dexerto.bsky.social", + "displayName": "Dexerto", + "avatar": "https://cdn.bsky.app/img/avatar/plain/did:plc:yfzqc42lkdurwcpslq7yp3ud/bafkreialdpbyzldcsbk2knioy5uoz35oomx7pgacfzgh34kizoxmoj2ypm", + "associated": { + "chat": { + "allowIncoming": "none", + "allowGroupInvites": "none" + }, + "activitySubscription": { + "allowSubscriptions": "followers" + } + }, + "labels": [], + "createdAt": "2023-04-27T17:42:28.330Z" + }, + "value": { + "$type": "app.bsky.feed.post", + "createdAt": "2026-07-10T09:03:00Z", + "embed": { + "$type": "app.bsky.embed.images", + "images": [ + { + "alt": "", + "aspectRatio": { + "height": 1600, + "width": 1400 + }, + "image": { + "$type": "blob", + "ref": { + "$link": "bafkreig4jkqjmn7mkivmhrk3rusqwbik4xgvkwxsux6wr7tbspz36nw4dy" + }, + "mimeType": "image/jpeg", + "size": 148913 + } + }, + { + "alt": "", + "aspectRatio": { + "height": 1600, + "width": 1400 + }, + "image": { + "$type": "blob", + "ref": { + "$link": "bafkreic6tc7c2kmawwk7yr5gtan2csfr3izt3qr5ffuqyjci33kfedk35i" + }, + "mimeType": "image/jpeg", + "size": 173646 + } + } + ] + }, + "text": "A Norway fan has gone viral for not taking part in their Row celebration after matches and won't do it if they win the World Cup\n\n\"It’s factually wrong; they didn’t row, they sailed over the Atlantic”" + }, + "labels": [], + "likeCount": 5056, + "replyCount": 256, + "repostCount": 612, + "quoteCount": 915, + "indexedAt": "2026-07-10T09:03:05.566Z", + "embeds": [ + { + "images": [ + { + "thumb": "https://cdn.bsky.app/img/feed_thumbnail/plain/did:plc:yfzqc42lkdurwcpslq7yp3ud/bafkreig4jkqjmn7mkivmhrk3rusqwbik4xgvkwxsux6wr7tbspz36nw4dy", + "fullsize": "https://cdn.bsky.app/img/feed_fullsize/plain/did:plc:yfzqc42lkdurwcpslq7yp3ud/bafkreig4jkqjmn7mkivmhrk3rusqwbik4xgvkwxsux6wr7tbspz36nw4dy", + "alt": "", + "aspectRatio": { + "height": 1600, + "width": 1400 + } + }, + { + "thumb": "https://cdn.bsky.app/img/feed_thumbnail/plain/did:plc:yfzqc42lkdurwcpslq7yp3ud/bafkreic6tc7c2kmawwk7yr5gtan2csfr3izt3qr5ffuqyjci33kfedk35i", + "fullsize": "https://cdn.bsky.app/img/feed_fullsize/plain/did:plc:yfzqc42lkdurwcpslq7yp3ud/bafkreic6tc7c2kmawwk7yr5gtan2csfr3izt3qr5ffuqyjci33kfedk35i", + "alt": "", + "aspectRatio": { + "height": 1600, + "width": 1400 + } + } + ], + "$type": "app.bsky.embed.images#view" + } + ], + "$type": "app.bsky.embed.record#viewRecord" + }, + "$type": "app.bsky.embed.record#view" + }, + "bookmarkCount": 116, + "replyCount": 152, + "repostCount": 545, + "likeCount": 6580, + "quoteCount": 14, + "indexedAt": "2026-07-10T11:46:12.181Z", + "labels": [] + }, + "reason": { + "by": { + "did": "did:plc:z72i7hdynmk6r22z27h6tvur", + "handle": "bsky.app", + "displayName": "Bluesky", + "avatar": "https://cdn.bsky.app/img/avatar/plain/did:plc:z72i7hdynmk6r22z27h6tvur/bafkreihwihm6kpd6zuwhhlro75p5qks5qtrcu55jp3gddbfjsieiv7wuka", + "associated": { + "chat": { + "allowIncoming": "none" + }, + "activitySubscription": { + "allowSubscriptions": "followers" + } + }, + "labels": [], + "createdAt": "2023-04-12T04:53:57.057Z", + "verification": { + "verifications": [], + "verifiedStatus": "none", + "trustedVerifierStatus": "valid" + } + }, + "uri": "at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.feed.repost/3mqcr3w3xru2d", + "cid": "bafyreieyetktyme4g7moqlzqv4t7kfff7xs7ykl6qbi7dmxb5ijdjffjxe", + "indexedAt": "2026-07-10T18:18:17.273Z", + "$type": "app.bsky.feed.defs#reasonRepost" + } + }, + { + "post": { + "uri": "at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.feed.post/3mqcp5qjdfs26", + "cid": "bafyreig6kgwgaixxenkfv62gkz43fgdyoe2ddbgqbfr2edkqv5ya5kginu", + "author": { + "did": "did:plc:z72i7hdynmk6r22z27h6tvur", + "handle": "bsky.app", + "displayName": "Bluesky", + "avatar": "https://cdn.bsky.app/img/avatar/plain/did:plc:z72i7hdynmk6r22z27h6tvur/bafkreihwihm6kpd6zuwhhlro75p5qks5qtrcu55jp3gddbfjsieiv7wuka", + "associated": { + "chat": { + "allowIncoming": "none" + }, + "activitySubscription": { + "allowSubscriptions": "followers" + } + }, + "labels": [], + "createdAt": "2023-04-12T04:53:57.057Z", + "verification": { + "verifications": [], + "verifiedStatus": "none", + "trustedVerifierStatus": "valid" + } + }, + "record": { + "$type": "app.bsky.feed.post", + "createdAt": "2026-07-10T17:43:30.972Z", + "embed": { + "$type": "app.bsky.embed.record", + "record": { + "cid": "bafyreigov2ns6elnyiksqufzr6ctfg4i4eiodharfupyoc5sq23tvyew4a", + "uri": "at://did:plc:cwf4mmm7mpzistinx3ox2zhj/app.bsky.feed.post/3mqcp425edfgx" + } + }, + "facets": [ + { + "$type": "app.bsky.richtext.facet", + "features": [ + { + "$type": "app.bsky.richtext.facet#mention", + "did": "did:plc:cwf4mmm7mpzistinx3ox2zhj" + } + ], + "index": { + "byteEnd": 31, + "byteStart": 16 + } + } + ], + "langs": [ + "en" + ], + "text": "Personnel news: @toni.bsky.team is dropping the “interim” from his title and will serve as our permanent CEO!" + }, + "embed": { + "record": { + "uri": "at://did:plc:cwf4mmm7mpzistinx3ox2zhj/app.bsky.feed.post/3mqcp425edfgx", + "cid": "bafyreigov2ns6elnyiksqufzr6ctfg4i4eiodharfupyoc5sq23tvyew4a", + "author": { + "did": "did:plc:cwf4mmm7mpzistinx3ox2zhj", + "handle": "toni.bsky.team", + "displayName": "Toni Schneider", + "avatar": "https://cdn.bsky.app/img/avatar/plain/did:plc:cwf4mmm7mpzistinx3ox2zhj/bafkreicixoq23lsimr6cabhtss53jimjzouzkstif4xdjmvt4rcr5bjsku", + "associated": { + "chat": { + "allowIncoming": "none" + }, + "activitySubscription": { + "allowSubscriptions": "followers" + } + }, + "labels": [], + "createdAt": "2023-11-28T19:15:32.965Z", + "verification": { + "verifications": [ + { + "issuer": "did:plc:z72i7hdynmk6r22z27h6tvur", + "issuerDisplayName": "Bluesky", + "issuerHandle": "bsky.app", + "uri": "at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.graph.verification/3mgngmgwfk62n", + "isValid": true, + "createdAt": "2026-03-09T17:58:00.876Z" + } + ], + "verifiedStatus": "valid", + "trustedVerifierStatus": "none" + } + }, + "value": { + "$type": "app.bsky.feed.post", + "createdAt": "2026-07-10T17:42:33.000Z", + "embed": { + "$type": "app.bsky.embed.external", + "external": { + "associatedRefs": [ + { + "$type": "com.atproto.repo.strongRef", + "cid": "bafyreihczgmq2lnsyjkj7rtpwd3stjbfqa2yxpncbktehkxzyodxklgwra", + "uri": "at://did:plc:cwf4mmm7mpzistinx3ox2zhj/site.standard.publication/3mhgursihlhgq" + }, + { + "$type": "com.atproto.repo.strongRef", + "cid": "bafyreic2nnco7omjpz3yf343rmwbcbepybtxhv3how5jmz25y2lk7ksksu", + "uri": "at://did:plc:cwf4mmm7mpzistinx3ox2zhj/site.standard.document/3mqcp424z5wgx" + } + ], + "description": "I'm four months into my interim CEO role at Bluesky, and it's time for an update. Most importantly, as of today, the interim part of the title is gone. I'm loving the mission and the job, and I'm all in as Bluesky’s official CEO. This job has been energizing from day one, and the Bluesky...", + "title": "Staying in the game", + "uri": "https://toni.org/2026/07/10/staying-in-the-game/" + } + }, + "facets": [ + { + "features": [ + { + "$type": "app.bsky.richtext.facet#link", + "uri": "https://toni.org/2026/07/10/staying-in-the-game/" + } + ], + "index": { + "byteEnd": 229, + "byteStart": 181 + } + } + ], + "langs": [ + "en" + ], + "tags": [ + "Bluesky" + ], + "text": "Staying in the game\n\nI'm four months into my interim CEO role at Bluesky, and it's time for an update. Most importantly, as of today, the interim part of the title is gone. I'm...\n\nhttps://toni.org/2026/07/10/staying-in-the-game/" + }, + "labels": [], + "likeCount": 722, + "replyCount": 102, + "repostCount": 95, + "quoteCount": 37, + "indexedAt": "2026-07-10T17:42:34.171Z", + "embeds": [ + { + "external": { + "uri": "https://toni.org/2026/07/10/staying-in-the-game/", + "title": "Staying in the game", + "description": "I'm four months into my interim CEO role at Bluesky, and it's time for an update. Most importantly, as of today, the interim part of the title is gone. I'm loving the mission and the job, and I'm all in as Bluesky’s official CEO. This job has been energizing from day one, and the Bluesky...", + "createdAt": "2026-07-10T17:42:33.000Z", + "readingTime": 3, + "source": { + "uri": "https://toni.org", + "icon": "https://cdn.bsky.app/img/avatar/plain/did:plc:cwf4mmm7mpzistinx3ox2zhj/bafkreidkmfyy5likpn4kvdb775rrviismf7lptkxhcbnkhimkzdgnazosq", + "title": "Toni.org", + "description": "Toni Schneider's blog", + "theme": { + "backgroundRGB": { + "r": 249, + "g": 249, + "b": 249, + "$type": "app.bsky.embed.external#colorRGB" + }, + "foregroundRGB": { + "r": 99, + "g": 99, + "b": 99, + "$type": "app.bsky.embed.external#colorRGB" + }, + "accentRGB": { + "r": 17, + "g": 17, + "b": 17, + "$type": "app.bsky.embed.external#colorRGB" + }, + "accentForegroundRGB": { + "r": 255, + "g": 255, + "b": 255, + "$type": "app.bsky.embed.external#colorRGB" + } + }, + "$type": "app.bsky.embed.external#viewExternalSource" + }, + "associatedProfiles": [ + { + "did": "did:plc:cwf4mmm7mpzistinx3ox2zhj", + "handle": "toni.bsky.team", + "displayName": "Toni Schneider", + "avatar": "https://cdn.bsky.app/img/avatar/plain/did:plc:cwf4mmm7mpzistinx3ox2zhj/bafkreicixoq23lsimr6cabhtss53jimjzouzkstif4xdjmvt4rcr5bjsku", + "associated": { + "chat": { + "allowIncoming": "none" + }, + "activitySubscription": { + "allowSubscriptions": "followers" + } + }, + "labels": [], + "createdAt": "2023-11-28T19:15:32.965Z", + "verification": { + "verifications": [ + { + "issuer": "did:plc:z72i7hdynmk6r22z27h6tvur", + "issuerDisplayName": "Bluesky", + "issuerHandle": "bsky.app", + "uri": "at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.graph.verification/3mgngmgwfk62n", + "isValid": true, + "createdAt": "2026-03-09T17:58:00.876Z" + } + ], + "verifiedStatus": "valid", + "trustedVerifierStatus": "none" + } + } + ], + "associatedRefs": [ + { + "$type": "com.atproto.repo.strongRef", + "cid": "bafyreihczgmq2lnsyjkj7rtpwd3stjbfqa2yxpncbktehkxzyodxklgwra", + "uri": "at://did:plc:cwf4mmm7mpzistinx3ox2zhj/site.standard.publication/3mhgursihlhgq" + }, + { + "$type": "com.atproto.repo.strongRef", + "cid": "bafyreic2nnco7omjpz3yf343rmwbcbepybtxhv3how5jmz25y2lk7ksksu", + "uri": "at://did:plc:cwf4mmm7mpzistinx3ox2zhj/site.standard.document/3mqcp424z5wgx" + } + ] + }, + "$type": "app.bsky.embed.external#view" + } + ], + "$type": "app.bsky.embed.record#viewRecord" + }, + "$type": "app.bsky.embed.record#view" + }, + "bookmarkCount": 43, + "replyCount": 99, + "repostCount": 182, + "likeCount": 1680, + "quoteCount": 31, + "indexedAt": "2026-07-10T17:43:31.572Z", + "labels": [] + } + }, + { + "post": { + "uri": "at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.feed.post/3mqafridzgk2e", + "cid": "bafyreibowb4zpdgv74mzhpn3w433q4affyql5jd6c43tluinucj4q2jdzy", + "author": { + "did": "did:plc:z72i7hdynmk6r22z27h6tvur", + "handle": "bsky.app", + "displayName": "Bluesky", + "avatar": "https://cdn.bsky.app/img/avatar/plain/did:plc:z72i7hdynmk6r22z27h6tvur/bafkreihwihm6kpd6zuwhhlro75p5qks5qtrcu55jp3gddbfjsieiv7wuka", + "associated": { + "chat": { + "allowIncoming": "none" + }, + "activitySubscription": { + "allowSubscriptions": "followers" + } + }, + "labels": [], + "createdAt": "2023-04-12T04:53:57.057Z", + "verification": { + "verifications": [], + "verifiedStatus": "none", + "trustedVerifierStatus": "valid" + } + }, + "record": { + "$type": "app.bsky.feed.post", + "createdAt": "2026-07-09T19:50:16.599Z", + "embed": { + "$type": "app.bsky.embed.images", + "images": [ + { + "alt": "A rendering of the new \"Filters\" button and option fields, which allow you to query and filter by keywords, people, date range, language and more.", + "aspectRatio": { + "height": 2000, + "width": 2140 + }, + "image": { + "$type": "blob", + "ref": { + "$link": "bafkreifvw4djmv7ney453nfyozyvoxvrwrlnpmbfhezarsn6plkrcvrw64" + }, + "mimeType": "image/jpeg", + "size": 1554074 + } + } + ] + }, + "langs": [ + "en" + ], + "text": "v1.127 is live! We're rolling out improvements to search over the coming weeks. You'll get more relevant results and more ways to find what you're looking for. \n\nThe new \"Filters\" button lets you query and filter by keywords, people, date range, language and more. You can even share your searches!" + }, + "embed": { + "images": [ + { + "thumb": "https://cdn.bsky.app/img/feed_thumbnail/plain/did:plc:z72i7hdynmk6r22z27h6tvur/bafkreifvw4djmv7ney453nfyozyvoxvrwrlnpmbfhezarsn6plkrcvrw64", + "fullsize": "https://cdn.bsky.app/img/feed_fullsize/plain/did:plc:z72i7hdynmk6r22z27h6tvur/bafkreifvw4djmv7ney453nfyozyvoxvrwrlnpmbfhezarsn6plkrcvrw64", + "alt": "A rendering of the new \"Filters\" button and option fields, which allow you to query and filter by keywords, people, date range, language and more.", + "aspectRatio": { + "height": 2000, + "width": 2140 + } + } + ], + "$type": "app.bsky.embed.images#view" + }, + "bookmarkCount": 162, + "replyCount": 253, + "repostCount": 579, + "likeCount": 3453, + "quoteCount": 183, + "indexedAt": "2026-07-09T19:50:21.327Z", + "labels": [] + } + }, + { + "post": { + "uri": "at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.feed.post/3mpok7nkjtc2o", + "cid": "bafyreib3xpze6k6lagrwhjn75wj4daeaj73hvhhjkmt3d5umm6qkqba5fa", + "author": { + "did": "did:plc:z72i7hdynmk6r22z27h6tvur", + "handle": "bsky.app", + "displayName": "Bluesky", + "avatar": "https://cdn.bsky.app/img/avatar/plain/did:plc:z72i7hdynmk6r22z27h6tvur/bafkreihwihm6kpd6zuwhhlro75p5qks5qtrcu55jp3gddbfjsieiv7wuka", + "associated": { + "chat": { + "allowIncoming": "none" + }, + "activitySubscription": { + "allowSubscriptions": "followers" + } + }, + "labels": [], + "createdAt": "2023-04-12T04:53:57.057Z", + "verification": { + "verifications": [], + "verifiedStatus": "none", + "trustedVerifierStatus": "valid" + } + }, + "record": { + "$type": "app.bsky.feed.post", + "createdAt": "2026-07-02T17:21:51.498Z", + "embed": { + "$type": "app.bsky.embed.images", + "images": [ + { + "alt": "two spidermen pointing at each other", + "aspectRatio": { + "height": 386, + "width": 686 + }, + "image": { + "$type": "blob", + "ref": { + "$link": "bafkreihlfz7eeyynraq3zqhozs4q3l3wabewbjevjb5jgophbtcjj7eijm" + }, + "mimeType": "image/jpeg", + "size": 161421 + } + } + ] + }, + "langs": [ + "en" + ], + "reply": { + "parent": { + "cid": "bafyreifzkundbxh3docpjjd5oyffcxyjk63ha7zoht6iy4y6bjvvct5zgy", + "uri": "at://did:plc:xxj5ugkba3k6ftpmcl67vv6r/app.bsky.feed.post/3mpog44tffs2r" + }, + "root": { + "cid": "bafyreigjxsxzw2ywmz3uu5pr33eg5j7xqwzarpjthc4nxmpn3cy4v3vopu", + "uri": "at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.feed.post/3mpoezgevvc23" + } + }, + "text": "wait a minute" + }, + "embed": { + "images": [ + { + "thumb": "https://cdn.bsky.app/img/feed_thumbnail/plain/did:plc:z72i7hdynmk6r22z27h6tvur/bafkreihlfz7eeyynraq3zqhozs4q3l3wabewbjevjb5jgophbtcjj7eijm", + "fullsize": "https://cdn.bsky.app/img/feed_fullsize/plain/did:plc:z72i7hdynmk6r22z27h6tvur/bafkreihlfz7eeyynraq3zqhozs4q3l3wabewbjevjb5jgophbtcjj7eijm", + "alt": "two spidermen pointing at each other", + "aspectRatio": { + "height": 386, + "width": 686 + } + } + ], + "$type": "app.bsky.embed.images#view" + }, + "bookmarkCount": 1, + "replyCount": 5, + "repostCount": 4, + "likeCount": 45, + "quoteCount": 0, + "indexedAt": "2026-07-02T17:21:53.024Z", + "labels": [] + }, + "reply": { + "root": { + "uri": "at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.feed.post/3mpoezgevvc23", + "cid": "bafyreigjxsxzw2ywmz3uu5pr33eg5j7xqwzarpjthc4nxmpn3cy4v3vopu", + "author": { + "did": "did:plc:z72i7hdynmk6r22z27h6tvur", + "handle": "bsky.app", + "displayName": "Bluesky", + "avatar": "https://cdn.bsky.app/img/avatar/plain/did:plc:z72i7hdynmk6r22z27h6tvur/bafkreihwihm6kpd6zuwhhlro75p5qks5qtrcu55jp3gddbfjsieiv7wuka", + "associated": { + "chat": { + "allowIncoming": "none" + }, + "activitySubscription": { + "allowSubscriptions": "followers" + } + }, + "labels": [], + "createdAt": "2023-04-12T04:53:57.057Z", + "verification": { + "verifications": [], + "verifiedStatus": "none", + "trustedVerifierStatus": "valid" + } + }, + "record": { + "$type": "app.bsky.feed.post", + "createdAt": "2026-07-02T15:48:53.937Z", + "embed": { + "$type": "app.bsky.embed.record", + "record": { + "cid": "bafyreicsssegckxhp5hvr5rgrpuol5yzgnk5gixexhrxsvta5ayigmdcai", + "uri": "at://did:plc:36qxymynrrlueovhitpmrfbu/app.bsky.feed.post/3mpmcyv2e6c2c" + } + }, + "langs": [ + "en" + ], + "text": "Pretty sure this is the first award issued for Bluesky posting 🎉" + }, + "embed": { + "record": { + "uri": "at://did:plc:36qxymynrrlueovhitpmrfbu/app.bsky.feed.post/3mpmcyv2e6c2c", + "cid": "bafyreicsssegckxhp5hvr5rgrpuol5yzgnk5gixexhrxsvta5ayigmdcai", + "author": { + "did": "did:plc:36qxymynrrlueovhitpmrfbu", + "handle": "melbuer.bsky.social", + "displayName": "Mel Buer", + "avatar": "https://cdn.bsky.app/img/avatar/plain/did:plc:36qxymynrrlueovhitpmrfbu/bafkreihyzz327efgou3hf7ucztdcf3xqf7vhhmrrlhdjjycb2c3omdjffq", + "associated": { + "chat": { + "allowIncoming": "all" + }, + "activitySubscription": { + "allowSubscriptions": "followers" + } + }, + "labels": [], + "createdAt": "2023-05-13T01:57:21.886Z", + "verification": { + "verifications": [ + { + "issuer": "did:plc:z72i7hdynmk6r22z27h6tvur", + "issuerDisplayName": "Bluesky", + "issuerHandle": "bsky.app", + "uri": "at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.graph.verification/3mpolodjvrt2a", + "isValid": true, + "createdAt": "2026-07-02T17:47:57.732Z" + } + ], + "verifiedStatus": "valid", + "trustedVerifierStatus": "none" + } + }, + "value": { + "$type": "app.bsky.feed.post", + "createdAt": "2026-07-01T20:07:28.798Z", + "embed": { + "$type": "app.bsky.embed.images", + "images": [ + { + "alt": "", + "aspectRatio": { + "height": 2048, + "width": 1536 + }, + "image": { + "$type": "blob", + "ref": { + "$link": "bafkreibgkiw34phdpi7m6z6lxqq43al62gi6oevuw3w2klyh3sgyv7fqwa" + }, + "mimeType": "image/jpeg", + "size": 1389286 + } + } + ] + }, + "langs": [ + "en" + ], + "text": "Oh yeah, I am now an award-winning independent journalist. Pretty cool!" + }, + "labels": [], + "likeCount": 2442, + "replyCount": 96, + "repostCount": 169, + "quoteCount": 29, + "indexedAt": "2026-07-01T20:07:33.830Z", + "embeds": [ + { + "images": [ + { + "thumb": "https://cdn.bsky.app/img/feed_thumbnail/plain/did:plc:36qxymynrrlueovhitpmrfbu/bafkreibgkiw34phdpi7m6z6lxqq43al62gi6oevuw3w2klyh3sgyv7fqwa", + "fullsize": "https://cdn.bsky.app/img/feed_fullsize/plain/did:plc:36qxymynrrlueovhitpmrfbu/bafkreibgkiw34phdpi7m6z6lxqq43al62gi6oevuw3w2klyh3sgyv7fqwa", + "alt": "", + "aspectRatio": { + "height": 2048, + "width": 1536 + } + } + ], + "$type": "app.bsky.embed.images#view" + } + ], + "$type": "app.bsky.embed.record#viewRecord" + }, + "$type": "app.bsky.embed.record#view" + }, + "bookmarkCount": 67, + "replyCount": 81, + "repostCount": 248, + "likeCount": 3226, + "quoteCount": 22, + "indexedAt": "2026-07-02T15:48:54.430Z", + "labels": [], + "$type": "app.bsky.feed.defs#postView" + }, + "parent": { + "uri": "at://did:plc:xxj5ugkba3k6ftpmcl67vv6r/app.bsky.feed.post/3mpog44tffs2r", + "cid": "bafyreifzkundbxh3docpjjd5oyffcxyjk63ha7zoht6iy4y6bjvvct5zgy", + "author": { + "did": "did:plc:xxj5ugkba3k6ftpmcl67vv6r", + "handle": "crimsongoose.bsky.social", + "displayName": "The Crimson Goose", + "avatar": "https://cdn.bsky.app/img/avatar/plain/did:plc:xxj5ugkba3k6ftpmcl67vv6r/bafkreihi5k6xdcq6vmely76y2xazpnumyc4mriyzt64ozy5nob7sos4uji", + "associated": { + "activitySubscription": { + "allowSubscriptions": "followers" + } + }, + "labels": [], + "createdAt": "2025-07-14T14:24:24.834Z" + }, + "record": { + "$type": "app.bsky.feed.post", + "createdAt": "2026-07-02T16:08:18.329Z", + "embed": { + "$type": "app.bsky.embed.images", + "images": [ + { + "alt": "", + "aspectRatio": { + "height": 3999, + "width": 3000 + }, + "image": { + "$type": "blob", + "ref": { + "$link": "bafkreiau2qppztxtyfyyfowe3f7p6m2gzgztybrjiulvdq5e5cupdp7uqm" + }, + "mimeType": "image/jpeg", + "size": 1999216 + } + } + ] + }, + "langs": [ + "en" + ], + "reply": { + "parent": { + "cid": "bafyreigjxsxzw2ywmz3uu5pr33eg5j7xqwzarpjthc4nxmpn3cy4v3vopu", + "uri": "at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.feed.post/3mpoezgevvc23" + }, + "root": { + "cid": "bafyreigjxsxzw2ywmz3uu5pr33eg5j7xqwzarpjthc4nxmpn3cy4v3vopu", + "uri": "at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.feed.post/3mpoezgevvc23" + } + }, + "text": "" + }, + "embed": { + "images": [ + { + "thumb": "https://cdn.bsky.app/img/feed_thumbnail/plain/did:plc:xxj5ugkba3k6ftpmcl67vv6r/bafkreiau2qppztxtyfyyfowe3f7p6m2gzgztybrjiulvdq5e5cupdp7uqm", + "fullsize": "https://cdn.bsky.app/img/feed_fullsize/plain/did:plc:xxj5ugkba3k6ftpmcl67vv6r/bafkreiau2qppztxtyfyyfowe3f7p6m2gzgztybrjiulvdq5e5cupdp7uqm", + "alt": "", + "aspectRatio": { + "height": 3999, + "width": 3000 + } + } + ], + "$type": "app.bsky.embed.images#view" + }, + "bookmarkCount": 0, + "replyCount": 1, + "repostCount": 1, + "likeCount": 13, + "quoteCount": 0, + "indexedAt": "2026-07-02T16:08:31.627Z", + "labels": [], + "$type": "app.bsky.feed.defs#postView" + }, + "grandparentAuthor": { + "did": "did:plc:z72i7hdynmk6r22z27h6tvur", + "handle": "bsky.app", + "displayName": "Bluesky", + "avatar": "https://cdn.bsky.app/img/avatar/plain/did:plc:z72i7hdynmk6r22z27h6tvur/bafkreihwihm6kpd6zuwhhlro75p5qks5qtrcu55jp3gddbfjsieiv7wuka", + "associated": { + "chat": { + "allowIncoming": "none" + }, + "activitySubscription": { + "allowSubscriptions": "followers" + } + }, + "labels": [], + "createdAt": "2023-04-12T04:53:57.057Z", + "verification": { + "verifications": [], + "verifiedStatus": "none", + "trustedVerifierStatus": "valid" + } + } + } + } + ], + "cursor": "2026-06-17T16:35:13.629Z" +} +\ No newline at end of file diff --git a/common/social/blueskyFetcher.test.ts b/common/social/blueskyFetcher.test.ts @@ -0,0 +1,99 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { readFileSync } from "node:fs"; +import path from "node:path"; +import { fileURLToPath } from "node:url"; +import { + didFromAtUri, + normalizeBlueskyItem, + rkeyFromAtUri, +} from "./blueskyFetcher"; +import { + monthShardFromCreatedAt, + parsePost, + uploadDateFromCreatedAt, +} from "../lib/posts"; + +// A recorded getAuthorFeed response — no network in tests. Captured live from +// public.api.bsky.app and trimmed to one repost, one reply and two plain posts. +const fixtureDir = path.join( + path.dirname(fileURLToPath(import.meta.url)), + "__fixtures__", +); +const feed = JSON.parse( + readFileSync(path.join(fixtureDir, "bluesky-author-feed.json"), "utf8"), +) as { feed: unknown[]; cursor?: string }; + +test("at-uri parsing pulls the rkey and did", () => { + const uri = "at://did:plc:z72i7hdynmk6r22z27h6tvur/app.bsky.feed.post/3mqc36slinc2m"; + assert.equal(rkeyFromAtUri(uri), "3mqc36slinc2m"); + assert.equal(didFromAtUri(uri), "did:plc:z72i7hdynmk6r22z27h6tvur"); + assert.equal(rkeyFromAtUri(undefined), null); + assert.equal(didFromAtUri("not-an-at-uri"), null); +}); + +test("normalizes every fixture item into a valid Post", () => { + const posts = feed.feed.map((item) => + normalizeBlueskyItem(item as never, "testchannel"), + ); + assert.equal(posts.length, 4); + for (const post of posts) { + assert.ok(post, "every fixture item should normalize"); + assert.equal(post!.platform, "bluesky"); + assert.equal(post!.channelSlug, "testchannel"); + assert.equal(post!.slug, `testchannel/${post!.id}`); + // A normalized post must survive the on-disk round trip unchanged. + assert.deepEqual(parsePost(JSON.parse(JSON.stringify(post))), post); + } +}); + +test("derives uploadDate and the permalink from the record", () => { + const post = normalizeBlueskyItem(feed.feed[0] as never, "testchannel")!; + assert.equal(post.uploadDate, uploadDateFromCreatedAt(post.createdAt)); + assert.match(post.uploadDate, /^\d{8}$/); + assert.equal( + post.url, + `https://bsky.app/profile/${post.author}/post/${post.id}`, + ); +}); + +test("marks the repost item and points repostOf at the original", () => { + const posts = feed.feed.map((i) => normalizeBlueskyItem(i as never, "c")!); + const repost = posts.find((p) => p.isRepost); + assert.ok(repost, "fixture contains a repost"); + assert.equal(repost!.repostOf?.platform, "bluesky"); + assert.equal(repost!.repostOf?.id, repost!.id); +}); + +test("marks the reply item and records its thread root", () => { + const posts = feed.feed.map((i) => normalizeBlueskyItem(i as never, "c")!); + const reply = posts.find((p) => p.isReply); + assert.ok(reply, "fixture contains a reply"); + assert.ok(reply!.replyTo?.id, "a reply carries a parent ref"); + assert.ok(reply!.threadId, "a reply carries a thread id"); + assert.notEqual(reply!.threadId, reply!.id); +}); + +test("a root post threads under its own id", () => { + const posts = feed.feed.map((i) => normalizeBlueskyItem(i as never, "c")!); + const root = posts.find((p) => !p.isReply); + assert.ok(root); + assert.equal(root!.threadId, root!.id); +}); + +test("rejects items with no resolvable id or timestamp", () => { + assert.equal(normalizeBlueskyItem({} as never, "c"), null); + assert.equal( + normalizeBlueskyItem({ post: { uri: "at://did/x/abc" } } as never, "c"), + null, + "no createdAt -> rejected", + ); +}); + +test("createdAt drives both the date filter field and the month shard", () => { + assert.equal(uploadDateFromCreatedAt("2026-07-10T11:46:11.858Z"), "20260710"); + assert.equal(monthShardFromCreatedAt("2026-07-10T11:46:11.858Z"), "2026-07"); + // Unparseable input degrades rather than poisoning the lexicographic bounds. + assert.equal(uploadDateFromCreatedAt("garbage"), ""); + assert.equal(monthShardFromCreatedAt("garbage"), "unknown"); +}); diff --git a/common/social/blueskyFetcher.ts b/common/social/blueskyFetcher.ts @@ -0,0 +1,486 @@ +// Bluesky ingest over the public AT Protocol — no auth, no external binary, +// no browser. Verified live against public.api.bsky.app: +// com.atproto.identity.resolveHandle -> { did } +// app.bsky.feed.getAuthorFeed -> { feed: [...], cursor } +// both 200 unauthenticated, with full post text, ISO-8601 ms `createdAt` and +// cursor paging. +// +// Bluesky leads the whole posts feature deliberately: it is free, complete and +// verified working, so the model / storage / index / search / AI pipeline can +// be proven end-to-end before touching anything as fragile as X. +// +// Note on PDS resolution: bsky.social 401s for accounts hosted on other PDSes, +// so any request that must hit the account's own PDS (the getRepo CAR backfill) +// resolves the service endpoint from the DID document at plc.directory first. +// The read-only AppView (public.api.bsky.app) needs no such resolution. + +import { + postPermalink, + uploadDateFromCreatedAt, + type Post, + type PostRef, +} from "../lib/posts"; +import { + registerSocialFetcher, + type PostFetchInput, + type PostFetchResult, + type SocialFetcher, + type SocialFetcherProbe, +} from "./fetchers"; + +const APPVIEW = "https://public.api.bsky.app"; +const PLC_DIRECTORY = "https://plc.directory"; + +// getAuthorFeed's documented ceiling. +const PAGE_LIMIT = 100; +// Safety valve so a first-run backfill of a prolific account can't page +// forever inside one job. +const MAX_PAGES = 200; + +// --------------------------------------------------------------------------- +// Wire types (only the fields we consume) +// --------------------------------------------------------------------------- + +type BskyAuthor = { + did?: string; + handle?: string; + displayName?: string; +}; + +type BskyFacetFeature = { + $type?: string; + uri?: string; +}; + +type BskyFacet = { + features?: BskyFacetFeature[]; +}; + +type BskyStrongRef = { uri?: string; cid?: string }; + +type BskyRecord = { + text?: string; + createdAt?: string; + langs?: string[]; + facets?: BskyFacet[]; + reply?: { root?: BskyStrongRef; parent?: BskyStrongRef }; + embed?: BskyEmbed; +}; + +type BskyEmbed = { + $type?: string; + images?: unknown[]; + media?: { images?: unknown[] }; + external?: { uri?: string }; + record?: BskyStrongRef & { record?: BskyStrongRef }; +}; + +type BskyPostView = { + uri?: string; + cid?: string; + author?: BskyAuthor; + record?: BskyRecord; + embed?: BskyEmbed; + replyCount?: number; + repostCount?: number; + likeCount?: number; + quoteCount?: number; + indexedAt?: string; +}; + +type BskyReplyRef = { + root?: BskyPostView; + parent?: BskyPostView; +}; + +type BskyReason = { + $type?: string; + by?: BskyAuthor; + indexedAt?: string; +}; + +type BskyFeedItem = { + post?: BskyPostView; + reply?: BskyReplyRef; + reason?: BskyReason; +}; + +type BskyFeedResponse = { + feed?: BskyFeedItem[]; + cursor?: string; +}; + +// --------------------------------------------------------------------------- +// Helpers +// --------------------------------------------------------------------------- + +// at://<did>/app.bsky.feed.post/<rkey> -> rkey. The rkey is the post's native +// id, the atproto analogue of a tweet id. +export function rkeyFromAtUri(uri: string | undefined): string | null { + if (!uri) return null; + const idx = uri.lastIndexOf("/"); + if (idx < 0 || idx === uri.length - 1) return null; + return uri.slice(idx + 1); +} + +export function didFromAtUri(uri: string | undefined): string | null { + if (!uri) return null; + const match = /^at:\/\/([^/]+)\//.exec(uri); + return match ? match[1] : null; +} + +function refFromPostView(view: BskyPostView | undefined): PostRef | null { + if (!view) return null; + const id = rkeyFromAtUri(view.uri); + if (!id) return null; + const author = view.author?.handle ?? ""; + const ref: PostRef = { platform: "bluesky", id }; + if (author) { + ref.author = author; + ref.url = postPermalink("bluesky", author, id); + } + return ref; +} + +// A strong ref carries only a URI, so the handle is unavailable; the DID +// works in a bsky.app profile URL, so the link still resolves. +function refFromStrongRef(ref: BskyStrongRef | undefined): PostRef | null { + if (!ref?.uri) return null; + const id = rkeyFromAtUri(ref.uri); + if (!id) return null; + const did = didFromAtUri(ref.uri); + const out: PostRef = { platform: "bluesky", id }; + if (did) { + out.author = did; + out.url = postPermalink("bluesky", did, id); + } + return out; +} + +// Outbound URLs: richtext link facets plus an external-embed card's target. +function linksFrom(record: BskyRecord | undefined, embed: BskyEmbed | undefined): string[] { + const links = new Set<string>(); + for (const facet of record?.facets ?? []) { + for (const feature of facet.features ?? []) { + if (feature.$type === "app.bsky.richtext.facet#link" && feature.uri) { + links.add(feature.uri); + } + } + } + const external = record?.embed?.external?.uri ?? embed?.external?.uri; + if (external) links.add(external); + return [...links]; +} + +// Media is counted, not archived (v1). Images can hang off either a plain +// image embed or the media half of a recordWithMedia embed. +function mediaCountFrom(embed: BskyEmbed | undefined): number { + if (!embed) return 0; + if (Array.isArray(embed.images)) return embed.images.length; + if (Array.isArray(embed.media?.images)) return embed.media!.images!.length; + return 0; +} + +function quotedRefFrom( + record: BskyRecord | undefined, + embed: BskyEmbed | undefined, +): PostRef | null { + // app.bsky.embed.record -> record is the strong ref; + // app.bsky.embed.recordWithMedia -> record.record is. + const candidate = record?.embed?.record ?? embed?.record; + if (!candidate) return null; + const inner = (candidate as { record?: BskyStrongRef }).record; + return refFromStrongRef(inner ?? (candidate as BskyStrongRef)); +} + +// Normalize one feed item into a Post. Returns null for items we can't +// identify (missing uri/text/createdAt). +// +// Reposts: for a repost the feed item's `post` IS the original post, and +// `reason.by` is the account that reposted it. We archive the original record +// with isRepost + repostOf set — i.e. "this post entered the archive because +// the tracked account reposted it" — which keeps attribution honest and lets +// the archive dedupe by the original id. +export function normalizeBlueskyItem( + item: BskyFeedItem, + channelSlug: string, +): Post | null { + const view = item.post; + if (!view) return null; + const id = rkeyFromAtUri(view.uri); + if (!id) return null; + const record = view.record; + const createdAt = record?.createdAt ?? view.indexedAt; + if (!createdAt) return null; + const text = typeof record?.text === "string" ? record.text : ""; + const author = view.author?.handle ?? ""; + + const isRepost = item.reason?.$type === "app.bsky.feed.defs#reasonRepost"; + const replyParent = + refFromPostView(item.reply?.parent) ?? refFromStrongRef(record?.reply?.parent); + const replyRoot = + refFromPostView(item.reply?.root) ?? refFromStrongRef(record?.reply?.root); + + const post: Post = { + id, + slug: `${channelSlug}/${id}`, + channelSlug, + author, + createdAt, + uploadDate: uploadDateFromCreatedAt(createdAt), + text, + url: postPermalink("bluesky", author, id), + platform: "bluesky", + isReply: Boolean(replyParent ?? replyRoot), + isRepost, + links: linksFrom(record, view.embed), + }; + + const displayName = view.author?.displayName; + if (displayName) post.authorName = displayName; + const lang = record?.langs?.[0]; + if (lang) post.lang = lang; + // A root post's thread id is its own id, so every post carries a usable + // grouping key without a second lookup. + post.threadId = replyRoot?.id ?? id; + if (replyParent) post.replyTo = replyParent; + const quoted = quotedRefFrom(record, view.embed); + if (quoted) post.quoted = quoted; + if (isRepost) { + post.repostOf = { + platform: "bluesky", + id, + ...(author + ? { author, url: postPermalink("bluesky", author, id) } + : {}), + }; + } + const media = mediaCountFrom(view.embed); + if (media > 0) post.mediaCount = media; + + const engagement: NonNullable<Post["engagement"]> = {}; + if (typeof view.likeCount === "number") engagement.likes = view.likeCount; + if (typeof view.repostCount === "number") engagement.reposts = view.repostCount; + if (typeof view.replyCount === "number") engagement.replies = view.replyCount; + if (typeof view.quoteCount === "number") engagement.quotes = view.quoteCount; + if (Object.keys(engagement).length > 0) post.engagement = engagement; + + return post; +} + +// --------------------------------------------------------------------------- +// API calls +// --------------------------------------------------------------------------- + +async function getJson<T>( + url: string, + signal: AbortSignal | undefined, +): Promise<T> { + const res = await fetch(url, { + signal, + headers: { accept: "application/json" }, + }); + if (!res.ok) { + throw new Error(`${res.status} ${res.statusText} for ${url}`); + } + return (await res.json()) as T; +} + +export async function resolveHandle( + handle: string, + signal?: AbortSignal, +): Promise<string> { + const url = `${APPVIEW}/xrpc/com.atproto.identity.resolveHandle?handle=${encodeURIComponent(handle)}`; + const body = await getJson<{ did?: string }>(url, signal); + if (!body.did) throw new Error(`No DID for handle ${handle}`); + return body.did; +} + +type BskyProfile = { + did?: string; + handle?: string; + displayName?: string; + description?: string; + avatar?: string; + postsCount?: number; +}; + +export async function getProfile( + actor: string, + signal?: AbortSignal, +): Promise<BskyProfile> { + const url = `${APPVIEW}/xrpc/app.bsky.actor.getProfile?actor=${encodeURIComponent(actor)}`; + return getJson<BskyProfile>(url, signal); +} + +// The account's own PDS endpoint, from its DID document. Required — and NOT +// optional — for any getRepo CAR backfill: bsky.social 401s for accounts hosted +// elsewhere. +export async function resolvePdsEndpoint( + did: string, + signal?: AbortSignal, +): Promise<string | null> { + if (!did.startsWith("did:plc:")) return null; + type DidDoc = { + service?: Array<{ id?: string; type?: string; serviceEndpoint?: string }>; + }; + const doc = await getJson<DidDoc>( + `${PLC_DIRECTORY}/${encodeURIComponent(did)}`, + signal, + ); + for (const service of doc.service ?? []) { + if ( + service.type === "AtprotoPersonalDataServer" && + typeof service.serviceEndpoint === "string" + ) { + return service.serviceEndpoint.replace(/\/$/, ""); + } + } + return null; +} + +async function getAuthorFeed( + actor: string, + cursor: string | undefined, + signal: AbortSignal | undefined, +): Promise<BskyFeedResponse> { + const params = new URLSearchParams({ + actor, + limit: String(PAGE_LIMIT), + filter: "posts_with_replies", + }); + if (cursor) params.set("cursor", cursor); + return getJson<BskyFeedResponse>( + `${APPVIEW}/xrpc/app.bsky.feed.getAuthorFeed?${params.toString()}`, + signal, + ); +} + +// --------------------------------------------------------------------------- +// Fetcher +// --------------------------------------------------------------------------- + +export const blueskyFetcher: SocialFetcher = { + id: "bluesky-atproto", + label: "Bluesky (public AT Protocol)", + platform: "bluesky", + fields: { limit: true }, + + detect(url: string): boolean { + try { + const host = new URL(url).hostname.toLowerCase(); + return host === "bsky.app" || host.endsWith(".bsky.app"); + } catch { + return false; + } + }, + + async probe(url: string, signal?: AbortSignal): Promise<SocialFetcherProbe> { + const handle = handleFromUrl(url); + if (!handle) { + return { ok: false, error: "Could not read a handle from that URL" }; + } + try { + const profile = await getProfile(handle, signal); + const resolved = profile.handle ?? handle; + return { + ok: true, + name: profile.displayName || resolved, + handle: resolved, + url: `https://bsky.app/profile/${resolved}`, + avatar: profile.avatar, + description: profile.description, + postCount: profile.postsCount, + }; + } catch (err) { + return { ok: false, error: (err as Error).message }; + } + }, + + async fetch(input: PostFetchInput): Promise<PostFetchResult> { + const { handle, channelSlug, since, seenIds, limit, signal, onLog } = input; + const posts: Post[] = []; + // Resuming an unfinished backfill: continue from the stored cursor and + // ignore the watermark, otherwise the first page (all already-archived) + // would stop the run before it reached the missing history. + let cursor: string | undefined = input.cursor; + const resuming = Boolean(input.cursor); + const watermark = resuming ? undefined : since; + let complete = false; + if (resuming) { + onLog?.(`Resuming backfill from cursor ${cursor}`); + } + + for (let page = 0; page < MAX_PAGES; page++) { + if (signal.aborted) break; + const body = await getAuthorFeed(handle, cursor, signal); + const items = body.feed ?? []; + if (items.length === 0) { + complete = true; + break; + } + + let sawKnown = false; + let sawOlderThanWatermark = false; + let newOnPage = 0; + for (const item of items) { + const post = normalizeBlueskyItem(item, channelSlug); + if (!post) continue; + if (seenIds.has(post.id)) { + sawKnown = true; + continue; + } + if (watermark && post.createdAt <= watermark) { + sawOlderThanWatermark = true; + continue; + } + posts.push(post); + newOnPage++; + } + + onLog?.( + `Bluesky page ${page + 1}: ${items.length} items, ${newOnPage} new` + + (sawKnown ? ", reached already-archived content" : "") + + (sawOlderThanWatermark ? ", reached the watermark" : ""), + ); + + cursor = body.cursor; + + // Stop conditions, in the same spirit as sync()'s "stop at the first + // page containing an already-archived entry": we still keep the new + // posts found on that page, we just don't page further back. + if (sawKnown || sawOlderThanWatermark) { + complete = true; + break; + } + if (!cursor) { + complete = true; + break; + } + if (limit && posts.length >= limit) { + posts.length = limit; + break; + } + } + + return { posts, complete, cursor }; + }, +}; + +function handleFromUrl(url: string): string | null { + const trimmed = url.trim(); + if (!trimmed) return null; + if (!/^https?:\/\//i.test(trimmed)) return trimmed.replace(/^@/, ""); + try { + const parsed = new URL(trimmed); + const segments = parsed.pathname.split("/").filter(Boolean); + if (segments[0] === "profile" && segments[1]) { + return decodeURIComponent(segments[1]).replace(/^@/, ""); + } + return null; + } catch { + return null; + } +} + +registerSocialFetcher(blueskyFetcher); diff --git a/common/social/fetchers.ts b/common/social/fetchers.ts @@ -0,0 +1,203 @@ +// First-class registry of social-post fetchers, modeled on +// common/lib/transcriptionApps.ts: each fetcher owns how it detects a URL it +// can handle, what config fields it surfaces in the editor form, how it probes +// an account cheaply, and how it pages a timeline into normalized `Post`s. +// +// The registry exists because every X retrieval path rots — cookies expire in +// days and X rotates its GraphQL query IDs every few weeks — so the SEAM +// matters more than any single implementation. Bluesky, by contrast, is a free +// and complete public API and needs no external binary at all. +// +// Server-side module (fetchers may spawn binaries / read paths). The editor +// form consumes `listSocialFetchers()` descriptors instead, exactly as the +// settings form consumes listTranscriptionApps() — so this module never lands +// in the client bundle. + +import type { Post, PostPlatform } from "../lib/posts"; + +export type PostFetchInput = { + // The account's canonical URL as configured on the channel. + accountUrl: string; + // The bare handle, without a leading "@" or URL wrapper. + handle: string; + // The channel slug the produced posts belong to. + channelSlug: string; + // ISO watermark: stop once posts older than this are reached. Absent on a + // first (full-backfill) run AND whenever a previous run stopped early — see + // `cursor`, which resumes the backfill instead. + since?: string; + // Opaque resume point from a previous run that did not finish (hit its + // limit, or was cancelled). When set, the fetcher continues paging from here + // rather than from the top of the timeline, so a capped first run does not + // leave a permanent hole in the archived history. + cursor?: string; + // Post ids already on disk. Paging stops at the first page containing one, + // mirroring how sync() stops at the first already-archived video. + seenIds: ReadonlySet<string>; + // Resolved cookie spec from resolveCookiePolicy(), when the fetcher needs + // credentials. Bluesky ignores it entirely. + cookies?: string; + // Soft cap on how many posts to return in one run. Undefined = no cap + // beyond the watermark/seen-id stop conditions. + limit?: number; + signal: AbortSignal; + // Progress/diagnostic sink, wired to the job log. + onLog?: (line: string) => void; +}; + +export type PostFetchResult = { + posts: Post[]; + // False when the run stopped early (limit hit, or aborted) and more history + // remains behind `cursor`. + complete: boolean; + cursor?: string; + // Set when the fetcher stopped because credentials are missing or expired. + // Maps onto the existing "needs cookies" snapshot bucket rather than + // inventing a new failure surface. + needsCookies?: boolean; +}; + +export type SocialFetcherProbe = { + ok: boolean; + // Display name for the account, used to autofill the channel form. + name?: string; + handle?: string; + // Canonical account URL, normalized. + url?: string; + avatar?: string; + description?: string; + postCount?: number; + error?: string; +}; + +// Which config fields this fetcher surfaces in the channel form. Mirrors +// TranscriptionApp.fields. +export type SocialFetcherFields = { + // Needs a browser cookie spec (cookiesFromBrowser / cookieMode). + cookies?: boolean; + // Needs an external binary whose path is configurable. + binPath?: boolean; + // Supports capping how many posts one run fetches. + limit?: boolean; +}; + +export type SocialFetcher = { + id: string; + label: string; + platform: PostPlatform; + // True when this fetcher can handle the given account URL. + detect(url: string): boolean; + fields: SocialFetcherFields; + // Cheap metadata lookup for the channel form's "Fetch details" button. + probe(url: string, signal?: AbortSignal): Promise<SocialFetcherProbe>; + fetch(input: PostFetchInput): Promise<PostFetchResult>; +}; + +// Populated by registerSocialFetcher() from each fetcher module. Indirection +// (rather than a literal object) keeps this module free of imports from the +// individual fetchers, so the heavier ones — which spawn binaries or launch a +// browser — are only pulled in by the server entry points that register them. +const REGISTRY = new Map<string, SocialFetcher>(); + +export function registerSocialFetcher(fetcher: SocialFetcher): void { + REGISTRY.set(fetcher.id, fetcher); +} + +export function getSocialFetcher( + id: string | undefined | null, +): SocialFetcher | undefined { + return id ? REGISTRY.get(id) : undefined; +} + +export function listSocialFetcherIds(): string[] { + return [...REGISTRY.keys()]; +} + +// The fetcher that claims this URL. When several match, the first registered +// wins — registration order therefore encodes preference, which is how +// x-gallery-dl stays the primary X path with x-playwright as its fallback. +export function detectSocialFetcher( + url: string | undefined | null, +): SocialFetcher | undefined { + if (!url) return undefined; + for (const fetcher of REGISTRY.values()) { + if (fetcher.detect(url)) return fetcher; + } + return undefined; +} + +// Fetchers for a platform, in registration (preference) order. +export function fetchersForPlatform( + platform: PostPlatform, +): SocialFetcher[] { + return [...REGISTRY.values()].filter((f) => f.platform === platform); +} + +// Resolve the fetcher for a channel: an explicit postFetcher wins, otherwise +// fall back to URL detection so an existing channel keeps working if its +// configured fetcher is later removed. +export function resolveSocialFetcher( + postFetcher: string | undefined | null, + accountUrl: string | undefined | null, +): SocialFetcher | undefined { + return getSocialFetcher(postFetcher) ?? detectSocialFetcher(accountUrl); +} + +// A client-safe view (no functions) for the channel form. +export type SocialFetcherDescriptor = { + id: string; + label: string; + platform: PostPlatform; + fields: SocialFetcherFields; +}; + +export function listSocialFetchers(): SocialFetcherDescriptor[] { + return [...REGISTRY.values()].map((f) => ({ + id: f.id, + label: f.label, + platform: f.platform, + fields: f.fields, + })); +} + +// --------------------------------------------------------------------------- +// Handle parsing +// --------------------------------------------------------------------------- + +// Extract a bare handle from an account URL or a raw "@handle" string. +// Returns null when nothing handle-shaped is present. +export function handleFromAccountUrl(input: string): string | null { + const trimmed = input.trim(); + if (!trimmed) return null; + if (!/^https?:\/\//i.test(trimmed)) { + // Bare handle form: "@name", "name", "name.bsky.social". + const bare = trimmed.replace(/^@/, ""); + return /^[A-Za-z0-9._-]+$/.test(bare) ? bare : null; + } + let url: URL; + try { + url = new URL(trimmed); + } catch { + return null; + } + const segments = url.pathname.split("/").filter(Boolean); + if (segments.length === 0) return null; + // bsky.app/profile/<handle-or-did> + if (segments[0] === "profile" && segments[1]) { + return decodeURIComponent(segments[1]).replace(/^@/, ""); + } + // x.com/<handle> — skip the reserved paths that are not accounts. + const first = decodeURIComponent(segments[0]).replace(/^@/, ""); + const RESERVED = new Set([ + "home", + "explore", + "search", + "settings", + "i", + "intent", + "notifications", + "messages", + ]); + if (RESERVED.has(first.toLowerCase())) return null; + return /^[A-Za-z0-9._-]+$/.test(first) ? first : null; +} diff --git a/common/social/xGalleryDlFetcher.ts b/common/social/xGalleryDlFetcher.ts @@ -0,0 +1,310 @@ +// The PRIMARY X/Twitter post fetcher: gallery-dl driven as a light headless +// subprocess. +// +// Why gallery-dl leads: it is exactly the shape this codebase already manages +// well for yt-dlp and whisper — spawn a binary, stream its stdout, parse +// records — so it reuses runManagedCommand's sibling (execa here, since we need +// the stdout payload rather than a job log), the job registry, and +// cookiePolicy.ts essentially unchanged. No browser process per fetch, no RAM +// overhead, and it runs anywhere including containers. +// +// Accepted weaknesses, both modelled rather than hidden: +// - X cookies expire in days. `cookieMode: "defer"` and the existing +// needsCookies snapshot bucket already describe "blocked pending +// credentials", so an auth failure sets `needsCookies` on the result rather +// than inventing a new failure surface. The Playwright session broker +// (xSessionBroker.ts) is the fix for the expiry itself. +// - gallery-dl hardcodes X's GraphQL query IDs, so it breaks when X rotates +// them (historically every 2–4 weeks) until the upstream extractor updates. +// That is what the x-playwright fallback exists for. +// +// NOT VERIFIED AGAINST LIVE X. gallery-dl is not installed in this environment +// and a live spike needs a throwaway account plus real cookies. The argv, +// parsing and failure mapping below are written to gallery-dl's documented +// `--dump-json` contract and are covered by a fixture binary in the editor e2e +// suite; treat the first real run as the spike. + +import { execa } from "execa"; +import { getPaths } from "../lib/paths"; +import type { Post } from "../lib/posts"; +import { normalizeXTweets, type XTweetRaw } from "./xNormalize"; +import { readXSessionStatus, xCookieFile } from "./xSessionBroker"; +import { + registerSocialFetcher, + type PostFetchInput, + type PostFetchResult, + type SocialFetcher, + type SocialFetcherProbe, +} from "./fetchers"; + +// Substrings in gallery-dl's stderr that mean "we are not authenticated" (or no +// longer are). Mapped to needsCookies so the channel lands in the existing +// Needs-cookies bucket instead of looking like a generic crash. +const AUTH_ERROR_PATTERNS = [ + "401", + "403", + "authorization", + "unauthorized", + "login required", + "requires authentication", + "no username", + "bad credentials", + "could not log in", +]; + +function looksLikeAuthFailure(text: string): boolean { + const lower = text.toLowerCase(); + return AUTH_ERROR_PATTERNS.some((p) => lower.includes(p)); +} + +// The argv for a text-only timeline read. Exported so a test can assert the +// flags without spawning anything. +export function buildGalleryDlArgs(opts: { + accountUrl: string; + cookies?: string; + // A cookies.txt exported by the Playwright session broker + // (xSessionBroker.ts). PREFERRED over `cookies`: a live browser profile keeps + // the session fresh, which is the fix for X cookies expiring in days. + cookieFile?: string; + limit?: number; +}): string[] { + const args = [ + // One JSON object per item on stdout — we want metadata, never files. + "--dump-json", + // Never write anything to disk: v1 archives text + links + thread + // structure only, no media. + "--no-download", + // Text-only tweets are skipped by default (gallery-dl is a media + // downloader); this is the flag that makes a timeline text archive work. + "-o", + "extractor.twitter.text-tweets=true", + // Retweets and replies are part of the archived record (they carry the + // thread structure), so keep both. + "-o", + "extractor.twitter.retweets=true", + "-o", + "extractor.twitter.replies=true", + // Media is counted, not archived. + "-o", + "extractor.twitter.videos=false", + "-o", + "extractor.twitter.cards=false", + ]; + if (opts.cookieFile) { + args.push("--cookies", opts.cookieFile); + } else if (opts.cookies) { + // Same browser-cookie spec yt-dlp uses (cookiePolicy.ts), same syntax. + args.push("--cookies-from-browser", opts.cookies); + } + if (opts.limit && opts.limit > 0) { + args.push("--range", `1-${Math.floor(opts.limit)}`); + } + args.push(opts.accountUrl); + return args; +} + +// gallery-dl's --dump-json emits either one object per line, or a top-level +// array. Tolerate both, plus interleaved non-JSON log noise. +export function parseGalleryDlOutput(stdout: string): XTweetRaw[] { + const trimmed = stdout.trim(); + if (!trimmed) return []; + // Whole-payload array first. + if (trimmed.startsWith("[")) { + try { + const parsed = JSON.parse(trimmed) as unknown; + if (Array.isArray(parsed)) return flattenRecords(parsed); + } catch { + // fall through to line mode + } + } + const out: XTweetRaw[] = []; + for (const line of trimmed.split("\n")) { + const t = line.trim(); + if (!t || (!t.startsWith("{") && !t.startsWith("["))) continue; + try { + out.push(...flattenRecords([JSON.parse(t) as unknown])); + } catch { + // Not JSON (a progress line) — ignore. + } + } + return out; +} + +// gallery-dl's dump entries are sometimes `[<type>, <url>, <metadata>]` tuples +// rather than bare objects. Unwrap to the metadata object either way. +function flattenRecords(items: unknown[]): XTweetRaw[] { + const out: XTweetRaw[] = []; + for (const item of items) { + if (Array.isArray(item)) { + for (const part of item) { + if (part && typeof part === "object" && !Array.isArray(part)) { + out.push(part as XTweetRaw); + } + } + continue; + } + if (item && typeof item === "object") out.push(item as XTweetRaw); + } + return out; +} + +export const xGalleryDlFetcher: SocialFetcher = { + id: "x-gallery-dl", + label: "X / Twitter (gallery-dl)", + platform: "twitter", + fields: { cookies: true, binPath: true, limit: true }, + + detect(url: string): boolean { + try { + const host = new URL(url).hostname.toLowerCase(); + return ( + host === "x.com" || + host.endsWith(".x.com") || + host === "twitter.com" || + host.endsWith(".twitter.com") + ); + } catch { + return false; + } + }, + + // A cheap liveness probe: read a single item. There is no free unauthenticated + // X metadata endpoint, so unlike Bluesky this genuinely spawns the binary. + async probe(url: string): Promise<SocialFetcherProbe> { + const bin = getPaths().galleryDlBin; + try { + const res = await execa( + bin, + buildGalleryDlArgs({ accountUrl: url, limit: 1 }), + { reject: false, timeout: 60_000 }, + ); + if (res.exitCode !== 0) { + const err = `${res.stderr ?? ""}`.trim().split("\n").slice(-3).join(" "); + return { + ok: false, + error: looksLikeAuthFailure(err) + ? `gallery-dl needs X credentials (configure cookies-from-browser): ${err}` + : err || `gallery-dl exited ${res.exitCode}`, + }; + } + const records = parseGalleryDlOutput(res.stdout ?? ""); + const first = records[0]; + const author = + first && typeof first === "object" + ? ((first.author as Record<string, unknown> | undefined) ?? undefined) + : undefined; + const handle = + (typeof author?.name === "string" ? author.name : undefined) ?? + (typeof first?.screen_name === "string" ? first.screen_name : undefined); + const nick = typeof author?.nick === "string" ? author.nick : undefined; + return { + ok: true, + name: nick ?? handle, + handle, + url: handle ? `https://x.com/${handle}` : url, + }; + } catch (err) { + const message = (err as Error).message; + return { + ok: false, + error: /ENOENT/.test(message) + ? `gallery-dl not found (looked for "${bin}"; set GALLERY_DL_BIN)` + : message, + }; + } + }, + + async fetch(input: PostFetchInput): Promise<PostFetchResult> { + const { accountUrl, channelSlug, since, seenIds, cookies, limit, signal, onLog } = + input; + const paths = getPaths(); + const bin = paths.galleryDlBin; + // Prefer the session broker's exported jar when one exists — it is kept + // fresh by a live browser profile, so it survives the cookie expiry that + // otherwise breaks gallery-dl within days. + const status = await readXSessionStatus(paths); + const cookieFile = + status.hasCookies && status.looksAuthenticated + ? xCookieFile(paths) + : undefined; + if (cookieFile) onLog?.("Using the X session broker's exported cookies."); + const args = buildGalleryDlArgs({ accountUrl, cookies, cookieFile, limit }); + onLog?.(`Running ${bin} for ${accountUrl}`); + + let stdout = ""; + let stderr = ""; + let exitCode: number | null = null; + try { + const res = await execa(bin, args, { + reject: false, + // gallery-dl walks the timeline newest-first; a long backfill is + // legitimately slow, so the cap is generous. + timeout: 30 * 60_000, + cancelSignal: signal, + }); + stdout = res.stdout ?? ""; + stderr = res.stderr ?? ""; + exitCode = res.exitCode ?? null; + } catch (err) { + const message = (err as Error).message; + if (/ENOENT/.test(message)) { + throw new Error( + `gallery-dl not found (looked for "${bin}"; set GALLERY_DL_BIN)`, + ); + } + throw err; + } + + if (exitCode !== 0) { + const tail = stderr.trim().split("\n").slice(-5).join("\n"); + if (looksLikeAuthFailure(tail)) { + onLog?.(`[auth] gallery-dl could not authenticate to X:\n${tail}`); + // Not a crash — a known, recoverable state the Needs-cookies bucket + // already models. Keep whatever it did manage to emit. + return { + posts: normalizeXTweets(parseGalleryDlOutput(stdout), channelSlug).filter( + (p) => !seenIds.has(p.id), + ), + complete: false, + needsCookies: true, + }; + } + throw new Error(`gallery-dl exited ${exitCode}: ${tail || "(no output)"}`); + } + + const records = parseGalleryDlOutput(stdout); + onLog?.(`gallery-dl returned ${records.length} record(s)`); + const all = normalizeXTweets(records, channelSlug); + + // Apply the same stop conditions the Bluesky fetcher uses. gallery-dl walks + // the whole timeline itself (there is no cursor to hand back), so we filter + // its output rather than stopping mid-stream. + const posts: Post[] = []; + let known = 0; + let older = 0; + for (const post of all) { + if (seenIds.has(post.id)) { + known++; + continue; + } + if (since && post.createdAt <= since) { + older++; + continue; + } + posts.push(post); + } + onLog?.( + `${posts.length} new post(s)` + + (known ? `, ${known} already archived` : "") + + (older ? `, ${older} older than the watermark` : ""), + ); + + // gallery-dl has no resumable cursor: it either walked the timeline it was + // given or it didn't. `complete` is therefore true whenever it exited + // cleanly — except when a --range cap may have cut it short. + const capped = Boolean(limit && all.length >= limit); + return { posts, complete: !capped }; + }, +}; + +registerSocialFetcher(xGalleryDlFetcher); diff --git a/common/social/xNormalize.test.ts b/common/social/xNormalize.test.ts @@ -0,0 +1,389 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + normalizeXTweet, + normalizeXTweets, + xCreatedAt, + xIdOf, + xLinks, + xText, +} from "./xNormalize"; +import { parsePost } from "../lib/posts"; +import { buildGalleryDlArgs, parseGalleryDlOutput } from "./xGalleryDlFetcher"; +import { isXCookie, toNetscapeCookieFile } from "./xSessionBroker"; +import { + collectTweetsFromGraphQL, + isTimelineResponseUrl, + xPlaywrightFetcher, +} from "./xPlaywrightFetcher"; + +// A gallery-dl `--dump-json` metadata record for a text tweet, in the shape its +// twitter extractor documents (flattened author/entities, "YYYY-MM-DD HH:MM:SS" +// dates). No network in tests — X is NEVER contacted by the suite. +const GALLERY_DL_RECORD = { + tweet_id: "1799999999999999999", + conversation_id: "1799999999999999999", + date: "2026-05-04 17:32:11", + content: "a plain text tweet about archives https://t.co/abc", + lang: "en", + author: { name: "someaccount", nick: "Some Account" }, + user: { name: "someaccount", nick: "Some Account" }, + favorite_count: 12, + retweet_count: 3, + reply_count: 1, + quote_count: 0, + entities: { + urls: [ + { url: "https://t.co/abc", expanded_url: "https://example.com/article" }, + ], + }, +}; + +// The legacy/raw GraphQL shape the Playwright fallback sees. Same normalizer. +const RAW_GRAPHQL_RECORD = { + id_str: "1800000000000000001", + conversation_id_str: "1799999999999999999", + created_at: "Wed Oct 10 20:19:24 +0000 2018", + full_text: "a reply in the same conversation", + lang: "en", + user: { name: "someaccount", nick: "Some Account" }, + in_reply_to_status_id_str: "1799999999999999999", + in_reply_to_screen_name: "someaccount", + favorite_count: 4, +}; + +test("normalizes a gallery-dl record into a valid Post", () => { + const post = normalizeXTweet(GALLERY_DL_RECORD, "xchan"); + assert.ok(post); + assert.equal(post!.platform, "twitter"); + assert.equal(post!.id, "1799999999999999999"); + assert.equal(post!.slug, "xchan/1799999999999999999"); + assert.equal(post!.author, "someaccount"); + assert.equal(post!.authorName, "Some Account"); + assert.equal(post!.uploadDate, "20260504"); + assert.equal(post!.url, "https://x.com/someaccount/status/1799999999999999999"); + assert.equal(post!.isReply, false); + assert.equal(post!.isRepost, false); + assert.deepEqual(post!.engagement, { + likes: 12, + reposts: 3, + replies: 1, + quotes: 0, + }); + // Survives the on-disk round trip unchanged. + assert.deepEqual(parsePost(JSON.parse(JSON.stringify(post))), post); +}); + +test("the SAME normalizer handles the raw GraphQL shape", () => { + // This is what makes the Playwright fallback a transport swap rather than a + // rewrite: both payloads land on identical Post fields. + const post = normalizeXTweet(RAW_GRAPHQL_RECORD, "xchan"); + assert.ok(post); + assert.equal(post!.id, "1800000000000000001"); + assert.equal(post!.isReply, true); + assert.equal(post!.replyTo?.id, "1799999999999999999"); + assert.equal(post!.threadId, "1799999999999999999"); + assert.equal(post!.uploadDate, "20181010"); +}); + +test("keeps 64-bit ids as strings (never through a JSON number)", () => { + // 1799999999999999999 > 2^53: parsing it as a number would silently corrupt + // the tweet id and break every permalink and dedupe. + assert.equal(xIdOf({ tweet_id: "1799999999999999999" }), "1799999999999999999"); + assert.ok(!Number.isSafeInteger(Number("1799999999999999999"))); +}); + +test("parses both date dialects to ISO-8601 UTC", () => { + assert.equal(xCreatedAt({ date: "2026-05-04 17:32:11" }), "2026-05-04T17:32:11.000Z"); + assert.equal( + xCreatedAt({ created_at: "Wed Oct 10 20:19:24 +0000 2018" }), + "2018-10-10T20:19:24.000Z", + ); + assert.equal(xCreatedAt({ date: "2026-05-04T17:32:11.000Z" }), "2026-05-04T17:32:11.000Z"); + assert.equal(xCreatedAt({}), undefined); +}); + +test("expands outbound links and drops t.co self-links", () => { + assert.deepEqual(xLinks(GALLERY_DL_RECORD), ["https://example.com/article"]); + assert.deepEqual( + xLinks({ entities: { urls: [{ url: "https://t.co/xyz" }] } }), + [], + ); +}); + +test("prefers a note-tweet body over the truncated text", () => { + const long = "x".repeat(400); + const rec = { + full_text: "x".repeat(140) + "…", + note_tweet: { note_tweet_results: { result: { text: long } } }, + }; + assert.equal(xText(rec), long); +}); + +test("marks a retweet and points repostOf at the original", () => { + const post = normalizeXTweet( + { + ...GALLERY_DL_RECORD, + retweeted_status: { + id_str: "1700000000000000000", + user: { name: "original" }, + }, + }, + "xchan", + ); + assert.ok(post); + assert.equal(post!.isRepost, true); + assert.equal(post!.repostOf?.id, "1700000000000000000"); + assert.equal(post!.repostOf?.author, "original"); +}); + +test("rejects records with no id or no timestamp", () => { + assert.equal(normalizeXTweet({}, "c"), null); + assert.equal(normalizeXTweet({ tweet_id: "1" }, "c"), null, "no date -> rejected"); + assert.equal( + normalizeXTweet({ date: "2026-05-04 17:32:11" }, "c"), + null, + "no id -> rejected", + ); +}); + +test("batch normalize dedupes by id and skips unparseable records", () => { + const posts = normalizeXTweets( + [GALLERY_DL_RECORD, GALLERY_DL_RECORD, {}, RAW_GRAPHQL_RECORD], + "xchan", + ); + assert.equal(posts.length, 2); +}); + +// ─── gallery-dl invocation contract ─── + +test("gallery-dl argv enables text-tweets and downloads nothing", () => { + const args = buildGalleryDlArgs({ + accountUrl: "https://x.com/someaccount", + cookies: "firefox", + limit: 50, + }); + const joined = args.join(" "); + // Text-only timeline extraction is the whole point — without this flag + // gallery-dl skips every tweet that has no media. + assert.match(joined, /extractor\.twitter\.text-tweets=true/); + assert.ok(args.includes("--no-download"), "v1 archives no media"); + assert.ok(args.includes("--dump-json")); + assert.deepEqual( + args.slice(args.indexOf("--cookies-from-browser"), args.indexOf("--cookies-from-browser") + 2), + ["--cookies-from-browser", "firefox"], + ); + assert.match(joined, /--range 1-50/); + assert.equal(args[args.length - 1], "https://x.com/someaccount"); +}); + +test("gallery-dl argv omits cookies when none are resolved", () => { + const args = buildGalleryDlArgs({ accountUrl: "https://x.com/a" }); + assert.ok(!args.includes("--cookies-from-browser")); + assert.ok(!args.join(" ").includes("--range")); +}); + +test("parses JSON-lines, whole-array and tuple dump forms", () => { + const lines = `{"tweet_id":"1","date":"2026-01-01 00:00:00"}\n{"tweet_id":"2","date":"2026-01-02 00:00:00"}`; + assert.equal(parseGalleryDlOutput(lines).length, 2); + + const arr = JSON.stringify([{ tweet_id: "3" }, { tweet_id: "4" }]); + assert.equal(parseGalleryDlOutput(arr).length, 2); + + // gallery-dl also emits [<type>, <url>, <metadata>] tuples. + const tuples = JSON.stringify([[3, "https://x.com/a/status/5", { tweet_id: "5" }]]); + const parsed = parseGalleryDlOutput(tuples); + assert.ok(parsed.some((r) => r.tweet_id === "5")); + + // Progress noise interleaved with JSON must not break parsing. + assert.equal( + parseGalleryDlOutput(`downloading...\n{"tweet_id":"6"}\ndone`).length, + 1, + ); + assert.deepEqual(parseGalleryDlOutput(""), []); +}); + +// ─── X session broker (cookie jar format) ─── + +test("exports cookies in the Netscape format gallery-dl reads", () => { + const jar = toNetscapeCookieFile([ + { + name: "auth_token", + value: "secret", + domain: ".x.com", + path: "/", + expires: 1893456000, + httpOnly: true, + secure: true, + }, + { + name: "ct0", + value: "csrf", + domain: "x.com", + path: "/", + // A session cookie (-1) must serialize as 0, not as a negative number. + expires: -1, + httpOnly: false, + secure: true, + }, + ]); + const lines = jar + .trim() + .split("\n") + .filter((l) => l.trim() !== "" && !l.startsWith("#")); + assert.equal(lines.length, 2); + assert.deepEqual(lines[0].split("\t"), [ + ".x.com", + "TRUE", + "/", + "TRUE", + "1893456000", + "auth_token", + "secret", + ]); + assert.deepEqual(lines[1].split("\t"), [ + "x.com", + "FALSE", + "/", + "TRUE", + "0", + "ct0", + "csrf", + ]); + assert.match(jar, /^# Netscape HTTP Cookie File/); +}); + +test("only X's own cookies are exported", () => { + const mk = (domain: string) => ({ + name: "n", + value: "v", + domain, + path: "/", + expires: 0, + httpOnly: false, + secure: true, + }); + assert.equal(isXCookie(mk(".x.com")), true); + assert.equal(isXCookie(mk("x.com")), true); + assert.equal(isXCookie(mk("api.twitter.com")), true); + // A profile can hold unrelated cookies; handing those to a subprocess would + // leak them for no benefit. + assert.equal(isXCookie(mk("google.com")), false); + assert.equal(isXCookie(mk("notx.com")), false); +}); + +test("gallery-dl prefers the broker's cookie file over browser extraction", () => { + const args = buildGalleryDlArgs({ + accountUrl: "https://x.com/a", + cookies: "firefox", + cookieFile: "/data/.x-session/cookies.txt", + }); + assert.ok(args.includes("--cookies")); + assert.ok( + !args.includes("--cookies-from-browser"), + "the live-profile jar wins — it is what survives X's short cookie expiry", + ); + assert.equal(args[args.indexOf("--cookies") + 1], "/data/.x-session/cookies.txt"); +}); + +// ─── X Playwright fallback (GraphQL payload extraction) ─── +// +// Driven entirely off a recorded-shape payload. The fallback is NEVER pointed +// at x.com in tests: the suite must stay deterministic and no run may risk the +// logged-in account. + +test("recognises the timeline GraphQL operations by substring", () => { + assert.equal( + isTimelineResponseUrl("https://x.com/i/api/graphql/AbC123/UserTweets?variables=%7B%7D"), + true, + ); + assert.equal(isTimelineResponseUrl("https://x.com/i/api/graphql/XyZ/UserTweetsAndReplies"), true); + // A query-id rotation must NOT break the match — that is the whole reason + // this fetcher exists. + assert.equal(isTimelineResponseUrl("https://x.com/i/api/graphql/TOTALLY-NEW-ID/UserTweets"), true); + assert.equal(isTimelineResponseUrl("https://x.com/i/api/2/notifications/all.json"), false); +}); + +test("extracts tweets from a nested GraphQL timeline payload", () => { + // The real payload buries tweets under + // data.user.result.timeline_v2.timeline.instructions[].entries[].content... + const payload = { + data: { + user: { + result: { + timeline_v2: { + timeline: { + instructions: [ + { + type: "TimelineAddEntries", + entries: [ + { + entryId: "tweet-1899000000000000001", + content: { + itemContent: { + tweet_results: { + result: { + __typename: "Tweet", + rest_id: "1899000000000000001", + core: { + user_results: { + result: { + legacy: { + screen_name: "someaccount", + name: "Some Account", + }, + }, + }, + }, + legacy: { + id_str: "1899000000000000001", + conversation_id_str: "1899000000000000001", + created_at: "Wed Oct 10 20:19:24 +0000 2018", + full_text: "a tweet captured through the fallback", + lang: "en", + favorite_count: 7, + }, + }, + }, + }, + }, + }, + { entryId: "cursor-bottom", content: { value: "DAABC" } }, + ], + }, + ], + }, + }, + }, + }, + }, + }; + + const tweets = collectTweetsFromGraphQL(payload); + assert.equal(tweets.length, 1); + + // The SHARED normalizer turns it into the same Post shape the gallery-dl + // path produces — this is what makes the fallback a transport swap. + const posts = normalizeXTweets(tweets, "xchan"); + assert.equal(posts.length, 1); + assert.equal(posts[0].id, "1899000000000000001"); + assert.equal(posts[0].author, "someaccount"); + assert.equal(posts[0].authorName, "Some Account"); + assert.equal(posts[0].text, "a tweet captured through the fallback"); + assert.equal(posts[0].platform, "twitter"); + assert.equal(posts[0].uploadDate, "20181010"); +}); + +test("GraphQL extraction survives cycles and ignores non-tweet nodes", () => { + const cyclic: Record<string, unknown> = { data: {} }; + cyclic.self = cyclic; + assert.deepEqual(collectTweetsFromGraphQL(cyclic), []); + assert.deepEqual(collectTweetsFromGraphQL(null), []); + assert.deepEqual(collectTweetsFromGraphQL({ data: { user: null } }), []); +}); + +test("the fallback never claims a URL by detection", () => { + // gallery-dl is the primary path; the fallback is opted into per channel. + assert.equal(xPlaywrightFetcher.detect("https://x.com/someaccount"), false); + assert.equal(xPlaywrightFetcher.platform, "twitter"); +}); diff --git a/common/social/xNormalize.ts b/common/social/xNormalize.ts @@ -0,0 +1,235 @@ +// The SHARED X/Twitter → Post normalizer. +// +// Both X fetchers produce the same underlying payload: gallery-dl parses X's +// GraphQL timeline responses and emits one metadata object per tweet, and the +// Playwright fallback re-issues that same GraphQL call from inside an +// authenticated page. So the fallback is a TRANSPORT swap, not a rewrite — +// this module is the single place tweet shape is understood. +// +// Deliberately tolerant: X's payload has drifted repeatedly (legacy vs. +// `__typename`-tagged results, `full_text` vs `text`, note-tweets for long +// posts), and gallery-dl flattens some of it. Every field is probed through a +// list of known aliases and a missing one degrades rather than throwing, so a +// partial payload still yields a usable Post. + +import { + postPermalink, + uploadDateFromCreatedAt, + type Post, + type PostRef, +} from "../lib/posts"; + +// A gallery-dl `--dump-json` / metadata-postprocessor record for one tweet, or +// an entry lifted out of a raw GraphQL timeline response. Everything optional. +export type XTweetRaw = Record<string, unknown>; + +function str(v: unknown): string | undefined { + return typeof v === "string" && v !== "" ? v : undefined; +} + +function num(v: unknown): number | undefined { + if (typeof v === "number" && Number.isFinite(v)) return v; + // X sometimes serializes ids and counts as strings. + if (typeof v === "string" && /^\d+$/.test(v)) return Number(v); + return undefined; +} + +function obj(v: unknown): Record<string, unknown> | undefined { + return v && typeof v === "object" && !Array.isArray(v) + ? (v as Record<string, unknown>) + : undefined; +} + +// First defined value among a record's alias keys. +function pick(rec: XTweetRaw, keys: string[]): unknown { + for (const k of keys) { + const v = rec[k]; + if (v !== undefined && v !== null && v !== "") return v; + } + return undefined; +} + +// X ids are 64-bit and MUST stay strings — JSON numbers lose precision above +// 2^53, which silently corrupts a tweet id. +export function xIdOf(rec: XTweetRaw): string | undefined { + const raw = pick(rec, ["tweet_id", "id_str", "rest_id", "id", "conversation_id"]); + if (typeof raw === "string" && raw !== "") return raw; + if (typeof raw === "number" && Number.isFinite(raw)) return String(raw); + return undefined; +} + +// gallery-dl emits `date` as "YYYY-MM-DD HH:MM:SS" (UTC); raw X emits +// `created_at` in the legacy "Wed Oct 10 20:19:24 +0000 2018" form. Normalize +// both to ISO-8601. +export function xCreatedAt(rec: XTweetRaw): string | undefined { + const raw = pick(rec, ["date", "created_at", "createdAt"]); + if (typeof raw !== "string" || !raw) return undefined; + // Already ISO. + if (/^\d{4}-\d{2}-\d{2}T/.test(raw)) { + const ms = Date.parse(raw); + return Number.isFinite(ms) ? new Date(ms).toISOString() : undefined; + } + // gallery-dl's "YYYY-MM-DD HH:MM:SS" is UTC but has no zone marker; adding + // one avoids it being read as local time (which would shift the date). + if (/^\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2}$/.test(raw)) { + const ms = Date.parse(raw.replace(" ", "T") + "Z"); + return Number.isFinite(ms) ? new Date(ms).toISOString() : undefined; + } + const ms = Date.parse(raw); + return Number.isFinite(ms) ? new Date(ms).toISOString() : undefined; +} + +// The author handle. gallery-dl nests it under `author` (the tweet's author) +// and `user` (the timeline's owner — different for a retweet). +export function xAuthor(rec: XTweetRaw): { handle: string; name?: string } { + const author = obj(rec.author) ?? obj(rec.user); + const handle = + str(author?.name) ?? + str(author?.screen_name) ?? + str(rec.author_name) ?? + str(rec.screen_name) ?? + ""; + const name = str(author?.nick) ?? str(author?.displayName) ?? str(author?.nickname); + return { handle, ...(name ? { name } : {}) }; +} + +// The tweet body. `content` is gallery-dl's field; raw X uses `full_text`, and +// a long post puts its untruncated body under a note-tweet. +export function xText(rec: XTweetRaw): string { + const note = obj(obj(obj(rec.note_tweet)?.note_tweet_results)?.result); + const noteText = str(note?.text); + if (noteText) return noteText; + return ( + str(pick(rec, ["content", "full_text", "text"])) ?? "" + ); +} + +// Expanded outbound URLs. gallery-dl flattens entities; raw X nests them. +export function xLinks(rec: XTweetRaw): string[] { + const out = new Set<string>(); + const entities = obj(rec.entities); + const urls = entities?.urls ?? (obj(rec.legacy)?.entities as Record<string, unknown> | undefined)?.urls; + if (Array.isArray(urls)) { + for (const u of urls) { + const e = obj(u); + const href = str(e?.expanded_url) ?? str(e?.url); + // Skip t.co self-links to the quoted tweet — the quote ref covers it. + if (href && !/^https?:\/\/t\.co\//.test(href)) out.add(href); + } + } + return [...out]; +} + +function refFrom( + id: string | undefined, + handle: string | undefined, +): PostRef | undefined { + if (!id) return undefined; + const ref: PostRef = { platform: "twitter", id }; + if (handle) { + ref.author = handle; + ref.url = postPermalink("twitter", handle, id); + } else { + ref.url = `https://x.com/i/status/${id}`; + } + return ref; +} + +// Normalize one tweet record into a Post. Returns null when the record has no +// resolvable id or timestamp — a shape we don't understand is skipped rather +// than stored as a corrupt entry. +export function normalizeXTweet( + rec: XTweetRaw, + channelSlug: string, +): Post | null { + const id = xIdOf(rec); + if (!id) return null; + const createdAt = xCreatedAt(rec); + if (!createdAt) return null; + + const { handle, name } = xAuthor(rec); + const text = xText(rec); + + const replyToId = + str(pick(rec, ["in_reply_to_status_id_str", "in_reply_to_tweet_id"])) ?? + (num(pick(rec, ["in_reply_to_status_id"])) !== undefined + ? String(num(pick(rec, ["in_reply_to_status_id"]))) + : undefined); + const replyToHandle = str( + pick(rec, ["in_reply_to_screen_name", "in_reply_to_user"]), + ); + + const quotedRec = obj(rec.quoted) ?? obj(rec.quoted_status); + const quotedId = quotedRec ? xIdOf(quotedRec) : str(rec.quote_id); + const quotedHandle = quotedRec ? xAuthor(quotedRec).handle : undefined; + + const retweetRec = obj(rec.retweeted_status) ?? obj(rec.retweet); + const retweetId = retweetRec ? xIdOf(retweetRec) : undefined; + const retweetHandle = retweetRec ? xAuthor(retweetRec).handle : undefined; + const isRepost = Boolean(retweetId) || rec.retweeted === true; + + // X's conversation_id IS the thread root, which is exactly our threadId. + const threadId = + str(pick(rec, ["conversation_id", "conversation_id_str"])) ?? id; + + const post: Post = { + id, + slug: `${channelSlug}/${id}`, + channelSlug, + author: handle, + createdAt, + uploadDate: uploadDateFromCreatedAt(createdAt), + text, + url: postPermalink("twitter", handle, id), + platform: "twitter", + isReply: Boolean(replyToId), + isRepost, + links: xLinks(rec), + threadId, + }; + if (name) post.authorName = name; + const lang = str(rec.lang); + if (lang && lang !== "und") post.lang = lang; + + const replyTo = refFrom(replyToId, replyToHandle); + if (replyTo) post.replyTo = replyTo; + const quoted = refFrom(quotedId, quotedHandle); + if (quoted) post.quoted = quoted; + const repostOf = refFrom(retweetId, retweetHandle); + if (repostOf) post.repostOf = repostOf; + + // Media is counted, not archived (v1). + const media = rec.media ?? obj(rec.extended_entities)?.media; + if (Array.isArray(media) && media.length > 0) post.mediaCount = media.length; + else { + const count = num(rec.count); + if (count && count > 0) post.mediaCount = count; + } + + const engagement: NonNullable<Post["engagement"]> = {}; + const likes = num(pick(rec, ["favorite_count", "like_count"])); + const reposts = num(pick(rec, ["retweet_count"])); + const replies = num(pick(rec, ["reply_count"])); + const quotes = num(pick(rec, ["quote_count"])); + if (likes !== undefined) engagement.likes = likes; + if (reposts !== undefined) engagement.reposts = reposts; + if (replies !== undefined) engagement.replies = replies; + if (quotes !== undefined) engagement.quotes = quotes; + if (Object.keys(engagement).length > 0) post.engagement = engagement; + + return post; +} + +// Normalize a batch, dropping unparseable records and deduping by id (a +// timeline page can repeat a tweet across a cursor boundary). +export function normalizeXTweets( + records: ReadonlyArray<XTweetRaw>, + channelSlug: string, +): Post[] { + const byId = new Map<string, Post>(); + for (const rec of records) { + const post = normalizeXTweet(rec, channelSlug); + if (post && !byId.has(post.id)) byId.set(post.id, post); + } + return [...byId.values()]; +} diff --git a/common/social/xPlaywrightFetcher.ts b/common/social/xPlaywrightFetcher.ts @@ -0,0 +1,251 @@ +// The X/Twitter FALLBACK fetcher: drive an authenticated Chromium and read the +// timeline from inside the page. +// +// Why it exists: gallery-dl hardcodes X's GraphQL query IDs, so it breaks every +// time X rotates them (historically every 2–4 weeks) until upstream catches up. +// A real browser is immune to that specific failure — X's own JavaScript +// computes the current query IDs, and the page's own `fetch` carries the bearer +// token, CSRF header and cookies. So this is a TRANSPORT swap, not a rewrite: +// the payload is the same tweet shape gallery-dl parses, and the `Post` +// normalizer (xNormalize.ts) is shared verbatim with the primary fetcher. +// +// Known costs, accepted deliberately: +// - Playwright Chromium's JA3/TLS fingerprint matches no real Chrome release +// (plus navigator.webdriver and CDP artifacts), so X CAN detect it. +// - Ban risk on the logged-in account is higher than an offline cookie read. +// - A browser process per fetch is slow and RAM-hungry. +// - It will NOT run in the minimal Docker build container — post fetching +// stays an editor-host concern, never a build-time one. +// Which is exactly why it is the FALLBACK: registered after x-gallery-dl, so +// URL detection prefers the cheap subprocess, and this is opted into per +// channel via `postFetcher: "x-playwright"`. +// +// Strategy: response INTERCEPTION (scroll the profile, capture UserTweets JSON) +// rather than re-issuing GraphQL by hand. It is the simpler of the two variants +// in the design and needs no knowledge of X's endpoint names beyond a substring +// match, so it survives more drift. +// +// NOT VERIFIED AGAINST LIVE X — it is tested against a local fixture page +// serving a recorded payload, never x.com, so the suite stays deterministic and +// no test run risks the account. + +import type { Post } from "../lib/posts"; +import { normalizeXTweets, type XTweetRaw } from "./xNormalize"; +import { xProfileDir } from "./xSessionBroker"; +import { getPaths } from "../lib/paths"; +import { + registerSocialFetcher, + type PostFetchInput, + type PostFetchResult, + type SocialFetcher, + type SocialFetcherProbe, +} from "./fetchers"; + +// The GraphQL operations that carry timeline tweets. Matched as substrings of +// the response URL, so a version suffix or a renamed query id still hits. +const TIMELINE_OPS = [ + "UserTweets", + "UserTweetsAndReplies", + "UserMedia", + "UserWithProfileTweetsQueryV2", +]; + +export function isTimelineResponseUrl(url: string): boolean { + return TIMELINE_OPS.some((op) => url.includes(op)); +} + +// Pull every tweet-shaped object out of a GraphQL timeline response. X nests +// them several layers deep and has moved them repeatedly, so rather than +// walking a fixed path this recurses and collects anything that looks like a +// tweet result. Tolerant by construction — the shape is a moving target. +export function collectTweetsFromGraphQL(payload: unknown): XTweetRaw[] { + const out: XTweetRaw[] = []; + const seen = new Set<unknown>(); + + const visit = (node: unknown): void => { + if (!node || typeof node !== "object") return; + if (seen.has(node)) return; + seen.add(node); + if (Array.isArray(node)) { + for (const item of node) visit(item); + return; + } + const rec = node as Record<string, unknown>; + + // A tweet result: `__typename: "Tweet"` with a legacy block, or a bare + // legacy object carrying the tweet fields. + const legacy = rec.legacy as Record<string, unknown> | undefined; + if (rec.__typename === "Tweet" && legacy && typeof legacy === "object") { + out.push(mergeTweet(rec, legacy)); + // Claim the legacy block so the bare-legacy branch below doesn't emit + // the same tweet a second time when the walk descends into it. + seen.add(legacy); + } else if ( + typeof rec.full_text === "string" && + (typeof rec.id_str === "string" || typeof rec.conversation_id_str === "string") + ) { + out.push(rec); + } + + for (const value of Object.values(rec)) visit(value); + }; + + visit(payload); + return out; +} + +// Flatten a `{ rest_id, legacy, core }` tweet result into the flat record shape +// the shared normalizer understands. +function mergeTweet( + rec: Record<string, unknown>, + legacy: Record<string, unknown>, +): XTweetRaw { + const merged: XTweetRaw = { ...legacy }; + if (typeof rec.rest_id === "string") merged.rest_id = rec.rest_id; + if (rec.note_tweet) merged.note_tweet = rec.note_tweet; + // The author lives under core.user_results.result.legacy.screen_name. + const core = rec.core as Record<string, unknown> | undefined; + const userResults = core?.user_results as Record<string, unknown> | undefined; + const userResult = userResults?.result as Record<string, unknown> | undefined; + const userLegacy = userResult?.legacy as Record<string, unknown> | undefined; + if (userLegacy) { + merged.user = { + name: userLegacy.screen_name, + nick: userLegacy.name, + }; + } + return merged; +} + +export const xPlaywrightFetcher: SocialFetcher = { + id: "x-playwright", + label: "X / Twitter (Playwright fallback)", + platform: "twitter", + fields: { limit: true }, + + // Never claims a URL by detection: gallery-dl is the primary path and is + // registered first, and this is opted into explicitly per channel. Returning + // false keeps it from ever being auto-selected. + detect(): boolean { + return false; + }, + + async probe(url: string): Promise<SocialFetcherProbe> { + // The probe would cost a full browser launch against x.com for very little + // information, and every such visit carries ban risk. The account URL alone + // is enough to configure the channel. + const handle = /(?:x|twitter)\.com\/([^/?#]+)/i.exec(url)?.[1]; + if (!handle) { + return { ok: false, error: "Could not read an X handle from that URL" }; + } + return { + ok: true, + name: handle, + handle, + url: `https://x.com/${handle}`, + }; + }, + + async fetch(input: PostFetchInput): Promise<PostFetchResult> { + const { accountUrl, channelSlug, since, seenIds, limit, signal, onLog } = input; + const paths = getPaths(); + + const { chromium } = await importPlaywright(); + onLog?.("Launching the authenticated browser profile (fallback path)."); + const context = await chromium.launchPersistentContext(xProfileDir(paths), { + headless: true, + }); + + const captured: XTweetRaw[] = []; + try { + const page = context.pages()[0] ?? (await context.newPage()); + + // Response interception: X's own JS issues the timeline queries with the + // CURRENT query ids, so nothing here needs to know them. + page.on("response", (res: { url: () => string; json: () => Promise<unknown> }) => { + const url = res.url(); + if (!isTimelineResponseUrl(url)) return; + void res + .json() + .then((body) => { + captured.push(...collectTweetsFromGraphQL(body)); + }) + .catch(() => { + /* a non-JSON or already-consumed response — ignore */ + }); + }); + + await page.goto(accountUrl, { + waitUntil: "domcontentloaded", + timeout: 60_000, + }); + + // Scroll until we stop learning anything new, hit the cap, or are + // cancelled. Bounded so a large account can't spin forever. + const MAX_SCROLLS = 60; + let lastCount = -1; + let idleRounds = 0; + for (let i = 0; i < MAX_SCROLLS; i++) { + if (signal.aborted) break; + if (limit && captured.length >= limit) break; + await page.evaluate("window.scrollBy(0, document.body.scrollHeight)"); + await page.waitForTimeout(1200); + if (captured.length === lastCount) { + if (++idleRounds >= 3) break; // timeline exhausted + } else { + idleRounds = 0; + lastCount = captured.length; + } + } + onLog?.(`Captured ${captured.length} tweet record(s) from the timeline.`); + } finally { + await context.close().catch(() => {}); + } + + // From here on it is identical to the gallery-dl path — same normalizer, + // same stop conditions. + const all = normalizeXTweets(captured, channelSlug); + const posts: Post[] = []; + for (const post of all) { + if (seenIds.has(post.id)) continue; + if (since && post.createdAt <= since) continue; + posts.push(post); + } + const capped = Boolean(limit && posts.length >= limit); + if (capped) posts.length = limit!; + return { posts, complete: !capped }; + }, +}; + +type PlaywrightPage = { + goto: (url: string, opts?: unknown) => Promise<unknown>; + evaluate: (fn: string) => Promise<unknown>; + waitForTimeout: (ms: number) => Promise<void>; + on: (event: string, cb: (res: never) => void) => void; +}; + +async function importPlaywright(): Promise<{ + chromium: { + launchPersistentContext: ( + dir: string, + opts: Record<string, unknown>, + ) => Promise<{ + pages: () => PlaywrightPage[]; + newPage: () => Promise<PlaywrightPage>; + close: () => Promise<void>; + }>; + }; +}> { + try { + // Assembled at runtime so TypeScript does not resolve it — `common` does + // not depend on Playwright (see xSessionBroker.ts for the same rationale). + const specifier = ["@playwright", "test"].join("/"); + return (await import(/* webpackIgnore: true */ specifier)) as never; + } catch { + throw new Error( + "Playwright is not available on this host — the X fallback fetcher needs it.", + ); + } +} + +registerSocialFetcher(xPlaywrightFetcher); diff --git a/common/social/xSessionBroker.ts b/common/social/xSessionBroker.ts @@ -0,0 +1,260 @@ +// The X session broker: a persistent Playwright profile that holds a logged-in +// X session and exports fresh cookies on demand. +// +// This is the highest-value half of the Playwright work, and it pays for itself +// immediately: gallery-dl's worst flaw is that X cookies expire in days, and a +// browser profile that stays logged in removes that problem WITHOUT changing +// the fetch path. The user logs in by hand once — 2FA and captcha included, +// which is precisely why this is a headed browser and not an automated login — +// and the profile then re-exports cookies whenever gallery-dl needs them. +// +// Deliberately isolated in its own module (and its own browser context) so the +// e2e harness — which is also Playwright — and the production fetcher never +// share state. Note the mild recursion: Playwright drives the test suite AND is +// a production dependency here; keeping the launch in one place is what keeps +// that untangled. +// +// Playwright is already a dependency of both editor and export with browsers +// installed, so this adds NO new dependency. It will NOT run in the minimal +// Docker build container — post fetching is an editor-host concern, never a +// build-time one. + +import path from "node:path"; +import { mkdir, readFile, rename, writeFile, rm, stat } from "node:fs/promises"; +import type { Paths } from "../lib/paths"; + +// Where the persistent browser profile lives. One profile per instance: a +// single X identity is all the archive needs. +export function xProfileDir(paths: Paths): string { + return path.join(paths.transcriptsDir, ".x-session", "profile"); +} + +// The exported cookie jar, in Netscape cookies.txt format — the format +// gallery-dl (and yt-dlp) read via `--cookies <file>`. +export function xCookieFile(paths: Paths): string { + return path.join(paths.transcriptsDir, ".x-session", "cookies.txt"); +} + +export type XSessionStatus = { + // A profile directory exists on disk. + hasProfile: boolean; + // A cookie jar has been exported. + hasCookies: boolean; + // When the cookie jar was last written. + cookiesUpdatedAt?: string; + // True when the exported jar carries the session cookie X actually + // authenticates with. A jar without it is present-but-useless. + looksAuthenticated: boolean; +}; + +// X's session cookie. Its presence is the only cheap local signal that a +// profile is actually logged in. +const AUTH_COOKIE = "auth_token"; + +export async function readXSessionStatus(paths: Paths): Promise<XSessionStatus> { + const profileDir = xProfileDir(paths); + const cookieFile = xCookieFile(paths); + const hasProfile = await exists(profileDir); + let hasCookies = false; + let cookiesUpdatedAt: string | undefined; + let looksAuthenticated = false; + try { + const st = await stat(cookieFile); + hasCookies = true; + cookiesUpdatedAt = new Date(st.mtimeMs).toISOString(); + const text = await readFile(cookieFile, "utf8"); + looksAuthenticated = text.includes(AUTH_COOKIE); + } catch { + /* no jar yet */ + } + return { hasProfile, hasCookies, cookiesUpdatedAt, looksAuthenticated }; +} + +async function exists(p: string): Promise<boolean> { + try { + await stat(p); + return true; + } catch { + return false; + } +} + +// A Playwright cookie, structurally typed so this module does not import +// @playwright/test at the top level (it must stay importable on a host that +// never installs browsers). +type BrowserCookie = { + name: string; + value: string; + domain: string; + path: string; + expires: number; // seconds since epoch; -1 for a session cookie + httpOnly: boolean; + secure: boolean; +}; + +// Serialize cookies to the Netscape cookies.txt format gallery-dl reads. +// Exported for testing — the format is fiddly (leading dot = include +// subdomains, TRUE/FALSE literals, tab separators) and getting it wrong fails +// silently as "not logged in". +export function toNetscapeCookieFile(cookies: ReadonlyArray<BrowserCookie>): string { + const lines = [ + "# Netscape HTTP Cookie File", + "# Written by the Archilyzer X session broker. Do not edit by hand.", + "", + ]; + for (const c of cookies) { + const includeSubdomains = c.domain.startsWith(".") ? "TRUE" : "FALSE"; + const expires = c.expires && c.expires > 0 ? Math.floor(c.expires) : 0; + lines.push( + [ + c.domain, + includeSubdomains, + c.path || "/", + c.secure ? "TRUE" : "FALSE", + String(expires), + c.name, + c.value, + ].join("\t"), + ); + } + return lines.join("\n") + "\n"; +} + +// Only X's own cookies are exported — the profile may hold unrelated ones and +// handing those to a subprocess would leak them for no benefit. +export function isXCookie(c: BrowserCookie): boolean { + const d = c.domain.replace(/^\./, "").toLowerCase(); + return d === "x.com" || d === "twitter.com" || d.endsWith(".x.com") || d.endsWith(".twitter.com"); +} + +// Launch a HEADED browser against the persistent profile so a human can log in +// (password, 2FA, captcha — all of it). Resolves once the caller closes the +// window, then exports the cookie jar. +// +// Playwright is imported dynamically: this module must stay importable on a +// host with no browsers installed (the Docker build image), where only the +// status/read helpers above are ever called. +export async function connectXAccount( + paths: Paths, + opts: { onLog?: (line: string) => void; timeoutMs?: number } = {}, +): Promise<XSessionStatus> { + const log = opts.onLog ?? (() => {}); + const profileDir = xProfileDir(paths); + await mkdir(profileDir, { recursive: true }); + + const { chromium } = await importPlaywright(); + log("Opening a browser window. Log in to X, then close the window."); + const context = await chromium.launchPersistentContext(profileDir, { + headless: false, + viewport: { width: 1280, height: 900 }, + }); + try { + const page = context.pages()[0] ?? (await context.newPage()); + await page.goto("https://x.com/login", { waitUntil: "domcontentloaded" }); + // Wait for the human. The window closing is the "done" signal — there is no + // reliable DOM marker for "logged in" that survives X's redesigns. + await new Promise<void>((resolve) => { + let settled = false; + const finish = () => { + if (settled) return; + settled = true; + resolve(); + }; + context.on("close", finish); + if (opts.timeoutMs) setTimeout(finish, opts.timeoutMs); + }); + const cookies = (await context.cookies()) as BrowserCookie[]; + await writeCookieJar(paths, cookies.filter(isXCookie), log); + } finally { + await context.close().catch(() => {}); + } + return readXSessionStatus(paths); +} + +// Re-export cookies from the stored profile WITHOUT any human interaction — +// this is the call gallery-dl's cookie refresh actually uses. Runs headless. +export async function refreshXCookies( + paths: Paths, + opts: { onLog?: (line: string) => void } = {}, +): Promise<XSessionStatus> { + const log = opts.onLog ?? (() => {}); + const profileDir = xProfileDir(paths); + if (!(await exists(profileDir))) { + throw new Error( + "No X session profile yet — run the headed 'Connect X account' flow first.", + ); + } + const { chromium } = await importPlaywright(); + const context = await chromium.launchPersistentContext(profileDir, { + headless: true, + }); + try { + // Touching the site lets X rotate/renew the session cookies before we read + // them, which is the whole point of keeping a live profile. + const page = context.pages()[0] ?? (await context.newPage()); + await page + .goto("https://x.com/home", { waitUntil: "domcontentloaded", timeout: 45_000 }) + .catch(() => { + log("Could not load x.com/home; exporting whatever the profile holds."); + }); + const cookies = (await context.cookies()) as BrowserCookie[]; + await writeCookieJar(paths, cookies.filter(isXCookie), log); + } finally { + await context.close().catch(() => {}); + } + return readXSessionStatus(paths); +} + +async function writeCookieJar( + paths: Paths, + cookies: BrowserCookie[], + log: (line: string) => void, +): Promise<void> { + const file = xCookieFile(paths); + await mkdir(path.dirname(file), { recursive: true }); + const body = toNetscapeCookieFile(cookies); + const tmp = `${file}.tmp-${process.pid}`; + await writeFile(tmp, body, { mode: 0o600 }); + await rename(tmp, file); + const authed = cookies.some((c) => c.name === AUTH_COOKIE); + log( + `Exported ${cookies.length} X cookie(s) to ${file}` + + (authed ? "" : " — WARNING: no auth_token, the session is not logged in"), + ); +} + +// Forget the stored session entirely (profile + exported jar). +export async function clearXSession(paths: Paths): Promise<void> { + await rm(path.dirname(xCookieFile(paths)), { recursive: true, force: true }); +} + +// Dynamic import so a host without Playwright can still import this module. +async function importPlaywright(): Promise<{ + chromium: { + launchPersistentContext: ( + dir: string, + opts: Record<string, unknown>, + ) => Promise<{ + pages: () => { goto: (u: string, o?: unknown) => Promise<unknown> }[]; + newPage: () => Promise<{ goto: (u: string, o?: unknown) => Promise<unknown> }>; + cookies: () => Promise<unknown[]>; + close: () => Promise<void>; + on: (event: string, cb: () => void) => void; + }>; + }; +}> { + try { + // The specifier is assembled at runtime so TypeScript does not try to + // resolve it: `common` deliberately does NOT depend on Playwright (the + // editor and export packages do, and they are what run this code). A + // static import would break `common`'s typecheck and its Docker install. + const specifier = ["@playwright", "test"].join("/"); + return (await import(/* webpackIgnore: true */ specifier)) as never; + } catch { + throw new Error( + "Playwright is not available on this host — the X session broker needs " + + "it (it ships with the editor; it is intentionally absent from the " + + "minimal Docker build image).", + ); + } +} diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,6 +1,9 @@ # Changelog ## [Unreleased] +- **Replace YouTube's auto-captions with transcripts of our own.** Most of the corpus rides on YouTube ASR captions, which are noticeably worse than what the transcription workers produce — no punctuation, rolling duplicate cues, `[Music]` filler — and they were *sticky*: `isVideoTranscribed()` counts any English VTT, so a video with only auto-captions was permanently invisible to every transcribe bucket and every transcribe job. There is now an opt-in, strictly-lowest-priority lane that finds those videos, downloads their audio, and transcribes them properly; whisper's `transcript.json` then wins the index pick automatically. Provenance is decided by a 4 KB sniff of the VTT itself (YouTube ASR marks ~96–100% of cues with `align:start position:N%` plus inline word timings; manual tracks mark 0%), with the 490 KB `metadata.info.json` parse kept only as a tie-breaker — so the per-regen cost is one small read per English-VTT-having video. Three snapshot buckets carry it: `autoSubsOnly` (needs audio) → `downloadedAutoSubsOnly` (needs whisper) → `supersededAutoSubs` (done, old VTT kept as a backup). Nothing is automatic by default: the auto-queue gains a per-runner **Replace YouTube auto-captions** switch that appends the bucket to the *tail* of the default union (real work always drains first), and a leaf can target the bucket by name for per-channel opt-in — the default unions are byte-identical to before, so existing setups are untouched. Manually, the Transcribe stage gains a two-step "YouTube auto-captions only" section and a single-video *Replace auto-captions* action, and each subtitle track is now labelled *YouTube auto-captions* / *manual captions*. A caption track whose provenance can't be proven machine-generated is **never** a candidate, and the original VTT is never deleted automatically — the Cleanup stage's purge button is the only thing that removes it (English ASR tracks only, honouring `do-not-clean`), so an AI-vs-YouTube comparison stays possible. See `common/lib/subtitleProvenance.ts`, `common/controller/purgeSupersededAutoSubs.ts`, `common/jobs/autoQueuePolicy.ts`, `editor/e2e/auto-subs-replace.spec.ts`. +- **Social posts are archived as a parallel corpus to video transcripts.** The archive can now ingest X/Twitter and Bluesky accounts from the same commentators and search them *together* with video transcripts — one corpus, one set of searches. A channel gains `sourceKind: "social"` (a separate axis from `handling`, so every existing `handling === "transcribe" ? … : …` branch stays binary and can never misroute), plus `postFetcher` and `socialHandle`. Posts are modelled on the live-chat layer, not the video layer: their own month-sharded JSONL on disk (`channels/<slug>/posts/YYYY-MM.jsonl` + a `posts-archive` of seen ids, so a re-run is a no-op), their own LMDB sub-DB keyed `[createdAt, channelSlug, id]` (ISO-8601 sorts chronologically, fixing the intra-day ordering the `YYYYMMDD` video key has), and their own `/posts/<slug>/{manifest,page-NNNN}.json` page tree. Ingest is a pluggable `SocialFetcher` registry (`common/social/fetchers.ts`) modelled on the transcription-app registry: **`bluesky-atproto`** (pure `fetch` against the public AT Protocol — no auth, no binary, verified end-to-end against a live account), **`x-gallery-dl`** (the primary X path: a light headless subprocess, cookies via the existing `cookiePolicy.ts`), and **`x-playwright`** (the fallback, immune to the GraphQL query-id rotations that periodically break gallery-dl). One `fetch-posts` job kind carries it, drainable and bookmarkable, routed to `platform:x.com` / `platform:bsky.app` by the existing queue keys. A social channel gets a minimal two-stage rail (Fetch → Index) instead of the six video stages, and the channel form hides every video-only control (audio format, download format, keep-source-video, extraction mode, saved-video dir). Auto-sync works unchanged — but the scheduler's *dispatch* now routes social channels to a post fetch rather than a yt-dlp video sync. See `common/lib/posts.ts`, `common/social/*`, `common/controller/fetchPosts.ts`, `editor/e2e/social-channel.spec.ts`. +- **Connect an X account once, instead of re-supplying cookies every few days.** X session cookies expire within days, which is gallery-dl's worst flaw as an archiving path. Settings gains an X-session broker: a headed "Connect X account" flow opens a browser on the editor host so the operator logs in by hand (2FA and captcha included — the login is deliberately never automated), storing a persistent browser profile. The fetcher then re-exports a fresh `cookies.txt` from that profile on demand and prefers it over `--cookies-from-browser`, so the session stops being the thing that breaks. Only X's own cookies are exported. See `common/social/xSessionBroker.ts`, `editor/app/settings/components/XSessionSection.tsx`. - **The dashboard is now a live mission-control cockpit.** The home page used to be a static SSR card stack (four stat tiles, a "needs attention" link, a bare channels table) — all the *live* operational density lived only in the opt-in `/widget` monitor. The dashboard is rebuilt as an information-first operations surface that updates in place, following the widget's proven pattern: `page.tsx` stays a server component that SSRs initial payloads and hands them to a client shell (`DashboardCockpit`) that polls the same `/api` endpoints (`/api/jobs/active`, `/api/workers`, `/api/widget/actionable`, `/api/widget/sync`). The hero is a full-width **Pipeline band** — a single instrument readout of running/queued jobs, worker-pool busy/total + pause state, sync heartbeat + scheduler on/off, the live job rows with progress bars, and the global controls (Pause transcriptions, Pause downloads, Sync all, + Add). Below it a **Needs-work** panel (top channels with ↓/✎ counts and inline Download/Transcribe actions, linking to `/actionable`) sits beside a **Quick-add / recent-changes** column, over an **enriched channels table** (relative "last sync" that ticks live, plus per-row inline Sync / Download-missing / Top-of-queue actions). Everything stays on the existing semantic "base" tokens (no new palette) and respects `prefers-reduced-motion`. The two ~1s poll hooks are extracted from the widget into `editor/app/widget/lib/usePolledPayload.ts` and the relative-time helpers into `.../relativeTime.ts`, shared by both surfaces. See `editor/app/page.tsx`, `editor/app/components/dashboard/*`, and `editor/e2e/dashboard.spec.ts`. - **URL-first channel onboarding.** The New-channel form now derives what it can from a pasted URL with **no network call** — platform (`detectPlatform`), handling (YouTube → subs, else transcribe), and a slug candidate (`@Veritasium` → `veritasium`, `/c/Some Name` → `some-name`) — and surfaces the **derived download queue** as a read-only hint (e.g. `Queue: platform:vimeo.com (new)` for a host the app doesn't recognize). An opt-in **Fetch details** button runs a lightweight single-entry yt-dlp probe (`probeChannelMeta` → `probeChannelUrlAction`) that fills the name and, crucially, lets a site **unknown to the app but known to yt-dlp** be created with its own `platform:<domain>` serial queue in the same action (no change to the closed `Platform` union). On create it now **always stores the playlist** ("Fetch playlist now", default on) so pending downloads populate immediately, with an optional **Add to top of auto-queue** that prepends a channel leaf at the head of the download policy tree and starts the runner (`prioritizeChannelDownloadAction`, also a one-click "Top of queue" on the dashboard channels table). Also fixes the missing **Kick** option in the platform select. See `editor/app/channels/{components/ChannelForm.tsx,actions.ts}`, `editor/app/auto-queue/actions.ts`, `common/ytdlp/runYtdlp.ts`, and `editor/e2e/new-channel-onboarding.spec.ts`. - **Pause is now first-class and survives restarts — for transcriptions and downloads.** Transcription pause was runtime-only (lost on restart) and buried on the Workers page; downloads had no pause at all. Both are now persisted in `settings.json` (`transcriptionsPaused` / `downloadsPaused`) and toggled from the dashboard Pipeline band (and, for parity, the monitor widget's controls). Pausing transcriptions persists the flag and re-pauses the worker pool at boot (`editor/instrumentation.ts`); pausing downloads idles the auto-download runner on its next loop iteration (re-read each tick, like the `enabled` flag) **and** makes manual download-bearing pipeline actions (`sync`, `download-from-playlist`, `download-missing`, `download-missing-subs`, `retry-bucket`) return a friendly "Downloads are paused" notice — store-playlist/enumeration stay allowed since they write no media. See `common/lib/settings.ts`, `editor/app/workers/actions.ts`, `editor/app/jobs/actions.ts`, `common/controller/autoRunner.ts`, and `editor/app/channels/[slug]/pipelineActions.ts`. diff --git a/editor/app/channels/[slug]/components/SocialChannelPanel.tsx b/editor/app/channels/[slug]/components/SocialChannelPanel.tsx @@ -0,0 +1,148 @@ +"use client"; + +// The social (posts) channel view: a minimal two-stage rail — Fetch → Index — +// instead of the six video stages and their 18 pathology buckets, none of which +// exist for an account that has no downloads, no audio and no transcripts. +// Its only failure surfaces are "the fetch failed" and "we need credentials", +// which is exactly what the state sidecar records. + +import { useTransition } from "react"; +import { fetchPostsAction } from "../socialActions"; + +export type SocialChannelState = { + postCount: number; + shards: string[]; + lastFetchedAt?: string; + lastFetchedCount?: number; + lastError?: string; + needsCookies?: boolean; + fetcherLabel: string; + handle: string; + accountUrl?: string; +}; + +export function SocialChannelPanel({ + slug, + state, +}: { + slug: string; + state: SocialChannelState; +}) { + const [pending, startTransition] = useTransition(); + + const run = (full: boolean) => + startTransition(async () => { + const res = await fetchPostsAction(slug, undefined, full); + // The job streams to the jobs page; nothing to consume here, but the + // stream must be released or its buffered chunks leak. + if (res.ok) void res.stream.cancel(); + }); + + return ( + <div className="flex flex-col gap-4"> + <section + data-social-panel="" + className="rounded-lg border border-border bg-card p-4" + > + <div className="flex flex-wrap items-baseline gap-x-3 gap-y-1"> + <h2 className="text-sm font-medium">Posts</h2> + <span className="text-xs text-muted-foreground"> + @{state.handle} · {state.fetcherLabel} + </span> + </div> + + <dl className="mt-3 grid grid-cols-2 gap-x-6 gap-y-1 text-sm sm:grid-cols-4"> + <Stat label="Archived posts" value={String(state.postCount)} /> + <Stat label="Month shards" value={String(state.shards.length)} /> + <Stat + label="Last fetch" + value={ + state.lastFetchedAt + ? new Date(state.lastFetchedAt).toLocaleString() + : "never" + } + /> + <Stat + label="Added last run" + value={ + state.lastFetchedCount === undefined + ? "—" + : String(state.lastFetchedCount) + } + /> + </dl> + + <div className="mt-4 flex flex-wrap gap-2"> + <button + type="button" + onClick={() => run(false)} + disabled={pending} + aria-label="fetch posts" + className="rounded-md bg-primary px-3 py-1 text-sm font-medium text-primary-foreground hover:opacity-90 disabled:opacity-50" + > + {pending ? "Fetching…" : "Fetch posts"} + </button> + <button + type="button" + onClick={() => run(true)} + disabled={pending} + aria-label="refetch full history" + title="Re-walk the account's whole history. Already-archived posts are deduped, so this repairs a gap rather than creating duplicates." + className="rounded-md border border-border px-3 py-1 text-sm hover:bg-muted disabled:opacity-50" + > + Re-fetch full history + </button> + </div> + + {/* The two failure surfaces a social channel actually has. */} + {state.needsCookies && ( + <p + role="status" + className="mt-3 rounded border border-warning/40 bg-warning-soft px-3 py-2 text-xs" + > + Needs credentials: the fetcher could not authenticate. Configure + cookies-from-browser for this channel (or globally) and re-run. + </p> + )} + {state.lastError && ( + <p + role="alert" + className="mt-3 rounded border border-destructive/30 bg-destructive-soft px-3 py-2 text-xs text-destructive" + > + Last fetch failed: {state.lastError} + </p> + )} + + {state.accountUrl && ( + <p className="mt-3 text-xs text-muted-foreground"> + Source:{" "} + <a + href={state.accountUrl} + target="_blank" + rel="noopener noreferrer" + className="underline" + > + {state.accountUrl} + </a> + </p> + )} + </section> + + <p className="text-xs text-muted-foreground"> + Posts are indexed into the search corpus by the normal build — there is + no per-post download, transcode or transcription stage. + </p> + </div> + ); +} + +function Stat({ label, value }: { label: string; value: string }) { + return ( + <div className="flex flex-col"> + <dt className="text-xs uppercase tracking-wide text-muted-foreground"> + {label} + </dt> + <dd className="font-mono text-sm">{value}</dd> + </div> + ); +} diff --git a/editor/app/channels/[slug]/page.tsx b/editor/app/channels/[slug]/page.tsx @@ -5,6 +5,17 @@ import type { ReactNode } from "react"; import type { Metadata } from "next"; import { notFound } from "next/navigation"; import { readChannelConfig } from "yt-dlp-transcript-common/controller/channels"; +import { isSocialChannel } from "yt-dlp-transcript-common/lib/channelConfig"; +import { + countPosts, + listPostShards, + readPostFetchState, +} from "yt-dlp-transcript-common/lib/posts-server"; +import { resolveSocialFetcher } from "yt-dlp-transcript-common/social/fetchers"; +import "yt-dlp-transcript-common/social/blueskyFetcher"; +import "yt-dlp-transcript-common/social/xGalleryDlFetcher"; +import "yt-dlp-transcript-common/social/xPlaywrightFetcher"; +import { SocialChannelPanel } from "./components/SocialChannelPanel"; import { excludedDownloadIdSet, generateChannelSnapshot, @@ -141,6 +152,57 @@ export default async function ChannelDetailPage({ const config = await readChannelConfig(paths, slug); if (!config) notFound(); + // A social (posts) channel short-circuits the whole video pipeline view: it + // has no downloads, audio, transcripts or availability, so none of the six + // stages or their pathology buckets apply. It gets a minimal Fetch → Index + // rail instead (see SocialChannelPanel). + if (isSocialChannel(config)) { + const channelRoot = path.join(paths.channelsDir, slug); + const [postCount, shards, fetchState] = await Promise.all([ + countPosts(channelRoot), + listPostShards(channelRoot), + readPostFetchState(channelRoot), + ]); + const fetcher = resolveSocialFetcher(config.postFetcher, config.url); + const socialRunningJobs = getRegistry() + .list() + .filter( + (j) => + j.channelSlug === slug && + (j.status === "running" || j.status === "queued"), + ); + return ( + <div className="flex flex-col gap-6"> + <h1 className="text-lg font-medium">{config.name ?? slug}</h1> + <RunningJobsList + jobs={socialRunningJobs.map((j) => ({ + id: j.id, + kind: j.kind, + status: j.status as "queued" | "running", + queueKey: j.queueKey, + channelSlug: j.channelSlug, + videoId: j.videoId, + }))} + hideChannelSlug + /> + <SocialChannelPanel + slug={slug} + state={{ + postCount, + shards, + lastFetchedAt: fetchState?.lastFetchedAt, + lastFetchedCount: fetchState?.lastFetchedCount, + lastError: fetchState?.lastError, + needsCookies: fetchState?.needsCookies, + fetcherLabel: fetcher?.label ?? "no fetcher configured", + handle: config.socialHandle ?? slug, + accountUrl: config.url, + }} + /> + </div> + ); + } + const update = updateChannelAction.bind(null, slug); const del = deleteChannelAction.bind(null, slug); const rename = renameChannelAction.bind(null, slug); @@ -287,7 +349,10 @@ export default async function ChannelDetailPage({ existingQueues={existingQueues} failedVideoIds={failedVideoIds} downloadedNoTranscriptIds={actionableDownloadedNoTranscriptIds} + autoSubsOnlyIds={buckets.autoSubsOnly} + downloadedAutoSubsOnlyIds={buckets.downloadedAutoSubsOnly} defaultQueueKey={TRANSCRIPTION_QUEUE} + downloadQueueKey={platformDefaultQueueKey} missingShard={transcribeMissingShard} /> ), @@ -296,6 +361,7 @@ export default async function ChannelDetailPage({ slug={slug} existingQueues={existingQueues} multipleAudioFormatIds={buckets.multipleAudioFormats} + supersededAutoSubsIds={buckets.supersededAutoSubs} foreignAudioIds={[ ...new Set([ ...buckets.untranscoded, diff --git a/editor/app/channels/[slug]/socialActions.ts b/editor/app/channels/[slug]/socialActions.ts @@ -0,0 +1,65 @@ +"use server"; + +// Server actions for social (posts) channels. Kept separate from +// pipelineActions.ts because runPipelineAction is video-shaped throughout — +// disk-space gates, download archives, yt-dlp handling overrides — none of +// which applies to a post fetch. A social channel has exactly two stages +// (Fetch → Index), and this file owns the first. + +import { revalidatePath } from "next/cache"; +import { getPaths } from "yt-dlp-transcript-common/lib/paths"; +import { getSettings } from "yt-dlp-transcript-common/lib/settings"; +import { + downloadQueueKey, + resolveQueueKey, +} from "yt-dlp-transcript-common/lib/queueKeys"; +import { readChannelConfig } from "yt-dlp-transcript-common/controller/channels"; +import { isSocialChannel } from "yt-dlp-transcript-common/lib/channelConfig"; +import { fetchPosts } from "yt-dlp-transcript-common/controller/fetchPosts"; +import { + runManagedFunction, + type StreamActionResult, +} from "yt-dlp-transcript-common/jobs/streamCommand"; + +export async function fetchPostsAction( + slug: string, + queueKey?: string, + full?: boolean, + limit?: number, +): Promise<StreamActionResult> { + const paths = getPaths(); + const config = await readChannelConfig(paths, slug); + if (!config) return { ok: false, error: `No such channel: ${slug}` }; + if (!isSocialChannel(config)) { + return { ok: false, error: `${slug} is not a social channel.` }; + } + + // queueKeyForUrl() already routes x.com / bsky.app to platform:x.com / + // platform:bsky.app, so per-platform serialization and the existing 429 + // backoff come free with no new queue plumbing. + const key = resolveQueueKey(downloadQueueKey(config), queueKey); + + return runManagedFunction({ + kind: "fetch-posts", + queueKey: key, + paths, + channelSlug: slug, + spec: { kind: "fetch-posts", slug, params: { queueKey, full, limit } }, + fn: async (onLog, signal) => { + const result = await fetchPosts({ + paths, + slug, + settings: getSettings(), + full, + limit, + onLog, + signal, + }); + revalidatePath(`/channels/${slug}`); + // Surface a failed fetch as a failed JOB (the managed wrapper turns a + // throw into status "failed"), so it shows up in the jobs list the same + // way a failed download does rather than silently logging. + if (!result.ok) throw new Error(result.error ?? "Post fetch failed"); + }, + }); +} diff --git a/editor/app/channels/actions.ts b/editor/app/channels/actions.ts @@ -10,6 +10,7 @@ import { type Platform, } from "yt-dlp-transcript-common/lib/platform"; import { probeChannelMeta } from "yt-dlp-transcript-common/ytdlp/runYtdlp"; +import { detectSocialFetcher } from "yt-dlp-transcript-common/social/fetchers"; import { channelExists, createChannel, @@ -39,10 +40,19 @@ import { planSiteMembershipWrites, } from "./lib/siteMemberships"; import { storePlaylistAction, syncAction } from "./[slug]/pipelineActions"; +import { fetchPostsAction } from "./[slug]/socialActions"; import { prioritizeChannelDownloadAction } from "../auto-queue/actions"; export type ActionResult = { error: string } | undefined; +// Import-for-side-effect, deferred to call time so the heavy fetcher modules +// never enter the module graph a client component imports. +async function registerBuiltinSocialFetchers(): Promise<void> { + await import("yt-dlp-transcript-common/social/blueskyFetcher"); + await import("yt-dlp-transcript-common/social/xGalleryDlFetcher"); + await import("yt-dlp-transcript-common/social/xPlaywrightFetcher"); +} + export type ProbeChannelResult = | { ok: true; @@ -59,6 +69,11 @@ export type ProbeChannelResult = // form hint. A successful probe on an unknown host is exactly the // "yt-dlp supports it, the app didn't know" case. queueKnown: boolean; + // Set when the URL resolved to a social account rather than a video + // channel: the form switches to its social branch and stores these. + sourceKind?: "social"; + postFetcher?: string; + socialHandle?: string; } | { ok: false; error: string }; @@ -79,6 +94,37 @@ export async function probeChannelUrlAction( } const platformDetected = detectPlatform(trimmed); const queueKey = queueKeyForUrl(trimmed); + + // Social accounts never go through yt-dlp: it cannot enumerate a timeline on + // either platform (no twitter:user extractor exists). Route them to the + // fetcher's own probe instead — for Bluesky that is a free, instant + // resolveHandle + getProfile. + // Register the built-in fetchers lazily, INSIDE the action body. A top-level + // side-effect import would put xGalleryDlFetcher (which pulls execa + node + // builtins) into this module's graph, and ChannelForm — a client component — + // imports this file for the action reference, which breaks the client bundle. + await registerBuiltinSocialFetchers(); + const socialFetcher = detectSocialFetcher(trimmed); + if (socialFetcher) { + const result = await socialFetcher.probe(trimmed); + if (!result.ok) { + return { + ok: false, + error: `${socialFetcher.label} probe failed: ${result.error ?? "unknown error"}`, + }; + } + return { + ok: true, + name: result.name ?? null, + platformDetected, + queueKey, + queueKnown: platformDetected !== null, + sourceKind: "social", + postFetcher: socialFetcher.id, + socialHandle: result.handle, + }; + } + try { const meta = await probeChannelMeta({ url: trimmed, paths: getPaths() }); return { @@ -154,6 +200,17 @@ export async function createChannelAction( /* best-effort — the channel was created regardless */ } } + // Social equivalent of "Fetch playlist now": kick off the first post fetch so + // the account's history starts archiving immediately. Same fire-and-forget + // shape — the channel exists either way. + if (config.url && formData.get("fetchPostsNow") != null) { + try { + const res = await fetchPostsAction(slug); + if (res.ok) void res.stream.cancel(); + } catch { + /* best-effort — the channel was created regardless */ + } + } // Optional "Add to top of auto-queue": prepend a channel leaf at the head of // the download policy tree and start the runner, for a channel you want // auto-downloading fast. diff --git a/editor/app/channels/components/ChannelForm.tsx b/editor/app/channels/components/ChannelForm.tsx @@ -5,7 +5,9 @@ import slugify from "@sindresorhus/slugify"; import type { ChannelConfig } from "yt-dlp-transcript-common/lib/channelConfig"; import { detectPlatform, + isSocialPlatform, queueKeyForUrl, + type Platform, } from "yt-dlp-transcript-common/lib/platform"; import { DOWNLOAD_FORMAT_LABELS, @@ -23,6 +25,7 @@ import { AUDIO_CHECK_MAX_ROLLBACKS_MIN, } from "yt-dlp-transcript-common/lib/channelConfig"; import { SYNC_INTERVAL_PRESETS } from "../../scheduler/intervalPresets"; +import { handleFromAccountUrl } from "yt-dlp-transcript-common/social/fetchers"; import { probeChannelUrlAction } from "../actions"; import { SiteMembershipsSection, @@ -118,6 +121,18 @@ export function ChannelForm({ const [handlingTouched, setHandlingTouched] = useState(false); const [platform, setPlatform] = useState<string>(c?.platform ?? ""); const [platformTouched, setPlatformTouched] = useState(false); + // Social sources are ingested into the posts corpus, not the video pipeline, + // so every video-only control below is hidden for them. In edit mode this is + // fixed by the stored config; in create mode it follows the URL/platform. + const [postFetcher, setPostFetcher] = useState(c?.postFetcher ?? ""); + const [socialHandle, setSocialHandle] = useState(c?.socialHandle ?? ""); + const [handleTouched, setHandleTouched] = useState(false); + const editSourceKind = c?.sourceKind ?? "video"; + const isSocial = isEdit + ? editSourceKind === "social" + : isSocialPlatform(platform as Platform) || + detectPlatform(url) === "twitter" || + detectPlatform(url) === "bluesky"; const [probe, setProbe] = useState<{ state: "idle" | "loading" | "done" | "error"; message?: string; @@ -138,7 +153,16 @@ export function ChannelForm({ const cand = slugCandidateFromUrl(url); if (cand) setSlug(cand); } - }, [url, isEdit, platformTouched, handlingTouched, slugTouched]); + // Social sources get their handle + fetcher offline too, from the URL + // alone — same no-network rule as platform/handling/slug. "Fetch details" + // only refines them (it can return the canonical-cased handle). + if (!handleTouched) { + setSocialHandle(handleFromAccountUrl(url) ?? ""); + } + if (detected === "twitter") setPostFetcher("x-gallery-dl"); + else if (detected === "bluesky") setPostFetcher("bluesky-atproto"); + else setPostFetcher(""); + }, [url, isEdit, platformTouched, handlingTouched, slugTouched, handleTouched]); // The derived download queue this channel will use. Offline from the URL; the // probe result (which may prove an unknown host) overrides it once fetched. @@ -164,6 +188,10 @@ export function ChannelForm({ if (!slugTouched) setSlug(slugify(res.name)); } if (!platformTouched) setPlatform(res.platformDetected ?? ""); + // A social URL resolves through the fetcher's own probe, which also + // reports which fetcher claimed it and the canonical handle. + if (res.postFetcher) setPostFetcher(res.postFetcher); + if (res.socialHandle) setSocialHandle(res.socialHandle); setProbe({ state: "done", message: res.name ?? undefined, @@ -281,6 +309,8 @@ export function ChannelForm({ <option value="odysee">Odysee</option> <option value="twitch">Twitch</option> <option value="kick">Kick</option> + <option value="twitter">X / Twitter (posts)</option> + <option value="bluesky">Bluesky (posts)</option> </select> <span className="text-xs text-muted-foreground"> Used as the default job queue, so all channels on the same @@ -344,6 +374,12 @@ export function ChannelForm({ </span> )} </label> + {isSocial ? ( + // Handling is a video-pipeline axis; a social source has no audio + // to transcribe. Emit a fixed value so the field stays present + // for parseChannelForm without offering a meaningless choice. + <input type="hidden" name="handling" value="transcribe" /> + ) : ( <fieldset className="flex flex-col gap-2"> <legend className="text-sm font-medium">Handling</legend> <label className="flex items-center gap-2 text-sm"> @@ -373,6 +409,7 @@ export function ChannelForm({ Transcribe (download audio, run whisper-cpp) </label> </fieldset> + )} <label className="flex flex-col gap-1 text-sm"> <span className="font-medium">Platform</span> <select @@ -391,6 +428,8 @@ export function ChannelForm({ <option value="odysee">Odysee</option> <option value="twitch">Twitch</option> <option value="kick">Kick</option> + <option value="twitter">X / Twitter (posts)</option> + <option value="bluesky">Bluesky (posts)</option> </select> <span className="text-xs text-muted-foreground"> Used as the default job queue, so all channels on the same @@ -402,6 +441,24 @@ export function ChannelForm({ <legend className="px-1 text-xs font-medium uppercase tracking-wide text-muted-foreground"> On create </legend> + {isSocial ? ( + <label className="flex items-start gap-2 text-sm"> + <input + type="checkbox" + name="fetchPostsNow" + defaultChecked + className="mt-1" + /> + <span className="flex flex-col gap-0.5"> + <span className="font-medium">Fetch posts now</span> + <span className="text-xs text-muted-foreground"> + Start a fetch-posts job immediately so the account&apos;s + history begins archiving. Requires a URL. + </span> + </span> + </label> + ) : ( + <> <label className="flex items-start gap-2 text-sm"> <input type="checkbox" @@ -432,11 +489,48 @@ export function ChannelForm({ </span> </span> </label> + </> + )} </fieldset> </> )} + {isSocial && ( + <> + {/* A social source is ingested into the posts corpus, so it carries + its fetcher + handle instead of any of the video controls. */} + <input type="hidden" name="sourceKind" value="social" /> + <input type="hidden" name="postFetcher" value={postFetcher} /> + <label className="flex flex-col gap-1 text-sm"> + <span className="font-medium">Account handle</span> + <input + type="text" + name="socialHandle" + value={socialHandle} + onChange={(e) => { + setSocialHandle(e.target.value); + setHandleTouched(true); + }} + readOnly={isEdit} + aria-label="account handle" + placeholder="example.bsky.social" + className="rounded border border-border bg-card px-2 py-1 text-sm" + /> + <span className="text-xs text-muted-foreground"> + The account whose posts are archived. Derived from the URL; + stored so a later URL-format change upstream can&apos;t + silently re-point ingest at a different account. + {postFetcher ? ` Fetcher: ${postFetcher}.` : ""} + </span> + </label> + </> + )} <SyncIntervalField value={c?.syncIntervalMinutes} /> </Section> + {/* Video-only sections. A social account has no audio, no download + format and no source video to retain, so they are hidden entirely + rather than shown as inert controls. */} + {!isSocial && ( + <> <Section title="Transcription"> <label className="flex flex-col gap-1 text-sm"> <span className="font-medium"> @@ -543,6 +637,8 @@ export function ChannelForm({ hint="Per-channel override for where this channel's persisted source videos live (e.g. a larger disk). Blank uses the global SAVED_VIDEOS_DIR default." /> </Section> + </> + )} <CollapsibleSection title="Advanced"> <label className="flex flex-col gap-1 text-sm"> <span className="font-medium"> diff --git a/editor/app/channels/components/parseChannelForm.ts b/editor/app/channels/components/parseChannelForm.ts @@ -15,8 +15,10 @@ import { import { PLATFORM_VALUES, detectPlatform, + isSocialPlatform, type Platform, } from "yt-dlp-transcript-common/lib/platform"; +import { handleFromAccountUrl } from "yt-dlp-transcript-common/social/fetchers"; import { isCookieMode } from "yt-dlp-transcript-common/lib/cookiePolicy"; import { isDownloadFormatPreset } from "yt-dlp-transcript-common/ytdlp/downloadFormat"; @@ -31,6 +33,9 @@ export type ParsedChannelForm = { // so a cleared input actually clears the stored value instead of inheriting // the previous one. New form fields just get added here. export const CHANNEL_FORM_FIELDS = [ + "sourceKind", + "postFetcher", + "socialHandle", "platform", "url", "audioFormat", @@ -67,6 +72,32 @@ export function parseChannelForm(formData: FormData): ParsedChannelForm { ) ? (platformRaw as Platform) : (detectPlatform(url) ?? undefined); + // Source kind. A social channel is ingested into the posts corpus instead of + // the video pipeline; the form hides every video-only control for it. An + // explicit "social" wins, otherwise infer from the platform so a social URL + // pasted into the form does the right thing without a hidden field. + const sourceKindRaw = String(formData.get("sourceKind") ?? "").trim(); + const sourceKind: ChannelConfig["sourceKind"] | undefined = + sourceKindRaw === "social" || (!sourceKindRaw && isSocialPlatform(platform)) + ? "social" + : sourceKindRaw === "video" + ? "video" + : undefined; + const postFetcher = + sourceKind === "social" ? stringOrUndef(formData, "postFetcher") : undefined; + const socialHandleRaw = + sourceKind === "social" ? stringOrUndef(formData, "socialHandle") : undefined; + const socialHandle = + socialHandleRaw?.replace(/^@/, "") || + (sourceKind === "social" && url + ? (handleFromAccountUrl(url) ?? undefined) + : undefined); + if (sourceKind === "social" && !socialHandle) { + throw new Error( + "A social source needs an account handle (could not derive one from the URL)", + ); + } + const audioFormatRaw = String(formData.get("audioFormat") ?? ""); const audioFormat: ChannelConfig["audioFormat"] | undefined = audioFormatRaw === "m4a" || @@ -215,6 +246,9 @@ export function parseChannelForm(formData: FormData): ParsedChannelForm { handling: handlingRaw, name, }; + if (sourceKind) config.sourceKind = sourceKind; + if (postFetcher) config.postFetcher = postFetcher; + if (socialHandle) config.socialHandle = socialHandle; if (platform) config.platform = platform; if (url) config.url = url; if (audioFormat) config.audioFormat = audioFormat; diff --git a/editor/app/jobs/jobReplayRegistry.ts b/editor/app/jobs/jobReplayRegistry.ts @@ -30,10 +30,13 @@ import { removeWrongFormatAudioAction, transcodeFailuresAction, transcodeUntranscodedAction, + purgeSupersededAutoSubsAction, + transcribeAutoSubsBucketAction, transcribeBucketAction, transcribeMissingAction, } from "../channels/[slug]/whisperActions"; import { persistKeptAction } from "../channels/[slug]/persistActions"; +import { fetchPostsAction } from "../channels/[slug]/socialActions"; import { redownloadIncompleteBucketAction, redownloadShortAudioBucketAction, @@ -126,6 +129,7 @@ export const JOB_REPLAY_HANDLERS: Record<string, ReplayHandler> = { str(p.handlingOverride), spec.bucket, Boolean(p.forceCookies), + Boolean(p.replaceAutoSubs), ); }, "download-from-playlist": (spec) => { @@ -150,6 +154,10 @@ export const JOB_REPLAY_HANDLERS: Record<string, ReplayHandler> = { const { queueKey } = params(spec); return syncAction(spec.slug, queueKey); }, + "fetch-posts": (spec) => { + const { p, queueKey } = params(spec); + return fetchPostsAction(spec.slug, queueKey, bool(p.full), num(p.limit)); + }, "download-missing-subs": (spec) => { const { p, queueKey } = params(spec); return downloadMissingSubsAction(spec.slug, queueKey, bool(p.abortOnError)); @@ -190,6 +198,29 @@ export const JOB_REPLAY_HANDLERS: Record<string, ReplayHandler> = { const { queueKey } = params(spec); return removeWrongFormatAudioAction(spec.slug, queueKey); }, + "whisper-bucket-auto-subs": async (spec) => { + const { p, queueKey } = params(spec); + if (!spec.bucket) return { ok: false, error: "Bookmark is missing its bucket." }; + const ids = await idsForBucket(spec.slug, spec.bucket); + if (ids.length === 0) { + return { + ok: false, + error: "No auto-caption-only videos with audio right now.", + info: true, + }; + } + return transcribeAutoSubsBucketAction( + spec.slug, + ids, + queueKey, + str(p.audioFormat) as AudioFormat | undefined, + bool(p.strictAudioFormat), + ); + }, + "purge-superseded-auto-subs": (spec) => { + const { queueKey } = params(spec); + return purgeSupersededAutoSubsAction(spec.slug, queueKey); + }, "clean-audio-transcribed": (spec) => { const { queueKey } = params(spec); return cleanAudioAction(spec.slug, queueKey); diff --git a/editor/app/scheduler/runTick.ts b/editor/app/scheduler/runTick.ts @@ -18,7 +18,9 @@ import { type SchedulerSkip, type SchedulerState, } from "yt-dlp-transcript-common/jobs/syncSchedulerState"; +import { isSocialChannel } from "yt-dlp-transcript-common/lib/channelConfig"; import { syncAction } from "../channels/[slug]/pipelineActions"; +import { fetchPostsAction } from "../channels/[slug]/socialActions"; import { checkKeptDeletedAction } from "../channels/[slug]/whisperActions"; import { backupSavedVideosAction } from "../saved-videos/backupActions"; @@ -120,8 +122,15 @@ export async function runSchedulerTick(): Promise<SchedulerTickResult> { } const queued: string[] = []; + const bySlug = new Map(channels.map((c) => [c.slug, c.config])); for (const slug of toQueue) { - const result = await syncAction(slug); + // The scheduler's ELIGIBILITY rules are source-agnostic (url + + // excludeFromSync + interval + lastSyncedAt), but the dispatch is not: a + // social channel must run a post fetch, not a yt-dlp video sync against + // its profile URL. + const result = isSocialChannel(bySlug.get(slug)) + ? await fetchPostsAction(slug) + : await syncAction(slug); if (!result.ok) { skipped.push({ slug, reason: result.error }); continue; diff --git a/editor/app/settings/components/XSessionSection.tsx b/editor/app/settings/components/XSessionSection.tsx @@ -0,0 +1,129 @@ +"use client"; + +// "Connect X account" — the operator-facing half of the X session broker. +// +// Why this exists: gallery-dl is the primary X fetcher and its worst flaw is +// that X cookies expire within days. A persistent, logged-in browser profile +// removes that problem without touching the fetch path at all — the fetcher +// just reads a jar this keeps fresh. + +import { useState, useTransition } from "react"; +import { + clearXSessionAction, + connectXAccountAction, + refreshXCookiesAction, +} from "../xSessionActions"; +import type { XSessionStatus } from "yt-dlp-transcript-common/social/xSessionBroker"; + +export function XSessionSection({ initial }: { initial: XSessionStatus }) { + const [status, setStatus] = useState<XSessionStatus>(initial); + const [error, setError] = useState<string | null>(null); + const [note, setNote] = useState<string | null>(null); + const [pending, startTransition] = useTransition(); + + const run = ( + fn: () => Promise< + { ok: true; status: XSessionStatus } | { ok: false; error: string } + >, + okNote: string, + ) => + startTransition(async () => { + setError(null); + setNote(null); + const res = await fn(); + if (res.ok) { + setStatus(res.status); + setNote(okNote); + } else { + setError(res.error); + } + }); + + return ( + <section + data-x-session="" + className="flex flex-col gap-3 rounded-lg border border-border bg-card p-4" + > + <div className="flex flex-wrap items-baseline gap-x-3"> + <h2 className="text-sm font-medium">X account session</h2> + <span + aria-label="x session state" + className={ + "text-xs " + + (status.looksAuthenticated ? "text-success" : "text-muted-foreground") + } + > + {status.looksAuthenticated + ? "connected" + : status.hasProfile + ? "profile present, not logged in" + : "not connected"} + </span> + </div> + + <p className="text-xs text-muted-foreground"> + X post archiving needs a logged-in session, and X cookies expire within + days. Connecting an account once stores a browser profile on this host; + the fetcher then re-exports fresh cookies from it automatically, so the + session stops being the thing that breaks. + </p> + + {status.cookiesUpdatedAt && ( + <p className="text-xs text-muted-foreground"> + Cookies last exported:{" "} + {new Date(status.cookiesUpdatedAt).toLocaleString()} + </p> + )} + + <div className="flex flex-wrap gap-2"> + <button + type="button" + onClick={() => + run(connectXAccountAction, "Connected. Cookies exported.") + } + disabled={pending} + aria-label="connect x account" + className="rounded-md bg-primary px-3 py-1 text-sm font-medium text-primary-foreground hover:opacity-90 disabled:opacity-50" + > + {pending ? "Working…" : "Connect X account"} + </button> + <button + type="button" + onClick={() => run(refreshXCookiesAction, "Cookies refreshed.")} + disabled={pending || !status.hasProfile} + aria-label="refresh x cookies" + className="rounded-md border border-border px-3 py-1 text-sm hover:bg-muted disabled:opacity-50" + > + Refresh cookies + </button> + <button + type="button" + onClick={() => run(clearXSessionAction, "Session cleared.")} + disabled={pending || !status.hasProfile} + aria-label="clear x session" + className="rounded-md border border-border px-3 py-1 text-sm hover:bg-muted disabled:opacity-50" + > + Forget session + </button> + </div> + + <p className="text-xs text-muted-foreground"> + Connecting opens a browser window <strong>on the machine running the + editor</strong>; complete the login (2FA and captcha included), then + close the window. The login is never automated — that is what gets + accounts flagged. + </p> + + {note && ( + <p role="status" className="text-xs text-success"> + {note} + </p> + )} + {error && ( + <p role="alert" className="text-xs text-destructive"> + {error} + </p> + )} + </section> + ); +} diff --git a/editor/app/settings/page.tsx b/editor/app/settings/page.tsx @@ -2,13 +2,15 @@ import type { Metadata } from "next"; import { getPaths } from "yt-dlp-transcript-common/lib/paths"; import { getSettings } from "yt-dlp-transcript-common/lib/settings"; import { listTranscriptionApps } from "yt-dlp-transcript-common/lib/transcriptionApps"; +import { readXSessionStatus } from "yt-dlp-transcript-common/social/xSessionBroker"; import { SettingsForm } from "./components/SettingsForm"; +import { XSessionSection } from "./components/XSessionSection"; export const dynamic = "force-dynamic"; export const metadata: Metadata = { title: "Settings" }; -export default function SettingsPage() { +export default async function SettingsPage() { const paths = getPaths(); const settings = getSettings(); @@ -22,8 +24,12 @@ export default function SettingsPage() { ["ytdlpBin", paths.ytdlpBin], ["whisperBin", paths.whisperBin], ["whisperModel", paths.whisperModel], + ["galleryDlBin", paths.galleryDlBin], ]; + // The X session broker's current state (see xSessionActions.ts). + const xSession = await readXSessionStatus(paths); + return ( <div className="flex flex-col gap-8"> <h1 className="text-2xl font-semibold">Settings</h1> @@ -40,6 +46,10 @@ export default function SettingsPage() { </section> <section className="flex flex-col gap-3 border-t border-border pt-6"> + <XSessionSection initial={xSession} /> + </section> + + <section className="flex flex-col gap-3 border-t border-border pt-6"> <details> <summary className="cursor-pointer text-lg font-semibold"> System paths diff --git a/editor/app/settings/xSessionActions.ts b/editor/app/settings/xSessionActions.ts @@ -0,0 +1,63 @@ +"use server"; + +// Server actions for the X session broker (common/social/xSessionBroker.ts). +// +// "Connect X account" opens a HEADED browser on the editor host so the operator +// can log in by hand — password, 2FA and captcha included. That is the whole +// point: automating an X login is what gets accounts flagged, and a human doing +// it once produces a profile that stays valid. +// +// Because it opens a window on the SERVER's display, this is deliberately an +// explicit operator action and never something a fetch triggers on its own. + +import { getPaths } from "yt-dlp-transcript-common/lib/paths"; +import { + clearXSession, + connectXAccount, + readXSessionStatus, + refreshXCookies, + type XSessionStatus, +} from "yt-dlp-transcript-common/social/xSessionBroker"; + +export type XSessionActionResult = + | { ok: true; status: XSessionStatus } + | { ok: false; error: string }; + +export async function xSessionStatusAction(): Promise<XSessionStatus> { + return readXSessionStatus(getPaths()); +} + +export async function connectXAccountAction(): Promise<XSessionActionResult> { + try { + const status = await connectXAccount(getPaths()); + if (!status.looksAuthenticated) { + return { + ok: false, + error: + "The browser closed without a logged-in X session (no auth_token " + + "cookie was exported). Try again and complete the login before " + + "closing the window.", + }; + } + return { ok: true, status }; + } catch (e) { + return { ok: false, error: (e as Error).message }; + } +} + +export async function refreshXCookiesAction(): Promise<XSessionActionResult> { + try { + return { ok: true, status: await refreshXCookies(getPaths()) }; + } catch (e) { + return { ok: false, error: (e as Error).message }; + } +} + +export async function clearXSessionAction(): Promise<XSessionActionResult> { + try { + await clearXSession(getPaths()); + return { ok: true, status: await readXSessionStatus(getPaths()) }; + } catch (e) { + return { ok: false, error: (e as Error).message }; + } +} diff --git a/editor/e2e/fixtures/bin/fake-gallery-dl.mjs b/editor/e2e/fixtures/bin/fake-gallery-dl.mjs @@ -0,0 +1,88 @@ +#!/usr/bin/env node +// E2E fake gallery-dl for the X/Twitter post fetcher. +// +// The real binary walks an X timeline and, with `--dump-json`, prints one JSON +// metadata object per tweet to stdout. This emits a small deterministic set in +// exactly that shape so the posts pipeline (normalize → JSONL → archive) can be +// driven end-to-end WITHOUT ever touching x.com — the suite must never depend +// on X's live defenses, and no test run may risk a real account. +// +// Deterministic by design: the same tweets every run, so a re-run proves the +// incremental (no-duplicate) guarantee. + +const args = process.argv.slice(2); + +// The account URL is the final positional argument. +const url = args[args.length - 1] ?? ""; +const handleMatch = /(?:x|twitter)\.com\/([^/?#]+)/i.exec(url); +const handle = handleMatch ? handleMatch[1] : "faketester"; + +// `--range 1-N` caps how many items the real binary emits; honour it so the +// limit path is exercised. +let limit = Infinity; +const rangeIdx = args.indexOf("--range"); +if (rangeIdx >= 0 && args[rangeIdx + 1]) { + const m = /^1-(\d+)$/.exec(args[rangeIdx + 1]); + if (m) limit = Number(m[1]); +} + +// Mirror the real failure the Needs-cookies bucket models: when the fixture is +// asked to simulate an expired session, exit non-zero with an auth-shaped +// message on stderr. +if (process.env.FAKE_GALLERY_DL_AUTH_FAIL === "1") { + process.stderr.write( + "[twitter][error] HTTP redirect to login page (401 Unauthorized)\n", + ); + process.exit(1); +} + +const AUTHOR = { name: handle, nick: "Fake Tester" }; + +// Ids are deliberately > 2^53 so any accidental numeric parse would corrupt +// them visibly rather than silently passing. +const TWEETS = [ + { + tweet_id: "1899000000000000001", + conversation_id: "1899000000000000001", + date: "2026-06-01 12:00:00", + content: "a fake tweet about zephyr archives", + lang: "en", + author: AUTHOR, + user: AUTHOR, + favorite_count: 5, + retweet_count: 1, + reply_count: 1, + quote_count: 0, + entities: { + urls: [ + { url: "https://t.co/fake", expanded_url: "https://example.com/fake" }, + ], + }, + }, + { + tweet_id: "1899000000000000002", + conversation_id: "1899000000000000001", + date: "2026-06-01 13:00:00", + content: "a fake reply in the same zephyr thread", + lang: "en", + author: AUTHOR, + user: AUTHOR, + in_reply_to_status_id_str: "1899000000000000001", + in_reply_to_screen_name: handle, + favorite_count: 2, + }, + { + tweet_id: "1899000000000000003", + conversation_id: "1899000000000000003", + date: "2026-05-20 09:30:00", + content: "an older standalone fake tweet", + lang: "en", + author: AUTHOR, + user: AUTHOR, + favorite_count: 0, + }, +]; + +for (const tweet of TWEETS.slice(0, limit)) { + process.stdout.write(JSON.stringify(tweet) + "\n"); +} diff --git a/editor/e2e/social-channel.spec.ts b/editor/e2e/social-channel.spec.ts @@ -0,0 +1,185 @@ +// Social (posts) channels end to end: create one through the real form, run a +// fetch-posts job, and assert the on-disk effect. +// +// X is NEVER contacted. The fetch runs against e2e/fixtures/bin/fake-gallery-dl.mjs +// (wired in via GALLERY_DL_BIN by `pnpm dev:test`), which prints the same +// `--dump-json` records the real binary does — so the suite stays deterministic +// and no test run can risk a real account. + +import { test, expect } from "@playwright/test"; +import { readFile } from "node:fs/promises"; +import { resetData, readJson, pathExists, resolvePath } from "./helpers"; + +type ChannelConfig = { + handling: string; + sourceKind?: string; + platform?: string; + postFetcher?: string; + socialHandle?: string; + url?: string; + name?: string; +}; + +const X_URL = "https://x.com/faketester"; +const BSKY_URL = "https://bsky.app/profile/someone.bsky.social"; + +async function readJsonl(relPath: string): Promise<Record<string, unknown>[]> { + const raw = await readFile(resolvePath(relPath), "utf8"); + return raw + .split("\n") + .filter((l) => l.trim()) + .map((l) => JSON.parse(l) as Record<string, unknown>); +} + +test("an X URL derives a social source offline and hides the video-only controls", async ({ + page, +}) => { + await resetData("empty"); + await page.goto("/channels/new"); + + await page.locator('input[name="url"]').fill(X_URL); + + // Offline autofill: detectPlatform() alone classifies the URL, no network. + await expect(page.locator('select[name="platform"]')).toHaveValue("twitter"); + // A social account has no audio to transcribe, so the video-only sections + // must not render at all (rather than showing inert controls). + await expect(page.locator('input[name="socialHandle"]')).toHaveValue("faketester"); + await expect(page.locator('select[name="audioFormat"]')).toHaveCount(0); + await expect(page.locator('select[name="downloadFormat"]')).toHaveCount(0); + await expect(page.locator('input[name="keepSourceVideo"]')).toHaveCount(0); + await expect(page.locator('input[name="savedVideosDir"]')).toHaveCount(0); + // Handling is a video axis — replaced by a fixed hidden value. + await expect(page.locator('input[name="handling"][type="radio"]')).toHaveCount(0); +}); + +test("a Bluesky URL is also recognised as a social source", async ({ page }) => { + await resetData("empty"); + await page.goto("/channels/new"); + + await page.locator('input[name="url"]').fill(BSKY_URL); + await expect(page.locator('select[name="platform"]')).toHaveValue("bluesky"); + await expect(page.locator('input[name="socialHandle"]')).toHaveValue( + "someone.bsky.social", + ); +}); + +test("a video URL still shows the video controls (no social regression)", async ({ + page, +}) => { + await resetData("empty"); + await page.goto("/channels/new"); + + await page.locator('input[name="url"]').fill("https://www.youtube.com/@Veritasium/videos"); + await expect(page.locator('select[name="platform"]')).toHaveValue("youtube"); + await expect(page.locator('select[name="audioFormat"]')).toHaveCount(1); + await expect(page.locator('input[name="handling"][type="radio"]')).toHaveCount(2); + await expect(page.locator('input[name="socialHandle"]')).toHaveCount(0); +}); + +test("creating a social channel round-trips through parseChannelConfig", async ({ + page, +}) => { + await resetData("empty"); + await page.goto("/channels/new"); + + await page.locator('input[name="url"]').fill(X_URL); + await page.locator('input[name="name"]').fill("Fake Tester"); + await page.locator('input[name="slug"]').fill("faketester"); + // Don't fetch on create — this test is about config persistence only. + await page.locator('input[name="fetchPostsNow"]').uncheck(); + await page.getByRole("button", { name: /create|add channel|save/i }).first().click(); + + await expect + .poll( + async () => + (await pathExists("test-transcripts/channels/faketester/config.json")) + ? await readJson<ChannelConfig>( + "test-transcripts/channels/faketester/config.json", + ) + : null, + { timeout: 15_000 }, + ) + .toMatchObject({ + sourceKind: "social", + platform: "twitter", + postFetcher: "x-gallery-dl", + socialHandle: "faketester", + url: X_URL, + }); +}); + +test("fetch-posts writes month-sharded JSONL + a posts-archive, and re-running is incremental", async ({ + page, +}) => { + await resetData("empty"); + await page.goto("/channels/new"); + + await page.locator('input[name="url"]').fill(X_URL); + await page.locator('input[name="name"]').fill("Fake Tester"); + await page.locator('input[name="slug"]').fill("faketester"); + // "Fetch posts now" is on by default — this is the create-time fetch path. + await page.getByRole("button", { name: /create|add channel|save/i }).first().click(); + + // The fake binary emits three tweets across two months. + await expect + .poll( + async () => + (await pathExists("test-transcripts/channels/faketester/posts-archive")) + ? ( + await readFile( + resolvePath("test-transcripts/channels/faketester/posts-archive"), + "utf8", + ) + ) + .split("\n") + .filter(Boolean).length + : 0, + { timeout: 30_000 }, + ) + .toBe(3); + + const june = await readJsonl( + "test-transcripts/channels/faketester/posts/2026-06.jsonl", + ); + const may = await readJsonl( + "test-transcripts/channels/faketester/posts/2026-05.jsonl", + ); + expect(june).toHaveLength(2); + expect(may).toHaveLength(1); + + // Ids stay strings: a 64-bit tweet id put through a JSON number would be + // silently corrupted. + expect(june[0].id).toBe("1899000000000000001"); + expect(june[0].platform).toBe("twitter"); + expect(june[0].uploadDate).toBe("20260601"); + expect(june[0].slug).toBe("faketester/1899000000000000001"); + expect(june[0].links).toEqual(["https://example.com/fake"]); + // The reply carries the thread structure. + expect(june[1].isReply).toBe(true); + expect(june[1].threadId).toBe("1899000000000000001"); + + // Re-run: the fake binary is deterministic, so every record comes back — and + // the archive must dedupe all of them rather than doubling the corpus. + await page.goto("/channels/faketester"); + await page.getByRole("button", { name: /fetch posts/i }).first().click(); + + await expect + .poll( + async () => { + const june2 = await readJsonl( + "test-transcripts/channels/faketester/posts/2026-06.jsonl", + ); + const archive = ( + await readFile( + resolvePath("test-transcripts/channels/faketester/posts-archive"), + "utf8", + ) + ) + .split("\n") + .filter(Boolean); + return { june: june2.length, archive: archive.length }; + }, + { timeout: 30_000 }, + ) + .toEqual({ june: 2, archive: 3 }); +}); diff --git a/editor/e2e/x-session.spec.ts b/editor/e2e/x-session.spec.ts @@ -0,0 +1,16 @@ +// The X session broker UI (common/social/xSessionBroker.ts). The headed login +// is never driven here — it would open a real browser and hit x.com, which the +// suite must never do. + +import { test, expect } from "@playwright/test"; +import { resetData } from "./helpers"; +test("settings page renders the X session section", async ({ page }) => { + await resetData("empty"); + await page.goto("/settings"); + await expect(page.locator("[data-x-session]")).toBeVisible(); + await expect(page.getByLabel("x session state")).toHaveText("not connected"); + // The headed login is an explicit operator action; the button exists but we + // never click it (it would open a real browser window and hit x.com). + await expect(page.getByLabel("connect x account")).toBeVisible(); + await expect(page.getByLabel("refresh x cookies")).toBeDisabled(); +}); diff --git a/editor/package.json b/editor/package.json @@ -5,8 +5,8 @@ "type": "module", "scripts": { "dev": "next dev --port ${EDITOR_PORT:-3001}", - "dev:test": "WORKER_TOKEN=test-worker-token TRANSCRIPTS_DIR=$(pwd)/test-transcripts EXPORT_PUBLIC_DIR=$(pwd)/test-transcripts/.export-public SETTINGS_FILE=$(pwd)/test-settings.json YTDLP_BIN=$(pwd)/e2e/fixtures/bin/fake-ytdlp.mjs WHISPER_BIN=$(pwd)/e2e/fixtures/bin/fake-whisper.mjs WHISPER_MODEL=/dev/null CHOUGH_BIN=$(pwd)/e2e/fixtures/bin/fake-chough.mjs CHOUGH_MODEL=/dev/null PARAKEET_STITCH_BIN=$(pwd)/e2e/fixtures/bin/fake-parakeet-stitch.mjs PARAKEET_CLI=/dev/null PARAKEET_MODEL=/dev/null FFMPEG_BIN=$(pwd)/e2e/fixtures/bin/fake-ffmpeg.mjs FFPROBE_BIN=$(pwd)/e2e/fixtures/bin/fake-ffprobe.mjs AUDIO_CHECK_INTERVAL_MS_OVERRIDE=300 AUDIO_CHECK_SIZE_GATE_OVERRIDE=4096 AUDIO_CHECK_INTERVAL_FLOOR_MS_OVERRIDE=50 AUDIO_CHECK_RECOVER_STEP_MS_OVERRIDE=100 AUDIO_CHECK_RECOVER_AFTER_OVERRIDE=2 next dev --port ${PORT:-3011}", - "start:test": "WORKER_TOKEN=test-worker-token TRANSCRIPTS_DIR=$(pwd)/test-transcripts EXPORT_PUBLIC_DIR=$(pwd)/test-transcripts/.export-public SETTINGS_FILE=$(pwd)/test-settings.json YTDLP_BIN=$(pwd)/e2e/fixtures/bin/fake-ytdlp.mjs WHISPER_BIN=$(pwd)/e2e/fixtures/bin/fake-whisper.mjs WHISPER_MODEL=/dev/null CHOUGH_BIN=$(pwd)/e2e/fixtures/bin/fake-chough.mjs CHOUGH_MODEL=/dev/null PARAKEET_STITCH_BIN=$(pwd)/e2e/fixtures/bin/fake-parakeet-stitch.mjs PARAKEET_CLI=/dev/null PARAKEET_MODEL=/dev/null FFMPEG_BIN=$(pwd)/e2e/fixtures/bin/fake-ffmpeg.mjs FFPROBE_BIN=$(pwd)/e2e/fixtures/bin/fake-ffprobe.mjs AUDIO_CHECK_INTERVAL_MS_OVERRIDE=300 AUDIO_CHECK_SIZE_GATE_OVERRIDE=4096 AUDIO_CHECK_INTERVAL_FLOOR_MS_OVERRIDE=50 AUDIO_CHECK_RECOVER_STEP_MS_OVERRIDE=100 AUDIO_CHECK_RECOVER_AFTER_OVERRIDE=2 next start --port ${PORT:-3011}", + "dev:test": "WORKER_TOKEN=test-worker-token TRANSCRIPTS_DIR=$(pwd)/test-transcripts EXPORT_PUBLIC_DIR=$(pwd)/test-transcripts/.export-public SETTINGS_FILE=$(pwd)/test-settings.json YTDLP_BIN=$(pwd)/e2e/fixtures/bin/fake-ytdlp.mjs GALLERY_DL_BIN=$(pwd)/e2e/fixtures/bin/fake-gallery-dl.mjs WHISPER_BIN=$(pwd)/e2e/fixtures/bin/fake-whisper.mjs WHISPER_MODEL=/dev/null CHOUGH_BIN=$(pwd)/e2e/fixtures/bin/fake-chough.mjs CHOUGH_MODEL=/dev/null PARAKEET_STITCH_BIN=$(pwd)/e2e/fixtures/bin/fake-parakeet-stitch.mjs PARAKEET_CLI=/dev/null PARAKEET_MODEL=/dev/null FFMPEG_BIN=$(pwd)/e2e/fixtures/bin/fake-ffmpeg.mjs FFPROBE_BIN=$(pwd)/e2e/fixtures/bin/fake-ffprobe.mjs AUDIO_CHECK_INTERVAL_MS_OVERRIDE=300 AUDIO_CHECK_SIZE_GATE_OVERRIDE=4096 AUDIO_CHECK_INTERVAL_FLOOR_MS_OVERRIDE=50 AUDIO_CHECK_RECOVER_STEP_MS_OVERRIDE=100 AUDIO_CHECK_RECOVER_AFTER_OVERRIDE=2 next dev --port ${PORT:-3011}", + "start:test": "WORKER_TOKEN=test-worker-token TRANSCRIPTS_DIR=$(pwd)/test-transcripts EXPORT_PUBLIC_DIR=$(pwd)/test-transcripts/.export-public SETTINGS_FILE=$(pwd)/test-settings.json YTDLP_BIN=$(pwd)/e2e/fixtures/bin/fake-ytdlp.mjs GALLERY_DL_BIN=$(pwd)/e2e/fixtures/bin/fake-gallery-dl.mjs WHISPER_BIN=$(pwd)/e2e/fixtures/bin/fake-whisper.mjs WHISPER_MODEL=/dev/null CHOUGH_BIN=$(pwd)/e2e/fixtures/bin/fake-chough.mjs CHOUGH_MODEL=/dev/null PARAKEET_STITCH_BIN=$(pwd)/e2e/fixtures/bin/fake-parakeet-stitch.mjs PARAKEET_CLI=/dev/null PARAKEET_MODEL=/dev/null FFMPEG_BIN=$(pwd)/e2e/fixtures/bin/fake-ffmpeg.mjs FFPROBE_BIN=$(pwd)/e2e/fixtures/bin/fake-ffprobe.mjs AUDIO_CHECK_INTERVAL_MS_OVERRIDE=300 AUDIO_CHECK_SIZE_GATE_OVERRIDE=4096 AUDIO_CHECK_INTERVAL_FLOOR_MS_OVERRIDE=50 AUDIO_CHECK_RECOVER_STEP_MS_OVERRIDE=100 AUDIO_CHECK_RECOVER_AFTER_OVERRIDE=2 next start --port ${PORT:-3011}", "build": "next build", "start": "next start --port ${EDITOR_PORT:-3001}", "lint": "eslint", diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md @@ -1,5 +1,10 @@ # Changelog +## [Unreleased] +- **One search now covers video transcripts *and* social posts.** Archived X/Twitter and Bluesky posts ship as a parallel corpus beside transcripts and live chat, and compose into the same boolean query tree — so `(transcripts:"foo" OR posts:"foo")` returns both kinds in one ranked, newest-first result set. `LayerScope` gains `"posts"` (whitelisted in `qt=` deserialization, so a shared link round-trips a posts leaf), the leaf scope selector gains **Posts**, and the filter row gains a **Posts** media kind beside Videos and Livestreams — a third kind, because a post is neither, and folding it into the video toggle would silently drop the whole corpus. Post and video slugs live in disjoint namespaces, partitioned per-leaf by the eval engine so a transcripts leaf never fetches a post and an AND across the two can't collapse to nothing. Post result cards drop what doesn't apply (no seek gutter, no livestream/age badges, no VOD expiry) and lead with the post body; opening one shows a new **PostModal** — a sibling of the transcript reader, not a generalization of it — with the post, its archived thread, its outbound links and its engagement counts. Date filters work unchanged: every post carries a derived `uploadDate`. Posts are cached and served under `/posts/`, offline-cached by the service worker, and CORS-readable so a federating hub merges them across origins. +- **"Ask AI" is grounded in posts as well as transcripts.** Retrieval adds a posts leaf per keyword alongside the transcript and metadata leaves, sharing the same term key so ranking still counts a keyword once rather than three times. Post excerpts render without a `[clock]` line and are cited as a bare `[n]` (never `[n @ mm:ss]` — a post has no timeline), the source list shows a date instead of a meaningless 0:00 seek button, and "load more context" on a post returns its **thread** rather than a time window. +- **MCP: posts are first-class.** `search_transcripts` gains `content_types` (defaulting to **both**, so existing agent flows pick posts up automatically) and renders post hits without moment links; new `get_post` and `get_thread` tools read one post or a whole conversation; `open_link` accepts a `posts` query scope; and the sweep prompt teaches the post citation form. All of it lands on the single `ShardSource` boundary, so local, remote and hub transports gain it at once. `corpus.json` bumps to spec 2 with a `postScheme` describing the new shards. + ## [0.8.2] - 2026-07-24 - **Admin: the editor dashboard is now a live mission-control cockpit, with URL-first channel onboarding and restart-surviving pause.** An operations-side change (it doesn't affect the published site): the editor's home page becomes a live, information-first operations surface — a full-width Pipeline band (running jobs, worker occupancy, sync heartbeat, global pause/sync controls) over a Needs-work / Quick-add split and an enriched channels table, all polling the same endpoints the monitor widget uses. Adding a channel is now URL-first (offline platform/handling/slug autofill + an opt-in yt-dlp "Fetch details" probe that can register an unknown platform's own queue, always store the playlist, and optionally jump the auto-queue), and pause is first-class for both transcriptions and downloads, persisted in `settings.json`. See the editor changelog for detail. - **"Ask AI" — the Report panel's citations are now clickable, precise to the exact moment cited.** The chat's answer bubbles already linked their `[1]`, `[2]`… markers to the cited transcript; the **Report** panel — the persistent document Report mode and whole-corpus sweeps maintain — rendered as plain text, so the citations a user works from most were dead. Now the report mirrors the chat: it renders each citation as an in-page link and lists its **numbered sources** beneath the document, and clicking one **opens the transcript modal seeked to the cited line**. Citations are also upgraded — in the report *and* the chat — to a precise-moment **`[n @ mm:ss]`** form (the source's number plus the specific line's timestamp, e.g. `[3 @ 12:34]`), so a citation jumps to the exact moment instead of the video's first matched snippet. A slightly-off model timestamp is **snapped to the nearest real transcript line**, and it stays graceful: a bare `[n]` still resolves to the first snippet, and an out-of-range number or unparseable time stays plain text rather than becoming a broken link. Making this work in the report (a single accumulated string, unlike a chat message that carries its own sources) required two new pieces: a **persisted report source registry** — every video the report cites, deduped and stored light (no excerpt text) — and **global citation numbering** so a report's `[n]` is stable across every batch/turn folded in, not renumbered per write. The registry is saved with the conversation and each saved chat, so a report's clickable citations survive a reload; New chat clears it; and on the federated hub /ask (no player) citations fall back to scroll-to-source, exactly like the chat. See `export/app/ask/citations.tsx` (new shared renderer — `linkifyCitations` + `CitationLink` + `CitationSources`, with `[n @ mm:ss]` parsing and nearest-snippet snapping), `export/app/ask/{MessageBubble,ReportPanel,AskChat,useAskChat,askChatStorage}.tsx/ts`, `export/app/lib/{askRetrieval,askConversation,searchAgent}.ts` (optional/global-registry numbering + `mergeReportSources` + the `[n @ mm:ss]` prompt updates), and `export/e2e/ask-chat.spec.ts`. diff --git a/export/app/(workspace)/SiteWorkspace.tsx b/export/app/(workspace)/SiteWorkspace.tsx @@ -12,6 +12,7 @@ import { PlayerProvider } from "yt-dlp-transcript-common/components/PlayerProvider"; import TranscriptModal from "yt-dlp-transcript-common/components/TranscriptModal"; +import PostModal from "yt-dlp-transcript-common/components/PostModal"; import { SingleSiteDataProvider } from "yt-dlp-transcript-common/components/SearchDataContext"; import { SearchSessionProvider } from "yt-dlp-transcript-common/components/SearchSessionContext"; import WorkspaceSearchBar from "yt-dlp-transcript-common/components/WorkspaceSearchBar"; @@ -36,6 +37,8 @@ export default function SiteWorkspace({ </SearchSessionProvider> </SingleSiteDataProvider> <TranscriptModal /> + {/* Sibling viewer for the social-post corpus (see PostModal.tsx). */} + <PostModal /> </PlayerProvider> ); } diff --git a/export/app/ask/citations.tsx b/export/app/ask/citations.tsx @@ -191,7 +191,15 @@ export function CitationSources({ )}{" "} <span className="text-muted-foreground/70"> — {s.channel} - {s.siteTitle ? ` · ${s.siteTitle}` : ""} ·{" "} + {s.siteTitle ? ` · ${s.siteTitle}` : ""} + {/* A post has no timeline, so the per-snippet moment button would + render a meaningless 0:00. Show the date instead and let the + source title itself be the link. */} + {s.isPost ? ( + <> · {s.uploadDate}</> + ) : ( + <> + {" · "} {s.snippets.map((sn, sni) => ( <span key={sn.seconds}> {sni > 0 ? ", " : ""} @@ -209,6 +217,8 @@ export function CitationSources({ )} </span> ))} + </> + )} </span> </li> ))} diff --git a/export/app/ask/useAskChat.ts b/export/app/ask/useAskChat.ts @@ -5,6 +5,8 @@ import { useSearchData } from "yt-dlp-transcript-common/components/SearchDataCon import { useSearchSession } from "yt-dlp-transcript-common/components/SearchSessionContext"; import type { SearchHandoff } from "yt-dlp-transcript-common/lib/aiHandoff"; import { fetchTranscript } from "yt-dlp-transcript-common/components/transcriptCache"; +import { useChannelPostsManifests } from "yt-dlp-transcript-common/components/postsCache"; +import { makeId, splitId } from "yt-dlp-transcript-common/components/originId"; import { cuesToSnippets, mergeSnippets, @@ -144,10 +146,38 @@ export function useAskChat() { const [searchMode, setSearchModeState] = useState<AgentMode>("auto"); const [markdownOn, setMarkdownOnState] = useState(true); - const { summariesState, channels, aliases } = useSearchData(); + const { summariesState, channels, aliases, postsManifest } = useSearchData(); const { summaries, summariesReady } = summariesState; const corpusError = summariesState.error?.message ?? null; + // The social-post corpus, so the chat's retrieval is grounded in BOTH + // datasets. Flattened from the per-channel posts manifests into the slug set + // the search engine scopes posts leaves to. Empty on a video-only site. + const postsRefs = useMemo( + () => + postsManifest + ? postsManifest.channels.map((c) => { + const { origin, slug } = splitId(c.slug); + return { origin, channelSlug: slug }; + }) + : [], + [postsManifest], + ); + const channelPostsQueries = useChannelPostsManifests(postsRefs); + const postScopeSlugs = useMemo<Set<string>>(() => { + const set = new Set<string>(); + for (let i = 0; i < channelPostsQueries.length; i++) { + const data = channelPostsQueries[i]?.data; + const ref = postsRefs[i]; + if (!data || !ref) continue; + for (const id of Object.keys(data.slugToPage)) { + set.add(makeId(ref.origin, `${ref.channelSlug}/${id}`)); + } + } + return set; + // eslint-disable-next-line react-hooks/exhaustive-deps + }, [postsRefs, channelPostsQueries.map((q) => (q.data ? 1 : 0)).join("")]); + // The live "active search" this workspace shares — the chat auto-grounds in it. // `activeGrounding` is the grounding for the ACTIVE mode (whole search ⇄ // selection); `getGrounding(target)` lazily builds a FULL-excerpt handoff over @@ -740,6 +770,7 @@ export function useAskChat() { question, prior, summaries, + postScopeSlugs, aliases, signal: ac.signal, onEvent, @@ -806,7 +837,17 @@ export function useAskChat() { abortRef.current = null; } }, - [provider, apiKey, model, searchMode, summaries, aliases, reportModeAvailable, pushDebug], + [ + provider, + apiKey, + model, + searchMode, + summaries, + postScopeSlugs, + aliases, + reportModeAvailable, + pushDebug, + ], ); const send = useCallback(async () => { diff --git a/export/app/components/hub/HubHome.tsx b/export/app/components/hub/HubHome.tsx @@ -8,6 +8,7 @@ import { useMemo } from "react"; import { PlayerProvider } from "yt-dlp-transcript-common/components/PlayerProvider"; import TranscriptModal from "yt-dlp-transcript-common/components/TranscriptModal"; +import PostModal from "yt-dlp-transcript-common/components/PostModal"; import TranscriptSearch from "yt-dlp-transcript-common/components/TranscriptSearch"; import { MultiSiteDataProvider, @@ -43,6 +44,8 @@ export default function HubHome() { </MultiSiteDataProvider> </div> <TranscriptModal /> + {/* Sibling viewer for the social-post corpus (see PostModal.tsx). */} + <PostModal /> </PlayerProvider> ); } diff --git a/export/app/duplicates/page.tsx b/export/app/duplicates/page.tsx @@ -1,6 +1,7 @@ import type { Metadata } from "next"; import { PlayerProvider } from "yt-dlp-transcript-common/components/PlayerProvider"; import TranscriptModal from "yt-dlp-transcript-common/components/TranscriptModal"; +import PostModal from "yt-dlp-transcript-common/components/PostModal"; import { DuplicatesClient } from "./DuplicatesClient"; export const metadata: Metadata = { title: "Duplicates" }; @@ -20,6 +21,8 @@ export default function DuplicatesPage() { <DuplicatesClient /> </div> <TranscriptModal /> + {/* Sibling viewer for the social-post corpus (see PostModal.tsx). */} + <PostModal /> </PlayerProvider> ); } diff --git a/export/app/lib/askConversation.ts b/export/app/lib/askConversation.ts @@ -420,14 +420,17 @@ export function gatherSystemPrompt( // the noise of every enabled alias. export function answerSystemPrompt(usedAliases: SearchAlias[]): string { const base = - "You are answering a question about a video-transcript archive. Base your " + - "answer on the transcript excerpts provided in this conversation. Excerpts " + + "You are answering a question about an archive of video transcripts AND " + + "social posts (X/Twitter, Bluesky) from the same commentators — one corpus, " + + "two kinds of source. Base your " + + "answer on the excerpts provided in this conversation. Excerpts " + "may appear in earlier turns, and a follow-up request (reformatting, " + "summarising, expanding) should reuse the relevant excerpts already " + "provided. Cite each excerpt you use as [n @ mm:ss]: its bracketed number " + "plus the specific line's timestamp (e.g. [1 @ 4:12]); a bare [n] is fine " + - "when no single line applies. Each excerpt shows a video title, channel, and " + - "timestamped lines. If the " + + "when no single line applies. A video excerpt shows a title, channel, and " + + "timestamped lines. A POST excerpt is marked `post by <author>` and has NO " + + "timestamps — always cite a post as a bare [n], never with @ mm:ss. If the " + "excerpts do not contain enough to answer, say so plainly rather than " + "guessing. Format your answer in GitHub-flavored Markdown."; const context = renderUsedAliasContext(usedAliases); diff --git a/export/app/lib/askRetrieval.ts b/export/app/lib/askRetrieval.ts @@ -13,6 +13,7 @@ import { type SearchAlias, } from "yt-dlp-transcript-common/lib/searchAliases"; import type { LayerHit } from "yt-dlp-transcript-common/components/searchPipeline"; +import { peekPost } from "yt-dlp-transcript-common/components/postsCache"; import type { DisplaySummary } from "yt-dlp-transcript-common/lib/transcripts"; import { formatDuration } from "yt-dlp-transcript-common/lib/format"; @@ -35,6 +36,9 @@ export type RetrievedVideo = { uploadDate: string; url?: string; snippets: { clock: string; seconds: number; text: string }[]; + // True for an entry from the social-post corpus. A post has no timeline, so + // its context lines carry no [clock] and its citations render no seek button. + isPost?: boolean; }; // Small English stopword set so common question words ("what", "does", "about") @@ -135,7 +139,36 @@ export function rankResults( const scored: { v: RetrievedVideo; score: number }[] = []; for (const slug of progress.slugs) { const summary = bySlug.get(slug); - if (!summary) continue; + if (!summary) { + // Not a video — it may be a post (a disjoint slug namespace). Posts are + // resolved from the posts page cache the retrieval pass already warmed. + const post = peekPost(slug); + if (!post) continue; + const hits = progress.hits.get(slug) ?? []; + const coverage = new Set(hits.map((h) => termKey(h.leafId))).size; + scored.push({ + score: coverage * 1000 + Math.min(hits.length, 500), + v: { + key: slug, + videoId: post.id, + title: post.text.slice(0, 120), + channel: post.authorName || post.author, + uploadDate: post.uploadDate, + url: post.url, + isPost: true, + // One snippet: the post body. It is already short, and the pipeline + // truncates hit text to a ±80-char window. + snippets: [ + { + clock: "", + seconds: 0, + text: post.text.trim().replace(/\s+/g, " ").slice(0, 480), + }, + ], + }, + }); + continue; + } const hits = progress.hits.get(slug) ?? []; const coverage = new Set(hits.map((h) => termKey(h.leafId))).size; const score = coverage * 1000 + Math.min(hits.length, 500); @@ -176,6 +209,9 @@ export type RetrieveOptions = RankOptions & { // query catches known AI-transcription misspellings (e.g. a name spelled // several ways). Empty/omitted = plain substring leaves, as before. aliases?: SearchAlias[]; + // Every post slug available to search (from the posts manifests). Omitted = + // no posts corpus, and every posts leaf resolves empty. + postScopeSlugs?: ReadonlySet<string> | null; // Total hit cap. Bounds how many transcript pages get fetched — the engine // stops once this many hits accumulate. NOT Infinity (that would scan the // whole ~800MB corpus). Capped results still work; they just aren't persisted @@ -214,6 +250,17 @@ export function buildSearchRoot( firedT.forEach((a, i) => children.push( newLeaf({ + id: `p#a${i}`, + query: a.suggestion, + scope: "posts", + useRegex: a.useRegex, + contributeHits: true, + }), + ), + ); + firedT.forEach((a, i) => + children.push( + newLeaf({ id: `t#a${i}`, query: a.suggestion, scope: "transcripts", @@ -238,6 +285,11 @@ export function buildSearchRoot( if (covered.has(kw)) return; children.push(newLeaf({ id: `t#${i}`, query: kw, scope: "transcripts", contributeHits: true })); children.push(newLeaf({ id: `m#${i}`, query: kw, scope: "metadata", contributeHits: true })); + // The social-post corpus is searched alongside transcripts, so the chat is + // grounded in BOTH datasets. The `p#<i>` id shares the term suffix with + // `t#<i>`/`m#<i>`, so rankResults' distinct-term coverage collapses all + // three scopes into one unit rather than triple-counting a keyword. + children.push(newLeaf({ id: `p#${i}`, query: kw, scope: "posts", contributeHits: true })); }); return { root: newGroup({ op: "OR", children }), firedT, firedM }; @@ -290,11 +342,19 @@ export function retrieve( fn(); }; + // The scope universe spans BOTH corpora: video slugs from summaries plus + // post slugs from the caller-supplied posts scope. searchEval partitions + // them per-leaf so a transcripts leaf never fetches a post and vice versa. + const postSlugs = opts.postScopeSlugs ?? null; const controller = runQueryTree({ root, - globalScope: summaries.map((s) => s.slug), + globalScope: [ + ...summaries.map((s) => s.slug), + ...(postSlugs ? Array.from(postSlugs) : []), + ], summaries, chatScopeSlugs: null, + postScopeSlugs: postSlugs, initialHitLimit: opts.hitLimit ?? 400, concurrency: 6, flushIntervalMs: 120, @@ -336,9 +396,13 @@ export function buildContext( return videos .map((v, i) => { const n = numberOf ? numberOf(v, i) : i + 1; - const head = `[${n}] "${v.title}" — ${v.channel}${v.siteTitle ? ` (${v.siteTitle})` : ""}`; + const head = v.isPost + ? `[${n}] post by ${v.channel}${v.siteTitle ? ` (${v.siteTitle})` : ""} — ${v.uploadDate}` + : `[${n}] "${v.title}" — ${v.channel}${v.siteTitle ? ` (${v.siteTitle})` : ""}`; + // A post has no timeline — a "[0:00]" prefix would be noise the model + // could mistake for a real timestamp and cite. const lines = v.snippets - .map((s) => ` [${s.clock}] ${s.text}`) + .map((s) => (v.isPost ? ` ${s.text}` : ` [${s.clock}] ${s.text}`)) .join("\n"); return `${head}\n${lines}`; }) diff --git a/export/app/lib/searchAgent.ts b/export/app/lib/searchAgent.ts @@ -42,6 +42,7 @@ import { } from "yt-dlp-transcript-common/lib/searchAliases"; import type { DisplaySummary } from "yt-dlp-transcript-common/lib/transcripts"; import { fetchTranscript } from "yt-dlp-transcript-common/components/transcriptCache"; +import { fetchThread } from "yt-dlp-transcript-common/components/postsCache"; import { cuesToSnippets, mergeSnippets, @@ -228,6 +229,9 @@ export type RunAskTurnOptions = { question: string; prior: UiMessage[]; summaries: DisplaySummary[]; + // Post slugs available to search, so the agent's retrieval covers the social + // corpus alongside video transcripts. Omitted = video-only, as before. + postScopeSlugs?: ReadonlySet<string> | null; aliases: SearchAlias[]; budget?: number; signal?: AbortSignal; @@ -431,6 +435,7 @@ export async function runAskTurn( const r = await retrieve({ question: query, summaries, + postScopeSlugs: opts.postScopeSlugs ?? null, aliases, signal, limit: PER_SEARCH_LIMIT, @@ -460,6 +465,34 @@ export async function runAskTurn( const ref = rawRef.trim(); const v = videos.get(ref); if (!v) return `No pinned video with ref "${ref}".`; + // For a POST, "more context" means the THREAD (parent + replies) — the + // natural analogue of a transcript's surrounding cue window, since a post + // has no timeline to window over. + if (v.isPost) { + onEvent({ type: "fetch_start", ref: v.key, label: v.title }); + let thread; + try { + thread = await fetchThread(v.key); + } catch { + onEvent({ type: "fetch_done", ref: v.key, count: 0 }); + return `Couldn't load the thread for this post.`; + } + const snips = thread.map((p) => ({ + clock: "", + seconds: 0, + text: `${p.authorName || p.author}: ${p.text.trim().replace(/\s+/g, " ")}`, + })); + const before = v.snippets.length; + v.snippets = mergeSnippets(v.snippets, snips, SNIPPETS_PER_VIDEO_CAP); + const added = v.snippets.length - before; + onEvent({ type: "fetch_done", ref: v.key, count: snips.length }); + if (snips.length === 0) return `No thread found for this post.`; + const body = snips.map((s) => s.text).join("\n"); + return ( + `Thread around the post by ${v.channel} ` + + `(${added} new post${added === 1 ? "" : "s"} added to your citable excerpts):\n${body}` + ); + } const center = typeof aroundSeconds === "number" ? aroundSeconds diff --git a/export/e2e/fixtures/data.ts b/export/e2e/fixtures/data.ts @@ -3,6 +3,14 @@ import { writeFileSync } from "node:fs"; export const CHANNEL = "Test Channel"; export const CHANNEL_SLUG = "test-channel"; +// A social channel + its posts, the parallel corpus to the video fixtures +// above. Kept on its own channel so channel-selection behaviour is separable. +export const POST_CHANNEL = "Test Social"; +export const POST_CHANNEL_SLUG = "test-social"; +export const POST_ROOT_ID = "post-root-1"; +export const POST_REPLY_ID = "post-reply-1"; +export const POST_OTHER_ID = "post-other-1"; + export const VIDEO_TRANSCRIPT_ONLY = "vid-transcript-only"; export const VIDEO_CHAT_SMALL = "vid-chat-small"; export const VIDEO_CHAT_LARGE = "vid-chat-large"; @@ -324,6 +332,84 @@ export function subsPage() { ]; } +// ─── Social-post corpus fixtures ─── +// "alpha" appears in BOTH the transcript cues and a post body, so a combined +// (transcripts OR posts) query is provably returning results from both corpora. + +function makePost( + id: string, + text: string, + extra: Record<string, unknown> = {}, +) { + return { + id, + slug: `${POST_CHANNEL_SLUG}/${id}`, + channelSlug: POST_CHANNEL_SLUG, + author: "tester.bsky.social", + authorName: "Tester", + createdAt: "2026-02-03T10:00:00.000Z", + uploadDate: "20260203", + text, + url: `https://bsky.app/profile/tester.bsky.social/post/${id}`, + platform: "bluesky" as const, + isReply: false, + isRepost: false, + links: [] as string[], + threadId: id, + ...extra, + }; +} + +export function postsManifest() { + return { + version: 1, + channels: [ + { + name: POST_CHANNEL, + slug: POST_CHANNEL_SLUG, + postCount: 3, + platform: "bluesky" as const, + }, + ], + totalCount: 3, + generatedAt: new Date().toISOString(), + }; +} + +export function channelPostsManifest() { + return { + version: 1, + channelSlug: POST_CHANNEL_SLUG, + pageCount: 1, + maxPageBytes: 8388608, + generatedAt: new Date().toISOString(), + slugToPage: { + [POST_ROOT_ID]: 0, + [POST_REPLY_ID]: 0, + [POST_OTHER_ID]: 0, + }, + }; +} + +export function postsPage() { + return [ + makePost(POST_ROOT_ID, "a post about alpha things", { + links: ["https://example.com/linked"], + engagement: { likes: 12, reposts: 3, replies: 1 }, + }), + makePost(POST_REPLY_ID, "replying about alpha again", { + isReply: true, + threadId: POST_ROOT_ID, + replyTo: { platform: "bluesky", id: POST_ROOT_ID }, + createdAt: "2026-02-03T11:00:00.000Z", + }), + makePost(POST_OTHER_ID, "an unrelated omega post", { + createdAt: "2026-02-02T09:00:00.000Z", + uploadDate: "20260202", + }), + ]; +} + // ─── Duplicate-shorts feature fixtures ─── // One site-filtered cluster of two members whose slugs resolve via the mocked // transcript routes, so a member click opens the real transcript modal. The diff --git a/export/e2e/helpers.ts b/export/e2e/helpers.ts @@ -7,6 +7,9 @@ import { statsPage, subsManifest, subsPage, + postsManifest, + channelPostsManifest, + postsPage, summaries, summariesManifest, transcriptPage, @@ -54,6 +57,15 @@ export async function installRoutes(page: Page) { await page.route(/\/subs\/[^/]+\/page-\d+\.json$/, async (route) => { await fulfillJson(route, subsPage()); }); + await page.route("**/posts/manifest.json", async (route) => { + await fulfillJson(route, postsManifest()); + }); + await page.route(/\/posts\/[^/]+\/manifest\.json$/, async (route) => { + await fulfillJson(route, channelPostsManifest()); + }); + await page.route(/\/posts\/[^/]+\/page-\d+\.json$/, async (route) => { + await fulfillJson(route, postsPage()); + }); // Search-alias dictionary — empty by default; alias-suggestion.spec overrides // this with a populated list. Kept here so other search tests get a clean // intercept instead of a real 404. diff --git a/export/e2e/posts-search.spec.ts b/export/e2e/posts-search.spec.ts @@ -0,0 +1,187 @@ +import { expect, test, type Page } from "@playwright/test"; +import { + CHANNEL_SLUG, + POST_CHANNEL_SLUG, + POST_OTHER_ID, + POST_REPLY_ID, + POST_ROOT_ID, + VIDEO_CHAT_LARGE, + VIDEO_CHAT_SMALL, + VIDEO_TRANSCRIPT_ONLY, +} from "./fixtures/data"; +import { installRoutes } from "./helpers"; + +// The social-post corpus as a PARALLEL dataset to video transcripts: one +// search, one result set, with a toggle-able mode. Seeded through `?qt=` +// (the composite query tree) exactly like query-tree.spec.ts. + +type SLeaf = { + k: "l"; + i?: string; + q: string; + s: "transcripts" | "chat" | "posts" | "metadata" | "description" | "tags"; + r?: 1; + h?: 0; + n?: 1; +}; +type SGroup = { + k: "g"; + i?: string; + o: "AND" | "OR"; + n?: 1; + c: (SLeaf | SGroup)[]; +}; + +function qt(root: SGroup): string { + return encodeURIComponent(JSON.stringify(root)); +} + +const TRANSCRIPT_ONLY_SLUG = `${CHANNEL_SLUG}/${VIDEO_TRANSCRIPT_ONLY}`; +const CHAT_SMALL_SLUG = `${CHANNEL_SLUG}/${VIDEO_CHAT_SMALL}`; +const CHAT_LARGE_SLUG = `${CHANNEL_SLUG}/${VIDEO_CHAT_LARGE}`; +const POST_ROOT_SLUG = `${POST_CHANNEL_SLUG}/${POST_ROOT_ID}`; +const POST_REPLY_SLUG = `${POST_CHANNEL_SLUG}/${POST_REPLY_ID}`; +const POST_OTHER_SLUG = `${POST_CHANNEL_SLUG}/${POST_OTHER_ID}`; + +async function expectResultSlugs(page: Page, slugs: string[]) { + const cards = page.locator("[data-card-header]"); + await expect(async () => { + const got = await cards.evaluateAll((els) => + els.map((e) => e.getAttribute("data-result-slug") ?? ""), + ); + expect(got.slice().sort()).toEqual(slugs.slice().sort()); + }).toPass({ timeout: 15_000 }); +} + +test.describe("social-post corpus — search", () => { + test.beforeEach(async ({ page }) => { + await installRoutes(page); + }); + + test("a posts-scope leaf returns post hits", async ({ page }) => { + const tree: SGroup = { + k: "g", + o: "AND", + c: [{ k: "l", q: "alpha", s: "posts" }], + }; + await page.goto(`/?qt=${qt(tree)}`); + await expectResultSlugs(page, [POST_ROOT_SLUG, POST_REPLY_SLUG]); + // The leaf section inside the card is labelled for the posts layer + // (scoped to the results list — "Posts" also appears in the scope select). + await expect( + page.locator("[data-leaf-section]").getByText("Posts", { exact: true }).first(), + ).toBeVisible(); + }); + + test("a combined (transcripts OR posts) query returns both kinds in one result set", async ({ + page, + }) => { + // This is the headline requirement: one search covering transcripts AND + // posts seamlessly, which falls out of the existing tree algebra. + const tree: SGroup = { + k: "g", + o: "OR", + c: [ + { k: "l", q: "alpha", s: "transcripts" }, + { k: "l", q: "alpha", s: "posts" }, + ], + }; + await page.goto(`/?qt=${qt(tree)}`); + await expectResultSlugs(page, [ + TRANSCRIPT_ONLY_SLUG, + CHAT_SMALL_SLUG, + CHAT_LARGE_SLUG, + POST_ROOT_SLUG, + POST_REPLY_SLUG, + ]); + }); + + test("post cards render no seek button", async ({ page }) => { + const tree: SGroup = { + k: "g", + o: "AND", + c: [{ k: "l", q: "alpha", s: "posts" }], + }; + await page.goto(`/?qt=${qt(tree)}`); + await expectResultSlugs(page, [POST_ROOT_SLUG, POST_REPLY_SLUG]); + // A post has no timeline, so its hit rows carry no mm:ss gutter — unlike a + // transcript hit row, which always does. + const postCard = page + .locator(`[data-card-header][data-result-slug="${POST_ROOT_SLUG}"]`) + .locator("xpath=ancestor::div[@data-card-header or @data-result-slug][1]"); + await expect(postCard.getByText(/^\d+:\d{2}$/)).toHaveCount(0); + }); + + test("qt= round-trips a posts leaf through the URL", async ({ page }) => { + const tree: SGroup = { + k: "g", + o: "AND", + c: [{ k: "l", q: "omega", s: "posts" }], + }; + await page.goto(`/?qt=${qt(tree)}`); + await expectResultSlugs(page, [POST_OTHER_SLUG]); + // The scope survived deserialize's explicit whitelist — a dropped leaf + // would have matched everything (empty tree) instead of just this post. + const select = page.locator('[data-testid^="leaf-scope-"]').first(); + await expect(select).toHaveValue("posts"); + }); + + test("opening a post result shows the post modal with its thread", async ({ + page, + }) => { + const tree: SGroup = { + k: "g", + o: "AND", + c: [{ k: "l", q: "alpha", s: "posts" }], + }; + await page.goto(`/?qt=${qt(tree)}`); + await expectResultSlugs(page, [POST_ROOT_SLUG, POST_REPLY_SLUG]); + + await page + .locator(`[data-card-header][data-result-slug="${POST_ROOT_SLUG}"]`) + .locator("[data-card-open]") + .click(); + + const modal = page.locator("[data-post-modal]"); + await expect(modal).toBeVisible(); + // Appears twice by design: as the primary post and again in its thread. + await expect( + modal.getByText("a post about alpha things").first(), + ).toBeVisible(); + // Thread context: the reply is archived under the same threadId. + await expect(modal.getByText(/Thread \(2 posts\)/)).toBeVisible(); + // Outbound links are surfaced, media is only counted. + await expect( + modal.getByRole("link", { name: "https://example.com/linked" }).first(), + ).toBeVisible(); + }); + + test("the Posts type toggle switches the corpus off", async ({ page }) => { + const tree: SGroup = { + k: "g", + o: "OR", + c: [ + { k: "l", q: "alpha", s: "transcripts" }, + { k: "l", q: "alpha", s: "posts" }, + ], + }; + await page.goto(`/?qt=${qt(tree)}`); + await expectResultSlugs(page, [ + TRANSCRIPT_ONLY_SLUG, + CHAT_SMALL_SLUG, + CHAT_LARGE_SLUG, + POST_ROOT_SLUG, + POST_REPLY_SLUG, + ]); + + // Posts are a third media kind beside Videos / Livestreams. + await page.getByRole("checkbox", { name: "Posts" }).uncheck(); + await page.getByTestId("search-submit").click(); + + await expectResultSlugs(page, [ + TRANSCRIPT_ONLY_SLUG, + CHAT_SMALL_SLUG, + CHAT_LARGE_SLUG, + ]); + }); +}); diff --git a/export/service-worker/site-sw.js b/export/service-worker/site-sw.js @@ -24,7 +24,7 @@ const PAGES = `pages-${VERSION}`; const META = `meta-${VERSION}`; // Matches /transcripts/<slug>/... and the parallel /subs, /summaries trees. -const SHARD_RE = /^\/(transcripts|subs|summaries)\/([^/]+)\/(.+)$/; +const SHARD_RE = /^\/(transcripts|subs|posts|summaries)\/([^/]+)\/(.+)$/; self.addEventListener("install", () => { // Activate immediately — no precache list (corpus is too large to bundle). diff --git a/export/service-worker/sw-hub.js b/export/service-worker/sw-hub.js @@ -23,7 +23,7 @@ const META = `meta-${VERSION}`; // Matches /transcripts/<slug>/... and the parallel /subs, /summaries trees, on // any origin (we test url.pathname, so it's origin-independent). -const SHARD_RE = /^\/(transcripts|subs|summaries)\/([^/]+)\/(.+)$/; +const SHARD_RE = /^\/(transcripts|subs|posts|summaries)\/([^/]+)\/(.+)$/; self.addEventListener("install", () => { self.skipWaiting(); diff --git a/mcp/src/search.test.ts b/mcp/src/search.test.ts @@ -8,6 +8,10 @@ import type { } from "yt-dlp-transcript-common/lib/manifest"; import type { TranscriptDetail } from "yt-dlp-transcript-common/lib/transcripts"; import type { SubsDetail } from "yt-dlp-transcript-common/lib/subs"; +import type { + ChannelPostsManifest, + Post, +} from "yt-dlp-transcript-common/lib/posts"; import type { SearchAlias } from "yt-dlp-transcript-common/lib/searchAliases"; import type { Cue } from "yt-dlp-transcript-common/lib/vtt"; import type { ChannelGroup } from "yt-dlp-transcript-common/lib/channelGroups"; @@ -20,6 +24,8 @@ import type { } from "./source"; import { searchTranscripts, + findPost, + getThread, getWindowedTranscript, buildMatcher, resolveSelectedChannels, @@ -149,6 +155,45 @@ const STUB_GROUPS: ChannelGroup[] = [ }, ]; +// A tiny social channel: one post sharing the "k cups" term with the video +// fixtures (so a default search proves BOTH corpora come back), plus a thread +// reply for get_thread. +const STUB_POSTS: Post[] = [ + { + id: "p1", + slug: "chan-b/p1", + channelSlug: "chan-b", + author: "tester.bsky.social", + authorName: "Tester", + createdAt: "2026-03-01T10:00:00.000Z", + uploadDate: "20260301", + text: "a zephyrpost mentioning something", + url: "https://bsky.app/profile/tester.bsky.social/post/p1", + platform: "bluesky", + isReply: false, + isRepost: false, + links: [], + threadId: "p1", + }, + { + id: "p2", + slug: "chan-b/p2", + channelSlug: "chan-b", + author: "tester.bsky.social", + authorName: "Tester", + createdAt: "2026-03-01T11:00:00.000Z", + uploadDate: "20260301", + text: "a zephyrpost reply in the same thread", + url: "https://bsky.app/profile/tester.bsky.social/post/p2", + platform: "bluesky", + isReply: true, + isRepost: false, + links: [], + threadId: "p1", + replyTo: { platform: "bluesky", id: "p1" }, + }, +]; + class StubSource implements ShardSource { readonly label = "stub"; constructor(private aliases: SearchAlias[] = [K_CUPS_ALIAS]) {} @@ -235,6 +280,25 @@ class StubSource implements ShardSource { ]; } + // Only chan-b ships posts; every other channel is video-only, which is + // what a real mixed corpus looks like. + async postsManifest(ch: ChannelRef): Promise<ChannelPostsManifest | null> { + if (ch.slug !== "chan-b") return null; + return { + version: 1, + channelSlug: ch.slug, + pageCount: 1, + maxPageBytes: 0, + generatedAt: "", + slugToPage: { p1: 0, p2: 0 }, + }; + } + + async postsPage(ch: ChannelRef, page: number): Promise<Post[]> { + if (ch.slug !== "chan-b" || page !== 0) return []; + return STUB_POSTS; + } + async availabilityMap(): Promise<Map<string, VideoAvailability>> { return new Map(Object.entries(AVAILABILITY)); } @@ -905,3 +969,114 @@ test("server: the sweep prompt honors parse_model", async () => { assert.ok(!/haiku/.test(text), "the default model name is fully replaced"); await client.close(); }); + +// ─── the social-post corpus ─── + +test("posts: default content_types searches BOTH corpora", async () => { + const src = new StubSource(); + // "zephyrpost" only exists in the posts fixtures, so a default search finding + // it proves posts are covered without an opt-in. + const r = await searchTranscripts(src, { query: "zephyrpost" }); + assert.equal(r.total, 2); + assert.ok(r.hits.every((h) => h.contentType === "post")); + assert.deepEqual( + r.hits.map((h) => h.videoId).sort(), + ["p1", "p2"], + ); +}); + +test("posts: a post hit carries author/date and no timestamps", async () => { + const src = new StubSource(); + const r = await searchTranscripts(src, { query: "zephyrpost mentioning" }); + assert.equal(r.total, 1); + const hit = r.hits[0]; + assert.equal(hit.contentType, "post"); + assert.equal(hit.author, "Tester"); + assert.equal(hit.uploadDate, "20260301"); + assert.equal(hit.platform, "bluesky"); + assert.match(hit.webpageUrl ?? "", /bsky\.app/); + // A post has no timeline: seconds stay 0 so nothing can render an @ mm:ss. + assert.ok(hit.snippets.every((sn) => sn.seconds === 0)); +}); + +test("posts: content_types can narrow to one corpus", async () => { + const src = new StubSource(); + const videosOnly = await searchTranscripts(src, { + query: "zephyrpost", + contentTypes: ["video"], + }); + assert.equal(videosOnly.total, 0); + + const postsOnly = await searchTranscripts(src, { + query: "coffee", + contentTypes: ["post"], + }); + assert.equal(postsOnly.total, 0, "coffee only appears in transcripts"); +}); + +test("posts: a posts-scope spec leaf matches post text", async () => { + const src = new StubSource(); + const root = newGroup({ + op: "AND", + children: [newLeaf({ id: "l1", query: "zephyrpost", scope: "posts" })], + }); + const channels = await src.listChannels(); + const r = await runSearchSpec(src, channels, { tree: root }); + assert.equal(r.total, 2); + assert.deepEqual(r.hits.map((h) => h.videoId).sort(), ["p1", "p2"]); +}); + +test("posts: findPost resolves by id and getThread returns the whole thread", async () => { + const src = new StubSource(); + const found = await findPost(src, "p2"); + assert.ok(found, "reply is findable by its own id"); + assert.equal(found!.post.id, "p2"); + assert.equal(found!.post.isReply, true); + + const thread = await getThread(src, found!.ch, found!.post); + // Root + reply, oldest first. + assert.deepEqual(thread.map((p) => p.id), ["p1", "p2"]); +}); + +test("server: get_post returns the post with no timestamps", async () => { + const client = await connectClient(new StubSource()); + const res = await client.callTool({ + name: "get_post", + arguments: { post_id: "p1" }, + }); + const out = firstText(res); + assert.match(out, /Post by Tester/); + assert.match(out, /a zephyrpost mentioning something/); + assert.match(out, /post_id: p1/); + // The ISO `posted:` line legitimately contains a time; what must NOT appear + // is a citable clock marker — a bracketed [m:ss] stamp or an `@ mm:ss`. + assert.doesNotMatch(out, /\[\d+:\d{2}/, "no bracketed cue stamps in a post"); + assert.doesNotMatch(out, /@\s*\d+:\d{2}/, "no @ mm:ss moment in a post"); + await client.close(); +}); + +test("server: get_thread returns parent + replies", async () => { + const client = await connectClient(new StubSource()); + const res = await client.callTool({ + name: "get_thread", + arguments: { post_id: "p1" }, + }); + const out = firstText(res); + assert.match(out, /Thread \(2 posts\)/); + assert.match(out, /a zephyrpost mentioning something/); + assert.match(out, /a zephyrpost reply in the same thread/); + await client.close(); +}); + +test("server: search_transcripts renders a post hit without a moment link", async () => { + const client = await connectClient(new StubSource()); + const res = await client.callTool({ + name: "search_transcripts", + arguments: { query: "zephyrpost", limit: 20 }, + }); + const out = firstText(res); + assert.match(out, /post by Tester/); + assert.match(out, /posted:/); + assert.doesNotMatch(out, /moment_base/); + await client.close(); +}); diff --git a/mcp/src/search.ts b/mcp/src/search.ts @@ -1,6 +1,7 @@ import { formatDuration } from "yt-dlp-transcript-common/lib/format"; import type { TranscriptDetail } from "yt-dlp-transcript-common/lib/transcripts"; import type { Cue } from "yt-dlp-transcript-common/lib/vtt"; +import type { Post } from "yt-dlp-transcript-common/lib/posts"; import type { Platform } from "yt-dlp-transcript-common/lib/platform"; import { forEachLeaf, @@ -135,6 +136,19 @@ export async function resolveSelectedChannels( }; } +// Which corpora a search covers. Defaults to BOTH, so existing agent flows +// pick up the social-post corpus automatically. +export type ContentType = "video" | "post"; +export const ALL_CONTENT_TYPES: ReadonlyArray<ContentType> = ["video", "post"]; + +export function parseContentTypes(raw: unknown): ContentType[] { + if (!Array.isArray(raw) || raw.length === 0) return [...ALL_CONTENT_TYPES]; + const out = raw.filter( + (t): t is ContentType => t === "video" || t === "post", + ); + return out.length > 0 ? out : [...ALL_CONTENT_TYPES]; +} + export type SearchHit = { videoId: string; // The video's slug (`<channelSlug>/<id>`) — the archilyzer viewer's `?v=` @@ -153,6 +167,11 @@ export type SearchHit = { webpageUrl?: string; matches: number; snippets: Snippet[]; + // Set on a hit from the social-post corpus. A post has no timeline, so its + // snippets carry seconds 0 and its citation takes no `@ mm:ss`. + contentType?: ContentType; + author?: string; + createdAt?: string; }; export type SearchResult = { @@ -273,6 +292,8 @@ export async function searchTranscripts( useAliases?: boolean; snippetsPerVideo?: number; aliases?: SearchAlias[]; + // Which corpora to search. Defaults to both video transcripts and posts. + contentTypes?: ContentType[]; }, ): Promise<SearchResult> { const limit = opts.limit ?? 20; @@ -305,8 +326,12 @@ export async function searchTranscripts( let pagesScanned = 0; let channelsScanned = 0; let truncated = false; + const contentTypes = opts.contentTypes ?? [...ALL_CONTENT_TYPES]; + const wantVideos = contentTypes.includes("video"); + const wantPosts = contentTypes.includes("post"); outer: for (const ch of channels) { + if (!wantVideos) break; let manifest; try { manifest = await source.transcriptsManifest(ch); @@ -364,6 +389,63 @@ export async function searchTranscripts( } } + // ── the social-post corpus ── + // A parallel pass over the same selected channels: only social channels ship + // a posts manifest, so a video-only site costs one 404 per channel and adds + // nothing to the result set. + if (wantPosts) { + postsOuter: for (const ch of channels) { + let pm; + try { + pm = await source.postsManifest(ch); + } catch { + continue; + } + if (!pm) continue; // not a social channel + for (let page = 0; page < pm.pageCount; page++) { + if (pagesScanned >= maxPages) { + truncated = true; + break postsOuter; + } + let posts: Post[]; + try { + posts = await source.postsPage(ch, page); + } catch { + continue; + } + pagesScanned++; + for (const post of posts) { + if (!match(post.text)) continue; + all.push({ + videoId: post.id, + slug: post.slug, + channelSlug: ch.slug, + channelName: ch.name, + ...(ch.siteTitle ? { siteTitle: ch.siteTitle } : {}), + ...(ch.siteUrl ? { siteUrl: ch.siteUrl } : {}), + platform: post.platform, + // A post has no title; its body stands in so a worklist row is + // still readable without pulling snippets. + title: truncate(post.text, 120), + uploadDate: post.uploadDate, + webpageUrl: post.url, + matches: 1, + contentType: "post", + author: post.authorName || post.author, + createdAt: post.createdAt, + snippets: includeSnippets + ? [{ clock: "", seconds: 0, text: truncate(post.text, 480) }] + : [], + }); + if (all.length >= HARD_VIDEO_CAP) { + truncated = true; + break postsOuter; + } + } + } + } + } + const total = all.length; const hits = all.slice(offset, offset + limit); return { @@ -390,6 +472,81 @@ export async function searchTranscripts( // `after` seconds), merge overlapping windows deduped by timestamp (capped), and // return timestamped excerpt lines plus the total match count. Bounded and // high-signal — the batch read the sweep prompt drives. +// Locate one archived post by id, mirroring findVideo. Only social channels +// ship a posts manifest, so a video-only corpus resolves this cheaply. +export async function findPost( + source: ShardSource, + postId: string, + channelHint?: string | string[], +): Promise<{ ch: ChannelRef; post: Post } | null> { + let channels = await source.listChannels(); + const hints = ( + Array.isArray(channelHint) ? channelHint : channelHint ? [channelHint] : [] + ) + .map((h) => (typeof h === "string" ? h.trim().toLowerCase() : "")) + .filter((h) => h !== ""); + if (hints.length > 0) { + const want = new Set(hints); + const filtered = channels.filter( + (c) => + want.has(c.slug.toLowerCase()) || + want.has(c.key.toLowerCase()) || + want.has(c.name.toLowerCase()), + ); + if (filtered.length > 0) channels = filtered; + } + for (const ch of channels) { + let manifest; + try { + manifest = await source.postsManifest(ch); + } catch { + continue; + } + if (!manifest) continue; + const page = manifest.slugToPage[postId]; + if (page === undefined) continue; + let posts: Post[]; + try { + posts = await source.postsPage(ch, page); + } catch { + continue; + } + const post = posts.find((p) => p.id === postId); + if (post) return { ch, post }; + } + return null; +} + +// Every archived post in the same thread as `post`, oldest first. A thread can +// straddle shard pages, so this walks the channel's whole (byte-capped) tree. +export async function getThread( + source: ShardSource, + ch: ChannelRef, + post: Post, +): Promise<Post[]> { + const threadId = post.threadId || post.id; + const manifest = await source.postsManifest(ch); + if (!manifest) return [post]; + const thread: Post[] = []; + for (let page = 0; page < manifest.pageCount; page++) { + let posts: Post[]; + try { + posts = await source.postsPage(ch, page); + } catch { + continue; + } + for (const p of posts) { + if ((p.threadId || p.id) === threadId) thread.push(p); + } + } + thread.sort((a, b) => + a.createdAt === b.createdAt + ? a.id.localeCompare(b.id) + : a.createdAt.localeCompare(b.createdAt), + ); + return thread.length > 0 ? thread : [post]; +} + export function getWindowedTranscript( record: TranscriptDetail, matcher: Matcher, @@ -599,6 +756,9 @@ type RecordCtx = { tags: string; cues: Cue[]; chatCues: Cue[]; + // The post body, when this record IS a post rather than a video. Empty for a + // video record, so a posts-scope leaf never matches one. + postText?: string; snippetsPerVideo: number; includeSnippets: boolean; }; @@ -638,6 +798,19 @@ function evalLeaf(leaf: QueryNode, m: LeafMatcher, ctx: RecordCtx): LeafOutcome }); } break; + case "posts": + // A post has no timeline: one hit, seconds 0 — the same convention the + // metadata / description / tags scopes already use. + if (ctx.postText && m.test(ctx.postText)) { + count++; + push({ + scope: "posts", + clock: clock(0), + seconds: 0, + text: truncate(ctx.postText), + }); + } + break; case "metadata": { const titleHit = m.test(ctx.title); const channelHit = m.test(ctx.channel); @@ -881,6 +1054,69 @@ export async function runSearchSpec( } } + // ── posts pass ── + // Only run when the tree actually has a posts leaf: a video-only spec must + // not pay a manifest probe per channel. Post records reuse the same evalNode + // with `postText` set and no cues, so AND/OR/negate semantics are identical. + const wantsPosts = [...matchers.values()].some((m) => m.scope === "posts"); + if (wantsPosts) { + postsOuter: for (const ch of channels) { + let pm; + try { + pm = await source.postsManifest(ch); + } catch { + continue; + } + if (!pm) continue; + for (let page = 0; page < pm.pageCount; page++) { + if (pagesScanned >= maxPages) { + truncated = true; + break postsOuter; + } + let posts: Post[]; + try { + posts = await source.postsPage(ch, page); + } catch { + continue; + } + pagesScanned++; + for (const post of posts) { + const ctx: RecordCtx = { + title: post.text, + channel: post.authorName || post.author, + description: "", + tags: "", + cues: [], + chatCues: [], + postText: post.text, + snippetsPerVideo, + includeSnippets, + }; + const r = evalNode(spec.tree, matchers, ctx); + if (!r.match) continue; + all.push({ + videoId: post.id, + slug: post.slug, + channelSlug: ch.slug, + channelName: ch.name, + ...(ch.siteTitle ? { siteTitle: ch.siteTitle } : {}), + ...(ch.siteUrl ? { siteUrl: ch.siteUrl } : {}), + platform: post.platform, + title: truncate(post.text, 120), + uploadDate: post.uploadDate, + webpageUrl: post.url, + matches: r.count || 1, + snippets: r.hits, + }); + if (all.length >= HARD_VIDEO_CAP) { + truncated = true; + break postsOuter; + } + } + } + } + } + const total = all.length; return { hits: all.slice(offset, offset + limit), diff --git a/mcp/src/server.ts b/mcp/src/server.ts @@ -25,8 +25,12 @@ import { } from "./source"; import type { SourceSpec } from "./sources"; import type { SourceController } from "./sourceController"; +import type { Post } from "yt-dlp-transcript-common/lib/posts"; import { searchTranscripts, + findPost, + getThread, + parseContentTypes, findVideo, buildMatcher, getWindowedTranscript, @@ -141,7 +145,11 @@ const TOOLS = [ { name: "search_transcripts", description: - "Search transcript captions for a term or phrase and return matching " + + "Search the archive for a term or phrase. The corpus holds video " + + "transcripts AND social posts (X/Twitter, Bluesky) from the same " + + "commentators; by default BOTH are searched and returned in one result " + + "set — use content_types to narrow. A post hit carries no timestamps " + + "(cite it as a bare source, never with @ mm:ss). Returns matching " + "videos with timestamped snippets. Substring match by default; set regex " + "to true for a case-insensitive regular expression. Scope is optional and " + "additive: restrict to one or more channels (channel / channels, by slug " + @@ -216,6 +224,14 @@ const TOOLS = [ "Max shard pages to scan before stopping (default 400). Reaching " + "it marks coverage partial.", }, + content_types: { + type: "array", + items: { type: "string", enum: ["video", "post"] }, + description: + "Which corpora to search: 'video' (transcripts) and/or 'post' " + + "(archived social posts). Defaults to BOTH. Posts have no " + + "timeline, so their hits carry no timestamps.", + }, link_style: { type: "string", enum: ["inline", "base"], @@ -254,6 +270,47 @@ const TOOLS = [ }, }, { + name: "get_post", + description: + "Fetch one archived social post by id: its full text, author, timestamp, " + + "outbound links and permalink. Posts have no timeline — cite them as a " + + "bare source with no @ mm:ss.", + inputSchema: { + type: "object", + properties: { + post_id: { type: "string", description: "The post id (tweet id / atproto rkey)." }, + channel: { + type: "string", + description: "Optional owning channel slug/name to skip the lookup.", + }, + }, + required: ["post_id"], + additionalProperties: false, + }, + }, + { + name: "get_thread", + description: + "Fetch the whole thread a post belongs to (its root and every archived " + + "reply), oldest first. This is the post-corpus analogue of reading the " + + "transcript around a cited moment.", + inputSchema: { + type: "object", + properties: { + post_id: { + type: "string", + description: "Any post id in the thread (root or a reply).", + }, + channel: { + type: "string", + description: "Optional owning channel slug/name to skip the lookup.", + }, + }, + required: ["post_id"], + additionalProperties: false, + }, + }, + { name: "get_transcripts", description: "Batch-read up to 20 videos' transcripts in one call. With a query, each " + @@ -310,6 +367,14 @@ const TOOLS = [ type: "number", description: "Seconds of context after each match (default 30).", }, + content_types: { + type: "array", + items: { type: "string", enum: ["video", "post"] }, + description: + "Which corpora to search: 'video' (transcripts) and/or 'post' " + + "(archived social posts). Defaults to BOTH. Posts have no " + + "timeline, so their hits carry no timestamps.", + }, link_style: { type: "string", enum: ["inline", "base"], @@ -489,8 +554,17 @@ const TOOLS = [ }, query_scope: { type: "string", - enum: ["transcripts", "chat", "metadata", "description", "tags"], - description: "Scope for the override query (default transcripts).", + enum: [ + "transcripts", + "chat", + "posts", + "metadata", + "description", + "tags", + ], + description: + "Scope for the override query (default transcripts). 'posts' " + + "targets the social-post corpus.", }, }, additionalProperties: false, @@ -587,6 +661,10 @@ export function createServer(sourceOrController: ShardSource | SourceController) return await handleGetTranscript(source, args); case "get_transcripts": return await handleGetTranscripts(source, args); + case "get_post": + return await handleGetPost(source, args); + case "get_thread": + return await handleGetThread(source, args); case "get_video_metadata": return await handleGetMetadata(source, args); case "list_sources": @@ -698,6 +776,7 @@ async function handleSearch( includeSnippets, useAliases: args.use_aliases !== false, maxPages: typeof args.max_pages === "number" ? args.max_pages : undefined, + contentTypes: parseContentTypes(args.content_types), }); const aliasNote = describeFiredAliases(result.firedAliases); @@ -723,6 +802,18 @@ async function handleSearch( return text(`${head}${footer}`); } const blocks = result.hits.map((h) => { + // A POST has no timeline: no moment link, no [m:ss] stamps. Render it as a + // dated, authored block so the model can cite it as a bare source. + if (h.contentType === "post") { + const head = + `### post by ${h.author ?? h.channelName}\n` + + `- post_id: ${h.videoId} | channel: ${h.channelName}` + + (h.siteTitle ? ` | site: ${h.siteTitle}` : "") + + ` | posted: ${formatDate(h.uploadDate)} | platform: ${h.platform ?? "post"}` + + (includeSnippets && h.webpageUrl ? `\n- source: ${h.webpageUrl}` : ""); + const body = h.snippets.map((sn) => ` - ${sn.text}`).join("\n"); + return body ? `${head}\n${body}` : head; + } // Worklist mode (include_snippets:false) trims the source line too — the // caller only wants ids/titles/counts, so no per-video URLs at all. const baseUrl = base && includeSnippets ? momentBaseFor(source, h) : null; @@ -744,8 +835,15 @@ async function handleSearch( }); // Worklist mode has no stamps (or bases) to expand — skip the note too. const baseNote = base && includeSnippets ? `\n\n(${BASE_EXPANSION_NOTE})` : ""; + const postCount = result.hits.filter((h) => h.contentType === "post").length; + const label = + postCount === 0 + ? "video(s)" + : postCount === result.hits.length + ? "post(s)" + : "result(s) (videos + posts)"; return text( - `${result.total} video(s) matching "${query}":\n\n${blocks.join("\n\n")}${footer}${baseNote}`, + `${result.total} ${label} matching "${query}":\n\n${blocks.join("\n\n")}${footer}${baseNote}`, ); } @@ -914,6 +1012,83 @@ async function handleGetTranscripts( return text(`${blocks.join("\n\n---\n\n")}${footer}`); } +// Render one post as markdown. No timestamps anywhere — a post has no timeline, +// and emitting a 0:00 would invite a bogus `@ mm:ss` citation. +function postToMarkdown(post: Post, heading = true): string { + const lines: string[] = []; + if (heading) lines.push(`# Post by ${post.authorName || post.author}`); + lines.push( + `- author: ${post.authorName ? `${post.authorName} (@${post.author})` : `@${post.author}`}`, + ); + lines.push(`- posted: ${post.createdAt}`); + lines.push(`- platform: ${post.platform}`); + lines.push(`- post_id: ${post.id}`); + lines.push(`- source: ${post.url}`); + if (post.isRepost) lines.push(`- repost: yes`); + if (post.isReply) lines.push(`- reply: yes`); + if (post.threadId && post.threadId !== post.id) { + lines.push(`- thread_id: ${post.threadId}`); + } + if (post.mediaCount) { + lines.push(`- media: ${post.mediaCount} attachment(s) (not archived)`); + } + if (post.engagement) { + const e = post.engagement; + const parts = [ + e.likes != null ? `${e.likes} likes` : "", + e.reposts != null ? `${e.reposts} reposts` : "", + e.replies != null ? `${e.replies} replies` : "", + e.quotes != null ? `${e.quotes} quotes` : "", + ].filter(Boolean); + if (parts.length) lines.push(`- engagement: ${parts.join(", ")}`); + } + lines.push(""); + lines.push(post.text); + if (post.links.length > 0) { + lines.push(""); + lines.push("Links:"); + for (const l of post.links) lines.push(`- ${l}`); + } + return lines.join("\n"); +} + +async function handleGetPost( + source: ShardSource, + args: Record<string, unknown>, +): Promise<ToolResult> { + const postId = String(args.post_id ?? "").trim(); + if (!postId) return errorText("post_id is required"); + const found = await findPost( + source, + postId, + typeof args.channel === "string" ? args.channel : undefined, + ); + if (!found) return errorText(`post not found: ${postId}`); + return text(postToMarkdown(found.post)); +} + +async function handleGetThread( + source: ShardSource, + args: Record<string, unknown>, +): Promise<ToolResult> { + const postId = String(args.post_id ?? "").trim(); + if (!postId) return errorText("post_id is required"); + const found = await findPost( + source, + postId, + typeof args.channel === "string" ? args.channel : undefined, + ); + if (!found) return errorText(`post not found: ${postId}`); + const thread = await getThread(source, found.ch, found.post); + const head = + `# Thread (${thread.length} post${thread.length === 1 ? "" : "s"}) — ` + + `${found.post.authorName || found.post.author}\n`; + const body = thread + .map((p, i) => `## ${i + 1}. ${p.authorName || p.author}\n${postToMarkdown(p, false)}`) + .join("\n\n"); + return text(`${head}\n${body}`); +} + async function handleGetTranscript( source: ShardSource, args: Record<string, unknown>, @@ -1644,11 +1819,13 @@ function buildSweepPrompt(args: Record<string, unknown>) { `reference the returned fragment against the report so far and upsert ` + `findings — claims, and contradictions with earlier claims — into ` + `well-titled \`## sections\` of \`${reportPath}\` (Write/Edit). Cite ` + - `every finding as **\`[title @ mm:ss](<moment url>)\`**, expanding each ` + + `every VIDEO finding as **\`[title @ mm:ss](<moment url>)\`**, expanding each ` + `kept \`[mm:ss|seconds]\` stamp by appending the integer after the ` + `\`|\` to that video's moment_base (full link = ` + `\`<moment_base><seconds>\`; no moment_base → link the \`source:\` URL ` + - `instead). Then discard the fragment. Batches are independent, so you ` + + `instead). A POST finding has no timestamp — cite it as ` + + `**\`[post by <author>, <date>](<source url>)\`** instead, never with ` + + `\`@ mm:ss\`. Then discard the fragment. Batches are independent, so you ` + `may dispatch several subagents in parallel.\n` + ` - **Fallback:** if no subagent/Task tool is available, do the batch ` + `inline — call \`get_transcripts\` with \`link_style: "base"\` yourself, ` + @@ -1660,7 +1837,8 @@ function buildSweepPrompt(args: Record<string, unknown>) { `**Finish.** Repeat to the end of the worklist, then write a short summary ` + `section (the scope swept, how many videos covered, headline findings, any ` + `partial-coverage caveat) and tell me the report path. Keep every citation ` + - `a clickable \`[title @ mm:ss](url)\` link.`, + `a clickable link — \`[title @ mm:ss](url)\` for a video, ` + + `\`[post by <author>, <date>](url)\` for a post.`, ); const numbered = steps @@ -1673,10 +1851,13 @@ function buildSweepPrompt(args: Record<string, unknown>) { `\`${reportPath}\`. You are the sweep engine — work through the whole match ` + `set methodically, using the transcript MCP tools for evidence and your own ` + `Write/Edit tools for the report. The MCP is read-only; never try to change ` + - `the archive. **Cite every finding as a clickable ` + + `the archive. The corpus holds video transcripts AND archived social posts. ` + + `**Cite every video finding as a clickable ` + `\`[title @ mm:ss](<moment url>)\` link** (build each moment URL by ` + `appending the cited integer seconds to that video's \`moment_base\` from ` + - `the tool output).\n\n` + + `the tool output); **cite every post finding as ` + + `\`[post by <author>, <date>](<source url>)\`** — posts have no timeline, ` + + `so they never take a \`@ mm:ss\`.\n\n` + `Follow these steps:\n\n${numbered}`; return { diff --git a/mcp/src/source.ts b/mcp/src/source.ts @@ -14,6 +14,11 @@ import type { } from "yt-dlp-transcript-common/lib/transcripts"; import type { SubsDetail } from "yt-dlp-transcript-common/lib/subs"; import { + postsPageFileName, + type ChannelPostsManifest, + type Post, +} from "yt-dlp-transcript-common/lib/posts"; +import { coerceAliasConfig, type SearchAlias, } from "yt-dlp-transcript-common/lib/searchAliases"; @@ -140,6 +145,14 @@ export interface ShardSource { // A page of a channel's subs records (subs/<slug>/page-NNNN.json). Each record // inlines its per-track cues under `tracks` (e.g. `tracks.live_chat`). subsPage(ch: ChannelRef, page: number): Promise<SubsDetail[]>; + // A channel's social-posts manifest (posts/<slug>/manifest.json), or null + // when the channel ships no posts shards (i.e. it is a video channel). Same + // slugToPage/pageCount shape as the transcripts manifest. This is the MCP's + // single abstraction boundary, so adding it here yields the posts corpus on + // all three transports (local / remote / hub) at once. + postsManifest(ch: ChannelRef): Promise<ChannelPostsManifest | null>; + // A page of a channel's posts (posts/<slug>/page-NNNN.json). + postsPage(ch: ChannelRef, page: number): Promise<Post[]>; // A map of every video's availability (deleted/unlisted), keyed by the // member-local video slug (`<channelSlug>/<id>`), built from the summaries // shards. Fetched lazily and cached — only the `fav` availability filter needs @@ -204,6 +217,36 @@ export class LocalSource implements ShardSource { return JSON.parse(raw) as SubsDetail[]; } + // Cached per channel INCLUDING the negative answer: most channels are + // video-only, and a posts-covering search would otherwise re-probe every one + // of them on every query. + private postsManifests = new Map<string, ChannelPostsManifest | null>(); + + async postsManifest(ch: ChannelRef): Promise<ChannelPostsManifest | null> { + const hit = this.postsManifests.get(ch.slug); + if (hit !== undefined) return hit; + let out: ChannelPostsManifest | null = null; + try { + const raw = await readFile( + path.join(this.dir, "posts", ch.slug, "manifest.json"), + "utf8", + ); + out = JSON.parse(raw) as ChannelPostsManifest; + } catch { + out = null; // channel ships no posts shards + } + this.postsManifests.set(ch.slug, out); + return out; + } + + async postsPage(ch: ChannelRef, page: number): Promise<Post[]> { + const raw = await readFile( + path.join(this.dir, "posts", ch.slug, postsPageFileName(page)), + "utf8", + ); + return JSON.parse(raw) as Post[]; + } + async availabilityMap(): Promise<Map<string, VideoAvailability>> { if (this.availability) return this.availability; this.availability = await buildAvailabilityMap( @@ -336,6 +379,28 @@ export class RemoteSource implements ShardSource { return this.getJson(`/subs/${ch.slug}/${subsPageFileName(page)}`); } + // Cached per channel including the negative answer — otherwise every + // posts-covering search costs one 404 per video-only channel. + private postsManifests = new Map<string, ChannelPostsManifest | null>(); + + async postsManifest(ch: ChannelRef): Promise<ChannelPostsManifest | null> { + const hit = this.postsManifests.get(ch.slug); + if (hit !== undefined) return hit; + let out: ChannelPostsManifest | null = null; + try { + const res = await fetch(`${this.base}/posts/${ch.slug}/manifest.json`); + out = res.ok ? ((await res.json()) as ChannelPostsManifest) : null; + } catch { + out = null; + } + this.postsManifests.set(ch.slug, out); + return out; + } + + postsPage(ch: ChannelRef, page: number): Promise<Post[]> { + return this.getJson(`/posts/${ch.slug}/${postsPageFileName(page)}`); + } + async availabilityMap(): Promise<Map<string, VideoAvailability>> { if (this.availability) return this.availability; this.availability = await buildAvailabilityMap( @@ -533,6 +598,20 @@ export class HubSource implements ShardSource { return this.memberFor(ch.siteId).subsPage(ch, page); } + async postsManifest(ch: ChannelRef): Promise<ChannelPostsManifest | null> { + if (!ch.siteId) return null; + try { + return await this.memberFor(ch.siteId).postsManifest(ch); + } catch { + return null; // member not yet registered / unreachable + } + } + + postsPage(ch: ChannelRef, page: number): Promise<Post[]> { + if (!ch.siteId) throw new Error("hub channel ref missing siteId"); + return this.memberFor(ch.siteId).postsPage(ch, page); + } + // Merge each member's availability map. Keys are member-local slugs // (`<channelSlug>/<id>`) — the same slug a member's transcript page records // carry — so a per-record `fav` lookup joins correctly. Built lazily/cached. diff --git a/mcp/src/sourceController.test.ts b/mcp/src/sourceController.test.ts @@ -71,6 +71,12 @@ class FakeSource implements ShardSource { async subsPage(): Promise<[]> { return []; } + async postsManifest(): Promise<null> { + return null; + } + async postsPage(): Promise<[]> { + return []; + } async availabilityMap(): Promise<Map<string, VideoAvailability>> { return new Map(); }