Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit ac7c3cce5b99fb4aa86728111527b9dd285de6fe
parent 4c96685d42781a106646f86238138e365d4c0f8b
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Fri, 25 Sep 2026 20:47:33 -0400

Merge one-core/c2-federated-search — release 9 slice C2: federated search (scope chips, per-archive state, progressive readiness, 'N of M archives answered', archive named on every card, paged fetch per archive)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Diffstat:
Mcommon/components/SearchDataContext.tsx | 407++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++---------------
Mcommon/components/SearchResults.tsx | 59+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/components/SearchSessionContext.tsx | 18++++++++++++++++++
Meditor/CHANGELOG.md | 1+
Mexport/app/components/hub/HubHome.tsx | 20++++++++++++++++----
Aexport/app/components/hub/HubScope.tsx | 104+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mexport/app/components/hub/HubStats.tsx | 15++++++++++-----
Aexport/app/components/hub/useHubScope.ts | 65+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aexport/e2e-hub/federated-search.spec.ts | 275+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mexport/e2e-hub/federation.spec.ts | 4++--
Mplans/release-9.md | 120+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
11 files changed, 1003 insertions(+), 85 deletions(-)

diff --git a/common/components/SearchDataContext.tsx b/common/components/SearchDataContext.tsx @@ -14,7 +14,7 @@ import { useMemo, type ReactNode, } from "react"; -import { useQueries } from "@tanstack/react-query"; +import { useQueries, useQueryClient } from "@tanstack/react-query"; import { useSummaries, type SummariesState } from "./summariesCache"; import { useSubsManifest } from "./subsCache"; import { usePostsManifest } from "./postsCache"; @@ -77,6 +77,41 @@ export type SearchDataValue = { // both modes: a hub's counts are the hub's, and summing member documents here // would put a number in front of someone that no site can reproduce. curatedTags: PublishedTag[]; + // Hub mode only (absent single-site): the per-archive state of the + // federation — which archives are in scope, loading, ready or failed — and a + // Retry per archive. Drives the scope chips and the "N of M archives + // answered." line. + federation?: FederationState; + // Hub mode only: the source archive's title for a content origin, so a result + // card can name where it came from in text, not by its accent alone. + siteTitleOf?: (origin: string) => string | undefined; +}; + +// One federated archive's state, as the hub shows it on its scope chip. +// off — the visitor took it out of the search; nothing is fetched. +// loading — its manifest or pages are still arriving (or being retried). +// ready — every summaries page arrived; its records are in the search. +// failed — its manifest or a summaries page could not be read; its records +// are NOT in the search until a Retry succeeds. +export type FederatedSiteStatus = "loading" | "ready" | "failed" | "off"; + +export type FederatedSiteState = { + origin: string; + siteTitle: string; + accent?: string; + status: FederatedSiteStatus; + // The archive's transcript count (its summaries manifest's totalCount), once + // the manifest has arrived. + count?: number; + // Why it failed, when it did. + error?: string; +}; + +export type FederationState = { + // In registry order, off archives included. + sites: FederatedSiteState[]; + // Refetch everything of this archive that failed. + retry: (origin: string) => void; }; // Single-site mode has no provenance accents; a module constant keeps the @@ -172,13 +207,60 @@ export function SingleSiteDataProvider({ children }: { children: ReactNode }) { // A federated origin the hub reads from. `origin` is "" for the hub's own // same-origin pool, else a full origin ("https://x.com"). `siteTitle` labels -// its channel group; `accent` is carried for provenance in the UI. +// its channel group; `accent` is carried for provenance in the UI. `enabled` +// false takes it out of the search: none of its feeds are fetched and nothing +// already cached from it is merged (default true). export type FederatedSite = { origin: string; siteTitle: string; accent?: string; + enabled?: boolean; }; +// At most this many summaries-page fetches in flight per archive, so one big +// member (Jeralyzer: dozens of pages) cannot take every connection the browser +// has and starve the small ones. Per ORIGIN, not global: different origins do +// not share a connection pool, so a global cap would only slow the whole hub. +const PAGE_FETCHES_PER_SITE = 6; + +// The per-archive feeds this provider fetches — what a Retry refetches. +const FEDERATED_KEYS = new Set([ + "manifest", + "summaries-page", + "subs-manifest", + "posts-manifest", + "search-aliases", +]); + +type Gate = { active: number; queue: Array<() => void> }; +const pageGates = new Map<string, Gate>(); + +async function gated<T>(origin: string, fn: () => Promise<T>): Promise<T> { + let g = pageGates.get(origin); + if (!g) { + g = { active: 0, queue: [] }; + pageGates.set(origin, g); + } + if (g.active >= PAGE_FETCHES_PER_SITE) { + // The releasing fetch hands its slot straight to us (active unchanged). + await new Promise<void>((resolve) => g.queue.push(resolve)); + } else { + g.active++; + } + try { + return await fn(); + } finally { + const next = g.queue.shift(); + if (next) next(); + else g.active--; + } +} + +function errorText(e: unknown): string | undefined { + if (!e) return undefined; + return e instanceof Error ? e.message : String(e); +} + // Multi-site data source for the hub: fans out the same summaries/subs feeds // across many origins and merges them into ONE origin-qualified view, so // TranscriptSearch stays mode-agnostic (it reads exactly the same context shape @@ -189,56 +271,176 @@ export type FederatedSite = { // model, and downstream fetches (fetchTranscript decodes the origin back out) // never collide across sites. Grouping falls out of the existing grouped- // checkbox UI: one ChannelGroup per site (id = origin, label = siteTitle). +// +// Per archive, in order: its summaries manifest, then — only once that has +// arrived, and only while the archive is in scope — its pages, at most +// PAGE_FETCHES_PER_SITE at a time. An archive's records join the merged list +// only when ALL its pages are in (status "ready"), so the list grows one whole +// archive at a time and a search re-runs once per archive, not once per page. +// A failed archive contributes nothing and says so (`federation`). +// +// `progressive`: summariesReady turns true as soon as ONE in-scope archive is +// ready (the hub's search page, where results should not wait for the slowest +// member). Without it, summariesReady waits until every in-scope archive has +// settled, ready or failed (the /ask chat, which grounds in what it was given). export function MultiSiteDataProvider({ sites, + progressive = false, children, }: { sites: FederatedSite[]; + progressive?: boolean; children: ReactNode; }) { + const queryClient = useQueryClient(); + const inScope = (s: FederatedSite) => s.enabled !== false; + // 1. Per-origin summaries manifests (channels/groups/pageCount/freshness). const manifestQueries = useQueries({ queries: sites.map((s) => ({ queryKey: ["manifest", s.origin], queryFn: () => readerFor(s.origin).readSummariesManifest(), + enabled: inScope(s), })), }); - const manifestsSettled = manifestQueries.every( - (q) => q.isSuccess || q.isError, - ); - // 2. Flat page descriptors across every origin whose manifest loaded, so a - // single useQueries can fan out all pages regardless of per-site counts. + // 2. Page descriptors for every in-scope origin whose manifest has arrived. + // Keyed by a string of (origin, pageCount) so the list's identity moves + // only when an archive's manifest lands or the scope changes. + const pagePlanKey = sites + .map((s, i) => { + const m = manifestQueries[i]?.data; + return inScope(s) && m ? `${s.origin}\n${m.pageCount}` : ""; + }) + .join("\u0000"); const pageDescriptors = useMemo<{ origin: string; index: number }[]>(() => { const out: { origin: string; index: number }[] = []; sites.forEach((s, i) => { + if (!inScope(s)) return; const pc = manifestQueries[i]?.data?.pageCount ?? 0; for (let p = 0; p < pc; p++) out.push({ origin: s.origin, index: p }); }); return out; - // manifestQueries identity churns; key off settled + the site set. + // manifestQueries identity churns; pagePlanKey is its content. // eslint-disable-next-line react-hooks/exhaustive-deps - }, [sites, manifestsSettled]); + }, [pagePlanKey]); const pageQueries = useQueries({ queries: pageDescriptors.map((d) => ({ queryKey: ["summaries-page", d.origin, d.index], - queryFn: () => readerFor(d.origin).readSummariesPage(d.index), + // Consuming `signal` lets TanStack cancel a page whose archive was + // switched off mid-load; a page still waiting at the gate then bails + // before it is fetched. + queryFn: ({ signal }: { signal: AbortSignal }) => + gated(d.origin, () => { + signal.throwIfAborted(); + return readerFor(d.origin).readSummariesPage(d.index); + }), })), }); + // Each archive's subs + posts manifests. Fetched alongside its summaries, + // merged below only once the archive is ready, so all of an archive's + // contributions land in the search in one re-run. A posts 404 (no posts + // corpus) resolves to an empty manifest, not a failure. + const subsQueries = useQueries({ + queries: sites.map((s) => ({ + queryKey: ["subs-manifest", s.origin], + queryFn: () => readerFor(s.origin).readSubsSiteManifest(), + enabled: inScope(s), + })), + }); + const postsQueries = useQueries({ + queries: sites.map((s) => ({ + queryKey: ["posts-manifest", s.origin], + queryFn: () => + readerFor(s.origin) + .readPostsSiteManifest() + .catch( + (): PostsManifest => ({ + version: 0, + channels: [], + totalCount: 0, + generatedAt: "", + }), + ), + enabled: inScope(s), + })), + }); + // 3. Each archive's state, from its manifest, page, subs and posts queries. + const pagesByOrigin = new Map<string, (typeof pageQueries)[number][]>(); + pageQueries.forEach((q, i) => { + const origin = pageDescriptors[i]?.origin; + if (origin === undefined) return; + const list = pagesByOrigin.get(origin); + if (list) list.push(q); + else pagesByOrigin.set(origin, [q]); + }); + const rawStates: FederatedSiteState[] = sites.map((s, i) => { + const base = { + origin: s.origin, + siteTitle: s.siteTitle, + ...(s.accent ? { accent: s.accent } : {}), + }; + if (!inScope(s)) return { ...base, status: "off" }; + const m = manifestQueries[i]; + const pages = pagesByOrigin.get(s.origin) ?? []; + const count = m?.data?.totalCount; + const withCount = count === undefined ? base : { ...base, count }; + // A subs manifest that errored still counts as settled (the archive just + // has no live chat in the search); posts never error (404 → empty). + const settled = (q: { isSuccess: boolean; isError: boolean; isFetching: boolean } | undefined) => + !!q && (q.isSuccess || (q.isError && !q.isFetching)); + if ( + m?.isSuccess && + pages.length === m.data.pageCount && + pages.every((p) => p.isSuccess) && + settled(subsQueries[i]) && + settled(postsQueries[i]) + ) { + return { ...withCount, status: "ready" }; + } + const fetching = !!m?.isFetching || pages.some((p) => p.isFetching); + const failed = m?.isError ? m : pages.find((p) => p.isError); + if (failed && !fetching) { + return { + ...withCount, + status: "failed", + ...(errorText(failed.error) ? { error: errorText(failed.error) } : {}), + }; + } + return { ...withCount, status: "loading" }; + }); + const statusKey = rawStates + .map((s) => `${s.origin}|${s.siteTitle}|${s.accent ?? ""}|${s.status}|${s.count ?? ""}|${s.error ?? ""}`) + .join("\n"); + const siteStates = useMemo( + () => rawStates, + // eslint-disable-next-line react-hooks/exhaustive-deps + [statusKey], + ); + + const readyOrigins = useMemo( + () => + new Set(siteStates.filter((s) => s.status === "ready").map((s) => s.origin)), + [siteStates], + ); + const readyKey = Array.from(readyOrigins).join("\u0000"); + const scoped = siteStates.filter((s) => s.status !== "off"); + const allSettled = scoped.every( + (s) => s.status === "ready" || s.status === "failed", + ); + const summariesReady = allSettled || (progressive && readyOrigins.size > 0); + const loadedPages = pageQueries.filter((q) => q.data).length; - const pagesSettled = pageQueries.every((q) => q.isSuccess || q.isError); - // Ready once every manifest and every page has settled (a failing origin - // resolves to error rather than blocking the rest of the shelf). - const summariesReady = manifestsSettled && pagesSettled; - // 3. Merge summaries, rewriting ids to be origin-qualified. + // 4. Merge the READY archives' summaries, rewriting ids to be + // origin-qualified, newest first. const summaries = useMemo<DisplaySummary[]>(() => { const out: DisplaySummary[] = []; pageQueries.forEach((q, i) => { const origin = pageDescriptors[i]?.origin ?? ""; - if (!q.data) return; + if (!q.data || !readyOrigins.has(origin)) return; for (const t of q.data) { out.push( origin @@ -251,35 +453,40 @@ export function MultiSiteDataProvider({ ); } }); - if (summariesReady) { - out.sort( - (a, b) => - b.uploadDate.localeCompare(a.uploadDate) || - a.channelSlug.localeCompare(b.channelSlug) || - a.id.localeCompare(b.id), - ); - } + out.sort( + (a, b) => + b.uploadDate.localeCompare(a.uploadDate) || + a.channelSlug.localeCompare(b.channelSlug) || + a.id.localeCompare(b.id), + ); return out; + // Only the ready set: a ready archive's pages never change, so a manifest + // landing elsewhere (a new pageDescriptors) must not hand the session a new + // list — every new identity re-runs the current search. // eslint-disable-next-line react-hooks/exhaustive-deps - }, [loadedPages, summariesReady, pageDescriptors]); + }, [readyKey]); - // 4. One channel group per site; channels keyed by makeId(origin, slug). + // 5. One channel group per in-scope site; channels keyed by + // makeId(origin, slug), from every in-scope manifest that has arrived. const groups = useMemo<ChannelGroup[]>( () => - sites.map((s, i) => ({ + sites.filter(inScope).map((s, i) => ({ id: s.origin, name: s.siteTitle, selectedByDefault: true, order: i, ...(s.accent ? { accent: s.accent } : {}), })), + // eslint-disable-next-line react-hooks/exhaustive-deps [sites], ); - const defaultGroupId = sites[0]?.origin ?? DEFAULT_GROUP_FALLBACK_ID; + const defaultGroupId = + sites.find(inScope)?.origin ?? sites[0]?.origin ?? DEFAULT_GROUP_FALLBACK_ID; const channels = useMemo<ChannelOption[]>(() => { const out: ChannelOption[] = []; sites.forEach((s, i) => { + if (!inScope(s)) return; const manifest = manifestQueries[i]?.data; if (!manifest) return; for (const c of manifest.channels) { @@ -295,20 +502,20 @@ export function MultiSiteDataProvider({ (a, b) => a.groupId.localeCompare(b.groupId) || a.name.localeCompare(b.name), ); // eslint-disable-next-line react-hooks/exhaustive-deps - }, [sites, manifestsSettled]); + }, [sites, pagePlanKey]); - // 5. Merge subs manifests: concat channels (slug → origin-qualified) so the - // chat scope resolves per-origin; sum the live-chat/total counts. - const subsQueries = useQueries({ - queries: sites.map((s) => ({ - queryKey: ["subs-manifest", s.origin], - queryFn: () => readerFor(s.origin).readSubsSiteManifest(), - })), - }); - const subsSettled = subsQueries.every((q) => q.isSuccess || q.isError); + // 6. Merge subs manifests: concat channels (slug → origin-qualified) so the + // chat scope resolves per-origin; sum the live-chat/total counts. Merged + // over the READY archives only (a failed archive contributes nothing). + const subsKey = sites + .map((s, i) => (readyOrigins.has(s.origin) && subsQueries[i]?.data ? s.origin : "")) + .join("\u0000"); const subsManifest = useMemo<SubsManifest | null>(() => { const loaded = sites - .map((s, i) => ({ origin: s.origin, data: subsQueries[i]?.data })) + .map((s, i) => ({ + origin: s.origin, + data: readyOrigins.has(s.origin) ? subsQueries[i]?.data : undefined, + })) .filter((e): e is { origin: string; data: SubsManifest } => !!e.data); if (loaded.length === 0) return null; const channelsOut: SubsManifest["channels"] = []; @@ -331,31 +538,21 @@ export function MultiSiteDataProvider({ generatedAt, }; // eslint-disable-next-line react-hooks/exhaustive-deps - }, [sites, subsSettled]); + }, [subsKey]); - // 5b. Merge posts manifests, same origin-qualification as subs. A member site + // 6b. Merge posts manifests, same origin-qualification as subs. A member site // with no posts corpus 404s; treat that as an empty contribution so one - // video-only origin can't blank the hub's posts scope. - const postsQueries = useQueries({ - queries: sites.map((s) => ({ - queryKey: ["posts-manifest", s.origin], - queryFn: () => - readerFor(s.origin) - .readPostsSiteManifest() - .catch( - (): PostsManifest => ({ - version: 0, - channels: [], - totalCount: 0, - generatedAt: "", - }), - ), - })), - }); - const postsSettled = postsQueries.every((q) => q.isSuccess || q.isError); + // video-only origin can't blank the hub's posts scope (and it is not a + // failure of that archive). + const postsKey = sites + .map((s, i) => (readyOrigins.has(s.origin) && postsQueries[i]?.data ? s.origin : "")) + .join("\u0000"); const postsManifest = useMemo<PostsManifest | null>(() => { const loaded = sites - .map((s, i) => ({ origin: s.origin, data: postsQueries[i]?.data })) + .map((s, i) => ({ + origin: s.origin, + data: readyOrigins.has(s.origin) ? postsQueries[i]?.data : undefined, + })) .filter((e): e is { origin: string; data: PostsManifest } => !!e.data); if (loaded.length === 0) return null; const channelsOut: PostsManifest["channels"] = []; @@ -375,7 +572,7 @@ export function MultiSiteDataProvider({ generatedAt, }; // eslint-disable-next-line react-hooks/exhaustive-deps - }, [sites, postsSettled]); + }, [postsKey]); // Alias dictionaries per federated origin, merged into one list (later origins // shadow earlier ones by id). Missing files resolve to [] — additive only. @@ -384,28 +581,36 @@ export function MultiSiteDataProvider({ queryKey: ["search-aliases", s.origin], queryFn: () => fetchAliases(s.origin), staleTime: Infinity, + enabled: inScope(s), })), }); - const aliasesSettled = aliasQueries.every((q) => q.isSuccess || q.isError); + const aliasesKey = sites + .map((s, i) => (readyOrigins.has(s.origin) && aliasQueries[i]?.data ? s.origin : "")) + .join("\u0000"); const aliases = useMemo<SearchAlias[]>(() => { let merged: SearchAlias[] = []; - for (const q of aliasQueries) if (q.data) merged = mergeAliases(merged, q.data); + sites.forEach((s, i) => { + const data = readyOrigins.has(s.origin) ? aliasQueries[i]?.data : undefined; + if (data) merged = mergeAliases(merged, data); + }); return merged; // eslint-disable-next-line react-hooks/exhaustive-deps - }, [aliasesSettled]); + }, [aliasesKey]); // The hub's OWN /tags.json, not a merge of its members'. See the field note // on SearchDataValue: a federated count nobody can reproduce is worse than no // chip, so until a hub build writes one, hub mode offers no tag chips. const curatedTags = useCuratedTags(""); - // Synthetic merged summaries manifest. TranscriptSearch reads channels/groups - // from the context (above), not from here, but the field is part of the - // SummariesState contract, so provide a coherent merged view. + // Synthetic merged summaries manifest over the READY archives — what the + // search actually covers. TranscriptSearch reads channels/groups from the + // context (above), not from here, but the field is part of the SummariesState + // contract, so provide a coherent merged view. const mergedManifest = useMemo<Manifest | null>(() => { - if (!manifestsSettled) return null; - const loaded = manifestQueries - .map((q) => q.data) + const loaded = sites + .map((s, i) => + readyOrigins.has(s.origin) ? manifestQueries[i]?.data : undefined, + ) .filter((m): m is Manifest => !!m); if (loaded.length === 0) return null; return { @@ -422,7 +627,24 @@ export function MultiSiteDataProvider({ defaultGroupId, }; // eslint-disable-next-line react-hooks/exhaustive-deps - }, [manifestsSettled, groups, defaultGroupId, pageDescriptors]); + }, [readyKey, groups, defaultGroupId, pageDescriptors]); + + // Every in-scope archive failed: say so to the consumers that read `error` + // (the /ask chat). One failure among several is the federation line's job. + const allFailed = + scoped.length > 0 && scoped.every((s) => s.status === "failed"); + const summariesError = useMemo<Error | null>( + () => + allFailed + ? new Error( + scoped.length === 1 + ? `${scoped[0].siteTitle} did not answer.` + : "None of the archives answered.", + ) + : null, + // eslint-disable-next-line react-hooks/exhaustive-deps + [allFailed, statusKey], + ); const summariesState = useMemo<SummariesState>( () => ({ @@ -431,12 +653,19 @@ export function MultiSiteDataProvider({ loadedPages, pageCount: pageDescriptors.length, summariesReady, - error: null, + error: summariesError, }), - [mergedManifest, summaries, loadedPages, pageDescriptors, summariesReady], + [ + mergedManifest, + summaries, + loadedPages, + pageDescriptors, + summariesReady, + summariesError, + ], ); - // origin → accent lookup for provenance in results. + // origin → accent / title lookups for provenance in results. const accentByOrigin = useMemo(() => { const m = new Map<string, string>(); for (const s of sites) if (s.accent) m.set(s.origin, s.accent); @@ -446,6 +675,32 @@ export function MultiSiteDataProvider({ (origin: string) => accentByOrigin.get(origin), [accentByOrigin], ); + const titleByOrigin = useMemo( + () => new Map(sites.map((s) => [s.origin, s.siteTitle])), + [sites], + ); + const siteTitleOf = useCallback( + (origin: string) => titleByOrigin.get(origin), + [titleByOrigin], + ); + + // Retry: refetch every query of this origin that ended in error — its + // manifest, failed pages, and any other feed of it that failed. + const retry = useCallback( + (origin: string) => { + void queryClient.refetchQueries({ + predicate: (q) => + FEDERATED_KEYS.has(String(q.queryKey[0])) && + q.queryKey[1] === origin && + q.state.status === "error", + }); + }, + [queryClient], + ); + const federation = useMemo<FederationState>( + () => ({ sites: siteStates, retry }), + [siteStates, retry], + ); const value = useMemo<SearchDataValue>( () => ({ @@ -461,6 +716,8 @@ export function MultiSiteDataProvider({ accentOf, aliases, curatedTags, + federation, + siteTitleOf, }), [ summariesState, @@ -472,6 +729,8 @@ export function MultiSiteDataProvider({ accentOf, aliases, curatedTags, + federation, + siteTitleOf, ], ); diff --git a/common/components/SearchResults.tsx b/common/components/SearchResults.tsx @@ -190,6 +190,8 @@ export default function SearchResults() { </div> </div> + <FederationStatus /> + {view !== "chart" && resultGroups.length > 0 && ( <div data-testid="selection-toolbar" @@ -318,6 +320,54 @@ export default function SearchResults() { ); } +// Hub mode only: when an archive in scope failed to answer, one plain line +// above the results says how many did, names the ones that did not, and offers +// a Retry for them. The results from the archives that answered render as +// usual underneath — one member failing never empties the page. Nothing in +// single-site mode (no `federation`), nothing while every archive is fine, and +// nothing while any archive in scope is still loading (the chips show that), so +// "N of M" always names every archive it did not count. +function FederationStatus() { + const { federation } = useSearchData(); + if (!federation) return null; + const scoped = federation.sites.filter((s) => s.status !== "off"); + const failed = scoped.filter((s) => s.status === "failed"); + // Only once every archive in scope has settled, so the count is final and + // every archive not counted is named (the chips show what is still loading). + if (failed.length === 0 || scoped.some((s) => s.status === "loading")) { + return null; + } + const answered = scoped.filter((s) => s.status === "ready").length; + const names = failed.map((s) => s.siteTitle); + const list = + names.length === 1 + ? names[0] + : `${names.slice(0, -1).join(", ")} and ${names[names.length - 1]}`; + return ( + <p + role="status" + data-testid="hub-scope-status" + className="mb-2 flex flex-wrap items-baseline gap-x-2 text-sm text-muted-foreground" + > + <span> + {answered} of {scoped.length}{" "} + {scoped.length === 1 ? "archive" : "archives"} answered. {list}{" "} + did not, so{" "} + {names.length === 1 ? "its" : "their"} videos are not in these results. + </span> + <button + type="button" + onClick={() => { + for (const s of failed) federation.retry(s.origin); + }} + className="text-brand transition-colors hover:underline focus-visible:outline-2 focus-visible:-outline-offset-2 focus-visible:outline-ring" + > + Retry + </button> + </p> + ); +} + // ─── Result rendering ──────────────────────────────────────────────────────── type ResultRow = | { kind: "card"; cardIndex: number; group: ResultGroup } @@ -633,6 +683,15 @@ const ResultCard = memo(function ResultCard({ </> )} <span className="basis-full sm:basis-auto text-xs text-muted-foreground shrink-0"> + {/* Hub mode: the archive this came from, in text. */} + {group.source && ( + <> + <span data-testid="result-source" className="text-foreground/80"> + {group.source} + </span> + {" · "} + </> + )} {group.channel && `${group.channel} · `} {group.date} {group.hits.length > 0 && diff --git a/common/components/SearchSessionContext.tsx b/common/components/SearchSessionContext.tsx @@ -158,6 +158,9 @@ export type ResultGroup = { // Provenance accent of the source origin (hub mode only); undefined // single-site, so no marker renders. accent?: string; + // The source archive's title (hub mode only), named on the card beside the + // channel so provenance is not carried by colour alone. Absent single-site. + source?: string; // Set on rows from the social-post corpus. A post is not a video: it has no // timeline to seek, no livestream/age state and no VOD expiry, so the result // card renders a distinct variant rather than a degraded video card. @@ -486,6 +489,7 @@ function useSearchSessionState() { groups: manifestGroups, channelKeyOf, accentOf, + siteTitleOf, } = useSearchData(); const { summaries, loadedPages, pageCount, summariesReady } = summariesState; @@ -873,6 +877,16 @@ function useSearchSessionState() { // those slugs that survived the query tree. For each, attach the per-leaf // hits collected by searchEval. Hits are already sorted by start time // within each video. + // `{ source }` for a hub result, `{}` single-site — so a single-site group + // carries no `source` key at all. + const sourceOf = useCallback( + (slug: string): { source?: string } => { + const title = siteTitleOf?.(splitId(slug).origin); + return title ? { source: title } : {}; + }, + [siteTitleOf], + ); + const resultGroups = useMemo<ResultGroup[]>(() => { if (!transcripts) return []; if (!hasActiveQuery) { @@ -897,6 +911,7 @@ function useSearchSessionState() { ? { curatedTags: t.curatedTags } : {}), accent: accentOf(splitId(t.slug).origin), + ...sourceOf(t.slug), }); } return out; @@ -924,6 +939,7 @@ function useSearchSessionState() { ? { curatedTags: t.curatedTags } : {}), accent: accentOf(splitId(t.slug).origin), + ...sourceOf(t.slug), }); } // Matched POST slugs. They are not in `transcripts` (a disjoint namespace), @@ -958,6 +974,7 @@ function useSearchSessionState() { uploadDate: post.uploadDate, hits: hitsBySlug.get(slug) ?? [], accent: accentOf(splitId(slug).origin), + ...sourceOf(slug), post, }); } @@ -973,6 +990,7 @@ function useSearchSessionState() { treeProgress, passesFilter, accentOf, + sourceOf, committedDateFrom, committedDateTo, committedStates, diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,6 +1,7 @@ # Changelog ## [Unreleased] +- **The hub's search shows each archive's state, lets you choose which archives to search, and no longer waits for the slowest.** Under the line "Searching N archives …" is a row of chips, one per archive on the hub (official and added). Each says whether that archive is loading, how many videos it has in the search once it is in, or that it failed, with a Retry beside it. Pressing a chip takes that archive out of the search: nothing more is fetched from it and its results disappear. The choice is kept in this browser, and an archive it has never seen is searched. The search now runs as soon as one archive has loaded and runs again as each further one arrives, where it used to wait for all of them. When an archive does not answer, a line above the results says so once the others have loaded ("4 of 5 archives answered. Hasanalyzer did not, so its videos are not in these results.") with a Retry, and the other archives' results show as usual. None of that archive's videos, posts or live chat are searched until a Retry succeeds; before, it dropped out silently or in part. Each result names its archive in text before the channel ("Jeralyzer · TheQuartering · 2026-09-25"). The hub now loads at most six summaries pages at a time from each archive, so a large archive does not hold up the small ones. The line under the archives now counts "videos", as the chips do. `/ask` on the hub is unchanged: it searches every archive and waits for all of them. Published sites are unchanged. Needs a rebuild and deploy of the hub. - **The hub and the homepage drop their subtitles, and both carry a Ko-fi link.** The headings are in Title Case on both pages: **Official Instances**, **Archives You Added** (hub) and **What It Does** (homepage). The line "The archives I run. Anyone can run their own." is gone from both, and the official-instance cards no longer show the site's description under its name; the four figures stay. The footer of the homepage and of the hub has a plain link reading "Ko-fi" to `https://ko-fi.com/archilyzer`. A published site's footer does not: an archive someone else hosts never carries it. Needs a rebuild and deploy of the hub and the homepage. - **A job that waited in a queue no longer ends `failed` after doing its work.** A sync, download or other job queued behind another on the same platform ran its final page refresh outside any request, where Next refuses it, so the job read `failed` and `pnpm ops … --wait` exited 1 even though the work was done (the teamrcn sync on 2026-09-25). The refresh is now skipped there with one warning in the server log; the pages re-read disk on their next load anyway. - **Downloads no longer sleep after a video the download filter declined.** The "sleep between downloads" (30 s by default) ran after every video, including each one the channel's download filter declined before fetching anything. A filtered channel's download-missing slept 193 times for 14 downloads on 2026-09-25. It still sleeps after every real fetch and after every failure, per-video ones included. diff --git a/export/app/components/hub/HubHome.tsx b/export/app/components/hub/HubHome.tsx @@ -19,6 +19,8 @@ import { useRegistry } from "yt-dlp-transcript-common/components/siteRegistry"; import ArchiveShelf from "./ArchiveShelf"; import HubStats from "./HubStats"; import HubOfflineManager from "./HubOfflineManager"; +import HubScope from "./HubScope"; +import { useHubScope } from "./useHubScope"; // `transcriptDownloads` comes from the server parent's currentSite() (a client // component cannot read site.json): false hides the modal's per-video export @@ -29,24 +31,34 @@ export default function HubHome({ transcriptDownloads?: boolean; }) { const { sites } = useRegistry(); + // Which archives the search covers (every one unless this browser switched + // it off with its chip). + const { isOn, toggle } = useHubScope(); - // Carry each site's accent through to the merged search for provenance. + // Carry each site's accent through to the merged search for provenance, and + // its scope: an archive switched off is not fetched at all. const federated = useMemo<FederatedSite[]>( () => sites.map((s) => ({ origin: s.origin, siteTitle: s.siteTitle, accent: s.accent, + enabled: isOn(s.origin), })), - [sites], + [sites, isOn], ); return ( <PlayerProvider features={{ transcriptDownloads }}> <div className="flex flex-col gap-6"> <ArchiveShelf /> - <MultiSiteDataProvider sites={federated}> - <HubStats /> + {/* progressive: the search runs as soon as one archive is in, and + re-runs as each further archive arrives. */} + <MultiSiteDataProvider sites={federated} progressive> + <div className="flex flex-col gap-3"> + <HubStats /> + <HubScope onToggle={toggle} /> + </div> <TranscriptSearch /> <HubOfflineManager /> </MultiSiteDataProvider> diff --git a/export/app/components/hub/HubScope.tsx b/export/app/components/hub/HubScope.tsx @@ -0,0 +1,104 @@ +"use client"; + +// The row of scope chips above the query builder: one per archive on the hub, +// official and added. Pressing a chip takes that archive in or out of the +// search (useHubScope remembers it per browser). Each chip says what its +// archive is doing — loading, its record count once ready (its summaries +// manifest's totalCount, the number the results header counts), or failed with +// a Retry beside it — so a slow or dead member is visible, not a silent gap. + +import { cn } from "yt-dlp-transcript-common/lib/utils"; +import { + useSearchData, + type FederatedSiteState, +} from "yt-dlp-transcript-common/components/SearchDataContext"; + +function detail(s: FederatedSiteState): string { + switch (s.status) { + case "off": + return "off"; + case "loading": + return "loading…"; + case "failed": + return "failed"; + case "ready": + return s.count === undefined ? "ready" : s.count.toLocaleString(); + } +} + +export default function HubScope({ + onToggle, +}: { + onToggle: (origin: string) => void; +}) { + const { federation } = useSearchData(); + if (!federation || federation.sites.length === 0) return null; + const { sites, retry } = federation; + + return ( + <div + role="group" + aria-label="Archives to search" + className="flex flex-wrap items-center gap-2" + > + {sites.map((s) => { + const on = s.status !== "off"; + return ( + <span + key={s.origin} + data-testid={`hub-scope-chip-${s.origin}`} + data-status={s.status} + className={cn( + "inline-flex min-w-0 items-stretch overflow-hidden rounded border text-sm transition-colors", + on ? "border-border-strong bg-card" : "border-border bg-transparent", + s.status === "failed" && "border-destructive/60", + )} + > + <button + type="button" + aria-pressed={on} + title={ + on + ? `${s.status === "ready" && s.count !== undefined ? `${s.count.toLocaleString()} videos from ${s.siteTitle} are in the search. ` : ""}Press to leave ${s.siteTitle} out.` + : `Press to search ${s.siteTitle} again.` + } + onClick={() => onToggle(s.origin)} + className={cn( + "flex min-w-0 items-center gap-1.5 px-2.5 py-1.5 select-none hover:bg-accent hover:text-accent-foreground focus-visible:outline-2 focus-visible:-outline-offset-2 focus-visible:outline-ring", + on ? "text-foreground" : "text-muted-foreground", + )} + > + {s.accent && ( + <span + aria-hidden="true" + className="inline-block size-2 shrink-0 rounded-full" + style={{ background: s.accent }} + /> + )} + <span className="truncate">{s.siteTitle}</span> + <span + className={cn( + "shrink-0 text-xs tabular-nums", + s.status === "failed" ? "text-destructive" : "text-muted-foreground", + )} + > + {detail(s)} + </span> + </button> + {s.status === "failed" && ( + <button + type="button" + onClick={() => retry(s.origin)} + aria-label={`Retry ${s.siteTitle}`} + title={s.error ? `${s.error} — try again` : "Try again"} + className="border-l border-border px-2.5 py-1.5 text-xs text-brand hover:bg-accent focus-visible:outline-2 focus-visible:-outline-offset-2 focus-visible:outline-ring" + > + Retry + </button> + )} + </span> + ); + })} + </div> + ); +} diff --git a/export/app/components/hub/HubStats.tsx b/export/app/components/hub/HubStats.tsx @@ -5,7 +5,7 @@ // archive on the hub, the visitor's added ones included — so they are not the // cards' numbers and are not meant to match them. The official cards show the // build-time figures the homepage prints (hub-summary.json); this line shows -// what actually loaded. +// what actually loaded, from the archives this browser has in scope. import { useSearchData } from "yt-dlp-transcript-common/components/SearchDataContext"; import { useRegistry } from "yt-dlp-transcript-common/components/siteRegistry"; @@ -21,18 +21,23 @@ function part(n: number, one: string, many: string) { export default function HubStats() { const { sites } = useRegistry(); - const { summariesState, channels } = useSearchData(); - const transcripts = summariesState.manifest?.totalCount ?? 0; + const { summariesState, channels, federation } = useSearchData(); + // The archives in scope (the visitor's chips), and what has arrived from them: + // channels as each manifest lands, videos as each archive is ready. + const archives = + federation?.sites.filter((s) => s.status !== "off").length ?? sites.length; + // The ready archives' record counts — the same figures the chips show. + const videos = summariesState.manifest?.totalCount ?? 0; if (sites.length === 0) return null; return ( <p className="text-sm text-muted-foreground"> - Searching {part(sites.length, "archive", "archives")} + Searching {part(archives, "archive", "archives")} <span className="mx-2 text-muted-foreground/40">·</span> {part(channels.length, "channel", "channels")} <span className="mx-2 text-muted-foreground/40">·</span> - {part(transcripts, "transcript", "transcripts")} right now. + {part(videos, "video", "videos")} right now. </p> ); } diff --git a/export/app/components/hub/useHubScope.ts b/export/app/components/hub/useHubScope.ts @@ -0,0 +1,65 @@ +"use client"; + +// Which archives the hub's search covers, per browser. Every archive is in by +// default; the visitor takes one out with its scope chip. Stored as the set of +// origins switched OFF, so an archive the browser has never seen (a new +// official instance, a site added later) is searched until someone says not to. +// localStorage is a convenience here: when it is unavailable every archive is +// simply searched. + +import { useCallback, useEffect, useRef, useState } from "react"; + +const STORAGE_KEY = "ytdlp-tb:hub-scope"; + +function readOff(): ReadonlySet<string> { + if (typeof window === "undefined") return new Set(); + try { + const raw = window.localStorage.getItem(STORAGE_KEY); + const parsed: unknown = raw ? JSON.parse(raw) : null; + const off = + parsed && typeof parsed === "object" && Array.isArray((parsed as { off?: unknown }).off) + ? (parsed as { off: unknown[] }).off + : []; + return new Set(off.filter((o): o is string => typeof o === "string")); + } catch { + return new Set(); + } +} + +function writeOff(off: ReadonlySet<string>) { + try { + window.localStorage.setItem( + STORAGE_KEY, + JSON.stringify({ off: Array.from(off) }), + ); + } catch { + // Storage full / disabled — the choice still holds for this visit. + } +} + +export function useHubScope() { + // Read on the client's first render. The registry's server snapshot is empty, + // so nothing that depends on this renders during hydration. + const [off, setOff] = useState<ReadonlySet<string>>(readOff); + const first = useRef(true); + useEffect(() => { + if (first.current) { + first.current = false; + return; + } + writeOff(off); + }, [off]); + + const toggle = useCallback((origin: string) => { + setOff((prev) => { + const next = new Set(prev); + if (next.has(origin)) next.delete(origin); + else next.add(origin); + return next; + }); + }, []); + + const isOn = useCallback((origin: string) => !off.has(origin), [off]); + + return { isOn, toggle }; +} diff --git a/export/e2e-hub/federated-search.spec.ts b/export/e2e-hub/federated-search.spec.ts @@ -0,0 +1,275 @@ +import { expect, test, type Page, type Route } from "@playwright/test"; + +// The hub's federated search, per archive: two official members (hub-sites.json) +// served by route mocks WITH CORS, each with one video. Proves the scope chips +// (take an archive out; it is not fetched, and the choice survives a reload), +// per-archive state (a member whose pages cannot be read is "failed", the page +// says "1 of 2 archives answered" with a Retry, and the healthy member's results +// still render), Retry, progressive readiness (a slow member does not hold the +// fast one's results back), and a card naming its source archive in text. + +const ORIGIN_A = "http://localhost:4598"; +const ORIGIN_B = "http://localhost:4599"; +const CORS = { "access-control-allow-origin": "*" }; + +async function fulfillJson(route: Route, body: unknown) { + await route.fulfill({ + status: 200, + contentType: "application/json", + headers: CORS, + body: JSON.stringify(body), + }); +} + +type Member = { origin: string; title: string; channel: string; video: string }; +const A: Member = { origin: ORIGIN_A, title: "Origin A", channel: "Channel A", video: "Alpha Video" }; +const B: Member = { origin: ORIGIN_B, title: "Origin B", channel: "Channel B", video: "Bravo Video" }; +const C: Member = { origin: "http://localhost:4597", title: "Origin C", channel: "Channel C", video: "Charlie Video" }; + +function manifestOf(m: Member) { + return { + version: 3, + totalCount: 1, + pageSize: 1000, + pageCount: 1, + generatedAt: "2026-01-01T00:00:00.000Z", + channels: [{ name: m.channel, count: 1, slug: "chan", groupId: "g" }], + groups: [{ id: "g", name: m.title, selectedByDefault: true }], + defaultGroupId: "g", + }; +} + +function pageOf(m: Member) { + const id = m === A ? "vida1" : m === B ? "vidb1" : "vidc1"; + return [ + { + slug: `chan/${id}`, + id, + channelSlug: "chan", + title: m.video, + uploadDate: m === A ? "20260102" : "20260101", + date: m === A ? "2026-01-02" : "2026-01-01", + duration: "5:00", + channel: m.channel, + isLivestream: false, + ageRestricted: false, + isDeleted: false, + isUnlisted: false, + platform: "youtube" as const, + webpageUrl: `${m.origin}/${id}`, + }, + ]; +} + +// Every request to a member origin is counted; `pages` decides how its +// summaries page is answered. +type PageMode = "ok" | "abort" | "hold"; +type MemberMock = { + requests: string[]; + pages: PageMode; + held: Route[]; +}; + +async function mockMember(page: Page, m: Member): Promise<MemberMock> { + const mock: MemberMock = { requests: [], pages: "ok", held: [] }; + await page.route(`${m.origin}/**`, async (route) => { + const url = new URL(route.request().url()); + mock.requests.push(url.pathname); + if (url.pathname === "/summaries/manifest.json") { + return fulfillJson(route, manifestOf(m)); + } + if (/^\/summaries\/page-\d+\.json$/.test(url.pathname)) { + if (mock.pages === "abort") return route.abort(); + if (mock.pages === "hold") { + mock.held.push(route); + return; + } + return fulfillJson(route, pageOf(m)); + } + if (url.pathname === "/subs/manifest.json") { + return fulfillJson(route, { + version: 4, + channels: [], + totalCount: 0, + liveChatTotalCount: 0, + generatedAt: "2026-01-01T00:00:00.000Z", + }); + } + // No posts, no aliases: a clean 404 (a member without a posts corpus is + // not a failure). + return route.fulfill({ status: 404, headers: CORS, body: "" }); + }); + return mock; +} + +async function setup(page: Page, members: Member[] = [A, B]) { + await page.route("**/hub-sites.json", (r) => + fulfillJson( + r, + members.map((m, i) => ({ + siteId: `origin${i}`, + siteTitle: m.title, + siteUrl: m.origin, + pwa: false, + contract: 1, + })), + ), + ); + await page.route("**/hub-summary.json", (r) => + r.fulfill({ status: 404, body: "" }), + ); + const a = await mockMember(page, A); + const b = await mockMember(page, B); + const c = members.includes(C) ? await mockMember(page, C) : null; + return { a, b, c }; +} + +const chip = (page: Page, m: Member) => + page.getByTestId(`hub-scope-chip-${m.origin}`); +const resultFrom = (page: Page, m: Member) => + page.locator(`[data-result-slug^="${m.origin}"]`); +const status = (page: Page) => page.getByTestId("hub-scope-status"); + +test.describe("hub federated search — scope, per-archive state, attribution", () => { + test("both official archives are searched by default and each card names its archive", async ({ + page, + }) => { + await setup(page); + await page.goto("/"); + await expect(chip(page, A)).toHaveAttribute("data-status", "ready"); + await expect(chip(page, B)).toHaveAttribute("data-status", "ready"); + await expect(chip(page, B)).toContainText("Origin B"); + await expect(chip(page, B).getByRole("button", { name: /Origin B/ })).toHaveAttribute( + "aria-pressed", + "true", + ); + // Browse listing: one card per archive, each naming its source in text. + const b = resultFrom(page, B); + await expect(b).toHaveCount(1); + await expect(b.getByTestId("result-source")).toHaveText("Origin B"); + await expect(b).toContainText("Channel B"); + await expect(resultFrom(page, A).getByTestId("result-source")).toHaveText( + "Origin A", + ); + // Nothing failed, so no "N of M" line. + await expect(status(page)).toHaveCount(0); + }); + + test("toggling an archive off removes its results and stops its fetches, across a reload", async ({ + page, + }) => { + const mocks = await setup(page); + await page.goto("/"); + await expect(resultFrom(page, B)).toHaveCount(1); + await expect(resultFrom(page, A)).toHaveCount(1); + + await chip(page, B).getByRole("button", { name: /Origin B/ }).click(); + await expect(chip(page, B)).toHaveAttribute("data-status", "off"); + await expect( + chip(page, B).getByRole("button", { name: /Origin B/ }), + ).toHaveAttribute("aria-pressed", "false"); + await expect(resultFrom(page, B)).toHaveCount(0); + await expect(resultFrom(page, A)).toHaveCount(1); + await expect(page.getByText(/Searching 1 archive\b/)).toBeVisible(); + + // The choice is this browser's: after a reload Origin B stays off and is + // not fetched at all. + mocks.b.requests.length = 0; + await page.reload(); + await expect(chip(page, A)).toHaveAttribute("data-status", "ready"); + await expect(chip(page, B)).toHaveAttribute("data-status", "off"); + await expect(resultFrom(page, A)).toHaveCount(1); + await expect(resultFrom(page, B)).toHaveCount(0); + expect(mocks.b.requests).toEqual([]); + + // Back on: fetched, and its results return. + await chip(page, B).getByRole("button", { name: /Origin B/ }).click(); + await expect(chip(page, B)).toHaveAttribute("data-status", "ready"); + await expect(resultFrom(page, B)).toHaveCount(1); + expect(mocks.b.requests).toContain("/summaries/manifest.json"); + }); + + test("a member that cannot be read fails alone; Retry brings it back", async ({ + page, + }) => { + const mocks = await setup(page); + mocks.b.pages = "abort"; + await page.goto("/"); + + await expect(chip(page, B)).toHaveAttribute("data-status", "failed"); + await expect(chip(page, B)).toContainText("failed"); + await expect( + chip(page, B).getByRole("button", { name: "Retry Origin B" }), + ).toBeVisible(); + await expect(chip(page, A)).toHaveAttribute("data-status", "ready"); + await expect(status(page)).toHaveAttribute("role", "status"); + await expect(status(page)).toContainText("1 of 2 archives answered."); + await expect(status(page)).toContainText("Origin B did not"); + // The healthy member's results still render. + await expect(resultFrom(page, A)).toHaveCount(1); + await expect(resultFrom(page, B)).toHaveCount(0); + + // Fix the member, then Retry from the line. + mocks.b.pages = "ok"; + await status(page).getByRole("button", { name: "Retry", exact: true }).click(); + await expect(chip(page, B)).toHaveAttribute("data-status", "ready"); + await expect(status(page)).toHaveCount(0); + await expect(resultFrom(page, B)).toHaveCount(1); + await expect(resultFrom(page, A)).toHaveCount(1); + }); + + test("a slow member does not hold back a fast one, and a search re-runs as it arrives", async ({ + page, + }) => { + const mocks = await setup(page); + mocks.b.pages = "hold"; + await page.goto("/"); + + await expect(chip(page, A)).toHaveAttribute("data-status", "ready"); + await expect(chip(page, B)).toHaveAttribute("data-status", "loading"); + await expect(chip(page, B)).toContainText("loading"); + + // Search while Origin B is still loading: Origin A answers now. + await page.getByTestId("query-builder").waitFor(); + await page + .locator('[data-testid^="leaf-scope-"]') + .first() + .selectOption({ label: "Title / channel" }); + const input = page.locator('[data-testid^="leaf-query-"]').first(); + await input.click(); + await input.fill("video"); + await input.press("Enter"); + await page.waitForURL(/[?&]qt=/); + await expect(resultFrom(page, A)).toHaveCount(1); + await expect(resultFrom(page, B)).toHaveCount(0); + + // Origin B arrives: the same search now covers it. + await expect.poll(() => mocks.b.held.length).toBeGreaterThan(0); + for (const r of mocks.b.held.splice(0)) await fulfillJson(r, pageOf(B)); + await expect(chip(page, B)).toHaveAttribute("data-status", "ready"); + await expect(resultFrom(page, B)).toHaveCount(1); + await expect(resultFrom(page, A)).toHaveCount(1); + }); + + test("the answered line waits until every archive in scope has settled", async ({ + page, + }) => { + const mocks = await setup(page, [A, B, C]); + mocks.b.pages = "abort"; + mocks.c!.pages = "hold"; + await page.goto("/"); + + await expect(chip(page, B)).toHaveAttribute("data-status", "failed"); + await expect(chip(page, C)).toHaveAttribute("data-status", "loading"); + // Origin C is neither answered nor failed yet: no count that leaves it out. + await expect(status(page)).toHaveCount(0); + await expect(resultFrom(page, A)).toHaveCount(1); + + await expect.poll(() => mocks.c!.held.length).toBeGreaterThan(0); + for (const r of mocks.c!.held.splice(0)) await fulfillJson(r, pageOf(C)); + await expect(chip(page, C)).toHaveAttribute("data-status", "ready"); + await expect(status(page)).toContainText("2 of 3 archives answered."); + await expect(status(page)).toContainText("Origin B did not"); + await expect(resultFrom(page, C)).toHaveCount(1); + }); +}); diff --git a/export/e2e-hub/federation.spec.ts b/export/e2e-hub/federation.spec.ts @@ -139,7 +139,7 @@ test.describe("hub federation — cross-origin browse + search", () => { await expect(builtWith).toHaveAttribute("href", PROJECT_URL); }); - test("adds an archive by URL and shows it under Archives you added", async ({ page }) => { + test("adds an archive by URL and shows it under Archives You Added", async ({ page }) => { await mockOriginB(page); await addArchive(page, ORIGIN_B); // Spine carries the descriptor's siteTitle → the site was read cross-origin. @@ -148,7 +148,7 @@ test.describe("hub federation — cross-origin browse + search", () => { ).toBeVisible(); // An added archive gets its own section, and its own Remove control. await expect( - page.getByRole("heading", { level: 2, name: "Archives you added" }), + page.getByRole("heading", { level: 2, name: "Archives You Added", exact: true }), ).toBeVisible(); await expect( page.getByRole("button", { name: "Remove Origin B" }), diff --git a/plans/release-9.md b/plans/release-9.md @@ -463,6 +463,126 @@ editor unit, test:scripts and mcp not run: nothing they build or test changed be - Needs a rebuild + deploy of the hub AND the homepage; this slice did not run build-hub, deploy-hub or deploy homepage. +### Slice C2, as shipped — federated search (2026-09-25) + +Branch `one-core/c2-federated-search` off `main` `6820bd21`; merged `main` (`8eb55add`, slice C1b) as +`9d20d3fc`. **The operator-approved target (2026-09-25):** the hub searches every official instance +by default, a visitor can take any archive out, each archive's state is visible, the search runs as +soon as one archive is in, a failed member is named and retryable, a result names its archive in text, +and pages are fetched per archive behind its manifest. Scoping: `$T/r9-hub-scope.md` §3. + +**The provider** (`common/components/SearchDataContext.tsx`, `MultiSiteDataProvider`). `FederatedSite` +gains `enabled` (default true): false means none of that archive's feeds are fetched (`enabled: false` +on its manifest, subs, posts and alias queries; no page descriptors) and nothing already cached from it +is merged — TanStack returns cached `data` for a disabled query, so every merge checks scope itself. +Per archive, its state is derived from its manifest, page, subs and posts queries: `ready` (manifest +in, every page in, and its subs + posts manifests settled — a posts 404 is an empty manifest, a subs +error counts as settled), `failed` (manifest or a page errored and nothing of it is fetching), `loading` +(otherwise, including while a Retry is in flight), `off`. `SearchDataValue` gains two OPTIONAL fields, +set only by the multi-site provider: `federation: {sites: FederatedSiteState[], retry(origin)}` +(`{origin, siteTitle, accent?, status, count?, error?}`, `count` = the manifest's `totalCount`) and +`siteTitleOf(origin)`. `retry` is `refetchQueries` over this provider's five feeds (`manifest`, `summaries-page`, +`subs-manifest`, `posts-manifest`, `search-aliases`) of that origin whose status is `error`. `summariesState.error` is no longer always null: it is set when +EVERY in-scope archive failed (read by `/ask`'s corpus-error line). `SingleSiteDataProvider` is not +touched. +- **Merge.** An archive's records join the merged list only when it is `ready`, so the list grows one + whole archive at a time; the merged list's identity is keyed on the ready set alone, so the search + session (which re-runs its pipeline on every new `summaries` identity) re-runs once per archive, not + per page or per manifest. The subs, posts and alias merges are keyed on the same ready set, so every contribution of an +archive lands in one re-run and a failed archive contributes nothing — no videos, posts or chat +(before: whatever pages had loaded, and its posts regardless). + The synthetic merged manifest (HubStats' transcript figure) covers the ready archives. Channel groups + and channels cover the in-scope archives whose manifest has arrived. +- **Progressive.** A new `progressive` prop: `summariesReady` turns true at the first ready in-scope + archive. Without it (the `/ask` hub, `AskHub.tsx`, unchanged), it waits until every in-scope archive + has settled, ready or failed. HubHome passes it. The release-8 slice E rule holds unchanged: a + restored query is `runHeld` in `SearchSessionContext`, independent of readiness, so it still waits + for the visitor; a `qt=` link runs over the first archive and re-runs as each further one arrives. +- **Speed.** Pages are requested per archive only once its manifest has settled and only while it is in + scope (as before for the manifest gate; new for scope), and each page fetch goes through a per-origin + gate of **6 in flight** (`PAGE_FETCHES_PER_SITE`), so a big member cannot take the browser's + connections from the small ones. The page `queryFn` consumes TanStack's `signal`, so switching an + archive off mid-load cancels its queries and a page still waiting at the gate bails unfetched. No new build-time data. + +**The page** (`export/app/components/hub/`). `useHubScope.ts` keeps the set of origins switched OFF in +`localStorage["ytdlp-tb:hub-scope"]` (`{off: [...]}`, every access in try/catch), so an archive the +browser has not seen — a new official instance, an added one — is searched. `HubScope.tsx` is the row +of chips under the live line, above the query builder: one per archive on the page (official and +added), `data-testid="hub-scope-chip-<origin>"` + `data-status`, a toggle button (`aria-pressed`, the +archive's title and `loading…` / its record count / `failed` / `off`) and, when failed, a `Retry +<title>` button. `HubStats` counts the archives in scope and says "videos" (the chips' word for the same figure; was +"transcripts"). Every chip button and the line's Retry carry a literal `focus-visible:outline-2 +focus-visible:-outline-offset-2 focus-visible:outline-ring` (the chip's `overflow-hidden` clipped the +default outline). `common/components/SearchResults.tsx`: one line +above the results when an in-scope archive failed, `role="status"`, `data-testid="hub-scope-status"`: +"N of M archives answered. X did not, so its videos are not in these results. Retry" ("archive" when +M = 1) — only once NOTHING in scope is loading, so M = ready + failed and every archive not counted is +named; nothing while every archive is fine (the chips show loading). A result card's meta line reads +`<archive> · <channel> · <date>` (`data-testid="result-source"`, from `ResultGroup.source`, set via +`siteTitleOf` in `SearchSessionContext`); single-site groups carry no `source` key. The accent stripe +and the `data-result-slug` contract are unchanged. + +**Specs.** New `export/e2e-hub/federated-search.spec.ts` (+5), two official members route-mocked with +CORS: (1) both searched by default, both chips `ready`, each card names its archive, no status line; +(2) toggling Origin B off removes its card and after a reload it stays off with **zero requests** to its +origin, and back on it is fetched and returns; (3) Origin B's pages aborted → chip `failed` with `Retry +Origin B`, "1 of 2 archives answered." / "Origin B did not", Origin A's card renders; fixed + the line's +Retry → chip `ready`, line gone, both cards; (4) Origin B's page held → Origin A's search result renders +while B is `loading`, and the same search covers B once it is released. (5, review) three members, B failed and C +held → no line; C released → "2 of 3 archives answered." naming B. `federation.spec.ts`: "Archives +You Added" exact (C1b's Title Case). Every existing hub test id / accessible name is unchanged. + +| sha | what | +|---|---| +| `af8b30ed` | provider: scope (`enabled`), per-archive state + Retry (`federation`), ready-only merge, `progressive`, per-origin page gate of 6, `error` when all failed, `siteTitleOf` | +| `f7f7482f` | hub: `useHubScope`, `HubScope` chips, HubHome wiring (`progressive`), HubStats counts the archives in scope | +| `d0225a20` | a result names its archive (`ResultGroup.source`); the "N of M archives answered." line with Retry | +| `0b891b8f` | `e2e-hub/federated-search.spec.ts` (+4) | +| `9d20d3fc` | merge `main` (`8eb55add`, C1b) — no conflicts | +| `6a69c034` | `federation.spec.ts`: "Archives You Added", exact | +| `a0bd7b02` | a chip's title says what its number counts and what pressing it does | +| `40818bbb` | the merged list's identity keys on the ready set alone (no re-run when another manifest lands) | +| `d54e6814` | review fixes: ready needs subs + posts settled and their merges key on the ready set; the answered line waits for every archive (+ `role="status"`, singular); focus rings; page `signal`; "videos"; Retry scoped to the five feeds; spec +1 | + +**Gates.** tsc clean (per commit; full workspace on `9d20d3fc`+ and on `40818bbb`). Common tests +**1,845/1,845** (no new unit tests: the provider is a hook, covered by the e2e). test:scripts **162 passed ++ 1 skipped**. `pnpm --filter export exec next build` (site) ok, 29 s; `compose:hub` (worktree, no +corpus: "0 built-in pool site(s) …; hub-summary.json skipped") + `INSTANCE_MODE=hub … next build` ok, +32 s. e2e, queued, all on the merged tree: **`e2e:hub` 15 passed** (pre-merge, 48.9 s), **16 passed** +(merged, 53.2 s), **16 passed** (`40818bbb`, 47.0 s); **export full suite 195 passed**, 8.5 min (single +site unchanged); **`e2e:2origin`** (`TWO_ORIGIN_REBUILD=1`) **3 passed**, 46.9 s on `a0bd7b02`, and **3 passed**, +1.0 min on `40818bbb`. + +**Review (SHIP AFTER FIXES) re-gate** on `d54e6814`: tsc clean; **`e2e:hub` 17 passed** (16 + 1), 52.9 s. The seven compose-hub outputs + `sw.js` in the worktree's `export/public` were +swapped for copies before every build/e2e and relinked after; the primary's copies were not written by +this slice (one change seen, `hub-summary.json` at 20:18:20, is the parent's build-hub, which wrote the +other six at 20:17:53). Editor build/unit and mcp not run: nothing they build changed. **Numbers tools: +none.** + +**Seen against the live members** (the built hub served locally with the real five-site +`hub-sites.json`, service workers blocked so routes apply): first result card at 1.0–2.2 s; chips ready +Hasanalyzer/Rekietalyzer ~1.6–2.0 s, Bonnellyzer ~2–2.9 s, Anilyzer ~2–4.5 s, Jeralyzer ~3.1–5.1 s. +With Hasanalyzer's pages aborted: "4 of 5 archives answered. Hasanalyzer did not, …", its chip failed +with Retry, 72,360 records from the other four listed. Checked at 1280 dark and 390 dark. + +**Found and left:** +- The chip's number is the member's summaries `totalCount` — its record (video) count, the number the + results header counts ("All videos (N)") — not the card's "transcripts" figure, which is the + homepage's build-time `transcribed.total`. The chip's title says "N videos from X are in the search"; + the live line now says "videos" too. +- A member whose manifest loaded but whose pages failed keeps its channel group in the filter panel + (the channels are known); its records are not searched. Hiding the group on failure would churn the + groups on every Retry. +- **Follow-up:** the `/ask` hub has no scope chips and searches every archive on the page, as before; + the scope is the front page's. Honouring it is `useHubScope().isOn` in `AskHub`. It waits for every archive to settle (not progressive), so a chat never grounds in + a half-loaded federation silently. +- **Follow-up:** official cards have no `accent` (C1: they wear `seriesColor` on the page), so result + stripes and chip dots appear only for an archive that sets one — the text attribution is what names + the source. The fix is one shared fallback (a `shelfAccent(site, summary, i)` out of `ArchiveShelf`) + used by HubHome and AskHub too. +- `route()` does not see requests a service worker makes; a manual check of the built hub needs + `serviceWorkers: "block"`. The e2e-hub suite runs `next dev`, where no SW is registered. + ## Rollout ## Rollout 2026-09-25 (evening) — `0213f6c8` live on :3001 (third restart of the day)