Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 3073d30b974e9b70bab47b4f0b9cbb6ee74585e9
parent bdf9511f75e4531659bd3b1d61ee9dd77682ac43
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Mon,  6 Jul 2026 21:39:58 -0400

Add in-browser bring-your-own-key AI chat (/ask)

Layer 3 of bring-your-own-AI (client-side only; export stays fully static):
- lib/askProvider.ts: browser-direct streaming to Anthropic (with the
  dangerous-direct-browser-access header), OpenAI, and Google Gemini, behind one
  askStream() adapter. No inference hosted; the user's key goes only to the
  provider.
- lib/askRetrieval.ts: retrieval over the published shards via the corpus.json
  contract — same-origin for a site, fanning out over member origins for a hub.
- ask/AskChat.tsx: chat UI — provider/key/model (localStorage, opt-out),
  retrieval → context → streamed answer with cited sources (links + timestamps).
- ask/page.tsx: static shell, hub-aware header.

Verified: export tsc clean; dev server renders /ask/ (200) with the chat UI and
all three providers. Retrieval mirrors the MCP scan already verified on real
data; a full provider round-trip needs a live API key.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

Diffstat:
Mexport/CHANGELOG.md | 1+
Aexport/app/ask/AskChat.tsx | 384+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aexport/app/ask/page.tsx | 40++++++++++++++++++++++++++++++++++++++++
Aexport/app/lib/askProvider.ts | 243+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aexport/app/lib/askRetrieval.ts | 183+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
5 files changed, 851 insertions(+), 0 deletions(-)

diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md @@ -4,6 +4,7 @@ - **Bring-your-own-AI: the archive is now machine-navigable for AI tools.** Every site publishes a small fixed set of discovery files — `llms.txt` (an LLM-readable overview) and `corpus.json` (a documented index of the channels and *how to fetch any transcript* from the existing paginated JSON shards), plus `robots.txt` and a `sitemap.xml`. Nothing is generated per video (the shard scheme is documented instead), so the file count stays constant no matter how large the corpus grows. This lets Claude Code and other tools browse and answer questions about the archive by fetching a couple of URLs. The federated hub publishes an aggregate `corpus.json`/`llms.txt` spanning every member site. - **New MCP server (`mcp/`) for Claude Code, Cursor, and other MCP clients.** A local tool that exposes the archive as MCP tools — `list_channels`, `search_transcripts` (timestamped snippets), `get_transcript`, `get_video_metadata` — reading the same static shards over disk or HTTP. It can point at one site or federate a whole hub. It never changes or hosts the site; see `mcp/README.md` for setup. - **New "Use with AI" page + copy-for-AI buttons.** A `/use-with-ai` page (linked from the header and footer) explains the in-browser chat, the `llms.txt`/`corpus.json` discovery files, and the MCP server. The player toolbar gains a "Copy as Markdown" control that copies the current transcript or live chat as clean, timestamped markdown for pasting into any AI chat, and the search results header gains a "Copy for AI" button that copies the matched videos and their hit snippets as context. +- **New in-browser "Ask AI" chat (`/ask`), bring-your-own-key.** Ask a question and get an answer grounded in the transcripts, with citations back to the source videos and timestamps. It runs entirely in your browser: it searches the published transcript shards, sends the relevant excerpts to the AI provider you choose (Anthropic, OpenAI, or Google Gemini) using your own API key, and streams the reply. The key is stored only on your device (or just for the session) and requests go straight to the provider — this site hosts no AI and sees no key. On the federated hub the chat searches across every member site. ## [0.6.4] - 2026-07-06 - **Fixed: sites always opened in light mode until you re-picked a theme.** If you'd chosen dark (or left it on "system" with a dark device), the page still loaded light on every visit and only switched after you opened the theme menu again. The saved theme is now re-applied before the page paints, so your choice sticks across reloads — no flash, no re-toggling. diff --git a/export/app/ask/AskChat.tsx b/export/app/ask/AskChat.tsx @@ -0,0 +1,384 @@ +"use client"; + +import { useCallback, useEffect, useRef, useState } from "react"; +import { + PROVIDERS, + askStream, + type Provider, + type ChatMessage, +} from "../lib/askProvider"; +import { + loadCorpus, + retrieve, + buildContext, + type CorpusInfo, + type RetrievedVideo, +} from "../lib/askRetrieval"; + +const SYSTEM_PROMPT = + "You are a helpful assistant answering questions about a video-transcript " + + "archive. Base your answer ONLY on the transcript excerpts provided with the " + + "user's question. Cite the excerpts you use by their bracketed number, e.g. " + + "[1]. Each excerpt shows a video title, channel, and timestamped lines. If the " + + "excerpts do not contain enough to answer, say so plainly rather than guessing. " + + "Be concise."; + +const K_PROVIDER = "ytdlp-tb:ai:provider"; +const K_REMEMBER = "ytdlp-tb:ai:remember"; +const keyFor = (p: Provider) => `ytdlp-tb:ai:key:${p}`; +const modelFor = (p: Provider) => `ytdlp-tb:ai:model:${p}`; + +type UiMessage = { + role: "user" | "assistant"; + content: string; + sources?: RetrievedVideo[]; + truncated?: boolean; + error?: boolean; +}; + +export default function AskChat() { + const [provider, setProvider] = useState<Provider>("anthropic"); + const [apiKey, setApiKey] = useState(""); + const [model, setModel] = useState(PROVIDERS.anthropic.defaultModel); + const [remember, setRemember] = useState(true); + const [showKey, setShowKey] = useState(false); + + const [corpus, setCorpus] = useState<CorpusInfo | null>(null); + const [corpusError, setCorpusError] = useState<string | null>(null); + const [messages, setMessages] = useState<UiMessage[]>([]); + const [input, setInput] = useState(""); + const [busy, setBusy] = useState(false); + const abortRef = useRef<AbortController | null>(null); + const scrollRef = useRef<HTMLDivElement | null>(null); + + // Restore saved provider + (if remembered) that provider's key/model. + useEffect(() => { + try { + const savedProvider = localStorage.getItem(K_PROVIDER) as Provider | null; + const rememberSaved = localStorage.getItem(K_REMEMBER) !== "0"; + const p = + savedProvider && PROVIDERS[savedProvider] ? savedProvider : "anthropic"; + setProvider(p); + setRemember(rememberSaved); + loadProviderCreds(p, rememberSaved); + } catch { + /* storage unavailable */ + } + // eslint-disable-next-line react-hooks/exhaustive-deps + }, []); + + // Load the channel list once. + useEffect(() => { + const ac = new AbortController(); + loadCorpus(ac.signal) + .then(setCorpus) + .catch((e: Error) => setCorpusError(e.message)); + return () => ac.abort(); + }, []); + + useEffect(() => { + scrollRef.current?.scrollTo({ top: scrollRef.current.scrollHeight }); + }, [messages]); + + function loadProviderCreds(p: Provider, rememberOn: boolean) { + if (rememberOn) { + setApiKey(localStorage.getItem(keyFor(p)) ?? ""); + setModel(localStorage.getItem(modelFor(p)) ?? PROVIDERS[p].defaultModel); + } else { + setApiKey(""); + setModel(PROVIDERS[p].defaultModel); + } + } + + function onProviderChange(p: Provider) { + setProvider(p); + try { + localStorage.setItem(K_PROVIDER, p); + } catch { + /* ignore */ + } + loadProviderCreds(p, remember); + } + + function persistKey(p: Provider, key: string, mdl: string, rememberOn: boolean) { + try { + if (rememberOn) { + localStorage.setItem(keyFor(p), key); + localStorage.setItem(modelFor(p), mdl); + } else { + localStorage.removeItem(keyFor(p)); + localStorage.removeItem(modelFor(p)); + } + localStorage.setItem(K_REMEMBER, rememberOn ? "1" : "0"); + } catch { + /* ignore */ + } + } + + const send = useCallback(async () => { + const question = input.trim(); + if (!question || busy) return; + if (!apiKey.trim()) return; + if (!corpus) return; + + persistKey(provider, apiKey, model, remember); + + const priorTurns: ChatMessage[] = messages + .filter((m) => !m.error) + .map((m) => ({ role: m.role, content: m.content })); + + setInput(""); + setMessages((prev) => [ + ...prev, + { role: "user", content: question }, + { role: "assistant", content: "" }, + ]); + setBusy(true); + const ac = new AbortController(); + abortRef.current = ac; + + // Update the trailing assistant message immutably. + const patchLast = (fn: (m: UiMessage) => UiMessage) => + setMessages((prev) => { + const next = prev.slice(); + next[next.length - 1] = fn(next[next.length - 1]); + return next; + }); + + try { + const { videos, truncated } = await retrieve(corpus.channels, question, { + signal: ac.signal, + }); + patchLast((m) => ({ ...m, sources: videos, truncated })); + + const context = buildContext(videos); + const apiMessages: ChatMessage[] = [ + ...priorTurns, + { + role: "user", + content: `${question}\n\n---\nTranscript excerpts you may cite (by number):\n${context}`, + }, + ]; + + await askStream({ + provider, + apiKey: apiKey.trim(), + model: model.trim() || PROVIDERS[provider].defaultModel, + system: SYSTEM_PROMPT, + messages: apiMessages, + signal: ac.signal, + onDelta: (chunk) => + patchLast((m) => ({ ...m, content: m.content + chunk })), + }); + } catch (e) { + const err = e as Error; + if (err.name === "AbortError") { + patchLast((m) => ({ ...m, content: m.content + "\n\n_(stopped)_" })); + } else { + patchLast((m) => ({ + ...m, + content: m.content || err.message, + error: true, + })); + } + } finally { + setBusy(false); + abortRef.current = null; + } + }, [input, busy, apiKey, corpus, provider, model, remember, messages]); + + const stop = () => abortRef.current?.abort(); + + const info = PROVIDERS[provider]; + const channelCount = corpus?.channels.length ?? 0; + + return ( + <div className="flex flex-col gap-5"> + {/* Provider / key settings */} + <details className="rounded-lg border border-border bg-card/40" open={!apiKey}> + <summary className="cursor-pointer px-4 py-2.5 text-sm font-medium text-foreground"> + {apiKey ? `${info.label} · key set` : "Set up your AI provider"} + </summary> + <div className="flex flex-col gap-3 border-t border-border px-4 py-3"> + <div className="flex flex-wrap gap-3"> + <label className="flex flex-col gap-1 text-xs text-muted-foreground"> + Provider + <select + value={provider} + onChange={(e) => onProviderChange(e.target.value as Provider)} + className="rounded-md border border-border bg-background px-2 py-1.5 text-sm text-foreground" + > + {(Object.keys(PROVIDERS) as Provider[]).map((p) => ( + <option key={p} value={p}> + {PROVIDERS[p].label} + </option> + ))} + </select> + </label> + <label className="flex min-w-[10rem] flex-1 flex-col gap-1 text-xs text-muted-foreground"> + Model + <input + list="ask-models" + value={model} + onChange={(e) => setModel(e.target.value)} + className="rounded-md border border-border bg-background px-2 py-1.5 text-sm text-foreground" + /> + <datalist id="ask-models"> + {info.models.map((m) => ( + <option key={m} value={m} /> + ))} + </datalist> + </label> + </div> + <label className="flex flex-col gap-1 text-xs text-muted-foreground"> + API key + <div className="flex gap-2"> + <input + type={showKey ? "text" : "password"} + value={apiKey} + onChange={(e) => setApiKey(e.target.value)} + placeholder={info.keyHint} + autoComplete="off" + className="flex-1 rounded-md border border-border bg-background px-2 py-1.5 font-mono text-sm text-foreground" + /> + <button + type="button" + onClick={() => setShowKey((v) => !v)} + className="rounded-md border border-border px-2 text-xs text-muted-foreground hover:text-foreground" + > + {showKey ? "Hide" : "Show"} + </button> + </div> + </label> + <label className="flex items-center gap-2 text-xs text-muted-foreground"> + <input + type="checkbox" + checked={remember} + onChange={(e) => { + setRemember(e.target.checked); + persistKey(provider, apiKey, model, e.target.checked); + }} + /> + Remember my key in this browser + </label> + <p className="text-xs text-muted-foreground/80"> + Your key is stored only {remember ? "in this browser" : "for this page session"} and is sent + directly to {info.label} — never to this site. Requests are billed to + your own account.{" "} + <a href={info.keyUrl} target="_blank" rel="noopener noreferrer" className="text-brand hover:underline"> + Get a key → + </a> + </p> + </div> + </details> + + {corpusError && ( + <p className="text-sm text-warning"> + Couldn&apos;t load the corpus index ({corpusError}). This chat needs the + site&apos;s <code className="font-mono">corpus.json</code>. + </p> + )} + + {/* Conversation */} + <div + ref={scrollRef} + className="flex max-h-[60vh] min-h-[8rem] flex-col gap-4 overflow-y-auto" + > + {messages.length === 0 && ( + <p className="text-sm text-muted-foreground"> + Ask a question about the transcripts + {channelCount > 0 ? ` (${channelCount} channel${channelCount === 1 ? "" : "s"} indexed)` : ""}. + Answers cite the videos they draw from. + </p> + )} + {messages.map((m, i) => ( + <div key={i} className="flex flex-col gap-2"> + <div + className={ + m.role === "user" + ? "self-end rounded-lg bg-primary/10 px-3 py-2 text-sm text-foreground" + : `rounded-lg border border-border bg-card/40 px-3 py-2 text-sm ${m.error ? "text-warning" : "text-foreground"}` + } + > + <span className="whitespace-pre-wrap">{m.content || (busy && i === messages.length - 1 ? "…" : "")}</span> + </div> + {m.role === "assistant" && m.sources && m.sources.length > 0 && ( + <ol className="ml-1 flex flex-col gap-1 text-xs text-muted-foreground"> + {m.sources.map((s, si) => ( + <li key={s.key}> + <span className="font-mono text-brand">[{si + 1}]</span>{" "} + {s.url ? ( + <a href={s.url} target="_blank" rel="noopener noreferrer" className="hover:underline"> + {s.title} + </a> + ) : ( + s.title + )}{" "} + <span className="text-muted-foreground/70"> + — {s.channel} + {s.siteTitle ? ` · ${s.siteTitle}` : ""} · {s.snippets.map((sn) => sn.clock).join(", ")} + </span> + </li> + ))} + {m.truncated && ( + <li className="text-muted-foreground/60"> + (search was truncated; ask more specifically for better coverage) + </li> + )} + </ol> + )} + </div> + ))} + </div> + + {/* Composer */} + <form + onSubmit={(e) => { + e.preventDefault(); + void send(); + }} + className="flex flex-col gap-2" + > + <textarea + value={input} + onChange={(e) => setInput(e.target.value)} + onKeyDown={(e) => { + if (e.key === "Enter" && (e.metaKey || e.ctrlKey)) { + e.preventDefault(); + void send(); + } + }} + rows={2} + placeholder={ + corpus + ? "Ask about the transcripts… (⌘/Ctrl+Enter to send)" + : "Loading corpus…" + } + disabled={!corpus} + className="w-full resize-y rounded-md border border-border bg-background px-3 py-2 text-sm text-foreground" + /> + <div className="flex items-center gap-2"> + <button + type="submit" + disabled={busy || !corpus || !input.trim() || !apiKey.trim()} + className="rounded-md bg-primary px-4 py-2 text-sm font-medium text-primary-foreground transition-colors hover:bg-brand-strong disabled:opacity-50" + > + {busy ? "Thinking…" : "Ask"} + </button> + {busy && ( + <button + type="button" + onClick={stop} + className="rounded-md border border-border px-3 py-2 text-sm text-muted-foreground hover:text-foreground" + > + Stop + </button> + )} + {!apiKey.trim() && ( + <span className="text-xs text-muted-foreground">Set an API key above to start.</span> + )} + </div> + </form> + </div> + ); +} diff --git a/export/app/ask/page.tsx b/export/app/ask/page.tsx @@ -0,0 +1,40 @@ +import type { Metadata } from "next"; +import Link from "next/link"; +import { currentSite } from "../lib/site"; +import { instanceMode } from "../lib/mode"; +import AskChat from "./AskChat"; + +export const metadata: Metadata = { title: "Ask AI" }; + +// Static shell for the bring-your-own-key chat. All the work happens client-side +// in <AskChat/> — retrieval over the published shards + a direct call to the +// visitor's chosen provider. The export stays fully static. +export default function AskPage() { + const site = currentSite(); + const scope = instanceMode() === "hub" ? "the federation" : "the transcripts"; + + return ( + <div className="mx-auto flex max-w-3xl flex-col gap-6"> + <header className="flex flex-col gap-2 border-b border-border pb-5"> + <p className="font-mono text-xs uppercase tracking-[0.18em] text-brand"> + Ask AI · {site.headerTitle} + </p> + <h1 className="font-display text-3xl font-semibold leading-tight text-foreground"> + Ask a question about {scope} + </h1> + <p className="max-w-prose text-sm text-muted-foreground"> + Bring your own API key. The chat searches the transcripts in your + browser, sends the relevant excerpts to your chosen AI, and answers with + citations. Nothing is hosted here — your key and the requests stay + between your browser and the provider. See{" "} + <Link href="/use-with-ai" className="text-brand hover:underline"> + Use with AI + </Link>{" "} + for other ways to use this archive. + </p> + </header> + + <AskChat /> + </div> + ); +} diff --git a/export/app/lib/askProvider.ts b/export/app/lib/askProvider.ts @@ -0,0 +1,243 @@ +// Bring-your-own-key chat providers, called DIRECTLY from the browser. The +// export site is fully static — there is no backend and we host no inference. +// The visitor supplies their own API key; requests go browser → provider only. +// +// Each provider is CORS-enabled for browser use: +// - Anthropic: requires the `anthropic-dangerous-direct-browser-access` header +// - OpenAI: chat-completions is CORS-open with a Bearer key +// - Gemini: Generative Language API takes the key as a query param +// +// Everything streams so the answer renders as it arrives. + +export type Provider = "anthropic" | "openai" | "gemini"; + +export type ChatRole = "user" | "assistant"; +export type ChatMessage = { role: ChatRole; content: string }; + +export type ProviderInfo = { + id: Provider; + label: string; + defaultModel: string; + // A few known-good model ids to offer; the field stays free-text so a user can + // enter anything their key supports. + models: string[]; + keyHint: string; + keyUrl: string; +}; + +export const PROVIDERS: Record<Provider, ProviderInfo> = { + anthropic: { + id: "anthropic", + label: "Anthropic (Claude)", + defaultModel: "claude-haiku-4-5", + models: ["claude-haiku-4-5", "claude-sonnet-5", "claude-opus-4-8"], + keyHint: "sk-ant-…", + keyUrl: "https://console.anthropic.com/settings/keys", + }, + openai: { + id: "openai", + label: "OpenAI", + defaultModel: "gpt-4o-mini", + models: ["gpt-4o-mini", "gpt-4o"], + keyHint: "sk-…", + keyUrl: "https://platform.openai.com/api-keys", + }, + gemini: { + id: "gemini", + label: "Google Gemini", + defaultModel: "gemini-2.5-flash", + models: ["gemini-2.5-flash", "gemini-2.5-pro"], + keyHint: "AIza…", + keyUrl: "https://aistudio.google.com/app/apikey", + }, +}; + +export type AskOptions = { + provider: Provider; + apiKey: string; + model?: string; + system: string; + messages: ChatMessage[]; + maxTokens?: number; + signal?: AbortSignal; + onDelta?: (chunk: string) => void; +}; + +// Stream a completion from the chosen provider, invoking onDelta for each text +// chunk and resolving with the full concatenated answer. +export async function askStream(opts: AskOptions): Promise<string> { + switch (opts.provider) { + case "anthropic": + return askAnthropic(opts); + case "openai": + return askOpenAI(opts); + case "gemini": + return askGemini(opts); + default: + throw new Error(`unknown provider: ${opts.provider}`); + } +} + +// ─── shared SSE plumbing ─── + +async function* sseLines( + res: Response, + signal?: AbortSignal, +): AsyncGenerator<string> { + if (!res.body) throw new Error("no response body"); + const reader = res.body.getReader(); + const decoder = new TextDecoder(); + let buf = ""; + try { + while (true) { + if (signal?.aborted) throw new DOMException("aborted", "AbortError"); + const { done, value } = await reader.read(); + if (done) break; + buf += decoder.decode(value, { stream: true }); + // SSE events are separated by a blank line; yield each `data:` payload. + let idx: number; + while ((idx = buf.indexOf("\n")) >= 0) { + const line = buf.slice(0, idx).trim(); + buf = buf.slice(idx + 1); + if (line.startsWith("data:")) yield line.slice(5).trim(); + } + } + } finally { + reader.releaseLock(); + } +} + +async function ensureOk(res: Response, provider: string): Promise<void> { + if (res.ok) return; + let detail = ""; + try { + detail = await res.text(); + } catch { + /* ignore */ + } + throw new Error( + `${provider} request failed (${res.status}). ${detail.slice(0, 300)}`, + ); +} + +function emit(full: string[], chunk: string, onDelta?: (c: string) => void): void { + if (!chunk) return; + full.push(chunk); + onDelta?.(chunk); +} + +// ─── Anthropic ─── + +async function askAnthropic(opts: AskOptions): Promise<string> { + const res = await fetch("https://api.anthropic.com/v1/messages", { + method: "POST", + signal: opts.signal, + headers: { + "content-type": "application/json", + "x-api-key": opts.apiKey, + "anthropic-version": "2023-06-01", + "anthropic-dangerous-direct-browser-access": "true", + }, + body: JSON.stringify({ + model: opts.model || PROVIDERS.anthropic.defaultModel, + max_tokens: opts.maxTokens ?? 1024, + system: opts.system, + messages: opts.messages.map((m) => ({ role: m.role, content: m.content })), + stream: true, + }), + }); + await ensureOk(res, "Anthropic"); + const full: string[] = []; + for await (const data of sseLines(res, opts.signal)) { + if (!data || data === "[DONE]") continue; + let evt: unknown; + try { + evt = JSON.parse(data); + } catch { + continue; + } + const e = evt as { + type?: string; + delta?: { type?: string; text?: string }; + }; + if (e.type === "content_block_delta" && e.delta?.type === "text_delta") { + emit(full, e.delta.text ?? "", opts.onDelta); + } + } + return full.join(""); +} + +// ─── OpenAI ─── + +async function askOpenAI(opts: AskOptions): Promise<string> { + const res = await fetch("https://api.openai.com/v1/chat/completions", { + method: "POST", + signal: opts.signal, + headers: { + "content-type": "application/json", + authorization: `Bearer ${opts.apiKey}`, + }, + body: JSON.stringify({ + model: opts.model || PROVIDERS.openai.defaultModel, + stream: true, + messages: [ + { role: "system", content: opts.system }, + ...opts.messages.map((m) => ({ role: m.role, content: m.content })), + ], + }), + }); + await ensureOk(res, "OpenAI"); + const full: string[] = []; + for await (const data of sseLines(res, opts.signal)) { + if (!data || data === "[DONE]") continue; + let evt: unknown; + try { + evt = JSON.parse(data); + } catch { + continue; + } + const e = evt as { choices?: { delta?: { content?: string } }[] }; + const chunk = e.choices?.[0]?.delta?.content; + if (chunk) emit(full, chunk, opts.onDelta); + } + return full.join(""); +} + +// ─── Google Gemini ─── + +async function askGemini(opts: AskOptions): Promise<string> { + const model = opts.model || PROVIDERS.gemini.defaultModel; + const url = + `https://generativelanguage.googleapis.com/v1beta/models/` + + `${encodeURIComponent(model)}:streamGenerateContent?alt=sse&key=${encodeURIComponent(opts.apiKey)}`; + const res = await fetch(url, { + method: "POST", + signal: opts.signal, + headers: { "content-type": "application/json" }, + body: JSON.stringify({ + system_instruction: { parts: [{ text: opts.system }] }, + // Gemini uses role "model" for assistant turns. + contents: opts.messages.map((m) => ({ + role: m.role === "assistant" ? "model" : "user", + parts: [{ text: m.content }], + })), + }), + }); + await ensureOk(res, "Gemini"); + const full: string[] = []; + for await (const data of sseLines(res, opts.signal)) { + if (!data) continue; + let evt: unknown; + try { + evt = JSON.parse(data); + } catch { + continue; + } + const e = evt as { + candidates?: { content?: { parts?: { text?: string }[] } }[]; + }; + const parts = e.candidates?.[0]?.content?.parts ?? []; + for (const p of parts) emit(full, p.text ?? "", opts.onDelta); + } + return full.join(""); +} diff --git a/export/app/lib/askRetrieval.ts b/export/app/lib/askRetrieval.ts @@ -0,0 +1,183 @@ +import { + transcriptPageFileName, + type ChannelTranscriptsManifest, +} from "yt-dlp-transcript-common/lib/manifest"; +import type { TranscriptDetail } from "yt-dlp-transcript-common/lib/transcripts"; +import { formatDuration } from "yt-dlp-transcript-common/lib/format"; + +// Browser-side retrieval for the /ask chat. Reads the site's own published +// static shards via the documented corpus.json contract — same for a single +// site (same-origin) or a hub (fanning out over member origins, which serve +// CORS-* shards). No server, no new files. Shards are fetched force-cache so +// the service worker / HTTP cache absorbs repeat questions. + +export type ChannelRef = { + slug: string; + name: string; + base: string; // origin ("" = same-origin); member origin in hub mode + siteTitle?: string; +}; + +export type CorpusInfo = { channels: ChannelRef[]; hub: boolean; title: string }; + +export type RetrievedVideo = { + key: string; + videoId: string; + title: string; + channel: string; + siteTitle?: string; + uploadDate: string; + url?: string; + snippets: { clock: string; seconds: number; text: string }[]; +}; + +async function getJson<T>(url: string, signal?: AbortSignal): Promise<T> { + const res = await fetch(url, { cache: "force-cache", signal }); + if (!res.ok) throw new Error(`GET ${url} -> ${res.status}`); + return (await res.json()) as T; +} + +type CorpusChannelJson = { slug: string; name?: string }; +type SiteCorpusJson = { + channels?: CorpusChannelJson[]; + site?: { title?: string }; +}; +type HubCorpusJson = { + kind?: string; + hub?: { title?: string }; + sites?: { siteId: string; title: string; url: string }[]; +}; + +// Load the channel list from /corpus.json. On a hub, fetch every member's own +// corpus.json and tag each channel with its origin + site title. +export async function loadCorpus(signal?: AbortSignal): Promise<CorpusInfo> { + const root = await getJson<SiteCorpusJson & HubCorpusJson>( + "/corpus.json", + signal, + ); + if (root.kind === "hub" && Array.isArray(root.sites)) { + const channels: ChannelRef[] = []; + for (const site of root.sites) { + const base = site.url.replace(/\/+$/, ""); + try { + const c = await getJson<SiteCorpusJson>(`${base}/corpus.json`, signal); + for (const ch of c.channels ?? []) { + channels.push({ + slug: ch.slug, + name: ch.name ?? ch.slug, + base, + siteTitle: site.title, + }); + } + } catch { + // skip an unreachable member + } + } + return { channels, hub: true, title: root.hub?.title ?? "the federation" }; + } + const channels: ChannelRef[] = (root.channels ?? []).map((ch) => ({ + slug: ch.slug, + name: ch.name ?? ch.slug, + base: "", + })); + return { channels, hub: false, title: root.site?.title ?? "this archive" }; +} + +function clock(seconds: number): string { + const s = Math.max(0, Math.floor(seconds)); + return s === 0 ? "0:00" : formatDuration(s); +} + +// Scan the channels' transcript shards for the query and return the top matching +// videos with a few cue snippets each. Bounded by `limit` and `maxPages` so a +// question can't fetch an unbounded slice of a large (or hub-wide) corpus. +export async function retrieve( + channels: ChannelRef[], + query: string, + opts: { + limit?: number; + snippetsPerVideo?: number; + maxPages?: number; + signal?: AbortSignal; + } = {}, +): Promise<{ videos: RetrievedVideo[]; truncated: boolean }> { + const limit = opts.limit ?? 12; + const perVideo = opts.snippetsPerVideo ?? 4; + const maxPages = opts.maxPages ?? 60; + const q = query.toLowerCase(); + const terms = q.split(/\s+/).filter((t) => t.length >= 3); + const match = (text: string): boolean => { + const t = text.toLowerCase(); + return terms.length === 0 ? t.includes(q) : terms.some((term) => t.includes(term)); + }; + + const videos: RetrievedVideo[] = []; + let pages = 0; + let truncated = false; + + outer: for (const ch of channels) { + let manifest: ChannelTranscriptsManifest; + try { + manifest = await getJson( + `${ch.base}/transcripts/${ch.slug}/manifest.json`, + opts.signal, + ); + } catch { + continue; + } + for (let p = 0; p < manifest.pageCount; p++) { + if (pages >= maxPages) { + truncated = true; + break outer; + } + let records: TranscriptDetail[]; + try { + records = await getJson( + `${ch.base}/transcripts/${ch.slug}/${transcriptPageFileName(p)}`, + opts.signal, + ); + } catch { + continue; + } + pages++; + for (const rec of records) { + const snippets: RetrievedVideo["snippets"] = []; + for (const cue of rec.cues ?? []) { + if (!match(cue.text)) continue; + snippets.push({ + clock: clock(cue.start), + seconds: cue.start, + text: cue.text.trim().replace(/\s+/g, " ").slice(0, 240), + }); + if (snippets.length >= perVideo) break; + } + if (snippets.length === 0) continue; + videos.push({ + key: `${ch.base}|${rec.id}`, + videoId: rec.id, + title: rec.title, + channel: ch.name, + siteTitle: ch.siteTitle, + uploadDate: rec.uploadDate, + url: rec.webpageUrl, + snippets, + }); + if (videos.length >= limit) break outer; + } + } + } + return { videos, truncated }; +} + +// Assemble numbered retrieved excerpts into the context block that gets appended +// to the user's question, plus the matching system instruction to cite by index. +export function buildContext(videos: RetrievedVideo[]): string { + if (videos.length === 0) return "(no matching transcript excerpts were found)"; + return videos + .map((v, i) => { + const head = `[${i + 1}] "${v.title}" — ${v.channel}${v.siteTitle ? ` (${v.siteTitle})` : ""}`; + const lines = v.snippets.map((s) => ` [${s.clock}] ${s.text}`).join("\n"); + return `${head}\n${lines}`; + }) + .join("\n\n"); +}