Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit a62cf6a859873f94dc5bc1d4c7e3998f4b8adf55
parent 2c696c721ae0e0c06dc55c7833d89730487bac13
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Tue,  7 Jul 2026 11:19:17 -0400

Overhaul /ask: agentic multi-step retrieval, animated feedback, Markdown

Turn the bring-your-own-key /ask chat into an agentic assistant and fix the
follow-up-loses-context bug.

Retrieval is now a bounded agent loop (gather → answer): the model decides what
to search, reads results, and can refine and search again. Two interchangeable
gather transports behind one shared loop — a provider-agnostic SEARCH:/DONE
scripted protocol and native function-calling for Anthropic/OpenAI/Gemini — with
Auto-detect, session-cached fallback on a tools-capability error, and an
Auto/Native/Scripted toggle. A search budget bounds round-trips/spend.

Conversation correctness: prior user turns are replayed with their grounded
content (question + that turn's excerpts), so earlier excerpts survive and a
reformat follow-up ("put that on a timeline") reuses them instead of re-searching
its own wording. Alias awareness: retrieve() expands keywords that match a known
transcription-misspelling alias into the alias regex, and the gather/answer
prompts carry an alias glossary.

UI (split the AskChat monolith into a hook + components): a staged, terminal-
flavoured pipeline status (per-search steps, spinner→check, thinking dots,
streaming caret), Markdown answers via a new common/components/Markdown.tsx with a
Formatted/Plain toggle, message enter animations (tw-animate-css, reduced-motion
honored), suggested-prompt empty state, copy-answer button, and smart auto-scroll.

Tests: 25 unit (scripted parsing, transport selection, result formatting, message
assembly + grounding, alias glossary, alias-expanded leaves, native tool-call
parsers) + a 4-case Playwright spec (scripted search+cite+Markdown, follow-up with
no junk search, native tool loop, Format toggle).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

Diffstat:
Acommon/components/Markdown.tsx | 87+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mexport/CHANGELOG.md | 3+++
Mexport/app/ask/AskChat.tsx | 464++++++++++++++++++++-----------------------------------------------------------
Aexport/app/ask/Composer.tsx | 88+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aexport/app/ask/MessageBubble.tsx | 127+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aexport/app/ask/PipelineStatus.tsx | 62++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aexport/app/ask/ProviderSettings.tsx | 161+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aexport/app/ask/useAskChat.ts | 275+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mexport/app/globals.css | 51+++++++++++++++++++++++++++++++++++++++++++++++++++
Aexport/app/lib/askConversation.test.ts | 95+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aexport/app/lib/askConversation.ts | 138+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mexport/app/lib/askProvider.ts | 13+++++++++++++
Mexport/app/lib/askRetrieval.test.ts | 31++++++++++++++++++++++++++++++-
Mexport/app/lib/askRetrieval.ts | 56++++++++++++++++++++++++++++++++++++++++++++++++--------
Aexport/app/lib/nativeTools/anthropic.ts | 122+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aexport/app/lib/nativeTools/gemini.ts | 124+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aexport/app/lib/nativeTools/nativeTools.test.ts | 80+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aexport/app/lib/nativeTools/openai.ts | 115+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aexport/app/lib/nativeTools/shared.ts | 92+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aexport/app/lib/searchAgent.test.ts | 72++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aexport/app/lib/searchAgent.ts | 280+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aexport/e2e/ask-chat.spec.ts | 160+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
22 files changed, 2339 insertions(+), 357 deletions(-)

diff --git a/common/components/Markdown.tsx b/common/components/Markdown.tsx @@ -0,0 +1,87 @@ +// Generic Markdown renderer for arbitrary content (AI chat answers, notes, …). +// Mirrors the house pattern established by common/components/Changelog.tsx — +// markdown-to-jsx with per-element Tailwind overrides on semantic design tokens +// (no `prose` plugin) — but drops the changelog-specific h1/h2 anchor behaviour +// so it's a drop-in for any Markdown string. Tuned for chat density: smaller +// headings, tight vertical rhythm, token-based code/quote/table styling. +// +// Pure/presentational and safe in a client tree. markdown-to-jsx escapes raw +// HTML by default, so streaming model output can't inject markup. + +import MarkdownToJsx from "markdown-to-jsx"; + +const OVERRIDES = { + h1: { props: { className: "mt-4 mb-2 text-base font-semibold text-foreground" } }, + h2: { props: { className: "mt-4 mb-2 text-sm font-semibold text-foreground" } }, + h3: { props: { className: "mt-3 mb-1 text-sm font-semibold text-foreground" } }, + h4: { props: { className: "mt-3 mb-1 text-sm font-semibold text-foreground" } }, + p: { props: { className: "my-2 leading-relaxed" } }, + ul: { + props: { + className: + "list-disc pl-5 my-2 space-y-1 marker:text-muted-foreground", + }, + }, + ol: { + props: { + className: + "list-decimal pl-5 my-2 space-y-1 marker:text-muted-foreground", + }, + }, + li: { props: { className: "leading-relaxed" } }, + strong: { props: { className: "font-semibold text-foreground" } }, + em: { props: { className: "italic" } }, + code: { + props: { + className: "px-1 py-0.5 rounded bg-muted font-mono text-[0.85em]", + }, + }, + pre: { + props: { + className: + "my-2 p-3 rounded-md bg-muted overflow-x-auto text-xs font-mono", + }, + }, + a: { + props: { + className: "text-brand underline decoration-brand/40 hover:decoration-brand", + target: "_blank", + rel: "noopener noreferrer", + }, + }, + hr: { props: { className: "my-4 border-t border-border" } }, + blockquote: { + props: { + className: "my-2 pl-3 border-l-2 border-border text-muted-foreground", + }, + }, + table: { + props: { + className: "my-2 block w-full overflow-x-auto text-left border-collapse", + }, + }, + th: { + props: { + className: "border-b border-border px-2 py-1 font-semibold text-foreground", + }, + }, + td: { props: { className: "border-b border-border/50 px-2 py-1 align-top" } }, +}; + +// Render a Markdown string. `forceBlock` keeps single-line input block-level so +// spacing is consistent regardless of content shape. +export function Markdown({ + children, + className, +}: { + children: string; + className?: string; +}) { + return ( + <div className={className}> + <MarkdownToJsx options={{ overrides: OVERRIDES, forceBlock: true }}> + {children} + </MarkdownToJsx> + </div> + ); +} diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md @@ -1,5 +1,8 @@ # Changelog +## [Unreleased] +- **The "Ask AI" chat now searches like an assistant — and follow-ups work.** Instead of one keyword search per message, the chat now runs an *agentic* loop: it decides what to look up, reads the results, and can refine and search again before answering — you see each search as it happens. Crucially, a follow-up that builds on the last answer (e.g. "put that on a timeline", "summarise it") no longer throws away the earlier context and re-searches your wording — earlier excerpts stay in play, so it just reformats what it already found. Answers render as **Markdown** (with a *Format answers* toggle to fall back to plain text if anything looks off), a live status shows the search→answer pipeline, and the model is told the archive's known transcription-misspelling aliases so it searches for the right variants. A **Search mode** control (Auto / Native tools / Scripted) chooses between your model's native function-calling and a provider-agnostic protocol, with automatic fallback. See `export/app/ask/*`, `export/app/lib/{searchAgent,askConversation,nativeTools/*}.ts`, `common/components/Markdown.tsx`, and `export/e2e/ask-chat.spec.ts`. + ## [0.7.3] - 2026-07-07 - **Rebuilds now skip channels that haven't changed.** The export build used to redo almost everything from scratch every run — it re-zipped every channel's transcript and live-chat download bundle, and `rm -rf`'d and re-copied every channel's transcript/subs pages into the served tree — even when nothing about that channel had changed. Each of these is now gated on a cheap per-channel signature: archive zips are built once into a persistent shared cache (`transcripts/export/archives`) keyed by a content signature and reused build-to-build and across sites, and the per-channel page trees are reconciled in place (only changed channels are re-copied, removed channels pruned). A no-change recompose drops from re-zipping/re-copying gigabytes to a few seconds. Output is identical — the signature is over inputs (video mtimes recorded by `build:index`, channel config, and archive options), so a channel is only rebuilt when its content actually changes; a schema bump or settings change invalidates the cache. Also parallelizes the per-channel archive compressor across a few channels at once. See `common/lib/channelSignature.ts`, `common/controller/{archiveTranscripts,archiveLiveChat}.ts`, and `common/bin/compose-site.ts`. diff --git a/export/app/ask/AskChat.tsx b/export/app/ask/AskChat.tsx @@ -1,269 +1,79 @@ "use client"; -import { useCallback, useEffect, useRef, useState } from "react"; -import { - PROVIDERS, - askStream, - type Provider, - type ChatMessage, -} from "../lib/askProvider"; -import { useSearchData } from "yt-dlp-transcript-common/components/SearchDataContext"; -import { - retrieve, - buildContext, - type RetrievedVideo, -} from "../lib/askRetrieval"; - -const SYSTEM_PROMPT = - "You are a helpful assistant answering questions about a video-transcript " + - "archive. Base your answer ONLY on the transcript excerpts provided with the " + - "user's question. Cite the excerpts you use by their bracketed number, e.g. " + - "[1]. Each excerpt shows a video title, channel, and timestamped lines. If the " + - "excerpts do not contain enough to answer, say so plainly rather than guessing. " + - "Be concise."; - -const K_PROVIDER = "ytdlp-tb:ai:provider"; -const K_REMEMBER = "ytdlp-tb:ai:remember"; -const keyFor = (p: Provider) => `ytdlp-tb:ai:key:${p}`; -const modelFor = (p: Provider) => `ytdlp-tb:ai:model:${p}`; - -type UiMessage = { - role: "user" | "assistant"; - content: string; - sources?: RetrievedVideo[]; - truncated?: boolean; - error?: boolean; -}; +import { useEffect, useMemo, useRef, useState } from "react"; +import { ArrowDownIcon } from "lucide-react"; +import { useAskChat } from "./useAskChat"; +import { ProviderSettings } from "./ProviderSettings"; +import { MessageBubble } from "./MessageBubble"; +import { Composer } from "./Composer"; export default function AskChat() { - const [provider, setProvider] = useState<Provider>("anthropic"); - const [apiKey, setApiKey] = useState(""); - const [model, setModel] = useState(PROVIDERS.anthropic.defaultModel); - const [remember, setRemember] = useState(true); - const [showKey, setShowKey] = useState(false); + const s = useAskChat(); + const { + messages, + busy, + channels, + corpusError, + summariesReady, + apiKey, + input, + setInput, + markdownOn, + } = s; - const { summariesState, channels } = useSearchData(); - const { summaries, summariesReady } = summariesState; - const corpusError = summariesState.error?.message ?? null; - const [messages, setMessages] = useState<UiMessage[]>([]); - const [input, setInput] = useState(""); - const [busy, setBusy] = useState(false); - const abortRef = useRef<AbortController | null>(null); const scrollRef = useRef<HTMLDivElement | null>(null); - - // Restore saved provider + (if remembered) that provider's key/model. + const stuckRef = useRef(true); + const [showJump, setShowJump] = useState(false); + + // Smart auto-scroll: only stick to the bottom when the reader is already near + // it, so streaming text doesn't yank the view while they scroll back to read. + const onScroll = () => { + const el = scrollRef.current; + if (!el) return; + const nearBottom = el.scrollHeight - el.scrollTop - el.clientHeight < 60; + stuckRef.current = nearBottom; + setShowJump(!nearBottom && messages.length > 0); + }; useEffect(() => { - try { - const savedProvider = localStorage.getItem(K_PROVIDER) as Provider | null; - const rememberSaved = localStorage.getItem(K_REMEMBER) !== "0"; - const p = - savedProvider && PROVIDERS[savedProvider] ? savedProvider : "anthropic"; - setProvider(p); - setRemember(rememberSaved); - loadProviderCreds(p, rememberSaved); - } catch { - /* storage unavailable */ - } - // eslint-disable-next-line react-hooks/exhaustive-deps - }, []); - - useEffect(() => { - scrollRef.current?.scrollTo({ top: scrollRef.current.scrollHeight }); + const el = scrollRef.current; + if (el && stuckRef.current) el.scrollTo({ top: el.scrollHeight }); }, [messages]); + const jumpToLatest = () => { + const el = scrollRef.current; + if (!el) return; + el.scrollTo({ top: el.scrollHeight, behavior: "smooth" }); + stuckRef.current = true; + setShowJump(false); + }; - function loadProviderCreds(p: Provider, rememberOn: boolean) { - if (rememberOn) { - setApiKey(localStorage.getItem(keyFor(p)) ?? ""); - setModel(localStorage.getItem(modelFor(p)) ?? PROVIDERS[p].defaultModel); - } else { - setApiKey(""); - setModel(PROVIDERS[p].defaultModel); - } - } - - function onProviderChange(p: Provider) { - setProvider(p); - try { - localStorage.setItem(K_PROVIDER, p); - } catch { - /* ignore */ - } - loadProviderCreds(p, remember); - } - - function persistKey(p: Provider, key: string, mdl: string, rememberOn: boolean) { - try { - if (rememberOn) { - localStorage.setItem(keyFor(p), key); - localStorage.setItem(modelFor(p), mdl); - } else { - localStorage.removeItem(keyFor(p)); - localStorage.removeItem(modelFor(p)); - } - localStorage.setItem(K_REMEMBER, rememberOn ? "1" : "0"); - } catch { - /* ignore */ - } - } - - const send = useCallback(async () => { - const question = input.trim(); - if (!question || busy) return; - if (!apiKey.trim()) return; - if (!summariesReady) return; - - persistKey(provider, apiKey, model, remember); - - const priorTurns: ChatMessage[] = messages - .filter((m) => !m.error) - .map((m) => ({ role: m.role, content: m.content })); - - setInput(""); - setMessages((prev) => [ - ...prev, - { role: "user", content: question }, - { role: "assistant", content: "" }, - ]); - setBusy(true); - const ac = new AbortController(); - abortRef.current = ac; - - // Update the trailing assistant message immutably. - const patchLast = (fn: (m: UiMessage) => UiMessage) => - setMessages((prev) => { - const next = prev.slice(); - next[next.length - 1] = fn(next[next.length - 1]); - return next; - }); - - try { - const { videos, truncated } = await retrieve({ - question, - summaries, - signal: ac.signal, - }); - patchLast((m) => ({ ...m, sources: videos, truncated })); - - const context = buildContext(videos); - const apiMessages: ChatMessage[] = [ - ...priorTurns, - { - role: "user", - content: `${question}\n\n---\nTranscript excerpts you may cite (by number):\n${context}`, - }, - ]; - - await askStream({ - provider, - apiKey: apiKey.trim(), - model: model.trim() || PROVIDERS[provider].defaultModel, - system: SYSTEM_PROMPT, - messages: apiMessages, - signal: ac.signal, - onDelta: (chunk) => - patchLast((m) => ({ ...m, content: m.content + chunk })), - }); - } catch (e) { - const err = e as Error; - if (err.name === "AbortError") { - patchLast((m) => ({ ...m, content: m.content + "\n\n_(stopped)_" })); - } else { - patchLast((m) => ({ - ...m, - content: m.content || err.message, - error: true, - })); - } - } finally { - setBusy(false); - abortRef.current = null; - } - }, [input, busy, apiKey, summariesReady, summaries, provider, model, remember, messages]); - - const stop = () => abortRef.current?.abort(); - - const info = PROVIDERS[provider]; const channelCount = channels.length; + const suggestions = useMemo(() => { + const base = [ + "What are the main topics discussed?", + "Summarize the most recent developments.", + "Where is a specific person or case talked about?", + ]; + if (channels[0]) base.unshift(`What does ${channels[0].name} focus on?`); + return base.slice(0, 4); + }, [channels]); return ( <div className="flex flex-col gap-5"> - {/* Provider / key settings */} - <details className="rounded-lg border border-border bg-card/40" open={!apiKey}> - <summary className="cursor-pointer px-4 py-2.5 text-sm font-medium text-foreground"> - {apiKey ? `${info.label} · key set` : "Set up your AI provider"} - </summary> - <div className="flex flex-col gap-3 border-t border-border px-4 py-3"> - <div className="flex flex-wrap gap-3"> - <label className="flex flex-col gap-1 text-xs text-muted-foreground"> - Provider - <select - value={provider} - onChange={(e) => onProviderChange(e.target.value as Provider)} - className="rounded-md border border-border bg-background px-2 py-1.5 text-sm text-foreground" - > - {(Object.keys(PROVIDERS) as Provider[]).map((p) => ( - <option key={p} value={p}> - {PROVIDERS[p].label} - </option> - ))} - </select> - </label> - <label className="flex min-w-[10rem] flex-1 flex-col gap-1 text-xs text-muted-foreground"> - Model - <input - list="ask-models" - value={model} - onChange={(e) => setModel(e.target.value)} - className="rounded-md border border-border bg-background px-2 py-1.5 text-sm text-foreground" - /> - <datalist id="ask-models"> - {info.models.map((m) => ( - <option key={m} value={m} /> - ))} - </datalist> - </label> - </div> - <label className="flex flex-col gap-1 text-xs text-muted-foreground"> - API key - <div className="flex gap-2"> - <input - type={showKey ? "text" : "password"} - value={apiKey} - onChange={(e) => setApiKey(e.target.value)} - placeholder={info.keyHint} - autoComplete="off" - className="flex-1 rounded-md border border-border bg-background px-2 py-1.5 font-mono text-sm text-foreground" - /> - <button - type="button" - onClick={() => setShowKey((v) => !v)} - className="rounded-md border border-border px-2 text-xs text-muted-foreground hover:text-foreground" - > - {showKey ? "Hide" : "Show"} - </button> - </div> - </label> - <label className="flex items-center gap-2 text-xs text-muted-foreground"> - <input - type="checkbox" - checked={remember} - onChange={(e) => { - setRemember(e.target.checked); - persistKey(provider, apiKey, model, e.target.checked); - }} - /> - Remember my key in this browser - </label> - <p className="text-xs text-muted-foreground/80"> - Your key is stored only {remember ? "in this browser" : "for this page session"} and is sent - directly to {info.label} — never to this site. Requests are billed to - your own account.{" "} - <a href={info.keyUrl} target="_blank" rel="noopener noreferrer" className="text-brand hover:underline"> - Get a key → - </a> - </p> - </div> - </details> + <ProviderSettings + provider={s.provider} + setProvider={s.setProvider} + apiKey={s.apiKey} + setApiKey={s.setApiKey} + model={s.model} + setModel={s.setModel} + remember={s.remember} + setRemember={s.setRemember} + showKey={s.showKey} + setShowKey={s.setShowKey} + searchMode={s.searchMode} + setSearchMode={s.setSearchMode} + persistKey={s.persistKey} + /> {corpusError && ( <p className="text-sm text-warning"> @@ -272,106 +82,64 @@ export default function AskChat() { </p> )} - {/* Conversation */} - <div - ref={scrollRef} - className="flex max-h-[60vh] min-h-[8rem] flex-col gap-4 overflow-y-auto" - > - {messages.length === 0 && ( - <p className="text-sm text-muted-foreground"> - Ask a question about the transcripts - {channelCount > 0 ? ` (${channelCount} channel${channelCount === 1 ? "" : "s"} indexed)` : ""}. - Answers cite the videos they draw from. - </p> - )} - {messages.map((m, i) => ( - <div key={i} className="flex flex-col gap-2"> - <div - className={ - m.role === "user" - ? "self-end rounded-lg bg-primary/10 px-3 py-2 text-sm text-foreground" - : `rounded-lg border border-border bg-card/40 px-3 py-2 text-sm ${m.error ? "text-warning" : "text-foreground"}` - } - > - <span className="whitespace-pre-wrap">{m.content || (busy && i === messages.length - 1 ? "…" : "")}</span> - </div> - {m.role === "assistant" && m.sources && m.sources.length > 0 && ( - <ol className="ml-1 flex flex-col gap-1 text-xs text-muted-foreground"> - {m.sources.map((s, si) => ( - <li key={s.key}> - <span className="font-mono text-brand">[{si + 1}]</span>{" "} - {s.url ? ( - <a href={s.url} target="_blank" rel="noopener noreferrer" className="hover:underline"> - {s.title} - </a> - ) : ( - s.title - )}{" "} - <span className="text-muted-foreground/70"> - — {s.channel} - {s.siteTitle ? ` · ${s.siteTitle}` : ""} · {s.snippets.map((sn) => sn.clock).join(", ")} - </span> - </li> + <div className="relative"> + <div + ref={scrollRef} + onScroll={onScroll} + className="flex max-h-[60vh] min-h-[8rem] flex-col gap-4 overflow-y-auto" + > + {messages.length === 0 ? ( + <div className="flex flex-col gap-3"> + <p className="text-sm text-muted-foreground"> + Ask a question about the transcripts + {channelCount > 0 + ? ` (${channelCount} channel${channelCount === 1 ? "" : "s"} indexed)` + : ""} + . The assistant searches for what it needs — refining as it reads — + and answers with citations. + </p> + <div className="flex flex-wrap gap-2"> + {suggestions.map((q) => ( + <button + key={q} + type="button" + onClick={() => setInput(q)} + className="rounded-full border border-border bg-card/40 px-3 py-1.5 text-xs text-muted-foreground transition-colors hover:border-brand hover:text-foreground" + > + {q} + </button> ))} - {m.truncated && ( - <li className="text-muted-foreground/60"> - (search was truncated; ask more specifically for better coverage) - </li> - )} - </ol> - )} - </div> - ))} - </div> + </div> + </div> + ) : ( + messages.map((m, i) => ( + <MessageBubble key={i} message={m} markdownOn={markdownOn} /> + )) + )} + </div> - {/* Composer */} - <form - onSubmit={(e) => { - e.preventDefault(); - void send(); - }} - className="flex flex-col gap-2" - > - <textarea - value={input} - onChange={(e) => setInput(e.target.value)} - onKeyDown={(e) => { - if (e.key === "Enter" && (e.metaKey || e.ctrlKey)) { - e.preventDefault(); - void send(); - } - }} - rows={2} - placeholder={ - summariesReady - ? "Ask about the transcripts… (⌘/Ctrl+Enter to send)" - : "Loading transcripts…" - } - disabled={!summariesReady} - className="w-full resize-y rounded-md border border-border bg-background px-3 py-2 text-sm text-foreground" - /> - <div className="flex items-center gap-2"> + {showJump && ( <button - type="submit" - disabled={busy || !summariesReady || !input.trim() || !apiKey.trim()} - className="rounded-md bg-primary px-4 py-2 text-sm font-medium text-primary-foreground transition-colors hover:bg-brand-strong disabled:opacity-50" + type="button" + onClick={jumpToLatest} + className="absolute bottom-2 left-1/2 flex -translate-x-1/2 items-center gap-1 rounded-full border border-border bg-card px-3 py-1 text-xs text-muted-foreground shadow-sm transition-colors hover:text-foreground animate-in fade-in motion-reduce:animate-none" > - {busy ? "Thinking…" : "Ask"} + <ArrowDownIcon className="size-3.5" /> Jump to latest </button> - {busy && ( - <button - type="button" - onClick={stop} - className="rounded-md border border-border px-3 py-2 text-sm text-muted-foreground hover:text-foreground" - > - Stop - </button> - )} - {!apiKey.trim() && ( - <span className="text-xs text-muted-foreground">Set an API key above to start.</span> - )} - </div> - </form> + )} + </div> + + <Composer + input={input} + setInput={setInput} + send={s.send} + stop={s.stop} + busy={busy} + summariesReady={summariesReady} + hasKey={!!apiKey.trim()} + markdownOn={markdownOn} + setMarkdownOn={s.setMarkdownOn} + /> </div> ); } diff --git a/export/app/ask/Composer.tsx b/export/app/ask/Composer.tsx @@ -0,0 +1,88 @@ +"use client"; + +type Props = { + input: string; + setInput: (v: string) => void; + send: () => void; + stop: () => void; + busy: boolean; + summariesReady: boolean; + hasKey: boolean; + markdownOn: boolean; + setMarkdownOn: (v: boolean) => void; +}; + +export function Composer(props: Props) { + const { + input, + setInput, + send, + stop, + busy, + summariesReady, + hasKey, + markdownOn, + setMarkdownOn, + } = props; + + return ( + <form + onSubmit={(e) => { + e.preventDefault(); + send(); + }} + className="flex flex-col gap-2" + > + <textarea + value={input} + onChange={(e) => setInput(e.target.value)} + onKeyDown={(e) => { + if (e.key === "Enter" && (e.metaKey || e.ctrlKey)) { + e.preventDefault(); + send(); + } + }} + rows={2} + placeholder={ + summariesReady + ? "Ask about the transcripts… (⌘/Ctrl+Enter to send)" + : "Loading transcripts…" + } + disabled={!summariesReady} + className="w-full resize-y rounded-md border border-border bg-background px-3 py-2 text-sm text-foreground" + /> + <div className="flex flex-wrap items-center gap-2"> + <button + type="submit" + disabled={busy || !summariesReady || !input.trim() || !hasKey} + className="rounded-md bg-primary px-4 py-2 text-sm font-medium text-primary-foreground transition-colors hover:bg-brand-strong disabled:opacity-50" + > + {busy ? "Thinking…" : "Ask"} + </button> + {busy && ( + <button + type="button" + onClick={stop} + className="rounded-md border border-border px-3 py-2 text-sm text-muted-foreground hover:text-foreground" + > + Stop + </button> + )} + {!hasKey && ( + <span className="text-xs text-muted-foreground"> + Set an API key above to start. + </span> + )} + {/* Formatted / Plain escape hatch for Markdown rendering. */} + <label className="ml-auto flex items-center gap-1.5 text-xs text-muted-foreground"> + <input + type="checkbox" + checked={markdownOn} + onChange={(e) => setMarkdownOn(e.target.checked)} + /> + Format answers + </label> + </div> + </form> + ); +} diff --git a/export/app/ask/MessageBubble.tsx b/export/app/ask/MessageBubble.tsx @@ -0,0 +1,127 @@ +"use client"; + +import { useState } from "react"; +import { CheckIcon, CopyIcon } from "lucide-react"; +import { Markdown } from "yt-dlp-transcript-common/components/Markdown"; +import type { UiMessage } from "../lib/askConversation"; +import { PipelineStatus } from "./PipelineStatus"; + +function CopyButton({ text }: { text: string }) { + const [copied, setCopied] = useState(false); + return ( + <button + type="button" + onClick={() => { + void navigator.clipboard?.writeText(text).then(() => { + setCopied(true); + setTimeout(() => setCopied(false), 1500); + }); + }} + className="inline-flex items-center gap-1 rounded px-1.5 py-0.5 text-xs text-muted-foreground transition-colors hover:text-foreground" + aria-label="Copy answer" + > + {copied ? ( + <> + <CheckIcon className="size-3.5 text-brand" /> Copied + </> + ) : ( + <> + <CopyIcon className="size-3.5" /> Copy + </> + )} + </button> + ); +} + +// One chat message. Assistant turns show the live pipeline while gathering, +// stream as plain text with a caret (so partial Markdown never flickers), then +// re-render as Markdown once complete. Sources cite the videos the answer drew +// from, with the searches that found them. +export function MessageBubble({ + message, + markdownOn, +}: { + message: UiMessage; + markdownOn: boolean; +}) { + const isUser = message.role === "user"; + const streaming = message.phase === "answering" || message.phase === "streaming"; + const renderMarkdown = + !isUser && markdownOn && !message.error && message.phase === "done"; + + return ( + <div className="flex flex-col gap-2 animate-in fade-in slide-in-from-bottom-2 motion-reduce:animate-none"> + {!isUser && <PipelineStatus phase={message.phase} steps={message.searchSteps ?? []} />} + + {(message.content || streaming) && ( + <div + className={ + isUser + ? "self-end rounded-lg bg-primary/10 px-3 py-2 text-sm text-foreground" + : `rounded-lg border border-border bg-card/40 px-3 py-2 text-sm ${ + message.error ? "text-warning" : "text-foreground" + }` + } + > + {renderMarkdown ? ( + <Markdown className="leading-relaxed">{message.content}</Markdown> + ) : ( + <span className="whitespace-pre-wrap leading-relaxed"> + {message.content} + {streaming && <span className="ai-caret h-[1em] align-text-bottom" aria-hidden />} + </span> + )} + </div> + )} + + {!isUser && message.phase === "done" && (message.searchSteps?.length ?? 0) > 0 && ( + <p className="font-mono text-xs text-muted-foreground/70"> + Searched: {message.searchSteps!.map((s) => s.query).join(" · ")} + </p> + )} + + {!isUser && message.phase === "done" && message.content && ( + <div className="flex items-center"> + <CopyButton text={message.content} /> + </div> + )} + + {!isUser && message.sources && message.sources.length > 0 && ( + <ol className="ml-1 flex flex-col gap-1 text-xs text-muted-foreground"> + {message.sources.map((s, si) => ( + <li + key={s.key} + className="animate-in fade-in motion-reduce:animate-none" + style={{ animationDelay: `${Math.min(si, 8) * 40}ms`, animationFillMode: "both" }} + > + <span className="font-mono text-brand">[{si + 1}]</span>{" "} + {s.url ? ( + <a + href={s.url} + target="_blank" + rel="noopener noreferrer" + className="hover:underline" + > + {s.title} + </a> + ) : ( + s.title + )}{" "} + <span className="text-muted-foreground/70"> + — {s.channel} + {s.siteTitle ? ` · ${s.siteTitle}` : ""} ·{" "} + {s.snippets.map((sn) => sn.clock).join(", ")} + </span> + </li> + ))} + {message.truncated && ( + <li className="text-muted-foreground/60"> + (some searches were truncated; ask more specifically for fuller + coverage) + </li> + )} + </ol> + )} + </div> + ); +} diff --git a/export/app/ask/PipelineStatus.tsx b/export/app/ask/PipelineStatus.tsx @@ -0,0 +1,62 @@ +"use client"; + +import { CheckIcon, Loader2Icon } from "lucide-react"; +import type { AssistantPhase, SearchStep } from "../lib/askConversation"; + +// The retrieval→answer feedback, shown while a turn is in flight (before the +// first answer token). A restrained terminal-flavoured trace: each search the +// model runs appears as its own row — a spinner while running, a check with the +// match count when done — so multi-step "search, read, search again" is visible. +// Once the answer starts streaming this returns null and the text takes over. +export function PipelineStatus({ + phase, + steps, +}: { + phase?: AssistantPhase; + steps: SearchStep[]; +}) { + if (phase !== "gathering" && phase !== "answering") return null; + const thinking = phase === "gathering" && steps.length === 0; + + return ( + <div className="flex flex-col gap-1.5 font-mono text-xs text-muted-foreground"> + {thinking && ( + <div className="flex items-center gap-2"> + <span className="flex items-center gap-1" aria-hidden> + <span className="ai-dot inline-block size-1.5 rounded-full bg-current" /> + <span className="ai-dot inline-block size-1.5 rounded-full bg-current" /> + <span className="ai-dot inline-block size-1.5 rounded-full bg-current" /> + </span> + <span>Reading your question…</span> + </div> + )} + + {steps.map((s, i) => ( + <div + key={i} + className="flex items-center gap-2 animate-in fade-in slide-in-from-top-1 motion-reduce:animate-none" + > + {s.count === undefined ? ( + <Loader2Icon className="size-3.5 shrink-0 animate-spin text-brand motion-reduce:animate-none" /> + ) : ( + <CheckIcon className="size-3.5 shrink-0 text-brand" /> + )} + <span className="shrink-0">searching</span> + <span className="truncate text-foreground">“{s.query}”</span> + {s.count !== undefined && ( + <span className="shrink-0 text-muted-foreground/70"> + · {s.count} result{s.count === 1 ? "" : "s"} + </span> + )} + </div> + ))} + + {phase === "answering" && ( + <div className="flex items-center gap-2 animate-in fade-in slide-in-from-top-1 motion-reduce:animate-none"> + <Loader2Icon className="size-3.5 shrink-0 animate-spin text-brand motion-reduce:animate-none" /> + <span>Writing answer…</span> + </div> + )} + </div> + ); +} diff --git a/export/app/ask/ProviderSettings.tsx b/export/app/ask/ProviderSettings.tsx @@ -0,0 +1,161 @@ +"use client"; + +import { PROVIDERS, type Provider } from "../lib/askProvider"; +import type { AgentMode } from "../lib/searchAgent"; + +const MODES: { value: AgentMode; label: string }[] = [ + { value: "auto", label: "Auto" }, + { value: "native", label: "Native tools" }, + { value: "scripted", label: "Scripted" }, +]; + +type Props = { + provider: Provider; + setProvider: (p: Provider) => void; + apiKey: string; + setApiKey: (v: string) => void; + model: string; + setModel: (v: string) => void; + remember: boolean; + setRemember: (v: boolean) => void; + showKey: boolean; + setShowKey: (v: boolean) => void; + searchMode: AgentMode; + setSearchMode: (m: AgentMode) => void; + persistKey: (p: Provider, key: string, mdl: string, rememberOn: boolean) => void; +}; + +export function ProviderSettings(props: Props) { + const { + provider, + setProvider, + apiKey, + setApiKey, + model, + setModel, + remember, + setRemember, + showKey, + setShowKey, + searchMode, + setSearchMode, + persistKey, + } = props; + const info = PROVIDERS[provider]; + + return ( + <details className="rounded-lg border border-border bg-card/40" open={!apiKey}> + <summary className="cursor-pointer px-4 py-2.5 text-sm font-medium text-foreground"> + {apiKey ? `${info.label} · key set` : "Set up your AI provider"} + </summary> + <div className="flex flex-col gap-3 border-t border-border px-4 py-3"> + <div className="flex flex-wrap gap-3"> + <label className="flex flex-col gap-1 text-xs text-muted-foreground"> + Provider + <select + value={provider} + onChange={(e) => setProvider(e.target.value as Provider)} + className="rounded-md border border-border bg-background px-2 py-1.5 text-sm text-foreground" + > + {(Object.keys(PROVIDERS) as Provider[]).map((p) => ( + <option key={p} value={p}> + {PROVIDERS[p].label} + </option> + ))} + </select> + </label> + <label className="flex min-w-[10rem] flex-1 flex-col gap-1 text-xs text-muted-foreground"> + Model + <input + list="ask-models" + value={model} + onChange={(e) => setModel(e.target.value)} + className="rounded-md border border-border bg-background px-2 py-1.5 text-sm text-foreground" + /> + <datalist id="ask-models"> + {info.models.map((m) => ( + <option key={m} value={m} /> + ))} + </datalist> + </label> + </div> + + <label className="flex flex-col gap-1 text-xs text-muted-foreground"> + API key + <div className="flex gap-2"> + <input + type={showKey ? "text" : "password"} + value={apiKey} + onChange={(e) => setApiKey(e.target.value)} + placeholder={info.keyHint} + autoComplete="off" + className="flex-1 rounded-md border border-border bg-background px-2 py-1.5 font-mono text-sm text-foreground" + /> + <button + type="button" + onClick={() => setShowKey(!showKey)} + className="rounded-md border border-border px-2 text-xs text-muted-foreground hover:text-foreground" + > + {showKey ? "Hide" : "Show"} + </button> + </div> + </label> + + <label className="flex items-center gap-2 text-xs text-muted-foreground"> + <input + type="checkbox" + checked={remember} + onChange={(e) => { + setRemember(e.target.checked); + persistKey(provider, apiKey, model, e.target.checked); + }} + /> + Remember my key in this browser + </label> + + {/* Search mode: how the model drives multi-step retrieval. */} + <div className="flex flex-col gap-1 text-xs text-muted-foreground"> + <span>Search mode</span> + <div className="inline-flex w-fit overflow-hidden rounded-md border border-border"> + {MODES.map((m) => ( + <button + key={m.value} + type="button" + onClick={() => setSearchMode(m.value)} + aria-pressed={searchMode === m.value} + className={ + "px-2.5 py-1 text-xs transition-colors " + + (searchMode === m.value + ? "bg-brand text-brand-ink" + : "text-muted-foreground hover:text-foreground") + } + > + {m.label} + </button> + ))} + </div> + <span className="text-muted-foreground/70"> + Auto uses your model’s native tool-calling when supported and falls + back to a text protocol otherwise. Switch to Scripted if searches + misbehave. + </span> + </div> + + <p className="text-xs text-muted-foreground/80"> + Your key is stored only{" "} + {remember ? "in this browser" : "for this page session"} and is sent + directly to {info.label} — never to this site. Requests are billed to + your own account.{" "} + <a + href={info.keyUrl} + target="_blank" + rel="noopener noreferrer" + className="text-brand hover:underline" + > + Get a key → + </a> + </p> + </div> + </details> + ); +} diff --git a/export/app/ask/useAskChat.ts b/export/app/ask/useAskChat.ts @@ -0,0 +1,275 @@ +"use client"; + +import { useCallback, useEffect, useRef, useState } from "react"; +import { useSearchData } from "yt-dlp-transcript-common/components/SearchDataContext"; +import { PROVIDERS, type Provider } from "../lib/askProvider"; +import { runAskTurn, type AgentEvent, type AgentMode } from "../lib/searchAgent"; +import type { SearchStep, UiMessage } from "../lib/askConversation"; + +const K_PROVIDER = "ytdlp-tb:ai:provider"; +const K_REMEMBER = "ytdlp-tb:ai:remember"; +const K_SEARCHMODE = "ytdlp-tb:ai:searchmode"; +const K_MARKDOWN = "ytdlp-tb:ai:md"; +const keyFor = (p: Provider) => `ytdlp-tb:ai:key:${p}`; +const modelFor = (p: Provider) => `ytdlp-tb:ai:model:${p}`; + +// State + turn orchestration for the /ask chat. The view is split into small +// presentational components; this hook owns everything mutable and calls +// runAskTurn (gather → answer), mapping its event stream onto the pending +// assistant message. +export function useAskChat() { + const [provider, setProviderState] = useState<Provider>("anthropic"); + const [apiKey, setApiKey] = useState(""); + const [model, setModel] = useState(PROVIDERS.anthropic.defaultModel); + const [remember, setRemember] = useState(true); + const [showKey, setShowKey] = useState(false); + const [searchMode, setSearchModeState] = useState<AgentMode>("auto"); + const [markdownOn, setMarkdownOnState] = useState(true); + + const { summariesState, channels, aliases } = useSearchData(); + const { summaries, summariesReady } = summariesState; + const corpusError = summariesState.error?.message ?? null; + + const [messages, setMessages] = useState<UiMessage[]>([]); + const [input, setInput] = useState(""); + const [busy, setBusy] = useState(false); + const abortRef = useRef<AbortController | null>(null); + + // Restore saved preferences + (if remembered) the provider's key/model. + useEffect(() => { + /* eslint-disable react-hooks/set-state-in-effect -- one-time restore from + localStorage after hydration; it must run in an effect, not a lazy + useState initializer, because there is no localStorage during SSR. */ + try { + const savedProvider = localStorage.getItem(K_PROVIDER) as Provider | null; + const rememberSaved = localStorage.getItem(K_REMEMBER) !== "0"; + const p = + savedProvider && PROVIDERS[savedProvider] ? savedProvider : "anthropic"; + setProviderState(p); + setRemember(rememberSaved); + const savedMode = localStorage.getItem(K_SEARCHMODE) as AgentMode | null; + if (savedMode === "auto" || savedMode === "native" || savedMode === "scripted") { + setSearchModeState(savedMode); + } + setMarkdownOnState(localStorage.getItem(K_MARKDOWN) !== "0"); + loadProviderCreds(p, rememberSaved); + } catch { + /* storage unavailable */ + } + /* eslint-enable react-hooks/set-state-in-effect */ + }, []); + + function loadProviderCreds(p: Provider, rememberOn: boolean) { + if (rememberOn) { + setApiKey(localStorage.getItem(keyFor(p)) ?? ""); + setModel(localStorage.getItem(modelFor(p)) ?? PROVIDERS[p].defaultModel); + } else { + setApiKey(""); + setModel(PROVIDERS[p].defaultModel); + } + } + + function setProvider(p: Provider) { + setProviderState(p); + try { + localStorage.setItem(K_PROVIDER, p); + } catch { + /* ignore */ + } + loadProviderCreds(p, remember); + } + + function persistKey(p: Provider, key: string, mdl: string, rememberOn: boolean) { + try { + if (rememberOn) { + localStorage.setItem(keyFor(p), key); + localStorage.setItem(modelFor(p), mdl); + } else { + localStorage.removeItem(keyFor(p)); + localStorage.removeItem(modelFor(p)); + } + localStorage.setItem(K_REMEMBER, rememberOn ? "1" : "0"); + } catch { + /* ignore */ + } + } + + function setSearchMode(m: AgentMode) { + setSearchModeState(m); + try { + localStorage.setItem(K_SEARCHMODE, m); + } catch { + /* ignore */ + } + } + + function setMarkdownOn(on: boolean) { + setMarkdownOnState(on); + try { + localStorage.setItem(K_MARKDOWN, on ? "1" : "0"); + } catch { + /* ignore */ + } + } + + const send = useCallback(async () => { + const question = input.trim(); + if (!question || busy || !apiKey.trim() || !summariesReady) return; + + persistKey(provider, apiKey, model, remember); + const prior = messages; + + setInput(""); + setMessages((prev) => [ + ...prev, + { role: "user", content: question }, + { role: "assistant", content: "", phase: "gathering", searchSteps: [] }, + ]); + setBusy(true); + const ac = new AbortController(); + abortRef.current = ac; + + // Patch the trailing assistant / the preceding user message immutably. + const patchAssistant = (fn: (m: UiMessage) => UiMessage) => + setMessages((prev) => { + const next = prev.slice(); + next[next.length - 1] = fn(next[next.length - 1]); + return next; + }); + const patchUser = (fn: (m: UiMessage) => UiMessage) => + setMessages((prev) => { + const next = prev.slice(); + next[next.length - 2] = fn(next[next.length - 2]); + return next; + }); + + const onEvent = (e: AgentEvent) => { + switch (e.type) { + case "search_start": + patchAssistant((m) => ({ + ...m, + phase: "gathering", + searchSteps: [...(m.searchSteps ?? []), { query: e.query }], + })); + break; + case "search_done": + patchAssistant((m) => { + const steps = (m.searchSteps ?? []).slice(); + // Fill the most recent step for this query still awaiting a count. + for (let i = steps.length - 1; i >= 0; i--) { + if (steps[i].query === e.query && steps[i].count === undefined) { + steps[i] = { ...steps[i], count: e.count }; + break; + } + } + return { ...m, searchSteps: steps }; + }); + break; + case "answer_start": + patchAssistant((m) => ({ ...m, phase: "answering" })); + break; + case "delta": + patchAssistant((m) => ({ + ...m, + phase: "streaming", + content: m.content + e.text, + })); + break; + } + }; + + try { + const result = await runAskTurn({ + provider, + apiKey: apiKey.trim(), + model: model.trim() || PROVIDERS[provider].defaultModel, + mode: searchMode, + question, + prior, + summaries, + aliases, + signal: ac.signal, + onEvent, + }); + patchUser((m) => ({ ...m, groundedContent: result.groundedContent })); + patchAssistant((m) => ({ + ...m, + content: m.content || result.answer, + sources: result.videos, + truncated: result.truncated, + phase: "done", + })); + } catch (e) { + const err = e as Error; + if (err.name === "AbortError") { + patchAssistant((m) => ({ + ...m, + content: m.content + "\n\n_(stopped)_", + phase: "stopped", + })); + } else { + patchAssistant((m) => ({ + ...m, + content: m.content || err.message, + error: true, + phase: "error", + })); + } + } finally { + setBusy(false); + abortRef.current = null; + } + }, [ + input, + busy, + apiKey, + summariesReady, + summaries, + aliases, + provider, + model, + remember, + searchMode, + messages, + ]); + + const stop = () => abortRef.current?.abort(); + const reset = () => { + if (busy) return; + setMessages([]); + }; + + return { + // provider settings + provider, + setProvider, + apiKey, + setApiKey, + model, + setModel, + remember, + setRemember, + showKey, + setShowKey, + persistKey, + searchMode, + setSearchMode, + markdownOn, + setMarkdownOn, + // data + summariesReady, + corpusError, + channels, + // conversation + messages, + input, + setInput, + busy, + send, + stop, + reset, + }; +} + +export type AskChatState = ReturnType<typeof useAskChat>; +export type { SearchStep }; diff --git a/export/app/globals.css b/export/app/globals.css @@ -55,3 +55,54 @@ body { animation: none; } } + +/* Ask chat: a terminal-flavoured blinking caret at the tail of the streaming + answer, and a three-dot "thinking" pulse before the first token arrives. Both + yield to a reduced-motion preference. */ +@keyframes ai-caret { + 0%, + 100% { + opacity: 1; + } + 50% { + opacity: 0; + } +} +.ai-caret { + display: inline-block; + width: 0.5ch; + margin-left: 1px; + background: currentColor; + animation: ai-caret 1s step-end infinite; +} +@keyframes ai-dot { + 0%, + 80%, + 100% { + opacity: 0.25; + transform: translateY(0); + } + 40% { + opacity: 1; + transform: translateY(-2px); + } +} +.ai-dot { + animation: ai-dot 1.2s ease-in-out infinite; +} +.ai-dot:nth-child(2) { + animation-delay: 0.15s; +} +.ai-dot:nth-child(3) { + animation-delay: 0.3s; +} +@media (prefers-reduced-motion: reduce) { + .ai-caret { + animation: none; + opacity: 1; + } + .ai-dot { + animation: none; + opacity: 0.6; + } +} diff --git a/export/app/lib/askConversation.test.ts b/export/app/lib/askConversation.test.ts @@ -0,0 +1,95 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import type { SearchAlias } from "yt-dlp-transcript-common/lib/searchAliases"; +import { + buildApiMessages, + buildGroundedContent, + renderAliasGlossary, + gatherSystemPrompt, + answerSystemPrompt, + type UiMessage, +} from "./askConversation"; +import type { RetrievedVideo } from "./askRetrieval"; + +const ALIASES: SearchAlias[] = [ + { + id: "platner", + label: "Graham Platner", + triggers: ["platner", "plattner", "platter"], + suggestion: "plat+ner", + useRegex: true, + note: "often mis-transcribed", + }, +]; + +function video(over: Partial<RetrievedVideo> = {}): RetrievedVideo { + return { + key: "a", + videoId: "a", + title: "Platner interview", + channel: "Rekieta", + uploadDate: "20200101", + snippets: [{ clock: "0:05", seconds: 5, text: "he said" }], + ...over, + }; +} + +test("buildApiMessages replays prior user turns with their grounded content", () => { + const prior: UiMessage[] = [ + { + role: "user", + content: "tell me about Platner", + groundedContent: "tell me about Platner\n\n---\n[1] excerpt about platner", + }, + { role: "assistant", content: "He is ...", sources: [video()] }, + ]; + const msgs = buildApiMessages(prior); + assert.equal(msgs.length, 2); + // The excerpts from turn 1 survive into the replayed history (the core fix). + assert.match(msgs[0].content, /excerpt about platner/); + assert.equal(msgs[1].content, "He is ..."); +}); + +test("buildApiMessages drops error turns and empty pending assistants", () => { + const prior: UiMessage[] = [ + { role: "user", content: "q", groundedContent: "q" }, + { role: "assistant", content: "boom", error: true }, + { role: "assistant", content: "" }, // pending, no grounded content + ]; + assert.equal(buildApiMessages(prior).length, 1); +}); + +test("buildGroundedContent appends numbered excerpts, or just the question", () => { + assert.equal(buildGroundedContent("hi", []), "hi"); + const grounded = buildGroundedContent("who?", [video()]); + assert.match(grounded, /who\?/); + assert.match(grounded, /\[1\] "Platner interview" — Rekieta/); +}); + +test("renderAliasGlossary lists variants and notes; empty when none", () => { + const g = renderAliasGlossary(ALIASES); + assert.match(g, /Graham Platner: also transcribed as platner, plattner, platter/); + assert.match(g, /often mis-transcribed/); + assert.equal(renderAliasGlossary([]), ""); + assert.equal( + renderAliasGlossary([{ ...ALIASES[0], enabled: false }]), + "", + ); +}); + +test("gatherSystemPrompt tailors the tail per mode and includes the glossary", () => { + const scripted = gatherSystemPrompt(ALIASES, "scripted", 4); + assert.match(scripted, /SEARCH: <query>/); + assert.match(scripted, /at most 4 searches/); + assert.match(scripted, /Graham Platner/); + const native = gatherSystemPrompt(ALIASES, "native", 3); + assert.match(native, /search_transcripts tool/); + assert.match(native, /finish tool/); +}); + +test("answerSystemPrompt asks for Markdown + citations and carries the glossary", () => { + const p = answerSystemPrompt(ALIASES); + assert.match(p, /Markdown/); + assert.match(p, /\[1\]/); + assert.match(p, /Graham Platner/); +}); diff --git a/export/app/lib/askConversation.ts b/export/app/lib/askConversation.ts @@ -0,0 +1,138 @@ +// Conversation assembly + prompts for the /ask chat. +// +// This is the layer that fixes the "follow-up loses its grounding" bug: prior +// user turns are replayed with the EXACT text that was sent (question + the +// excerpts retrieved that turn), not the bare question — so excerpts from +// earlier turns stay in context and a follow-up like "format it in a timeline" +// can reuse them. The search agent (searchAgent.ts) drives retrieval; this +// module owns the message shapes and the system prompts (including an alias +// glossary so the model accounts for AI-transcription misspellings). +// +// Pure (no I/O, no React) so it can be unit-tested. + +import type { ChatMessage } from "./askProvider"; +import { buildContext, type RetrievedVideo } from "./askRetrieval"; +import type { SearchAlias } from "yt-dlp-transcript-common/lib/searchAliases"; + +// One search the agent ran this turn (shown live in the pipeline UI). `count` +// is undefined while the search is in flight, then set to the match count. +export type SearchStep = { query: string; count?: number }; + +// Top-level stage of a pending assistant turn, for the status indicator. +export type AssistantPhase = + | "gathering" // agent is deciding / running searches + | "answering" // gather done, awaiting the first answer token + | "streaming" // answer tokens are arriving + | "done" + | "error" + | "stopped"; + +export type UiMessage = { + role: "user" | "assistant"; + content: string; + // User turns: the exact string sent to the model (question + this turn's + // excerpts). Replayed verbatim so earlier grounding persists across turns. + groundedContent?: string; + // Assistant turns: + sources?: RetrievedVideo[]; + searchSteps?: SearchStep[]; + phase?: AssistantPhase; + truncated?: boolean; + error?: boolean; +}; + +// Replay completed prior turns for the model. User turns use their persisted +// grounded content (with excerpts); assistant turns use their text. Error turns +// are dropped. The CURRENT turn is appended by the caller (bare question for the +// gather phase, grounded question for the answer phase). +export function buildApiMessages(prior: UiMessage[]): ChatMessage[] { + return prior + .filter((m) => !m.error && (m.content.trim() !== "" || m.groundedContent)) + .map((m) => ({ + role: m.role, + content: + m.role === "user" ? m.groundedContent ?? m.content : m.content, + })); +} + +// Assemble the grounded user content for a turn: the question plus this turn's +// retrieved excerpts (numbered for citation). No excerpts → just the question, +// so a no-search follow-up leans on excerpts already in earlier turns. +export function buildGroundedContent( + question: string, + videos: RetrievedVideo[], +): string { + if (videos.length === 0) return question; + return ( + `${question}\n\n---\n` + + `Transcript excerpts you may cite (by number):\n${buildContext(videos)}` + ); +} + +// A compact glossary of known terms and their common (mis)transcriptions, so the +// model forms misspelling-aware searches and reconciles variant spellings it +// sees in excerpts. Empty when there are no usable aliases (block is omitted). +export function renderAliasGlossary(aliases: SearchAlias[], max = 40): string { + const usable = aliases.filter((a) => a.enabled !== false); + if (usable.length === 0) return ""; + const shown = usable.slice(0, max); + const lines = shown.map((a) => { + const variants = a.triggers.join(", "); + const note = a.note ? ` (${a.note})` : ""; + return `- ${a.label}: also transcribed as ${variants}${note}`; + }); + const overflow = + usable.length > max ? `\n…and ${usable.length - max} more.` : ""; + return ( + "Known terms in this archive and how AI transcription commonly mangles " + + "them — account for these variant spellings when you search and when you " + + "read excerpts:\n" + + lines.join("\n") + + overflow + ); +} + +function withGlossary(base: string, aliases: SearchAlias[]): string { + const glossary = renderAliasGlossary(aliases); + return glossary ? `${base}\n\n${glossary}` : base; +} + +// System prompt for the gather (search-decision) phase. `mode` tailors the +// closing instruction: native tool-calling vs. the scripted text protocol. +export function gatherSystemPrompt( + aliases: SearchAlias[], + mode: "native" | "scripted", + budget: number, +): string { + const base = + "You are helping answer a question about a video-transcript archive. " + + "Before answering you may search the transcripts to gather relevant " + + "excerpts. Use focused keyword or name queries. Read each result set and, " + + "if it helps, search again with a refined query — for example correcting a " + + "spelling, trying an alternate name, or narrowing to a specific event. " + + "Only search when you need transcript evidence: if the latest message just " + + "asks to reformat, summarise, translate, or expand on the previous answer, " + + `do not search. You may run at most ${budget} searches.`; + const tail = + mode === "native" + ? "Call the search_transcripts tool to search. Call the finish tool as " + + "soon as you have enough excerpts, or immediately if no search is needed." + : "Reply with EXACTLY one line and nothing else: either " + + "`SEARCH: <query>` to run a search, or `DONE` when you have enough " + + "excerpts (or need no search)."; + return withGlossary(`${base}\n\n${tail}`, aliases); +} + +// System prompt for the final answer phase. +export function answerSystemPrompt(aliases: SearchAlias[]): string { + const base = + "You are answering a question about a video-transcript archive. Base your " + + "answer on the transcript excerpts provided in this conversation. Excerpts " + + "may appear in earlier turns, and a follow-up request (reformatting, " + + "summarising, expanding) should reuse the relevant excerpts already " + + "provided. Cite the excerpts you use by their bracketed number, e.g. [1]. " + + "Each excerpt shows a video title, channel, and timestamped lines. If the " + + "excerpts do not contain enough to answer, say so plainly rather than " + + "guessing. Format your answer in GitHub-flavored Markdown."; + return withGlossary(base, aliases); +} diff --git a/export/app/lib/askProvider.ts b/export/app/lib/askProvider.ts @@ -78,6 +78,13 @@ export async function askStream(opts: AskOptions): Promise<string> { } } +// Non-streaming single-shot completion, used by the scripted search loop to get +// a short SEARCH:/DONE decision. Reuses the streaming machinery and collects the +// full text; callers pass a small `maxTokens` to keep decision turns cheap. +export function askOnce(opts: AskOptions): Promise<string> { + return askStream({ ...opts, onDelta: undefined }); +} + // ─── shared SSE plumbing ─── async function* sseLines( @@ -180,6 +187,9 @@ async function askOpenAI(opts: AskOptions): Promise<string> { body: JSON.stringify({ model: opts.model || PROVIDERS.openai.defaultModel, stream: true, + // Only cap when asked (the scripted search loop passes a small value); + // otherwise let the provider default so answers aren't truncated. + ...(opts.maxTokens ? { max_tokens: opts.maxTokens } : {}), messages: [ { role: "system", content: opts.system }, ...opts.messages.map((m) => ({ role: m.role, content: m.content })), @@ -221,6 +231,9 @@ async function askGemini(opts: AskOptions): Promise<string> { role: m.role === "assistant" ? "model" : "user", parts: [{ text: m.content }], })), + ...(opts.maxTokens + ? { generationConfig: { maxOutputTokens: opts.maxTokens } } + : {}), }), }); await ensureOk(res, "Gemini"); diff --git a/export/app/lib/askRetrieval.test.ts b/export/app/lib/askRetrieval.test.ts @@ -3,7 +3,14 @@ import assert from "node:assert/strict"; import type { DisplaySummary } from "yt-dlp-transcript-common/lib/transcripts"; import type { LayerHit } from "yt-dlp-transcript-common/components/searchPipeline"; import type { TreeProgress } from "yt-dlp-transcript-common/lib/searchEval"; -import { extractKeywords, rankResults, buildContext } from "./askRetrieval"; +import type { SearchAlias } from "yt-dlp-transcript-common/lib/searchAliases"; +import { + extractKeywords, + rankResults, + buildContext, + buildSearchRoot, +} from "./askRetrieval"; +import { isLeaf } from "yt-dlp-transcript-common/lib/searchQuery"; function summary(over: Partial<DisplaySummary>): DisplaySummary { return { @@ -137,6 +144,28 @@ test("rankResults picks diverse snippets and formats timestamps", () => { assert.equal(v.snippets[0].text, "first hit"); }); +test("buildSearchRoot alias-expands a matching keyword into a regex leaf", () => { + const aliases: SearchAlias[] = [ + { + id: "loli", + label: "loli", + triggers: ["loli", "lolly", "loly"], + suggestion: "\\blol(i|ly)", + useRegex: true, + }, + ]; + const root = buildSearchRoot(["lolly", "banana"], aliases); + const leaves = root.children.filter(isLeaf); + const t0 = leaves.find((l) => l.id === "t#0")!; + // "lolly" matches the alias trigger → transcript leaf uses the regex pattern. + assert.equal(t0.query, "\\blol(i|ly)"); + assert.equal(t0.useRegex, true); + const t1 = leaves.find((l) => l.id === "t#1")!; + // "banana" has no alias → plain substring leaf. + assert.equal(t1.query, "banana"); + assert.equal(t1.useRegex, false); +}); + test("buildContext numbers videos and indents snippets", () => { const ctx = buildContext([ { diff --git a/export/app/lib/askRetrieval.ts b/export/app/lib/askRetrieval.ts @@ -2,7 +2,8 @@ import { runQueryTree, type TreeProgress, } from "yt-dlp-transcript-common/lib/searchEval"; -import { newGroup, newLeaf } from "yt-dlp-transcript-common/lib/searchQuery"; +import { newGroup, newLeaf, type LayerScope } from "yt-dlp-transcript-common/lib/searchQuery"; +import { matchAliases, type SearchAlias } from "yt-dlp-transcript-common/lib/searchAliases"; import type { LayerHit } from "yt-dlp-transcript-common/components/searchPipeline"; import type { DisplaySummary } from "yt-dlp-transcript-common/lib/transcripts"; import { formatDuration } from "yt-dlp-transcript-common/lib/format"; @@ -162,6 +163,11 @@ export type RetrieveOptions = RankOptions & { question: string; summaries: DisplaySummary[]; signal?: AbortSignal; + // Known search aliases. When a keyword matches an alias trigger, the leaf is + // built from the alias's replacement pattern (`suggestion` + `useRegex`) so a + // query catches known AI-transcription misspellings (e.g. a name spelled + // several ways). Empty/omitted = plain substring leaves, as before. + aliases?: SearchAlias[]; // Total hit cap. Bounds how many transcript pages get fetched — the engine // stops once this many hits accumulate. NOT Infinity (that would scan the // whole ~800MB corpus). Capped results still work; they just aren't persisted @@ -169,6 +175,45 @@ export type RetrieveOptions = RankOptions & { hitLimit?: number; }; +// Build a search leaf for one keyword in one scope, applying an alias +// replacement (regex pattern) when the keyword matches a known trigger. +function leafFor( + keyword: string, + scope: LayerScope, + id: string, + aliases: SearchAlias[], +) { + const hit = aliases.length ? matchAliases(keyword, scope, aliases)[0] : undefined; + if (hit) { + return newLeaf({ + id, + query: hit.suggestion, + scope, + useRegex: hit.useRegex, + contributeHits: true, + }); + } + return newLeaf({ id, query: keyword, scope, contributeHits: true }); +} + +// Build the OR query root for a set of keywords: a transcript-cue leaf and a +// fetch-free metadata leaf per keyword, each alias-expanded when the keyword +// matches a known trigger. Exported for unit testing. Leaf ids encode the +// keyword index after "#" so ranking collapses a keyword's two scopes into one +// unit of term coverage (see termKey() in rankResults). +export function buildSearchRoot( + keywords: string[], + aliases: SearchAlias[] = [], +) { + return newGroup({ + op: "OR", + children: keywords.flatMap((kw, i) => [ + leafFor(kw, "transcripts", `t#${i}`, aliases), + leafFor(kw, "metadata", `m#${i}`, aliases), + ]), + }); +} + // Run the question through the shared search engine and return ranked videos. // Browser-only: runQueryTree depends on window timers + fetch. Resolves once the // engine reports done; rejects with an AbortError if the signal fires first. @@ -176,6 +221,7 @@ export function retrieve( opts: RetrieveOptions, ): Promise<{ videos: RetrievedVideo[]; truncated: boolean }> { const { question, summaries, signal } = opts; + const aliases = opts.aliases ?? []; const keywords = extractKeywords(question); return new Promise((resolve, reject) => { @@ -194,13 +240,7 @@ export function retrieve( // Leaf ids encode the keyword index after a "#" so ranking can collapse a // keyword's transcript + metadata leaves into ONE unit of term coverage // (see termKey() in rankResults) rather than double-counting the two scopes. - const root = newGroup({ - op: "OR", - children: keywords.flatMap((kw, i) => [ - newLeaf({ id: `t#${i}`, query: kw, scope: "transcripts", contributeHits: true }), - newLeaf({ id: `m#${i}`, query: kw, scope: "metadata", contributeHits: true }), - ]), - }); + const root = buildSearchRoot(keywords, aliases); let settled = false; const finish = (fn: () => void) => { diff --git a/export/app/lib/nativeTools/anthropic.ts b/export/app/lib/nativeTools/anthropic.ts @@ -0,0 +1,122 @@ +// Anthropic native tool-calling gather. Non-streaming /v1/messages loop using +// tool_use / tool_result content blocks. + +import { PROVIDERS } from "../askProvider"; +import { + abortError, + EMPTY_PARAMS, + FINISH_TOOL_DESCRIPTION, + FINISH_TOOL_NAME, + postJson, + SEARCH_PARAMS, + SEARCH_TOOL_DESCRIPTION, + SEARCH_TOOL_NAME, + type NativeGatherContext, +} from "./shared"; + +const TOOLS = [ + { + name: SEARCH_TOOL_NAME, + description: SEARCH_TOOL_DESCRIPTION, + input_schema: SEARCH_PARAMS, + }, + { + name: FINISH_TOOL_NAME, + description: FINISH_TOOL_DESCRIPTION, + input_schema: EMPTY_PARAMS, + }, +]; + +// Pure: extract tool_use blocks from an Anthropic message response. +export function parseAnthropicToolUses( + json: unknown, +): { id: string; name: string; query: string }[] { + const content = (json as { content?: unknown }).content; + if (!Array.isArray(content)) return []; + const out: { id: string; name: string; query: string }[] = []; + for (const block of content) { + const b = block as { + type?: string; + id?: string; + name?: string; + input?: { query?: unknown }; + }; + if (b.type === "tool_use" && typeof b.name === "string") { + out.push({ + id: b.id ?? "", + name: b.name, + query: typeof b.input?.query === "string" ? b.input.query : "", + }); + } + } + return out; +} + +export async function anthropicGather(ctx: NativeGatherContext): Promise<void> { + const messages: { role: string; content: unknown }[] = [ + ...ctx.history.map((m) => ({ role: m.role, content: m.content })), + { role: "user", content: ctx.question }, + ]; + let searchesRun = 0; + + for (let round = 0; round <= ctx.budget + 1; round++) { + if (ctx.signal?.aborted) throw abortError(); + const json = await postJson( + "https://api.anthropic.com/v1/messages", + { + "x-api-key": ctx.apiKey, + "anthropic-version": "2023-06-01", + "anthropic-dangerous-direct-browser-access": "true", + }, + { + model: ctx.model || PROVIDERS.anthropic.defaultModel, + max_tokens: 512, + system: ctx.system, + tools: TOOLS, + messages, + }, + "Anthropic", + ctx.signal, + ); + + const calls = parseAnthropicToolUses(json); + const searches = calls.filter( + (c) => c.name === SEARCH_TOOL_NAME && c.query.trim() !== "", + ); + // No search requested → the model finished (or answered directly). Done. + if (searches.length === 0) return; + + // Replay the assistant's tool_use turn verbatim, then answer each tool_use. + messages.push({ + role: "assistant", + content: (json as { content?: unknown[] }).content ?? [], + }); + const toolResults: unknown[] = []; + for (const call of calls) { + if (call.name === SEARCH_TOOL_NAME && call.query.trim() !== "") { + const content = + searchesRun < ctx.budget + ? await ((): Promise<string> => { + searchesRun += 1; + return ctx.runSearch(call.query); + })() + : "Search budget reached — answer with the excerpts gathered so far."; + toolResults.push({ + type: "tool_result", + tool_use_id: call.id, + content, + }); + } else { + // finish (or any other tool): acknowledge so the block is satisfied. + toolResults.push({ + type: "tool_result", + tool_use_id: call.id, + content: "Acknowledged.", + }); + } + } + messages.push({ role: "user", content: toolResults }); + + if (searchesRun >= ctx.budget) return; + } +} diff --git a/export/app/lib/nativeTools/gemini.ts b/export/app/lib/nativeTools/gemini.ts @@ -0,0 +1,124 @@ +// Gemini native tool-calling gather. Non-streaming :generateContent loop using +// functionDeclarations + functionCall / functionResponse parts. + +import { PROVIDERS } from "../askProvider"; +import { + abortError, + EMPTY_PARAMS, + FINISH_TOOL_DESCRIPTION, + FINISH_TOOL_NAME, + postJson, + SEARCH_PARAMS, + SEARCH_TOOL_DESCRIPTION, + SEARCH_TOOL_NAME, + type NativeGatherContext, +} from "./shared"; + +const TOOLS = [ + { + functionDeclarations: [ + { + name: SEARCH_TOOL_NAME, + description: SEARCH_TOOL_DESCRIPTION, + parameters: SEARCH_PARAMS, + }, + { + name: FINISH_TOOL_NAME, + description: FINISH_TOOL_DESCRIPTION, + parameters: EMPTY_PARAMS, + }, + ], + }, +]; + +// Pure: extract functionCall parts from a Gemini generateContent response. +export function parseGeminiFunctionCalls( + json: unknown, +): { name: string; query: string }[] { + const parts = + ( + json as { + candidates?: { content?: { parts?: unknown[] } }[]; + } + ).candidates?.[0]?.content?.parts ?? []; + const out: { name: string; query: string }[] = []; + for (const part of parts) { + const fc = (part as { functionCall?: { name?: string; args?: { query?: unknown } } }) + .functionCall; + if (fc?.name) { + out.push({ + name: fc.name, + query: typeof fc.args?.query === "string" ? fc.args.query : "", + }); + } + } + return out; +} + +export async function geminiGather(ctx: NativeGatherContext): Promise<void> { + const model = ctx.model || PROVIDERS.gemini.defaultModel; + const url = + `https://generativelanguage.googleapis.com/v1beta/models/` + + `${encodeURIComponent(model)}:generateContent?key=${encodeURIComponent(ctx.apiKey)}`; + const contents: unknown[] = [ + ...ctx.history.map((m) => ({ + role: m.role === "assistant" ? "model" : "user", + parts: [{ text: m.content }], + })), + { role: "user", parts: [{ text: ctx.question }] }, + ]; + let searchesRun = 0; + + for (let round = 0; round <= ctx.budget + 1; round++) { + if (ctx.signal?.aborted) throw abortError(); + const json = await postJson( + url, + {}, + { + system_instruction: { parts: [{ text: ctx.system }] }, + tools: TOOLS, + contents, + }, + "Gemini", + ctx.signal, + ); + + const calls = parseGeminiFunctionCalls(json); + const searches = calls.filter( + (c) => c.name === SEARCH_TOOL_NAME && c.query.trim() !== "", + ); + if (searches.length === 0) return; // finish or a plain answer → done + + // Replay the model's functionCall turn, then send functionResponse parts. + const modelParts = calls.map((c) => ({ + functionCall: { + name: c.name, + args: c.name === SEARCH_TOOL_NAME ? { query: c.query } : {}, + }, + })); + contents.push({ role: "model", parts: modelParts }); + + const responseParts: unknown[] = []; + for (const call of calls) { + let result: string; + if (call.name === SEARCH_TOOL_NAME && call.query.trim() !== "") { + result = + searchesRun < ctx.budget + ? await (() => { + searchesRun += 1; + return ctx.runSearch(call.query); + })() + : "Search budget reached — answer with the excerpts gathered so far."; + } else { + result = "Acknowledged."; + } + responseParts.push({ + functionResponse: { name: call.name, response: { result } }, + }); + } + // Function responses are sent back as a user-role turn. + contents.push({ role: "user", parts: responseParts }); + + if (searchesRun >= ctx.budget) return; + } +} diff --git a/export/app/lib/nativeTools/nativeTools.test.ts b/export/app/lib/nativeTools/nativeTools.test.ts @@ -0,0 +1,80 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { parseAnthropicToolUses } from "./anthropic"; +import { parseOpenAIToolCalls } from "./openai"; +import { parseGeminiFunctionCalls } from "./gemini"; + +test("parseAnthropicToolUses extracts tool_use blocks", () => { + const json = { + content: [ + { type: "text", text: "let me search" }, + { type: "tool_use", id: "tu_1", name: "search_transcripts", input: { query: "platner" } }, + { type: "tool_use", id: "tu_2", name: "finish", input: {} }, + ], + }; + assert.deepEqual(parseAnthropicToolUses(json), [ + { id: "tu_1", name: "search_transcripts", query: "platner" }, + { id: "tu_2", name: "finish", query: "" }, + ]); + assert.deepEqual(parseAnthropicToolUses({ content: "no tools" }), []); +}); + +test("parseOpenAIToolCalls parses JSON arguments", () => { + const json = { + choices: [ + { + message: { + tool_calls: [ + { + id: "call_1", + function: { name: "search_transcripts", arguments: '{"query":"senate campaign"}' }, + }, + { id: "call_2", function: { name: "finish", arguments: "{}" } }, + ], + }, + }, + ], + }; + assert.deepEqual(parseOpenAIToolCalls(json), [ + { id: "call_1", name: "search_transcripts", query: "senate campaign" }, + { id: "call_2", name: "finish", query: "" }, + ]); + // Malformed arguments degrade to an empty query, not a throw. + const bad = { + choices: [ + { + message: { + tool_calls: [ + { + id: "x", + function: { name: "search_transcripts", arguments: "{not json" }, + }, + ], + }, + }, + ], + }; + assert.deepEqual(parseOpenAIToolCalls(bad), [ + { id: "x", name: "search_transcripts", query: "" }, + ]); + assert.deepEqual(parseOpenAIToolCalls({ choices: [{ message: {} }] }), []); +}); + +test("parseGeminiFunctionCalls extracts functionCall parts", () => { + const json = { + candidates: [ + { + content: { + parts: [ + { text: "searching" }, + { functionCall: { name: "search_transcripts", args: { query: "timeline" } } }, + ], + }, + }, + ], + }; + assert.deepEqual(parseGeminiFunctionCalls(json), [ + { name: "search_transcripts", query: "timeline" }, + ]); + assert.deepEqual(parseGeminiFunctionCalls({ candidates: [] }), []); +}); diff --git a/export/app/lib/nativeTools/openai.ts b/export/app/lib/nativeTools/openai.ts @@ -0,0 +1,115 @@ +// OpenAI native tool-calling gather. Non-streaming /v1/chat/completions loop +// using function tool_calls + role:"tool" result messages. + +import { PROVIDERS } from "../askProvider"; +import { + abortError, + EMPTY_PARAMS, + FINISH_TOOL_DESCRIPTION, + FINISH_TOOL_NAME, + postJson, + SEARCH_PARAMS, + SEARCH_TOOL_DESCRIPTION, + SEARCH_TOOL_NAME, + type NativeGatherContext, +} from "./shared"; + +const TOOLS = [ + { + type: "function", + function: { + name: SEARCH_TOOL_NAME, + description: SEARCH_TOOL_DESCRIPTION, + parameters: SEARCH_PARAMS, + }, + }, + { + type: "function", + function: { + name: FINISH_TOOL_NAME, + description: FINISH_TOOL_DESCRIPTION, + parameters: EMPTY_PARAMS, + }, + }, +]; + +type RawToolCall = { + id?: string; + function?: { name?: string; arguments?: string }; +}; + +// Pure: extract tool calls from an OpenAI chat-completion response. +export function parseOpenAIToolCalls( + json: unknown, +): { id: string; name: string; query: string }[] { + const msg = (json as { choices?: { message?: { tool_calls?: unknown } }[] }) + .choices?.[0]?.message; + const calls = (msg as { tool_calls?: unknown })?.tool_calls; + if (!Array.isArray(calls)) return []; + return (calls as RawToolCall[]).map((c) => { + let query = ""; + try { + const args = JSON.parse(c.function?.arguments ?? "{}"); + if (typeof args.query === "string") query = args.query; + } catch { + /* malformed args → empty query */ + } + return { id: c.id ?? "", name: c.function?.name ?? "", query }; + }); +} + +export async function openaiGather(ctx: NativeGatherContext): Promise<void> { + const messages: unknown[] = [ + { role: "system", content: ctx.system }, + ...ctx.history.map((m) => ({ role: m.role, content: m.content })), + { role: "user", content: ctx.question }, + ]; + let searchesRun = 0; + + for (let round = 0; round <= ctx.budget + 1; round++) { + if (ctx.signal?.aborted) throw abortError(); + const json = await postJson( + "https://api.openai.com/v1/chat/completions", + { authorization: `Bearer ${ctx.apiKey}` }, + { + model: ctx.model || PROVIDERS.openai.defaultModel, + max_tokens: 512, + tools: TOOLS, + tool_choice: "auto", + messages, + }, + "OpenAI", + ctx.signal, + ); + + const rawMsg = (json as { choices?: { message?: unknown }[] }).choices?.[0] + ?.message; + const calls = parseOpenAIToolCalls(json); + const searches = calls.filter( + (c) => c.name === SEARCH_TOOL_NAME && c.query.trim() !== "", + ); + // No tool call (or only a plain answer) → done gathering. + if (calls.length === 0 || searches.length === 0) return; + + // Replay the assistant message with its tool_calls, then answer EVERY call + // (OpenAI requires a tool response for each tool_call id). + messages.push(rawMsg); + for (const call of calls) { + let content: string; + if (call.name === SEARCH_TOOL_NAME && call.query.trim() !== "") { + content = + searchesRun < ctx.budget + ? await (() => { + searchesRun += 1; + return ctx.runSearch(call.query); + })() + : "Search budget reached — answer with the excerpts gathered so far."; + } else { + content = "Acknowledged."; + } + messages.push({ role: "tool", tool_call_id: call.id, content }); + } + + if (searchesRun >= ctx.budget) return; + } +} diff --git a/export/app/lib/nativeTools/shared.ts b/export/app/lib/nativeTools/shared.ts @@ -0,0 +1,92 @@ +// Shared plumbing for the native function-calling gather transports. Each +// provider file (anthropic/openai/gemini) implements a non-streaming tool loop: +// call the model with a `search_transcripts` tool + a `finish` tool, run the +// searches it requests, feed results back, and stop when it finishes (or the +// search budget is reached). Only the request/response shapes differ per +// provider; this module holds the common types, tool text, and POST helper. + +import type { ChatMessage } from "../askProvider"; + +export type NativeGatherContext = { + apiKey: string; + model: string; + // Full gather system prompt (built by the caller via gatherSystemPrompt). + system: string; + // Prior turns (already grounded), replayed for context. + history: ChatMessage[]; + question: string; + // Max searches this turn. + budget: number; + // Runs one search and returns the results text to feed back to the model. + runSearch: (query: string) => Promise<string>; + signal?: AbortSignal; +}; + +export const SEARCH_TOOL_NAME = "search_transcripts"; +export const FINISH_TOOL_NAME = "finish"; + +export const SEARCH_TOOL_DESCRIPTION = + "Search the video-transcript archive for excerpts relevant to a query. " + + "Returns matching videos with timestamped snippet lines. Use focused keyword " + + "or name queries; call it again to refine based on what you find."; + +export const FINISH_TOOL_DESCRIPTION = + "Call this when you have gathered enough excerpts (or none are needed) and " + + "are ready to answer."; + +// JSON-Schema for the search tool's input (provider files wrap this in their +// own tool envelope). +export const SEARCH_PARAMS = { + type: "object", + properties: { + query: { + type: "string", + description: "Keywords or a name/phrase to search for.", + }, + }, + required: ["query"], +} as const; + +export const EMPTY_PARAMS = { type: "object", properties: {} } as const; + +// Thrown when a native tool-calling request fails in a way that suggests the +// model/endpoint doesn't support tools (HTTP 400/404). The agent catches this +// and retries the turn with the scripted transport. +export class ToolsUnavailableError extends Error {} + +export function abortError(): DOMException { + return new DOMException("Aborted", "AbortError"); +} + +// POST JSON and parse the response. 400/404 → ToolsUnavailableError (so the agent +// can fall back to scripted); other non-OK statuses throw a normal Error the chat +// surfaces (e.g. 401 bad key). +export async function postJson( + url: string, + headers: Record<string, string>, + body: unknown, + provider: string, + signal?: AbortSignal, +): Promise<unknown> { + if (signal?.aborted) throw abortError(); + const res = await fetch(url, { + method: "POST", + signal, + headers: { "content-type": "application/json", ...headers }, + body: JSON.stringify(body), + }); + if (!res.ok) { + let detail = ""; + try { + detail = await res.text(); + } catch { + /* ignore */ + } + const msg = `${provider} request failed (${res.status}). ${detail.slice(0, 300)}`; + if (res.status === 400 || res.status === 404) { + throw new ToolsUnavailableError(msg); + } + throw new Error(msg); + } + return res.json(); +} diff --git a/export/app/lib/searchAgent.test.ts b/export/app/lib/searchAgent.test.ts @@ -0,0 +1,72 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + parseScriptedDecision, + supportsNativeTools, + pickTransport, + formatResultsForModel, +} from "./searchAgent"; +import type { RetrievedVideo } from "./askRetrieval"; + +test("parseScriptedDecision reads a SEARCH line and strips quotes", () => { + assert.deepEqual(parseScriptedDecision("SEARCH: graham platner"), { + kind: "search", + query: "graham platner", + }); + assert.deepEqual(parseScriptedDecision('SEARCH: "travel ban"'), { + kind: "search", + query: "travel ban", + }); +}); + +test("parseScriptedDecision tolerates code fences and surrounding prose", () => { + const reply = "Sure, let me look.\n```\nSEARCH: senate campaign\n```"; + assert.deepEqual(parseScriptedDecision(reply), { + kind: "search", + query: "senate campaign", + }); +}); + +test("parseScriptedDecision treats DONE or plain prose as done", () => { + assert.deepEqual(parseScriptedDecision("DONE"), { kind: "done" }); + assert.deepEqual(parseScriptedDecision("I have enough to answer."), { + kind: "done", + }); + // A degenerate "SEARCH: done" is not a real query. + assert.deepEqual(parseScriptedDecision("SEARCH: done"), { kind: "done" }); +}); + +test("supportsNativeTools recognises tool-capable model families", () => { + assert.equal(supportsNativeTools("anthropic", "claude-haiku-4-5"), true); + assert.equal(supportsNativeTools("openai", "gpt-4o-mini"), true); + assert.equal(supportsNativeTools("gemini", "gemini-2.5-flash"), true); + assert.equal(supportsNativeTools("openai", "text-babbage-001"), false); + assert.equal(supportsNativeTools("gemini", "palm-2"), false); +}); + +test("pickTransport: auto by capability, explicit modes override", () => { + assert.equal(pickTransport("openai", "gpt-4o", "auto"), "native"); + assert.equal(pickTransport("openai", "mystery-model", "auto"), "scripted"); + assert.equal(pickTransport("openai", "gpt-4o", "scripted"), "scripted"); + assert.equal(pickTransport("gemini", "palm-2", "native"), "native"); +}); + +test("formatResultsForModel numbers videos with channel, date, and clocks", () => { + const videos: RetrievedVideo[] = [ + { + key: "a", + videoId: "a", + title: "Timeline of the trial", + channel: "Rekieta", + uploadDate: "20200101", + snippets: [{ clock: "0:05", seconds: 5, text: "opening" }], + }, + ]; + const out = formatResultsForModel(videos); + assert.match(out, /1\. "Timeline of the trial" — Rekieta \[20200101\]/); + assert.match(out, /0:05 opening/); + assert.equal( + formatResultsForModel([]), + "No transcript excerpts matched that query.", + ); +}); diff --git a/export/app/lib/searchAgent.ts b/export/app/lib/searchAgent.ts @@ -0,0 +1,280 @@ +// The /ask search agent: one shared loop, two gather transports. +// +// A turn is: GATHER (let the model search the transcripts — possibly several +// times, refining as it reads results) → ANSWER (stream a cited answer over the +// excerpts it gathered). The gather transport is either provider-agnostic +// "scripted" (a SEARCH:/DONE text protocol) or provider-native function-calling; +// everything else — running searches, accumulating/deduping videos, the answer +// phase, the event stream the UI renders — is shared. Auto-detect picks native +// where supported and falls back to scripted on a capability error. + +import { + askOnce, + askStream, + type ChatMessage, + type Provider, +} from "./askProvider"; +import { retrieve, type RetrievedVideo } from "./askRetrieval"; +import { + answerSystemPrompt, + buildApiMessages, + buildGroundedContent, + gatherSystemPrompt, + type UiMessage, +} from "./askConversation"; +import { anthropicGather } from "./nativeTools/anthropic"; +import { openaiGather } from "./nativeTools/openai"; +import { geminiGather } from "./nativeTools/gemini"; +import { + abortError, + ToolsUnavailableError, + type NativeGatherContext, +} from "./nativeTools/shared"; +import type { SearchAlias } from "yt-dlp-transcript-common/lib/searchAliases"; +import type { DisplaySummary } from "yt-dlp-transcript-common/lib/transcripts"; + +export type AgentMode = "auto" | "native" | "scripted"; + +export type AgentEvent = + | { type: "search_start"; query: string } + | { type: "search_done"; query: string; count: number } + | { type: "answer_start" } + | { type: "delta"; text: string }; + +// Default search budget per turn. Bounds round-trips (and spend on the user's +// key). Round-trips ≈ searches + 1 (answer), + gather decision turns. +export const DEFAULT_BUDGET = 4; +// Per-search result cap and overall context cap (bounds tokens sent to the model). +const PER_SEARCH_LIMIT = 8; +const MAX_CONTEXT_VIDEOS = 15; + +// ─── transport selection ─── + +// Provider+model pairs that failed a native tool call this session; Auto avoids +// re-trying native for them. +const forcedScripted = new Set<string>(); +const pairKey = (p: Provider, m: string) => `${p}:${m.toLowerCase()}`; + +export function supportsNativeTools(provider: Provider, model: string): boolean { + const m = model.toLowerCase(); + switch (provider) { + case "anthropic": + return m.startsWith("claude-"); + case "openai": + return /^(gpt-4o|gpt-4\.1|gpt-4-turbo|gpt-5|o1|o3|o4)/.test(m); + case "gemini": + return /gemini-(1\.5|2)/.test(m) || m.includes("2.0") || m.includes("2.5"); + default: + return false; + } +} + +export function pickTransport( + provider: Provider, + model: string, + mode: AgentMode, +): "native" | "scripted" { + if (mode === "scripted") return "scripted"; + if (mode === "native") return "native"; // explicit choice ignores the cache + if (forcedScripted.has(pairKey(provider, model))) return "scripted"; + return supportsNativeTools(provider, model) ? "native" : "scripted"; +} + +// ─── scripted transport ─── + +// Parse one scripted decision line. Lenient: accepts a SEARCH: line anywhere +// (stripping code fences/quotes); anything else — including "DONE" or prose — is +// treated as done gathering. +export function parseScriptedDecision( + reply: string, +): { kind: "search"; query: string } | { kind: "done" } { + const cleaned = reply.replace(/```[a-z]*/gi, "").replace(/```/g, "").trim(); + const m = cleaned.match(/SEARCH\s*:\s*(.+)/i); + if (m) { + const q = m[1] + .split(/\r?\n/)[0] + .trim() + .replace(/^["'`]+|["'`]+$/g, "") + .trim(); + if (q && !/^done$/i.test(q)) return { kind: "search", query: q }; + } + return { kind: "done" }; +} + +async function scriptedGather( + ctx: NativeGatherContext, + provider: Provider, +): Promise<void> { + const convo: ChatMessage[] = [ + ...ctx.history, + { role: "user", content: ctx.question }, + ]; + for (let i = 0; i < ctx.budget; i++) { + if (ctx.signal?.aborted) throw abortError(); + const reply = await askOnce({ + provider, + apiKey: ctx.apiKey, + model: ctx.model, + system: ctx.system, + messages: convo, + maxTokens: 120, + signal: ctx.signal, + }); + const decision = parseScriptedDecision(reply); + if (decision.kind === "done") return; + convo.push({ role: "assistant", content: `SEARCH: ${decision.query}` }); + const results = await ctx.runSearch(decision.query); + convo.push({ + role: "user", + content: + `Results for "${decision.query}":\n${results}\n\n` + + `Reply "SEARCH: <query>" to search again, or "DONE" to answer.`, + }); + } +} + +// ─── result formatting fed back to the model during gather ─── + +export function formatResultsForModel(videos: RetrievedVideo[]): string { + if (videos.length === 0) return "No transcript excerpts matched that query."; + return videos + .map((v, i) => { + const site = v.siteTitle ? ` (${v.siteTitle})` : ""; + const date = v.uploadDate ? ` [${v.uploadDate}]` : ""; + const head = `${i + 1}. "${v.title}" — ${v.channel}${site}${date}`; + const lines = v.snippets.map((s) => ` ${s.clock} ${s.text}`).join("\n"); + return lines ? `${head}\n${lines}` : head; + }) + .join("\n\n"); +} + +// ─── the turn ─── + +export type RunAskTurnOptions = { + provider: Provider; + apiKey: string; + model: string; + mode: AgentMode; + question: string; + prior: UiMessage[]; + summaries: DisplaySummary[]; + aliases: SearchAlias[]; + budget?: number; + signal?: AbortSignal; + onEvent: (e: AgentEvent) => void; +}; + +export type AskTurnResult = { + answer: string; + videos: RetrievedVideo[]; + groundedContent: string; + truncated: boolean; +}; + +export async function runAskTurn( + opts: RunAskTurnOptions, +): Promise<AskTurnResult> { + const { + provider, + apiKey, + model, + mode, + question, + prior, + summaries, + aliases, + signal, + onEvent, + } = opts; + const budget = opts.budget ?? DEFAULT_BUDGET; + const history = buildApiMessages(prior); + const hasPriorGrounding = prior.some( + (m) => m.role === "assistant" && (m.sources?.length ?? 0) > 0, + ); + + const videos = new Map<string, RetrievedVideo>(); + const queries: string[] = []; + let truncated = false; + + const runSearch = async (rawQuery: string): Promise<string> => { + const query = rawQuery.trim(); + if (!query) return "No query provided."; + onEvent({ type: "search_start", query }); + const r = await retrieve({ + question: query, + summaries, + aliases, + signal, + limit: PER_SEARCH_LIMIT, + }); + for (const v of r.videos) videos.set(v.key, v); + if (r.truncated) truncated = true; + queries.push(query); + onEvent({ type: "search_done", query, count: r.videos.length }); + return formatResultsForModel(r.videos); + }; + + const gatherCtx: NativeGatherContext = { + apiKey, + model, + system: "", // set per transport below + history, + question, + budget, + runSearch, + signal, + }; + + const doGather = async (kind: "native" | "scripted"): Promise<void> => { + const ctx: NativeGatherContext = { + ...gatherCtx, + system: gatherSystemPrompt(aliases, kind, budget), + }; + if (kind === "scripted") return scriptedGather(ctx, provider); + switch (provider) { + case "anthropic": + return anthropicGather(ctx); + case "openai": + return openaiGather(ctx); + case "gemini": + return geminiGather(ctx); + } + }; + + const kind = pickTransport(provider, model, mode); + try { + await doGather(kind); + } catch (e) { + if ((e as Error).name === "AbortError") throw e; + if (kind === "native" && e instanceof ToolsUnavailableError) { + // This model/endpoint doesn't support tools — remember and retry scripted. + forcedScripted.add(pairKey(provider, model)); + await doGather("scripted"); + } else { + throw e; + } + } + + // Safety net: never answer a fresh question with zero grounding just because a + // model ignored the protocol / declined to search. + if (queries.length === 0 && !hasPriorGrounding) { + await runSearch(question); + } + + const finalVideos = [...videos.values()].slice(0, MAX_CONTEXT_VIDEOS); + const groundedContent = buildGroundedContent(question, finalVideos); + + onEvent({ type: "answer_start" }); + const answer = await askStream({ + provider, + apiKey, + model, + system: answerSystemPrompt(aliases), + messages: [...history, { role: "user", content: groundedContent }], + maxTokens: 2048, + signal, + onDelta: (chunk) => onEvent({ type: "delta", text: chunk }), + }); + + return { answer, videos: finalVideos, groundedContent, truncated }; +} diff --git a/export/e2e/ask-chat.spec.ts b/export/e2e/ask-chat.spec.ts @@ -0,0 +1,160 @@ +import { expect, test, type Page } from "@playwright/test"; +import { installRoutes } from "./helpers"; + +// The /ask agentic chat. Retrieval runs client-side over the mocked transcript +// fixtures (installRoutes); the AI provider is mocked here. We drive the +// provider-agnostic "scripted" mode so a single Anthropic SSE mock covers every +// call (gather decisions via askOnce + the streamed answer): the mock reads the +// request to decide what to return — SEARCH on a fresh question, DONE once +// results are in or when the message is a reformat follow-up, and a Markdown +// answer for the answer phase. + +const CORS = { + "access-control-allow-origin": "*", + "access-control-allow-headers": "*", + "access-control-allow-methods": "*", +}; + +function sse(text: string): string { + return ( + `data: ${JSON.stringify({ + type: "content_block_delta", + delta: { type: "text_delta", text }, + })}\n\n` + `data: ${JSON.stringify({ type: "message_stop" })}\n\n` + ); +} + +function toolUse(name: string, input: Record<string, unknown>) { + return JSON.stringify({ + content: [{ type: "tool_use", id: "tu_1", name, input }], + stop_reason: "tool_use", + }); +} + +async function mockAnthropic(page: Page) { + await page.route("https://api.anthropic.com/**", async (route) => { + if (route.request().method() === "OPTIONS") { + await route.fulfill({ status: 204, headers: CORS }); + return; + } + const body = route.request().postDataJSON() as { + system?: string; + tools?: unknown[]; + messages?: { role: string; content: unknown }[]; + }; + const system = body.system ?? ""; + const msgs = body.messages ?? []; + const lastUser = [...msgs].reverse().find((m) => m.role === "user"); + const lastText = + typeof lastUser?.content === "string" ? lastUser.content : ""; + + // Native tool-calling gather turn → respond with a tool_use JSON block. + if (Array.isArray(body.tools) && !system.includes("Markdown")) { + const searchedAlready = Array.isArray(lastUser?.content); // tool_result turn + if (searchedAlready || /timeline|format/i.test(lastText)) { + await route.fulfill({ + status: 200, + headers: { ...CORS, "content-type": "application/json" }, + body: toolUse("finish", {}), + }); + } else { + await route.fulfill({ + status: 200, + headers: { ...CORS, "content-type": "application/json" }, + body: toolUse("search_transcripts", { query: "alpha" }), + }); + } + return; + } + + // Scripted gather + answer phase → stream SSE text. + let text: string; + if (system.includes("Markdown")) { + text = "Here is the answer with a list:\n\n- point one [1]\n- point two"; + } else if (lastText.includes('Results for "')) { + text = "DONE"; // we've searched once — stop + } else if (/timeline|format/i.test(lastText)) { + text = "DONE"; // reformat follow-up — no search needed + } else { + text = "SEARCH: alpha"; // fresh question — search a term in the fixtures + } + + await route.fulfill({ + status: 200, + headers: { ...CORS, "content-type": "text/event-stream" }, + body: sse(text), + }); + }); +} + +async function setup(page: Page, mode: "Scripted" | "Native tools" = "Scripted") { + await installRoutes(page); + await mockAnthropic(page); + await page.goto("/ask/"); + // Pick the search mode while the settings panel is still open — filling the + // key collapses it (open={!apiKey}). + await page.getByRole("button", { name: mode }).click(); + await page.locator('input[placeholder^="sk-ant"]').fill("sk-ant-test"); + // Guard against the one-time restore effect clearing the freshly-typed key. + await expect(page.locator('input[placeholder^="sk-ant"]')).toHaveValue( + "sk-ant-test", + ); +} + +async function ask(page: Page, text: string) { + const box = page.getByPlaceholder(/Ask about the transcripts/); + await box.fill(text); + await page.getByRole("button", { name: "Ask", exact: true }).click(); +} + +test.describe("ask chat", () => { + test("searches, answers with citations, and renders Markdown", async ({ + page, + }) => { + await setup(page); + await ask(page, "tell me about the alpha discussion"); + + // The turn ran a search (persistent trace) for the query the model chose. + await expect(page.getByText(/Searched:\s*alpha/)).toBeVisible(); + // Citations reference the fixture videos. + await expect(page.getByText("Transcript only", { exact: false }).first()).toBeVisible(); + // The answer rendered as Markdown (a real list item). + await expect(page.locator("li", { hasText: "point one" }).first()).toBeVisible(); + }); + + test("a reformat follow-up reuses context without a junk search", async ({ + page, + }) => { + await setup(page); + await ask(page, "tell me about the alpha discussion"); + // The "Searched:" trace only appears once the turn completes. + await expect(page.getByText(/Searched:\s*alpha/)).toBeVisible(); + await expect(page.locator("li", { hasText: "point one" })).toHaveCount(1); + + await ask(page, "format it in a timeline"); + // A second answer rendered (turn 2 completed). + await expect(page.locator("li", { hasText: "point one" })).toHaveCount(2); + // …but the follow-up ran NO new search — only turn 1's trace exists. + await expect(page.getByText(/^Searched:/)).toHaveCount(1); + }); + + test("native function-calling drives the same search loop", async ({ + page, + }) => { + await setup(page, "Native tools"); + await ask(page, "tell me about the alpha discussion"); + // The native tool_use loop searched and then answered. + await expect(page.getByText(/Searched:\s*alpha/)).toBeVisible(); + await expect(page.locator("li", { hasText: "point one" }).first()).toBeVisible(); + }); + + test("the Format answers toggle switches to plain text", async ({ page }) => { + await setup(page); + await ask(page, "tell me about the alpha discussion"); + await expect(page.locator("li", { hasText: "point one" }).first()).toBeVisible(); + + await page.getByLabel("Format answers").uncheck(); + // Plain mode shows the raw Markdown source, dashes and all. + await expect(page.getByText("- point one [1]", { exact: false })).toBeVisible(); + }); +});