Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 03a8f821f01def5da39e7b924b99336095b07b04
parent b35da36055ad5b3192a2a587df47c79455ca8ba9
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Tue,  7 Jul 2026 16:03:51 -0400

Ask chat: alias-aware search + saved/resumable/editable context

Two follow-on improvements to the /ask chat.

Alias-aware search (the "searched Graham instead of the graham-platner alias
regex" bug): retrieval matched aliases one keyword at a time, so a multi-word
phrase trigger ("graham platner") never fired. buildSearchRoot now phrase-matches
aliases against the whole query string and adds the alias's regex leaf (dropping
the covered plain keywords), and both the gather and answer prompts now carry an
explicit per-alias directive — search the FULL aliased term, not a fragment, plus
the alias note — so the model forms alias-aware queries. Built on main's
phrase-aware matchAliases.

Saved, resumable, editable context:
- Persistence: the conversation (messages + each turn's grounded excerpts +
  search steps) and any context override are saved to localStorage, debounced,
  and restored on load; a New chat control clears them.
- Retry: runAskTurn now surfaces grounding at answer_start (before the stream) so
  it survives a failed answer, and takes a precomputedGrounding option to skip
  gather. A Retry button on a failed turn re-runs it reusing the excerpts already
  found — no re-search.
- Editable context: a Context panel serializes the effective context
  (USER:/ASSISTANT: blocks, excerpts inline) into an uncontrolled textarea the
  human can prune; Apply parses it back and continues the conversation from the
  trimmed seed (also the "new session from context" path). runAskTurn gains a
  historyOverride; send/runTurn read messages + override from refs to avoid a
  stale-closure send right after applying an edit.

Tests: +4 unit (query-level alias fire on full phrase / not on a fragment;
directive glossary; context serialize/parse round-trip + fallback) = 29 total;
+3 e2e (Retry-after-503 reuses excerpts, reload persists the conversation, editing
the context changes the next request) = 7 total. tsc + lint clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

Diffstat:
Mexport/CHANGELOG.md | 4++++
Mexport/app/ask/AskChat.tsx | 38++++++++++++++++++++++++++++++++++++--
Aexport/app/ask/ContextPanel.tsx | 107+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mexport/app/ask/MessageBubble.tsx | 21++++++++++++++++++++-
Mexport/app/ask/useAskChat.ts | 367++++++++++++++++++++++++++++++++++++++++++++++++++++++-------------------------
Mexport/app/lib/askConversation.test.ts | 34++++++++++++++++++++++++++++++++--
Mexport/app/lib/askConversation.ts | 77++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++---------
Mexport/app/lib/askRetrieval.test.ts | 67++++++++++++++++++++++++++++++++++++++++++++++++-------------------
Mexport/app/lib/askRetrieval.ts | 111+++++++++++++++++++++++++++++++++++++++++++++++++------------------------------
Mexport/app/lib/searchAgent.ts | 35++++++++++++++++++++++++++++++++---
Mexport/e2e/ask-chat.spec.ts | 132+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
11 files changed, 801 insertions(+), 192 deletions(-)

diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md @@ -1,5 +1,9 @@ # Changelog +## [Unreleased] +- **The "Ask AI" chat now uses your curated name aliases when it searches.** Asking about a term with a known alias (e.g. *Graham Platner*, which AI transcription often mangles to "Grand Platina") now applies the alias's regex, so misspelled mentions are found instead of missed. Previously the chat matched aliases one word at a time and a multi-word alias never fired; it now matches the whole search phrase, and each prompt explicitly tells the model to search the full aliased term (not a fragment like "Graham") and why. See `export/app/lib/askRetrieval.ts` (`buildSearchRoot`) and `export/app/lib/askConversation.ts` (`renderAliasGlossary`). +- **The chat conversation is now saved, resumable, and editable.** Your conversation and the excerpts it gathered are kept in your browser, so a reload no longer loses them, and a *New chat* button clears them. If a request fails partway (e.g. a provider 503), the turn shows a **Retry** button that re-runs it **reusing the excerpts already found** — no re-searching. And a new **Context** panel lets you view and prune exactly what gets sent to the model (conversation + excerpts) as editable text, for a leaner, cheaper starting point — mid-conversation or as the seed for a fresh session. See `export/app/ask/{useAskChat.ts,ContextPanel.tsx,MessageBubble.tsx}`, `export/app/lib/{searchAgent,askConversation}.ts`, and `export/e2e/ask-chat.spec.ts`. + ## [0.7.4] - 2026-07-07 - **The export build can now fan out across sites in parallel (Docker mode).** Archive-zip generation is hoisted into a new shared host step (`build:archives`) that warms the persistent cache once for the union of all sites' channels, so the per-site `compose:site` can run read-only over that cache (`ARCHIVES_READONLY=1`) without re-zipping or racing — which is what lets several sites build at once. No change to a site's output; this is a build-pipeline change. See `common/bin/build-archives.ts`, `common/controller/{archiveTranscripts,archiveLiveChat}.ts` (`readOnly`), and `common/bin/compose-site.ts`. (Orchestration + UI live in the editor changelog.) - **The "Ask AI" chat now searches like an assistant — and follow-ups work.** Instead of one keyword search per message, the chat now runs an *agentic* loop: it decides what to look up, reads the results, and can refine and search again before answering — you see each search as it happens. Crucially, a follow-up that builds on the last answer (e.g. "put that on a timeline", "summarise it") no longer throws away the earlier context and re-searches your wording — earlier excerpts stay in play, so it just reformats what it already found. Answers render as **Markdown** (with a *Format answers* toggle to fall back to plain text if anything looks off), a live status shows the search→answer pipeline, and the model is told the archive's known transcription-misspelling aliases so it searches for the right variants. A **Search mode** control (Auto / Native tools / Scripted) chooses between your model's native function-calling and a provider-agnostic protocol, with automatic fallback. See `export/app/ask/*`, `export/app/lib/{searchAgent,askConversation,nativeTools/*}.ts`, `common/components/Markdown.tsx`, and `export/e2e/ask-chat.spec.ts`. diff --git a/export/app/ask/AskChat.tsx b/export/app/ask/AskChat.tsx @@ -1,9 +1,10 @@ "use client"; import { useEffect, useMemo, useRef, useState } from "react"; -import { ArrowDownIcon } from "lucide-react"; +import { ArrowDownIcon, PlusIcon } from "lucide-react"; import { useAskChat } from "./useAskChat"; import { ProviderSettings } from "./ProviderSettings"; +import { ContextPanel } from "./ContextPanel"; import { MessageBubble } from "./MessageBubble"; import { Composer } from "./Composer"; @@ -82,6 +83,34 @@ export default function AskChat() { </p> )} + {(messages.length > 0 || s.contextOverride) && ( + <div className="flex items-center justify-between"> + <span className="text-xs text-muted-foreground/70"> + {s.contextOverride + ? "Continuing from an edited context." + : "Conversation is saved in this browser."} + </span> + <button + type="button" + onClick={s.reset} + disabled={busy} + className="inline-flex items-center gap-1 rounded-md border border-border px-2 py-1 text-xs text-muted-foreground transition-colors hover:text-foreground disabled:opacity-50" + > + <PlusIcon className="size-3.5" /> New chat + </button> + </div> + )} + + {(messages.length > 0 || s.contextOverride) && ( + <ContextPanel + contextText={s.contextText} + overrideActive={!!s.contextOverride} + busy={busy} + onApply={s.applyContext} + onClear={s.clearContextOverride} + /> + )} + <div className="relative"> <div ref={scrollRef} @@ -113,7 +142,12 @@ export default function AskChat() { </div> ) : ( messages.map((m, i) => ( - <MessageBubble key={i} message={m} markdownOn={markdownOn} /> + <MessageBubble + key={i} + message={m} + markdownOn={markdownOn} + onRetry={m.role === "assistant" ? () => s.retry(i) : undefined} + /> )) )} </div> diff --git a/export/app/ask/ContextPanel.tsx b/export/app/ask/ContextPanel.tsx @@ -0,0 +1,107 @@ +"use client"; + +import { useRef, useState } from "react"; + +// A view/edit surface for the context that will be sent to the model — the +// replayed conversation + retrieved excerpts, serialized as USER:/ASSISTANT: +// blocks. The human can prune it for a leaner, cheaper starting point (mid- +// conversation) and "apply" it: the edit becomes the seed the conversation +// continues from (the display resets to it). This is also the "forward to a new +// session" path after an error. +// +// The textarea is uncontrolled (seeded via defaultValue + a remount key that +// changes only when the underlying context changes) so a human's edits are never +// clobbered by a re-render. +export function ContextPanel({ + contextText, + overrideActive, + busy, + onApply, + onClear, +}: { + contextText: string; + overrideActive: boolean; + busy: boolean; + onApply: (text: string) => void; + onClear: () => void; +}) { + const [open, setOpen] = useState(false); + const taRef = useRef<HTMLTextAreaElement>(null); + // Bump to force the uncontrolled textarea to remount and reseed from the + // current context (on open, and on an explicit "Reload"). + const [seed, setSeed] = useState(0); + + const onToggle = (e: React.SyntheticEvent<HTMLDetailsElement>) => { + const isOpen = e.currentTarget.open; + setOpen(isOpen); + if (isOpen) setSeed((n) => n + 1); + }; + + const empty = contextText.trim() === ""; + + return ( + <details + className="rounded-lg border border-border bg-card/40" + onToggle={onToggle} + open={open} + > + <summary className="flex cursor-pointer items-center gap-2 px-4 py-2.5 text-sm font-medium text-foreground"> + Context + {overrideActive && ( + <span className="rounded bg-brand-soft px-1.5 py-0.5 font-mono text-[0.7rem] text-brand"> + edited + </span> + )} + <span className="ml-auto font-mono text-xs text-muted-foreground/70"> + {contextText.length.toLocaleString()} chars + </span> + </summary> + {open && ( + <div className="flex flex-col gap-3 border-t border-border px-4 py-3"> + <p className="text-xs text-muted-foreground"> + This is what the model receives as prior context on your next question. + Prune it for a leaner, cheaper starting point, then apply — the + conversation continues from the trimmed context. + </p> + <textarea + key={`${seed}:${contextText.length}`} + ref={taRef} + defaultValue={contextText} + rows={10} + spellCheck={false} + placeholder={empty ? "(no context yet — ask a question first)" : ""} + className="w-full resize-y rounded-md border border-border bg-background px-3 py-2 font-mono text-xs leading-relaxed text-foreground" + /> + <div className="flex flex-wrap items-center gap-2"> + <button + type="button" + disabled={busy} + onClick={() => onApply(taRef.current?.value ?? "")} + className="rounded-md bg-primary px-3 py-1.5 text-xs font-medium text-primary-foreground transition-colors hover:bg-brand-strong disabled:opacity-50" + > + Apply as starting point + </button> + <button + type="button" + disabled={busy} + onClick={() => setSeed((n) => n + 1)} + className="rounded-md border border-border px-2.5 py-1.5 text-xs text-muted-foreground transition-colors hover:text-foreground disabled:opacity-50" + > + Reload from conversation + </button> + {overrideActive && ( + <button + type="button" + disabled={busy} + onClick={onClear} + className="rounded-md border border-border px-2.5 py-1.5 text-xs text-muted-foreground transition-colors hover:text-foreground disabled:opacity-50" + > + Clear edit + </button> + )} + </div> + </div> + )} + </details> + ); +} diff --git a/export/app/ask/MessageBubble.tsx b/export/app/ask/MessageBubble.tsx @@ -1,7 +1,7 @@ "use client"; import { useState } from "react"; -import { CheckIcon, CopyIcon } from "lucide-react"; +import { CheckIcon, CopyIcon, RotateCwIcon } from "lucide-react"; import { Markdown } from "yt-dlp-transcript-common/components/Markdown"; import type { UiMessage } from "../lib/askConversation"; import { PipelineStatus } from "./PipelineStatus"; @@ -40,9 +40,11 @@ function CopyButton({ text }: { text: string }) { export function MessageBubble({ message, markdownOn, + onRetry, }: { message: UiMessage; markdownOn: boolean; + onRetry?: () => void; }) { const isUser = message.role === "user"; const streaming = message.phase === "answering" || message.phase === "streaming"; @@ -86,6 +88,23 @@ export function MessageBubble({ </div> )} + {!isUser && message.error && onRetry && ( + <div className="flex items-center gap-2 text-xs"> + <button + type="button" + onClick={onRetry} + className="inline-flex items-center gap-1 rounded-md border border-border px-2 py-1 text-muted-foreground transition-colors hover:text-foreground" + > + <RotateCwIcon className="size-3.5" /> Retry + </button> + {(message.sources?.length ?? 0) > 0 && ( + <span className="text-muted-foreground/70"> + reuses the excerpts already found + </span> + )} + </div> + )} + {!isUser && message.sources && message.sources.length > 0 && ( <ol className="ml-1 flex flex-col gap-1 text-xs text-muted-foreground"> {message.sources.map((s, si) => ( diff --git a/export/app/ask/useAskChat.ts b/export/app/ask/useAskChat.ts @@ -1,22 +1,30 @@ "use client"; -import { useCallback, useEffect, useRef, useState } from "react"; +import { useCallback, useEffect, useMemo, useRef, useState } from "react"; import { useSearchData } from "yt-dlp-transcript-common/components/SearchDataContext"; import { PROVIDERS, type Provider } from "../lib/askProvider"; import { runAskTurn, type AgentEvent, type AgentMode } from "../lib/searchAgent"; -import type { SearchStep, UiMessage } from "../lib/askConversation"; +import { + buildApiMessages, + parseContext, + serializeContext, + type SearchStep, + type UiMessage, +} from "../lib/askConversation"; const K_PROVIDER = "ytdlp-tb:ai:provider"; const K_REMEMBER = "ytdlp-tb:ai:remember"; const K_SEARCHMODE = "ytdlp-tb:ai:searchmode"; const K_MARKDOWN = "ytdlp-tb:ai:md"; +const K_CONVO = "ytdlp-tb:ai:conversation"; const keyFor = (p: Provider) => `ytdlp-tb:ai:key:${p}`; const modelFor = (p: Provider) => `ytdlp-tb:ai:model:${p}`; -// State + turn orchestration for the /ask chat. The view is split into small -// presentational components; this hook owns everything mutable and calls -// runAskTurn (gather → answer), mapping its event stream onto the pending -// assistant message. +type PersistedConvo = { messages: UiMessage[]; contextOverride: string | null }; + +// State + turn orchestration for the /ask chat. Owns the conversation (persisted +// across reloads), the retrieval agent turns (gather → answer), a Retry path that +// reuses already-gathered excerpts, and a human-editable context override. export function useAskChat() { const [provider, setProviderState] = useState<Provider>("anthropic"); const [apiKey, setApiKey] = useState(""); @@ -31,11 +39,22 @@ export function useAskChat() { const corpusError = summariesState.error?.message ?? null; const [messages, setMessages] = useState<UiMessage[]>([]); + // A human-pruned context that seeds the conversation: prepended to the replayed + // history on every turn. Set by applying an edit in the context panel. + const [contextOverride, setContextOverride] = useState<string | null>(null); const [input, setInput] = useState(""); const [busy, setBusy] = useState(false); const abortRef = useRef<AbortController | null>(null); - // Restore saved preferences + (if remembered) the provider's key/model. + // Always-current mirrors so send()/runTurn() never act on a stale snapshot + // (e.g. sending immediately after applying an edited context). + const messagesRef = useRef(messages); + messagesRef.current = messages; + const contextOverrideRef = useRef(contextOverride); + contextOverrideRef.current = contextOverride; + + // Restore saved preferences + (if remembered) the provider's key/model, plus a + // persisted conversation. useEffect(() => { /* eslint-disable react-hooks/set-state-in-effect -- one-time restore from localStorage after hydration; it must run in an effect, not a lazy @@ -53,12 +72,42 @@ export function useAskChat() { } setMarkdownOnState(localStorage.getItem(K_MARKDOWN) !== "0"); loadProviderCreds(p, rememberSaved); + const rawConvo = localStorage.getItem(K_CONVO); + if (rawConvo) { + const parsed = JSON.parse(rawConvo) as PersistedConvo; + if (Array.isArray(parsed.messages)) setMessages(parsed.messages); + if (typeof parsed.contextOverride === "string") { + setContextOverride(parsed.contextOverride); + } + } } catch { - /* storage unavailable */ + /* storage unavailable / corrupt */ } /* eslint-enable react-hooks/set-state-in-effect */ }, []); + // Persist the conversation (messages + override), debounced so streaming deltas + // don't thrash storage. Survives reloads and hard errors. + const persistTimer = useRef<ReturnType<typeof setTimeout> | null>(null); + useEffect(() => { + if (persistTimer.current) clearTimeout(persistTimer.current); + persistTimer.current = setTimeout(() => { + try { + if (messages.length === 0 && !contextOverride) { + localStorage.removeItem(K_CONVO); + } else { + const payload: PersistedConvo = { messages, contextOverride }; + localStorage.setItem(K_CONVO, JSON.stringify(payload)); + } + } catch { + /* ignore */ + } + }, 400); + return () => { + if (persistTimer.current) clearTimeout(persistTimer.current); + }; + }, [messages, contextOverride]); + function loadProviderCreds(p: Provider, rememberOn: boolean) { if (rememberOn) { setApiKey(localStorage.getItem(keyFor(p)) ?? ""); @@ -112,131 +161,215 @@ export function useAskChat() { } } + // Patch one message by index, immutably. + const patchAt = (index: number, fn: (m: UiMessage) => UiMessage) => + setMessages((prev) => { + if (index < 0 || index >= prev.length) return prev; + const next = prev.slice(); + next[index] = fn(next[index]); + return next; + }); + + // Core turn runner shared by send() and retry(). Patches the assistant message + // at `assistantIndex` (and its user turn at `userIndex`) as the agent streams. + const runTurn = useCallback( + async (args: { + question: string; + prior: UiMessage[]; + userIndex: number; + assistantIndex: number; + precomputedGrounding?: { + videos: UiMessage["sources"]; + groundedContent: string; + }; + }) => { + const { question, prior, userIndex, assistantIndex } = args; + setBusy(true); + const ac = new AbortController(); + abortRef.current = ac; + + const override = contextOverrideRef.current; + const historyOverride = override + ? [...parseContext(override), ...buildApiMessages(prior)] + : undefined; + + const onEvent = (e: AgentEvent) => { + switch (e.type) { + case "search_start": + patchAt(assistantIndex, (m) => ({ + ...m, + phase: "gathering", + searchSteps: [...(m.searchSteps ?? []), { query: e.query }], + })); + break; + case "search_done": + patchAt(assistantIndex, (m) => { + const steps = (m.searchSteps ?? []).slice(); + for (let i = steps.length - 1; i >= 0; i--) { + if (steps[i].query === e.query && steps[i].count === undefined) { + steps[i] = { ...steps[i], count: e.count }; + break; + } + } + return { ...m, searchSteps: steps }; + }); + break; + case "answer_start": + // Store grounding NOW (before streaming) so a failed stream can be + // retried without re-searching, and the excerpts persist. + patchAt(assistantIndex, (m) => ({ + ...m, + phase: "answering", + sources: e.videos, + groundedContent: e.groundedContent, + })); + patchAt(userIndex, (m) => ({ + ...m, + groundedContent: e.groundedContent, + })); + break; + case "delta": + patchAt(assistantIndex, (m) => ({ + ...m, + phase: "streaming", + content: m.content + e.text, + })); + break; + } + }; + + try { + const result = await runAskTurn({ + provider, + apiKey: apiKey.trim(), + model: model.trim() || PROVIDERS[provider].defaultModel, + mode: searchMode, + question, + prior, + summaries, + aliases, + signal: ac.signal, + onEvent, + historyOverride, + precomputedGrounding: args.precomputedGrounding?.groundedContent + ? { + videos: args.precomputedGrounding.videos ?? [], + groundedContent: args.precomputedGrounding.groundedContent, + } + : undefined, + }); + patchAt(userIndex, (m) => ({ ...m, groundedContent: result.groundedContent })); + patchAt(assistantIndex, (m) => ({ + ...m, + content: m.content || result.answer, + sources: result.videos, + truncated: result.truncated, + error: false, + phase: "done", + })); + } catch (e) { + const err = e as Error; + if (err.name === "AbortError") { + patchAt(assistantIndex, (m) => ({ + ...m, + content: m.content + "\n\n_(stopped)_", + phase: "stopped", + })); + } else { + // Preserve any grounding already captured at answer_start so Retry can + // skip the search. + patchAt(assistantIndex, (m) => ({ + ...m, + content: m.content || err.message, + error: true, + phase: "error", + })); + } + } finally { + setBusy(false); + abortRef.current = null; + } + }, + [provider, apiKey, model, searchMode, summaries, aliases], + ); + const send = useCallback(async () => { const question = input.trim(); if (!question || busy || !apiKey.trim() || !summariesReady) return; - persistKey(provider, apiKey, model, remember); - const prior = messages; + const prior = messagesRef.current; + const userIndex = prior.length; + const assistantIndex = prior.length + 1; setInput(""); setMessages((prev) => [ ...prev, { role: "user", content: question }, { role: "assistant", content: "", phase: "gathering", searchSteps: [] }, ]); - setBusy(true); - const ac = new AbortController(); - abortRef.current = ac; - - // Patch the trailing assistant / the preceding user message immutably. - const patchAssistant = (fn: (m: UiMessage) => UiMessage) => - setMessages((prev) => { - const next = prev.slice(); - next[next.length - 1] = fn(next[next.length - 1]); - return next; - }); - const patchUser = (fn: (m: UiMessage) => UiMessage) => - setMessages((prev) => { - const next = prev.slice(); - next[next.length - 2] = fn(next[next.length - 2]); - return next; - }); + await runTurn({ question, prior, userIndex, assistantIndex }); + }, [input, busy, apiKey, summariesReady, provider, model, remember, runTurn]); - const onEvent = (e: AgentEvent) => { - switch (e.type) { - case "search_start": - patchAssistant((m) => ({ - ...m, - phase: "gathering", - searchSteps: [...(m.searchSteps ?? []), { query: e.query }], - })); - break; - case "search_done": - patchAssistant((m) => { - const steps = (m.searchSteps ?? []).slice(); - // Fill the most recent step for this query still awaiting a count. - for (let i = steps.length - 1; i >= 0; i--) { - if (steps[i].query === e.query && steps[i].count === undefined) { - steps[i] = { ...steps[i], count: e.count }; - break; - } - } - return { ...m, searchSteps: steps }; - }); - break; - case "answer_start": - patchAssistant((m) => ({ ...m, phase: "answering" })); - break; - case "delta": - patchAssistant((m) => ({ - ...m, - phase: "streaming", - content: m.content + e.text, - })); - break; - } - }; - - try { - const result = await runAskTurn({ - provider, - apiKey: apiKey.trim(), - model: model.trim() || PROVIDERS[provider].defaultModel, - mode: searchMode, - question, - prior, - summaries, - aliases, - signal: ac.signal, - onEvent, - }); - patchUser((m) => ({ ...m, groundedContent: result.groundedContent })); - patchAssistant((m) => ({ + // Re-run a failed assistant turn. Reuses the grounding captured before the + // answer stream failed (no re-search); falls back to a full turn otherwise. + const retry = useCallback( + async (assistantIndex: number) => { + if (busy || !apiKey.trim() || !summariesReady) return; + const current = messagesRef.current; + const userIndex = assistantIndex - 1; + const userMsg = current[userIndex]; + const failed = current[assistantIndex]; + if (!userMsg || userMsg.role !== "user" || !failed) return; + const prior = current.slice(0, userIndex); + const precomputedGrounding = failed.groundedContent + ? { videos: failed.sources, groundedContent: failed.groundedContent } + : undefined; + // Reset the failed message to a pending state (keep captured grounding). + patchAt(assistantIndex, (m) => ({ ...m, - content: m.content || result.answer, - sources: result.videos, - truncated: result.truncated, - phase: "done", + content: "", + error: false, + phase: "gathering", + searchSteps: precomputedGrounding ? m.searchSteps : [], })); - } catch (e) { - const err = e as Error; - if (err.name === "AbortError") { - patchAssistant((m) => ({ - ...m, - content: m.content + "\n\n_(stopped)_", - phase: "stopped", - })); - } else { - patchAssistant((m) => ({ - ...m, - content: m.content || err.message, - error: true, - phase: "error", - })); - } - } finally { - setBusy(false); - abortRef.current = null; - } - }, [ - input, - busy, - apiKey, - summariesReady, - summaries, - aliases, - provider, - model, - remember, - searchMode, - messages, - ]); + await runTurn({ + question: userMsg.content, + prior, + userIndex, + assistantIndex, + precomputedGrounding, + }); + }, + [busy, apiKey, summariesReady, runTurn], + ); const stop = () => abortRef.current?.abort(); + const reset = () => { if (busy) return; setMessages([]); + setContextOverride(null); + }; + + // The effective context that will be sent next turn, serialized for the panel: + // the override (if any) prepended to the replayed conversation. + const contextText = useMemo(() => { + const base = contextOverride ? parseContext(contextOverride) : []; + return serializeContext([...base, ...buildApiMessages(messages)]); + }, [contextOverride, messages]); + + // Apply an edited context as the new starting point: it becomes the override + // (prepended to future turns) and the visible conversation resets to it. + const applyContext = (text: string) => { + if (busy) return; + const trimmed = text.trim(); + setContextOverride(trimmed ? trimmed : null); + setMessages([]); + }; + + const clearContextOverride = () => { + if (busy) return; + setContextOverride(null); }; return { @@ -268,6 +401,12 @@ export function useAskChat() { send, stop, reset, + retry, + // context editing + contextText, + contextOverride, + applyContext, + clearContextOverride, }; } diff --git a/export/app/lib/askConversation.test.ts b/export/app/lib/askConversation.test.ts @@ -7,6 +7,8 @@ import { renderAliasGlossary, gatherSystemPrompt, answerSystemPrompt, + serializeContext, + parseContext, type UiMessage, } from "./askConversation"; import type { RetrievedVideo } from "./askRetrieval"; @@ -66,9 +68,14 @@ test("buildGroundedContent appends numbered excerpts, or just the question", () assert.match(grounded, /\[1\] "Platner interview" — Rekieta/); }); -test("renderAliasGlossary lists variants and notes; empty when none", () => { +test("renderAliasGlossary directs searching the full term + carries the note", () => { const g = renderAliasGlossary(ALIASES); - assert.match(g, /Graham Platner: also transcribed as platner, plattner, platter/); + // Directive to search the full term, not a fragment. + assert.match(g, /Graham Platner: when your search concerns this, search the full term "platner"/); + assert.match(g, /NOT a fragment/); + // Extra triggers surface as variant spellings. + assert.match(g, /also written: plattner, platter/); + // The alias note is included. assert.match(g, /often mis-transcribed/); assert.equal(renderAliasGlossary([]), ""); assert.equal( @@ -87,6 +94,29 @@ test("gatherSystemPrompt tailors the tail per mode and includes the glossary", ( assert.match(native, /finish tool/); }); +test("serializeContext + parseContext round-trip through role markers", () => { + const msgs = [ + { role: "user" as const, content: "tell me about Platner\n\n[1] excerpt" }, + { role: "assistant" as const, content: "He is ..." }, + ]; + const text = serializeContext(msgs); + assert.match(text, /^USER:\n/); + assert.match(text, /\nASSISTANT:\n/); + assert.deepEqual(parseContext(text), msgs); +}); + +test("parseContext survives pruning and falls back to a single user block", () => { + // A human deletes the assistant turn — still parses to the remaining user turn. + assert.deepEqual(parseContext("USER:\nonly the question"), [ + { role: "user", content: "only the question" }, + ]); + // No markers at all → treat the whole blob as one user message. + assert.deepEqual(parseContext("just some pasted notes"), [ + { role: "user", content: "just some pasted notes" }, + ]); + assert.deepEqual(parseContext(" "), []); +}); + test("answerSystemPrompt asks for Markdown + citations and carries the glossary", () => { const p = answerSystemPrompt(ALIASES); assert.match(p, /Markdown/); diff --git a/export/app/lib/askConversation.ts b/export/app/lib/askConversation.ts @@ -55,6 +55,51 @@ export function buildApiMessages(prior: UiMessage[]): ChatMessage[] { })); } +// ─── editable context (view / prune / forward to a new session) ─── + +const ROLE_MARKER = /^(USER|ASSISTANT):\s*$/i; + +// Serialize a message list into an editable transcript: a `USER:` / `ASSISTANT:` +// header line per turn followed by its content (excerpts inline). This is what +// the human sees and prunes in the context panel. +export function serializeContext(messages: ChatMessage[]): string { + return messages + .map((m) => `${m.role.toUpperCase()}:\n${m.content}`) + .join("\n\n"); +} + +// Parse an edited context transcript back into messages, splitting on the +// `USER:` / `ASSISTANT:` header lines. Text before any marker (or with no markers +// at all) becomes a single user message, so a hand-pasted blob still works. +export function parseContext(text: string): ChatMessage[] { + const lines = text.split(/\r?\n/); + const out: ChatMessage[] = []; + let role: "user" | "assistant" = "user"; + let buf: string[] = []; + let sawMarker = false; + const flush = () => { + const content = buf.join("\n").trim(); + if (content) out.push({ role, content }); + buf = []; + }; + for (const line of lines) { + const m = line.match(ROLE_MARKER); + if (m) { + flush(); + role = m[1].toLowerCase() === "assistant" ? "assistant" : "user"; + sawMarker = true; + } else { + buf.push(line); + } + } + flush(); + if (!sawMarker) { + const content = text.trim(); + return content ? [{ role: "user", content }] : []; + } + return out; +} + // Assemble the grounded user content for a turn: the question plus this turn's // retrieved excerpts (numbered for citation). No excerpts → just the question, // so a no-search follow-up leans on excerpts already in earlier turns. @@ -69,24 +114,38 @@ export function buildGroundedContent( ); } -// A compact glossary of known terms and their common (mis)transcriptions, so the -// model forms misspelling-aware searches and reconciles variant spellings it -// sees in excerpts. Empty when there are no usable aliases (block is omitted). +// A directive glossary of known terms that AI transcription mangles. It tells +// the model, per alias, to SEARCH THE FULL TERM (not a fragment) so the archive +// can apply the curated match, and to treat the variant spellings as the same +// entity when reading excerpts. Empty when there are no usable aliases. +// +// This matters because the transcript search only applies an alias's regex when +// the search query contains the alias's full trigger phrase — searching a +// fragment (e.g. "Graham" instead of "graham platner") silently misses the +// curated misspelling coverage. The instruction below is what steers the model +// to search the whole term. Each line also carries the alias `note` so the model +// understands why the term is special. export function renderAliasGlossary(aliases: SearchAlias[], max = 40): string { const usable = aliases.filter((a) => a.enabled !== false); if (usable.length === 0) return ""; const shown = usable.slice(0, max); const lines = shown.map((a) => { - const variants = a.triggers.join(", "); - const note = a.note ? ` (${a.note})` : ""; - return `- ${a.label}: also transcribed as ${variants}${note}`; + const term = a.triggers[0] ?? a.label; + const alts = + a.triggers.length > 1 + ? ` (also written: ${a.triggers.slice(1).join(", ")})` + : ""; + const note = a.note ? ` — ${a.note}` : ""; + return `- ${a.label}: when your search concerns this, search the full term "${term}"${alts}, NOT a fragment; the archive then automatically broadens it to catch mis-transcriptions${note}`; }); const overflow = usable.length > max ? `\n…and ${usable.length - max} more.` : ""; return ( - "Known terms in this archive and how AI transcription commonly mangles " + - "them — account for these variant spellings when you search and when you " + - "read excerpts:\n" + + "KNOWN TERMS — this archive has curated aliases for names/terms that AI " + + "transcription commonly mis-spells. When a search concerns one of these, " + + "search the FULL term shown below (never an abbreviated fragment) so the " + + "curated match applies; and when reading excerpts, treat the listed variant " + + "spellings as the same entity:\n" + lines.join("\n") + overflow ); diff --git a/export/app/lib/askRetrieval.test.ts b/export/app/lib/askRetrieval.test.ts @@ -144,26 +144,55 @@ test("rankResults picks diverse snippets and formats timestamps", () => { assert.equal(v.snippets[0].text, "first hit"); }); -test("buildSearchRoot alias-expands a matching keyword into a regex leaf", () => { - const aliases: SearchAlias[] = [ - { - id: "loli", - label: "loli", - triggers: ["loli", "lolly", "loly"], - suggestion: "\\blol(i|ly)", - useRegex: true, - }, - ]; - const root = buildSearchRoot(["lolly", "banana"], aliases); +const LOLI_ALIAS: SearchAlias = { + id: "loli", + label: "loli", + triggers: ["loli", "lolly", "loly"], + suggestion: "\\blol(i|ly)", + useRegex: true, +}; +const PLATNER_ALIAS: SearchAlias = { + id: "graham-platner", + label: "Graham Platner", + triggers: ["graham platner"], + suggestion: "\\bgra\\w+ plat\\w+", + useRegex: true, +}; + +test("buildSearchRoot: a single-word alias fires and drops the plain leaf", () => { + const root = buildSearchRoot("lolly banana", ["lolly", "banana"], [LOLI_ALIAS]); + const leaves = root.children.filter(isLeaf); + const aliasLeaf = leaves.find((l) => l.id === "t#a0")!; + // "lolly" matches the alias trigger → a regex leaf uses the suggestion. + assert.equal(aliasLeaf.query, "\\blol(i|ly)"); + assert.equal(aliasLeaf.useRegex, true); + // The covered "lolly" keyword is dropped; only "banana" remains as a plain leaf. + assert.ok(!leaves.some((l) => l.query === "lolly")); + const banana = leaves.find((l) => l.query === "banana")!; + assert.equal(banana.useRegex, false); +}); + +test("buildSearchRoot: a MULTI-WORD alias fires on the full phrase (the bug)", () => { + // The whole-query phrase contains "graham platner" → the regex leaf appears. + const root = buildSearchRoot( + "graham platner", + ["graham", "platner"], + [PLATNER_ALIAS], + ); + const leaves = root.children.filter(isLeaf); + const aliasLeaf = leaves.find((l) => l.useRegex); + assert.ok(aliasLeaf, "expected a regex alias leaf"); + assert.equal(aliasLeaf!.query, "\\bgra\\w+ plat\\w+"); + // Both trigger tokens are covered → no bare "graham"/"platner" plain leaves. + assert.ok(!leaves.some((l) => l.query === "graham" || l.query === "platner")); +}); + +test("buildSearchRoot: a fragment does NOT fire a multi-word alias", () => { + // Searching just "graham" can't satisfy the two-token trigger → no regex leaf. + const root = buildSearchRoot("graham", ["graham"], [PLATNER_ALIAS]); const leaves = root.children.filter(isLeaf); - const t0 = leaves.find((l) => l.id === "t#0")!; - // "lolly" matches the alias trigger → transcript leaf uses the regex pattern. - assert.equal(t0.query, "\\blol(i|ly)"); - assert.equal(t0.useRegex, true); - const t1 = leaves.find((l) => l.id === "t#1")!; - // "banana" has no alias → plain substring leaf. - assert.equal(t1.query, "banana"); - assert.equal(t1.useRegex, false); + assert.ok(!leaves.some((l) => l.useRegex)); + assert.ok(leaves.some((l) => l.query === "graham")); }); test("buildContext numbers videos and indents snippets", () => { diff --git a/export/app/lib/askRetrieval.ts b/export/app/lib/askRetrieval.ts @@ -2,8 +2,16 @@ import { runQueryTree, type TreeProgress, } from "yt-dlp-transcript-common/lib/searchEval"; -import { newGroup, newLeaf, type LayerScope } from "yt-dlp-transcript-common/lib/searchQuery"; -import { matchAliases, type SearchAlias } from "yt-dlp-transcript-common/lib/searchAliases"; +import { + newGroup, + newLeaf, + type QueryNode, +} from "yt-dlp-transcript-common/lib/searchQuery"; +import { + matchAliases, + tokenizeQuery, + type SearchAlias, +} from "yt-dlp-transcript-common/lib/searchAliases"; import type { LayerHit } from "yt-dlp-transcript-common/components/searchPipeline"; import type { DisplaySummary } from "yt-dlp-transcript-common/lib/transcripts"; import { formatDuration } from "yt-dlp-transcript-common/lib/format"; @@ -175,43 +183,64 @@ export type RetrieveOptions = RankOptions & { hitLimit?: number; }; -// Build a search leaf for one keyword in one scope, applying an alias -// replacement (regex pattern) when the keyword matches a known trigger. -function leafFor( - keyword: string, - scope: LayerScope, - id: string, - aliases: SearchAlias[], -) { - const hit = aliases.length ? matchAliases(keyword, scope, aliases)[0] : undefined; - if (hit) { - return newLeaf({ - id, - query: hit.suggestion, - scope, - useRegex: hit.useRegex, - contributeHits: true, - }); - } - return newLeaf({ id, query: keyword, scope, contributeHits: true }); -} - -// Build the OR query root for a set of keywords: a transcript-cue leaf and a -// fetch-free metadata leaf per keyword, each alias-expanded when the keyword -// matches a known trigger. Exported for unit testing. Leaf ids encode the -// keyword index after "#" so ranking collapses a keyword's two scopes into one -// unit of term coverage (see termKey() in rankResults). +// Build the OR query root for a search. Aliases are matched against the WHOLE +// query string (phrase-aware — matchAliases fires on single- OR multi-word +// triggers like "graham platner" when their tokens appear in order), and each +// fired alias contributes a regex leaf (`suggestion` + `useRegex`) so the search +// catches known AI-transcription misspellings (e.g. "Grand Platina"). The plain +// per-keyword leaves cover everything else — but a keyword already covered by a +// fired alias's trigger is dropped so a broad bare leaf (e.g. "graham") doesn't +// dilute ranking. Leaf ids encode the term after "#" so ranking collapses a +// term's transcript + metadata leaves into one unit of coverage (see termKey()). +// +// NOTE: alias firing depends on the *query string* containing the trigger phrase. +// The gather prompt steers the model to search the full aliased term (not a +// fragment like "graham"), which is what lets a phrase alias fire here. export function buildSearchRoot( + query: string, keywords: string[], aliases: SearchAlias[] = [], -) { - return newGroup({ - op: "OR", - children: keywords.flatMap((kw, i) => [ - leafFor(kw, "transcripts", `t#${i}`, aliases), - leafFor(kw, "metadata", `m#${i}`, aliases), - ]), +): ReturnType<typeof newGroup> { + const children: QueryNode[] = []; + const firedT = aliases.length ? matchAliases(query, "transcripts", aliases) : []; + const firedM = aliases.length ? matchAliases(query, "metadata", aliases) : []; + + // Tokens covered by any fired alias's trigger — skip plain leaves for these. + const covered = new Set<string>(); + for (const a of [...firedT, ...firedM]) { + for (const t of a.triggers) for (const tok of tokenizeQuery(t)) covered.add(tok); + } + + firedT.forEach((a, i) => + children.push( + newLeaf({ + id: `t#a${i}`, + query: a.suggestion, + scope: "transcripts", + useRegex: a.useRegex, + contributeHits: true, + }), + ), + ); + firedM.forEach((a, i) => + children.push( + newLeaf({ + id: `m#a${i}`, + query: a.suggestion, + scope: "metadata", + useRegex: a.useRegex, + contributeHits: true, + }), + ), + ); + + keywords.forEach((kw, i) => { + if (covered.has(kw)) return; + children.push(newLeaf({ id: `t#${i}`, query: kw, scope: "transcripts", contributeHits: true })); + children.push(newLeaf({ id: `m#${i}`, query: kw, scope: "metadata", contributeHits: true })); }); + + return newGroup({ op: "OR", children }); } // Run the question through the shared search engine and return ranked videos. @@ -234,13 +263,11 @@ export function retrieve( return; } - // OR of the keywords. Each term gets a transcript-cue leaf (contributes - // timestamped snippets) plus a fetch-free metadata leaf (title/channel) for - // cheap title recall. contributeHits stays true so both surface citations. - // Leaf ids encode the keyword index after a "#" so ranking can collapse a - // keyword's transcript + metadata leaves into ONE unit of term coverage - // (see termKey() in rankResults) rather than double-counting the two scopes. - const root = buildSearchRoot(keywords, aliases); + // OR of the query's alias-regex leaves (phrase-matched on the full query) + // and its keyword leaves. Each term gets a transcript-cue leaf (timestamped + // snippets) plus a fetch-free metadata leaf (title/channel) for cheap title + // recall. contributeHits stays true so both surface citations. + const root = buildSearchRoot(question, keywords, aliases); let settled = false; const finish = (fn: () => void) => { diff --git a/export/app/lib/searchAgent.ts b/export/app/lib/searchAgent.ts @@ -38,7 +38,10 @@ export type AgentMode = "auto" | "native" | "scripted"; export type AgentEvent = | { type: "search_start"; query: string } | { type: "search_done"; query: string; count: number } - | { type: "answer_start" } + // Fired once gather is done, BEFORE the answer streams — carries the grounding + // so the UI can persist it and a failed answer stream can be retried without + // re-searching. + | { type: "answer_start"; videos: RetrievedVideo[]; groundedContent: string } | { type: "delta"; text: string }; // Default search budget per turn. Bounds round-trips (and spend on the user's @@ -162,6 +165,13 @@ export type RunAskTurnOptions = { budget?: number; signal?: AbortSignal; onEvent: (e: AgentEvent) => void; + // Retry path: reuse a previous attempt's grounding and SKIP the whole gather + // phase — only (re)stream the answer. Set when retrying a turn whose gather + // succeeded but whose answer stream failed (e.g. a 503 mid-answer). + precomputedGrounding?: { videos: RetrievedVideo[]; groundedContent: string }; + // Human-edited context: replaces the history reconstructed from `prior` (used + // when the user has pruned the context in the context panel). + historyOverride?: ChatMessage[]; }; export type AskTurnResult = { @@ -187,7 +197,26 @@ export async function runAskTurn( onEvent, } = opts; const budget = opts.budget ?? DEFAULT_BUDGET; - const history = buildApiMessages(prior); + const history = opts.historyOverride ?? buildApiMessages(prior); + + // Retry path: grounding already gathered on a prior attempt — skip gather and + // go straight to (re)streaming the answer over the same excerpts. + if (opts.precomputedGrounding) { + const { videos: pv, groundedContent } = opts.precomputedGrounding; + onEvent({ type: "answer_start", videos: pv, groundedContent }); + const answer = await askStream({ + provider, + apiKey, + model, + system: answerSystemPrompt(aliases), + messages: [...history, { role: "user", content: groundedContent }], + maxTokens: 2048, + signal, + onDelta: (chunk) => onEvent({ type: "delta", text: chunk }), + }); + return { answer, videos: pv, groundedContent, truncated: false }; + } + const hasPriorGrounding = prior.some( (m) => m.role === "assistant" && (m.sources?.length ?? 0) > 0, ); @@ -264,7 +293,7 @@ export async function runAskTurn( const finalVideos = [...videos.values()].slice(0, MAX_CONTEXT_VIDEOS); const groundedContent = buildGroundedContent(question, finalVideos); - onEvent({ type: "answer_start" }); + onEvent({ type: "answer_start", videos: finalVideos, groundedContent }); const answer = await askStream({ provider, apiKey, diff --git a/export/e2e/ask-chat.spec.ts b/export/e2e/ask-chat.spec.ts @@ -157,4 +157,136 @@ test.describe("ask chat", () => { // Plain mode shows the raw Markdown source, dashes and all. await expect(page.getByText("- point one [1]", { exact: false })).toBeVisible(); }); + + test("Retry re-runs a failed answer reusing the gathered excerpts", async ({ + page, + }) => { + await installRoutes(page); + let answerAttempts = 0; + await page.route("https://api.anthropic.com/**", async (route) => { + if (route.request().method() === "OPTIONS") { + await route.fulfill({ status: 204, headers: CORS }); + return; + } + const body = route.request().postDataJSON() as { + system?: string; + messages?: { role: string; content: unknown }[]; + }; + const system = body.system ?? ""; + const msgs = body.messages ?? []; + const lastUser = [...msgs].reverse().find((m) => m.role === "user"); + const lastText = + typeof lastUser?.content === "string" ? lastUser.content : ""; + if (system.includes("Markdown")) { + answerAttempts += 1; + if (answerAttempts === 1) { + // First answer attempt fails transiently. + await route.fulfill({ + status: 503, + headers: { ...CORS, "content-type": "application/json" }, + body: JSON.stringify({ type: "error", error: { message: "overloaded" } }), + }); + return; + } + await route.fulfill({ + status: 200, + headers: { ...CORS, "content-type": "text/event-stream" }, + body: sse("Recovered answer [1]"), + }); + return; + } + const text = lastText.includes('Results for "') ? "DONE" : "SEARCH: alpha"; + await route.fulfill({ + status: 200, + headers: { ...CORS, "content-type": "text/event-stream" }, + body: sse(text), + }); + }); + await page.goto("/ask/"); + await page.getByRole("button", { name: "Scripted" }).click(); + await page.locator('input[placeholder^="sk-ant"]').fill("sk-ant-test"); + await ask(page, "tell me about the alpha discussion"); + + // The answer failed → an error turn with a Retry action. + const retry = page.getByRole("button", { name: "Retry" }); + await expect(retry).toBeVisible(); + await retry.click(); + + // Retry succeeds… + await expect(page.getByText(/Recovered answer/)).toBeVisible(); + // …and it reused the excerpts — only the ONE original search ever ran. + await expect(page.getByText(/Searched:\s*alpha/)).toHaveCount(1); + }); + + test("the conversation persists across a reload", async ({ page }) => { + await setup(page); + await ask(page, "tell me about the alpha discussion"); + await expect(page.locator("li", { hasText: "point one" }).first()).toBeVisible(); + // Let the debounced persistence flush, then reload. + await page.waitForTimeout(600); + await page.reload(); + + // Restored from localStorage — question + answer are back without a new call. + await expect(page.getByText("tell me about the alpha discussion")).toBeVisible(); + await expect(page.getByText("point one", { exact: false }).first()).toBeVisible(); + }); + + test("editing the context changes what the next turn sends", async ({ + page, + }) => { + await installRoutes(page); + let sawMarker = false; + await page.route("https://api.anthropic.com/**", async (route) => { + if (route.request().method() === "OPTIONS") { + await route.fulfill({ status: 204, headers: CORS }); + return; + } + const body = route.request().postDataJSON() as { + system?: string; + messages?: { role: string; content: unknown }[]; + }; + if (JSON.stringify(body.messages ?? []).includes("PRUNED_MARKER")) { + sawMarker = true; + } + const system = body.system ?? ""; + const msgs = body.messages ?? []; + const lastUser = [...msgs].reverse().find((m) => m.role === "user"); + const lastText = + typeof lastUser?.content === "string" ? lastUser.content : ""; + let text: string; + if (system.includes("Markdown")) text = "answer with a point [1]"; + else if (lastText.includes('Results for "')) text = "DONE"; + else text = "SEARCH: alpha"; + await route.fulfill({ + status: 200, + headers: { ...CORS, "content-type": "text/event-stream" }, + body: sse(text), + }); + }); + await page.goto("/ask/"); + await page.getByRole("button", { name: "Scripted" }).click(); + await page.locator('input[placeholder^="sk-ant"]').fill("sk-ant-test"); + await ask(page, "first question"); + await expect(page.getByText("answer with a point", { exact: false }).first()).toBeVisible(); + + // Open the Context panel, prune to a marker, and apply it as the new base. + await page.locator("summary", { hasText: "Context" }).click(); + const ta = page.locator("details textarea"); + await ta.fill("USER:\nPRUNED_MARKER only this context"); + await expect(ta).toHaveValue(/PRUNED_MARKER/); + await page.getByRole("button", { name: "Apply as starting point" }).click(); + // Wait for the override to take effect before sending (avoids a state race). + await expect(page.getByText(/Continuing from an edited context/)).toBeVisible(); + // The display reset to empty (messages cleared) → the suggestion chips return. + await expect( + page.getByRole("button", { name: "What are the main topics discussed?" }), + ).toBeVisible(); + + // The next turn must carry the pruned context, not the original conversation. + await ask(page, "second question"); + // Scope to the rendered answer paragraph (the open context textarea also + // contains this text). + await expect(page.locator("p", { hasText: "answer with a point" }).first()).toBeVisible(); + expect(sawMarker).toBe(true); + }); });