Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit a906f1341598f10d3ce8915f7e36660b33013b00
parent 7790cfcdd0bf25535cca7850dce56c3531a79399
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Thu, 16 Jul 2026 21:26:56 -0400

Ask chat: opt-in Report mode (model-maintained living document)

An opt-in mode (native tool-capable providers only) where the model maintains
a persistent markdown report via an update_report tool while the user keeps
chatting, and the conversation is compacted into that report instead of
replaying every prior turn.

- update_report tool (section upsert) in shared.ts + all 3 native providers
  via buildTools(includeFetch, includeReport); parsers extract section/content
  conditionally (existing deep-equal assertions hold); independent per-turn
  counter, off the search budget.
- applyReportPatch (pure, tested): case-insensitive ##-level section upsert,
  preserves preamble/other sections, collapses blank runs.
- searchAgent: report/reportMode options, report in the result, report_update
  event, runUpdateReport executor; COMPACTION — when report mode is on the
  replayed history becomes a single "Report so far" message (the excerpt pool
  still flows into grounding), rebuilt post-gather for the answer phase.
- useAskChat: report/reportMode/reportUpdating state, persisted + restored,
  gated to tool-capable transports; ReportPanel (collapsible, Markdown, toggle
  disabled on scripted with a hint, updating shimmer).

Off by default → all existing paths byte-for-byte unchanged.

Verified: tsc clean (export + common); +7 unit (applyReportPatch ×4,
update_report parse ×3) = 30 green; full e2e 108/108 (native report-mode +
compaction assertion + scripted-disabled cases); no new lint errors.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

Diffstat:
Mexport/CHANGELOG.md | 1+
Mexport/app/ask/AskChat.tsx | 10++++++++++
Aexport/app/ask/ReportPanel.tsx | 95+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mexport/app/ask/useAskChat.ts | 83+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++----
Mexport/app/lib/askConversation.test.ts | 46++++++++++++++++++++++++++++++++++++++++++++++
Mexport/app/lib/askConversation.ts | 55+++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mexport/app/lib/nativeTools/anthropic.ts | 61+++++++++++++++++++++++++++++++++++++++++++++++++++++++------
Mexport/app/lib/nativeTools/gemini.ts | 59+++++++++++++++++++++++++++++++++++++++++++++++++++++------
Mexport/app/lib/nativeTools/openai.ts | 58+++++++++++++++++++++++++++++++++++++++++++++++++++++-----
Mexport/app/lib/nativeTools/parsers.test.ts | 74+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-
Mexport/app/lib/nativeTools/shared.ts | 35++++++++++++++++++++++++++++++++++-
Mexport/app/lib/searchAgent.ts | 61+++++++++++++++++++++++++++++++++++++++++++++++++++++++++----
Mexport/e2e/ask-chat.spec.ts | 93+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
13 files changed, 704 insertions(+), 27 deletions(-)

diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md @@ -1,6 +1,7 @@ # Changelog ## [Unreleased] +- **"Ask AI" chat — an opt-in Report mode for long research sessions.** Turn on **Report mode** (in the new **Report** panel; tool-capable providers only) and the assistant maintains a persistent Markdown **report** document — a running canvas it updates via an `update_report` tool as you keep chatting, upserting well-titled sections that persist across turns. Crucially, while it's on the conversation is **compacted into the report** instead of replaying every prior question and answer: each turn sends a single *"report so far"* summary plus the usual deduplicated excerpt pool, so a long back-and-forth stays bounded in tokens instead of growing every turn. The report renders live in the panel (with an *updating…* shimmer while a turn writes to it) and is saved with the conversation, so it survives reloads; *New chat* clears it. On the Scripted transport there's no tool to drive it, so the toggle is disabled with a hint. See `export/app/lib/nativeTools/{shared,anthropic,openai,gemini}.ts` (the `update_report` tool), `export/app/lib/askConversation.ts` (`applyReportPatch`), `export/app/lib/searchAgent.ts` (compaction + executor), `export/app/ask/{useAskChat.ts,ReportPanel.tsx,AskChat.tsx}`, and `export/e2e/ask-chat.spec.ts`. - **"Ask AI" now grounds in your *current* search automatically.** With search and chat sharing one workspace, you no longer click "Ask AI about these results" to hand a frozen copy of your results to the chat — the chat reads the **live** search directly. Run a search, switch to Chat, and it's already grounded in exactly those results (with the same *answer only from these / may also search* toggle); change the search and the grounding follows. No active search → the chat searches on its own, as before. A **Detach** control lets you ask a free-form question without the current search grounding it (and a **Ground in my search** button re-attaches). The old "Ask AI about these results" button and its one-shot hand-off are retired. See `common/components/SearchSessionContext.tsx` (`liveGrounding`), `export/app/ask/useAskChat.ts`, `export/app/ask/{AskChat,PinnedResultsPanel}.tsx`, and `common/components/SearchResults.tsx`. - **Search and "Ask AI" are now one workspace — the search bar stays put when you switch between them.** Previously `/` (search) and `/ask` (chat) were separate pages, and navigating from one to the other threw away your search. They now share a single shell: the search bar and its results live in a persistent layout, with a **Results ⇄ Chat** switch between the two views. Run a search, flip to Chat to ask about it, flip back — your search is exactly where you left it. Under the hood the search state was lifted out of the monolithic search component into a shared `SearchSession` (both views read the same committed search), so it's the one source of truth features build on. Hub mode is unchanged. See `common/components/{SearchSessionContext,WorkspaceSearchBar,SearchResults,TranscriptSearch}.tsx`, `export/app/(workspace)/*`, and `export/e2e/workspace-shell.spec.ts`. - **Handing a search to "Ask AI" now passes the *whole* result set, not just the first 20.** Previously only the top 20 videos reached the chat, so it literally couldn't see the rest. Now every match is handed off (up to a generous cap) and presented in tiers: the model gets a compact **index of all matching videos** (title, channel, date, hit count) plus **full excerpts for the top ~12** — and it can pull excerpts for any other result on demand via the existing *fetch_context* tool, or you can click **Load context** on any of them. This keeps the payload bounded while letting the assistant reason over the complete set. See `common/lib/aiHandoff.ts` (tiered `buildSearchHandoff`) and `export/app/lib/askConversation.ts` (`buildTieredGrounding`). diff --git a/export/app/ask/AskChat.tsx b/export/app/ask/AskChat.tsx @@ -5,6 +5,7 @@ import { ArrowDownIcon, PlusIcon } from "lucide-react"; import { useAskChat } from "./useAskChat"; import { ProviderSettings } from "./ProviderSettings"; import { ContextPanel } from "./ContextPanel"; +import { ReportPanel } from "./ReportPanel"; import { PinnedResultsPanel } from "./PinnedResultsPanel"; import { MessageBubble } from "./MessageBubble"; import { Composer } from "./Composer"; @@ -160,6 +161,15 @@ export default function AskChat() { /> )} + <ReportPanel + report={s.report} + reportMode={s.reportMode} + setReportMode={s.setReportMode} + available={s.reportModeAvailable} + updating={s.reportUpdating} + busy={busy} + /> + <div className="relative"> <div ref={scrollRef} diff --git a/export/app/ask/ReportPanel.tsx b/export/app/ask/ReportPanel.tsx @@ -0,0 +1,95 @@ +"use client"; + +import { useState } from "react"; +import { FileTextIcon, Loader2Icon } from "lucide-react"; +import { Markdown } from "yt-dlp-transcript-common/components/Markdown"; + +// The running "report" surface for report mode: a persistent markdown document +// the model maintains via the update_report tool while the user keeps chatting. +// The header carries the on/off toggle (disabled on transports without native +// tools, since there is no tool to write the report), a live "updating…" shimmer +// while a turn is writing, and the rendered report (or an empty-state hint). +// Collapsible like the Context panel so it stays out of the way until wanted. +export function ReportPanel({ + report, + reportMode, + setReportMode, + available, + updating, + busy, +}: { + report: string; + reportMode: boolean; + setReportMode: (on: boolean) => void; + available: boolean; + updating: boolean; + busy: boolean; +}) { + const [open, setOpen] = useState(false); + const hasReport = report.trim() !== ""; + + return ( + <details + className="rounded-lg border border-border bg-card/40" + open={open} + onToggle={(e) => setOpen((e.currentTarget as HTMLDetailsElement).open)} + > + <summary className="flex cursor-pointer items-center gap-2 px-4 py-2.5 text-sm font-medium text-foreground"> + <FileTextIcon className="size-4 shrink-0 text-brand" /> + Report + {updating && ( + <span className="inline-flex items-center gap-1 text-xs font-normal text-brand"> + <Loader2Icon className="size-3 animate-spin motion-reduce:animate-none" /> + updating… + </span> + )} + {/* The toggle lives in the header so it's reachable without expanding. */} + <label + className="ml-auto flex items-center gap-1.5 text-xs font-normal text-muted-foreground" + title={ + available + ? "Maintain a persistent report; the conversation is compacted into it." + : "Report mode needs a tool-capable provider (switch off Scripted)." + } + onClick={(e) => e.stopPropagation()} + > + <input + type="checkbox" + checked={reportMode && available} + disabled={!available || busy} + onChange={(e) => setReportMode(e.target.checked)} + className="accent-brand" + /> + Report mode + </label> + </summary> + + {open && ( + <div className="flex flex-col gap-3 border-t border-border px-4 py-3"> + {!available && ( + <p className="text-xs text-warning"> + Report mode needs a tool-capable provider. Switch the search mode off + Scripted (Auto or Native tools) to enable it. + </p> + )} + {hasReport ? ( + <div + aria-live="polite" + className={ + "rounded-md border border-border bg-background px-3 py-2 text-sm text-foreground transition-opacity " + + (updating ? "opacity-60" : "opacity-100") + } + > + <Markdown className="leading-relaxed">{report}</Markdown> + </div> + ) : ( + <p className="text-xs text-muted-foreground"> + Turn on report mode and ask a question — the assistant will build a + report here, keeping it current as the conversation grows. + </p> + )} + </div> + )} + </details> + ); +} diff --git a/export/app/ask/useAskChat.ts b/export/app/ask/useAskChat.ts @@ -11,7 +11,12 @@ import { windowCues, } from "yt-dlp-transcript-common/lib/transcriptWindow"; import { PROVIDERS, type Provider } from "../lib/askProvider"; -import { runAskTurn, type AgentEvent, type AgentMode } from "../lib/searchAgent"; +import { + runAskTurn, + supportsNativeTools, + type AgentEvent, + type AgentMode, +} from "../lib/searchAgent"; import type { RetrievedVideo } from "../lib/askRetrieval"; import { buildApiMessages, @@ -43,6 +48,10 @@ type PersistedConvo = { // Per-video extra transcript pulled in via "Load context", overlaid onto the // live search grounding (keyed by video slug). enrichments?: Record<string, Snip[]>; + // Report mode: the running report document + whether the mode is on. Persisted + // so a long research session's report survives reloads. + report?: string; + reportMode?: boolean; }; const NON_TERMINAL_PHASES = new Set(["gathering", "answering", "streaming"]); @@ -109,6 +118,15 @@ export function useAskChat() { // as a starting point the AI may expand with its own searches. const [strictGrounding, setStrictGroundingState] = useState(true); + // Report mode: an opt-in mode (native tool-capable providers only) where the + // model maintains a persistent markdown report via update_report, and the + // conversation is compacted into that report instead of replaying every turn. + const [report, setReport] = useState(""); + const [reportMode, setReportModeState] = useState(false); + // True while a turn is writing to the report (drives the panel shimmer); + // cleared when the turn ends. + const [reportUpdating, setReportUpdating] = useState(false); + // The effective pinned grounding = the live search, overlaid with any per-video // "Load context" enrichments, unless the user detached. const pinned = useMemo<SearchHandoff | null>(() => { @@ -153,6 +171,10 @@ export function useAskChat() { pinnedRef.current = pinned; const strictRef = useRef(strictGrounding); strictRef.current = strictGrounding; + const reportRef = useRef(report); + reportRef.current = report; + const reportModeRef = useRef(reportMode); + reportModeRef.current = reportMode; // Restore saved preferences + (if remembered) the provider's key/model, plus a // persisted conversation. @@ -189,6 +211,10 @@ export function useAskChat() { if (parsed.enrichments && typeof parsed.enrichments === "object") { setEnrichments(parsed.enrichments); } + if (typeof parsed.report === "string") setReport(parsed.report); + if (typeof parsed.reportMode === "boolean") { + setReportModeState(parsed.reportMode); + } } } catch { /* storage unavailable / corrupt */ @@ -208,7 +234,9 @@ export function useAskChat() { messages.length === 0 && !contextOverride && !detached && - !hasEnrichments + !hasEnrichments && + !report && + !reportMode ) { localStorage.removeItem(K_CONVO); setStorageWarning(false); @@ -226,6 +254,8 @@ export function useAskChat() { strictGrounding, detached, enrichments, + report, + reportMode, }) ) { setStorageWarning(msgs.length < messages.length); @@ -249,7 +279,15 @@ export function useAskChat() { return () => { if (persistTimer.current) clearTimeout(persistTimer.current); }; - }, [messages, contextOverride, strictGrounding, detached, enrichments]); + }, [ + messages, + contextOverride, + strictGrounding, + detached, + enrichments, + report, + reportMode, + ]); function loadProviderCreds(p: Provider, rememberOn: boolean) { // Prefer a model the user typed for this provider earlier this session, so a @@ -311,6 +349,17 @@ export function useAskChat() { } } + // Report mode is only meaningful on a tool-capable transport: Scripted has no + // tool to maintain the report, so the toggle is disabled there. (Persistence is + // handled by the conversation-persist effect, which lists reportMode as a dep.) + const reportModeAvailable = + searchMode !== "scripted" && + (searchMode === "native" || supportsNativeTools(provider, model)); + + function setReportMode(on: boolean) { + setReportModeState(on); + } + // Patch one message by index, immutably. const patchAt = (index: number, fn: (m: UiMessage) => UiMessage) => setMessages((prev) => { @@ -399,6 +448,12 @@ export function useAskChat() { sources: e.videos, })); break; + case "report_update": + // The model wrote a section — flag the live "updating report" state + // (cleared when the turn ends). The updated report text arrives with + // the turn's result below. + setReportUpdating(true); + break; case "delta": patchAt(assistantIndex, (m) => ({ ...m, @@ -409,6 +464,11 @@ export function useAskChat() { } }; + // Report mode only takes effect on a tool-capable transport; on scripted it + // must stay off (there's no update_report tool, and dropping the replayed + // history would lose context with nothing to compact into). + const effectiveReportMode = reportModeRef.current && reportModeAvailable; + try { const result = await runAskTurn({ provider, @@ -423,6 +483,8 @@ export function useAskChat() { onEvent, historyOverride, seedVideos: args.seedVideos, + report: reportRef.current, + reportMode: effectiveReportMode, precomputedGrounding: args.precomputedGrounding?.groundedContent ? { videos: args.precomputedGrounding.videos ?? [], @@ -431,6 +493,10 @@ export function useAskChat() { } : undefined, }); + // Persist the (possibly updated) report from this turn. + if (effectiveReportMode && result.report !== reportRef.current) { + setReport(result.report); + } patchAt(assistantIndex, (m) => ({ ...m, content: m.content || result.answer, @@ -463,11 +529,12 @@ export function useAskChat() { } } finally { setBusy(false); + setReportUpdating(false); sendingRef.current = false; abortRef.current = null; } }, - [provider, apiKey, model, searchMode, summaries, aliases], + [provider, apiKey, model, searchMode, summaries, aliases, reportModeAvailable], ); const send = useCallback(async () => { @@ -581,6 +648,8 @@ export function useAskChat() { setDetached(false); setEnrichments({}); setStrictGroundingState(true); + setReport(""); + setReportModeState(false); }; // Detach from the live search → ask a free-form question (the AI searches on @@ -690,6 +759,12 @@ export function useAskChat() { setSearchMode, markdownOn, setMarkdownOn, + // report mode + report, + reportMode, + setReportMode, + reportModeAvailable, + reportUpdating, // data summariesReady, corpusError, diff --git a/export/app/lib/askConversation.test.ts b/export/app/lib/askConversation.test.ts @@ -2,6 +2,7 @@ import { test } from "node:test"; import assert from "node:assert/strict"; import type { SearchAlias } from "yt-dlp-transcript-common/lib/searchAliases"; import { + applyReportPatch, buildApiMessages, buildGroundedContent, collectPriorPool, @@ -164,6 +165,51 @@ test("parseContext survives pruning and falls back to a single user block", () = assert.deepEqual(parseContext(" "), []); }); +test("applyReportPatch appends a new section to the end, preserving others", () => { + const before = "## Findings\n\nAlpha is discussed."; + const after = applyReportPatch(before, "Open Questions", "Who said beta?"); + // The existing section is preserved… + assert.match(after, /## Findings\n\nAlpha is discussed\./); + // …and the new one is appended with a `## ` heading and its content. + assert.match(after, /## Open Questions\n\nWho said beta\?/); + // New section comes AFTER the existing one. + assert.ok(after.indexOf("## Findings") < after.indexOf("## Open Questions")); +}); + +test("applyReportPatch replaces an existing section's body, leaving neighbours", () => { + const before = + "## Findings\n\nOld body.\n\n## Timeline\n\n- 2020: a thing"; + const after = applyReportPatch(before, "Findings", "New body with more detail."); + assert.match(after, /## Findings\n\nNew body with more detail\./); + // The old body is gone… + assert.doesNotMatch(after, /Old body/); + // …and the untouched neighbouring section survives intact. + assert.match(after, /## Timeline\n\n- 2020: a thing/); +}); + +test("applyReportPatch preserves a preamble and appends into empty reports", () => { + // First write into an empty report is just the section block. + const first = applyReportPatch("", "Findings", "Alpha."); + assert.equal(first, "## Findings\n\nAlpha."); + // A preamble ahead of any heading is kept when a new section is appended. + const withPreamble = applyReportPatch( + "Research notes on the archive.\n\n## Findings\n\nAlpha.", + "Sources", + "[1] video", + ); + assert.match(withPreamble, /^Research notes on the archive\./); + assert.match(withPreamble, /## Sources\n\n\[1\] video/); +}); + +test("applyReportPatch matches the heading case-insensitively", () => { + const before = "## Findings\n\nOld."; + const after = applyReportPatch(before, "findings", "Updated."); + // Matched despite the case difference → replaced, not appended (one heading). + assert.equal((after.match(/## Findings/g) ?? []).length, 1); + assert.match(after, /## Findings\n\nUpdated\./); + assert.doesNotMatch(after, /Old\./); +}); + test("answerSystemPrompt asks for Markdown + citations and carries the glossary", () => { const p = answerSystemPrompt(ALIASES); assert.match(p, /Markdown/); diff --git a/export/app/lib/askConversation.ts b/export/app/lib/askConversation.ts @@ -90,6 +90,53 @@ export function collectPriorPool( return order.map((k) => byKey.get(k)!); } +// ─── report mode (persistent markdown document, maintained via update_report) ─── + +// Upsert a `## <section>` block into the running report. If a section with that +// heading already exists (matched case-insensitively at the `##` level), its body +// is replaced with `content`; otherwise a new `## <section>` block is appended to +// the end. Any preamble and the other sections are preserved, and runs of 3+ +// newlines are collapsed to a single blank line so repeated edits stay tidy. +// Pure + testable — the executor in searchAgent applies it as the model calls the +// tool, threading the result through the turn. +export function applyReportPatch( + report: string, + section: string, + content: string, +): string { + const heading = /^##\s+(.+?)\s*$/; + const target = section.trim().toLowerCase(); + const body = content.trim(); + const lines = report.split("\n"); + + // Index every `##`-level section heading (exactly two hashes: `### …` and + // deeper don't match, so sub-headings inside a section body are left intact). + const heads: { line: number; title: string }[] = []; + lines.forEach((l, i) => { + const m = heading.exec(l); + if (m) heads.push({ line: i, title: m[1] }); + }); + + const hit = heads.findIndex((h) => h.title.trim().toLowerCase() === target); + + let out: string; + if (hit === -1) { + // Append a new section, keeping any existing preamble/sections ahead of it. + const block = `## ${section.trim()}\n\n${body}`; + out = report.trim() === "" ? block : `${report.replace(/\s+$/, "")}\n\n${block}`; + } else { + // Replace the matched section's body, preserving its original heading line. + const start = heads[hit].line; + const end = hit + 1 < heads.length ? heads[hit + 1].line : lines.length; + const before = lines.slice(0, start); + const after = lines.slice(end); + const block = `${lines[start]}\n\n${body}`; + out = [...before, ...block.split("\n"), ...after].join("\n"); + } + + return out.replace(/\n{3,}/g, "\n\n").trim(); +} + // ─── editable context (view / prune / forward to a new session) ─── const ROLE_MARKER = /^(USER|ASSISTANT):\s*$/i; @@ -241,6 +288,7 @@ export function gatherSystemPrompt( mode: "native" | "scripted", budget: number, canFetch = false, + canReport = false, ): string { const base = "You are helping answer a question about a video-transcript archive. " + @@ -256,10 +304,17 @@ export function gatherSystemPrompt( "fetch_context tool with the video's ref (and optionally a timestamp in " + "seconds) to read the surrounding transcript before answering." : ""; + const reportLine = canReport + ? " You are maintaining a running report document that persists across turns. " + + "After gathering evidence, call the update_report tool to record findings in " + + "well-titled sections (pass the section heading and its full markdown " + + "content); keep the report the source of truth. Then answer briefly." + : ""; const tail = mode === "native" ? "Call the search_transcripts tool to search." + fetchLine + + reportLine + " Call the finish tool as soon as you have enough excerpts, or " + "immediately if no search is needed." : "Reply with EXACTLY one line and nothing else: either " + diff --git a/export/app/lib/nativeTools/anthropic.ts b/export/app/lib/nativeTools/anthropic.ts @@ -13,6 +13,10 @@ import { FINISH_TOOL_NAME, type ParsedToolCall, postJson, + REPORT_BUDGET, + REPORT_PARAMS, + REPORT_TOOL_DESCRIPTION, + REPORT_TOOL_NAME, SEARCH_PARAMS, SEARCH_TOOL_DESCRIPTION, SEARCH_TOOL_NAME, @@ -20,8 +24,9 @@ import { } from "./shared"; // The fetch_context tool is offered only when the turn can read transcripts -// (expand mode with a pinned set); search + finish are always present. -function buildTools(includeFetch: boolean) { +// (expand mode with a pinned set); update_report only in report mode; search + +// finish are always present. +function buildTools(includeFetch: boolean, includeReport: boolean) { const tools: { name: string; description: string; input_schema: unknown }[] = [ { name: SEARCH_TOOL_NAME, @@ -36,6 +41,13 @@ function buildTools(includeFetch: boolean) { input_schema: FETCH_PARAMS, }); } + if (includeReport) { + tools.push({ + name: REPORT_TOOL_NAME, + description: REPORT_TOOL_DESCRIPTION, + input_schema: REPORT_PARAMS, + }); + } tools.push({ name: FINISH_TOOL_NAME, description: FINISH_TOOL_DESCRIPTION, @@ -54,7 +66,13 @@ export function parseAnthropicToolUses(json: unknown): ParsedToolCall[] { type?: string; id?: string; name?: string; - input?: { query?: unknown; video?: unknown; aroundSeconds?: unknown }; + input?: { + query?: unknown; + video?: unknown; + aroundSeconds?: unknown; + section?: unknown; + content?: unknown; + }; }; if (b.type === "tool_use" && typeof b.name === "string") { out.push({ @@ -66,6 +84,12 @@ export function parseAnthropicToolUses(json: unknown): ParsedToolCall[] { typeof b.input?.aroundSeconds === "number" ? b.input.aroundSeconds : undefined, + ...(typeof b.input?.section === "string" + ? { section: b.input.section } + : {}), + ...(typeof b.input?.content === "string" + ? { content: b.input.content } + : {}), }); } } @@ -74,14 +98,20 @@ export function parseAnthropicToolUses(json: unknown): ParsedToolCall[] { export async function anthropicGather(ctx: NativeGatherContext): Promise<void> { const canFetch = !!ctx.runFetchContext; - const tools = buildTools(canFetch); + const canReport = !!ctx.runUpdateReport; + const tools = buildTools(canFetch, canReport); const messages: { role: string; content: unknown }[] = [ ...ctx.history.map((m) => ({ role: m.role, content: m.content })), { role: "user", content: ctx.question }, ]; let searchesRun = 0; let fetchesRun = 0; - const maxRounds = ctx.budget + 1 + (canFetch ? FETCH_BUDGET : 0); + let reportsRun = 0; + const maxRounds = + ctx.budget + + 1 + + (canFetch ? FETCH_BUDGET : 0) + + (canReport ? REPORT_BUDGET : 0); for (let round = 0; round <= maxRounds; round++) { if (ctx.signal?.aborted) throw abortError(); @@ -110,8 +140,13 @@ export async function anthropicGather(ctx: NativeGatherContext): Promise<void> { const wantsFetch = canFetch && calls.some((c) => c.name === FETCH_TOOL_NAME && c.video.trim() !== ""); + const wantsReport = + canReport && + calls.some( + (c) => c.name === REPORT_TOOL_NAME && (c.section ?? "").trim() !== "", + ); // No actionable tool call → the model finished (or answered directly). Done. - if (!wantsSearch && !wantsFetch) return; + if (!wantsSearch && !wantsFetch && !wantsReport) return; // Replay the assistant's tool_use turn verbatim, then answer each tool_use. messages.push({ @@ -139,6 +174,20 @@ export async function anthropicGather(ctx: NativeGatherContext): Promise<void> { } else { content = "Reached the limit on transcript reads this turn."; } + } else if ( + canReport && + call.name === REPORT_TOOL_NAME && + (call.section ?? "").trim() !== "" + ) { + if (reportsRun < REPORT_BUDGET) { + reportsRun += 1; + content = await ctx.runUpdateReport!( + call.section!, + call.content ?? "", + ); + } else { + content = "Reached the limit on report updates this turn."; + } } else { // finish (or any other tool): acknowledge so the block is satisfied. content = "Acknowledged."; diff --git a/export/app/lib/nativeTools/gemini.ts b/export/app/lib/nativeTools/gemini.ts @@ -13,13 +13,17 @@ import { FINISH_TOOL_NAME, type ParsedToolCall, postJson, + REPORT_BUDGET, + REPORT_PARAMS, + REPORT_TOOL_DESCRIPTION, + REPORT_TOOL_NAME, SEARCH_PARAMS, SEARCH_TOOL_DESCRIPTION, SEARCH_TOOL_NAME, type NativeGatherContext, } from "./shared"; -function buildTools(includeFetch: boolean) { +function buildTools(includeFetch: boolean, includeReport: boolean) { const functionDeclarations: { name: string; description: string; @@ -38,6 +42,13 @@ function buildTools(includeFetch: boolean) { parameters: FETCH_PARAMS, }); } + if (includeReport) { + functionDeclarations.push({ + name: REPORT_TOOL_NAME, + description: REPORT_TOOL_DESCRIPTION, + parameters: REPORT_PARAMS, + }); + } functionDeclarations.push({ name: FINISH_TOOL_NAME, description: FINISH_TOOL_DESCRIPTION, @@ -60,7 +71,13 @@ export function parseGeminiFunctionCalls(json: unknown): ParsedToolCall[] { part as { functionCall?: { name?: string; - args?: { query?: unknown; video?: unknown; aroundSeconds?: unknown }; + args?: { + query?: unknown; + video?: unknown; + aroundSeconds?: unknown; + section?: unknown; + content?: unknown; + }; }; } ).functionCall; @@ -74,6 +91,12 @@ export function parseGeminiFunctionCalls(json: unknown): ParsedToolCall[] { typeof fc.args?.aroundSeconds === "number" ? fc.args.aroundSeconds : undefined, + ...(typeof fc.args?.section === "string" + ? { section: fc.args.section } + : {}), + ...(typeof fc.args?.content === "string" + ? { content: fc.args.content } + : {}), }); } } @@ -82,7 +105,8 @@ export function parseGeminiFunctionCalls(json: unknown): ParsedToolCall[] { export async function geminiGather(ctx: NativeGatherContext): Promise<void> { const canFetch = !!ctx.runFetchContext; - const tools = buildTools(canFetch); + const canReport = !!ctx.runUpdateReport; + const tools = buildTools(canFetch, canReport); const model = ctx.model || PROVIDERS.gemini.defaultModel; const url = `https://generativelanguage.googleapis.com/v1beta/models/` + @@ -96,7 +120,12 @@ export async function geminiGather(ctx: NativeGatherContext): Promise<void> { ]; let searchesRun = 0; let fetchesRun = 0; - const maxRounds = ctx.budget + 1 + (canFetch ? FETCH_BUDGET : 0); + let reportsRun = 0; + const maxRounds = + ctx.budget + + 1 + + (canFetch ? FETCH_BUDGET : 0) + + (canReport ? REPORT_BUDGET : 0); for (let round = 0; round <= maxRounds; round++) { if (ctx.signal?.aborted) throw abortError(); @@ -119,7 +148,12 @@ export async function geminiGather(ctx: NativeGatherContext): Promise<void> { const wantsFetch = canFetch && calls.some((c) => c.name === FETCH_TOOL_NAME && c.video.trim() !== ""); - if (!wantsSearch && !wantsFetch) return; // finish or a plain answer → done + const wantsReport = + canReport && + calls.some( + (c) => c.name === REPORT_TOOL_NAME && (c.section ?? "").trim() !== "", + ); + if (!wantsSearch && !wantsFetch && !wantsReport) return; // finish/answer → done // Replay the model's functionCall turn, then send functionResponse parts. const modelParts = calls.map((c) => ({ @@ -135,7 +169,9 @@ export async function geminiGather(ctx: NativeGatherContext): Promise<void> { ? { aroundSeconds: c.aroundSeconds } : {}), } - : {}, + : c.name === REPORT_TOOL_NAME + ? { section: c.section ?? "", content: c.content ?? "" } + : {}, }, })); contents.push({ role: "model", parts: modelParts }); @@ -161,6 +197,17 @@ export async function geminiGather(ctx: NativeGatherContext): Promise<void> { } else { result = "Reached the limit on transcript reads this turn."; } + } else if ( + canReport && + call.name === REPORT_TOOL_NAME && + (call.section ?? "").trim() !== "" + ) { + if (reportsRun < REPORT_BUDGET) { + reportsRun += 1; + result = await ctx.runUpdateReport!(call.section!, call.content ?? ""); + } else { + result = "Reached the limit on report updates this turn."; + } } else { result = "Acknowledged."; } diff --git a/export/app/lib/nativeTools/openai.ts b/export/app/lib/nativeTools/openai.ts @@ -13,13 +13,17 @@ import { FINISH_TOOL_NAME, type ParsedToolCall, postJson, + REPORT_BUDGET, + REPORT_PARAMS, + REPORT_TOOL_DESCRIPTION, + REPORT_TOOL_NAME, SEARCH_PARAMS, SEARCH_TOOL_DESCRIPTION, SEARCH_TOOL_NAME, type NativeGatherContext, } from "./shared"; -function buildTools(includeFetch: boolean) { +function buildTools(includeFetch: boolean, includeReport: boolean) { const tools: { type: "function"; function: { name: string; description: string; parameters: unknown }; @@ -43,6 +47,16 @@ function buildTools(includeFetch: boolean) { }, }); } + if (includeReport) { + tools.push({ + type: "function", + function: { + name: REPORT_TOOL_NAME, + description: REPORT_TOOL_DESCRIPTION, + parameters: REPORT_PARAMS, + }, + }); + } tools.push({ type: "function", function: { @@ -69,21 +83,34 @@ export function parseOpenAIToolCalls(json: unknown): ParsedToolCall[] { let query = ""; let video = ""; let aroundSeconds: number | undefined; + let section: string | undefined; + let content: string | undefined; try { const args = JSON.parse(c.function?.arguments ?? "{}"); if (typeof args.query === "string") query = args.query; if (typeof args.video === "string") video = args.video; if (typeof args.aroundSeconds === "number") aroundSeconds = args.aroundSeconds; + if (typeof args.section === "string") section = args.section; + if (typeof args.content === "string") content = args.content; } catch { /* malformed args → empty */ } - return { id: c.id ?? "", name: c.function?.name ?? "", query, video, aroundSeconds }; + return { + id: c.id ?? "", + name: c.function?.name ?? "", + query, + video, + aroundSeconds, + ...(section !== undefined ? { section } : {}), + ...(content !== undefined ? { content } : {}), + }; }); } export async function openaiGather(ctx: NativeGatherContext): Promise<void> { const canFetch = !!ctx.runFetchContext; - const tools = buildTools(canFetch); + const canReport = !!ctx.runUpdateReport; + const tools = buildTools(canFetch, canReport); const messages: unknown[] = [ { role: "system", content: ctx.system }, ...ctx.history.map((m) => ({ role: m.role, content: m.content })), @@ -91,7 +118,12 @@ export async function openaiGather(ctx: NativeGatherContext): Promise<void> { ]; let searchesRun = 0; let fetchesRun = 0; - const maxRounds = ctx.budget + 1 + (canFetch ? FETCH_BUDGET : 0); + let reportsRun = 0; + const maxRounds = + ctx.budget + + 1 + + (canFetch ? FETCH_BUDGET : 0) + + (canReport ? REPORT_BUDGET : 0); for (let round = 0; round <= maxRounds; round++) { if (ctx.signal?.aborted) throw abortError(); @@ -118,8 +150,13 @@ export async function openaiGather(ctx: NativeGatherContext): Promise<void> { const wantsFetch = canFetch && calls.some((c) => c.name === FETCH_TOOL_NAME && c.video.trim() !== ""); + const wantsReport = + canReport && + calls.some( + (c) => c.name === REPORT_TOOL_NAME && (c.section ?? "").trim() !== "", + ); // No actionable tool call → done gathering. - if (!wantsSearch && !wantsFetch) return; + if (!wantsSearch && !wantsFetch && !wantsReport) return; // Replay the assistant message with its tool_calls, then answer EVERY call // (OpenAI requires a tool response for each tool_call id). @@ -144,6 +181,17 @@ export async function openaiGather(ctx: NativeGatherContext): Promise<void> { } else { content = "Reached the limit on transcript reads this turn."; } + } else if ( + canReport && + call.name === REPORT_TOOL_NAME && + (call.section ?? "").trim() !== "" + ) { + if (reportsRun < REPORT_BUDGET) { + reportsRun += 1; + content = await ctx.runUpdateReport!(call.section!, call.content ?? ""); + } else { + content = "Reached the limit on report updates this turn."; + } } else { content = "Acknowledged."; } diff --git a/export/app/lib/nativeTools/parsers.test.ts b/export/app/lib/nativeTools/parsers.test.ts @@ -3,7 +3,7 @@ import assert from "node:assert/strict"; import { parseAnthropicToolUses } from "./anthropic"; import { parseOpenAIToolCalls } from "./openai"; import { parseGeminiFunctionCalls } from "./gemini"; -import { FETCH_TOOL_NAME, SEARCH_TOOL_NAME } from "./shared"; +import { FETCH_TOOL_NAME, REPORT_TOOL_NAME, SEARCH_TOOL_NAME } from "./shared"; test("parseAnthropicToolUses reads search and fetch_context calls", () => { const calls = parseAnthropicToolUses({ @@ -69,6 +69,78 @@ test("parseOpenAIToolCalls tolerates malformed arguments", () => { assert.equal(calls[0].video, ""); }); +test("parseAnthropicToolUses reads update_report calls (section + content)", () => { + const calls = parseAnthropicToolUses({ + content: [ + { + type: "tool_use", + id: "r1", + name: REPORT_TOOL_NAME, + input: { section: "Findings", content: "Alpha is discussed." }, + }, + ], + }); + assert.equal(calls.length, 1); + assert.equal(calls[0].name, REPORT_TOOL_NAME); + assert.equal(calls[0].section, "Findings"); + assert.equal(calls[0].content, "Alpha is discussed."); + // A search call carries no section/content keys (deep-equality with the older + // shape still holds). + const search = parseAnthropicToolUses({ + content: [{ type: "tool_use", id: "s", name: SEARCH_TOOL_NAME, input: { query: "x" } }], + }); + assert.equal("section" in search[0], false); + assert.equal("content" in search[0], false); +}); + +test("parseOpenAIToolCalls reads update_report calls (section + content)", () => { + const calls = parseOpenAIToolCalls({ + choices: [ + { + message: { + tool_calls: [ + { + id: "r1", + function: { + name: REPORT_TOOL_NAME, + arguments: JSON.stringify({ + section: "Timeline", + content: "- 2020: a thing", + }), + }, + }, + ], + }, + }, + ], + }); + assert.equal(calls[0].name, REPORT_TOOL_NAME); + assert.equal(calls[0].section, "Timeline"); + assert.equal(calls[0].content, "- 2020: a thing"); +}); + +test("parseGeminiFunctionCalls reads update_report calls (section + content)", () => { + const calls = parseGeminiFunctionCalls({ + candidates: [ + { + content: { + parts: [ + { + functionCall: { + name: REPORT_TOOL_NAME, + args: { section: "Sources", content: "[1] video" }, + }, + }, + ], + }, + }, + ], + }); + assert.equal(calls[0].name, REPORT_TOOL_NAME); + assert.equal(calls[0].section, "Sources"); + assert.equal(calls[0].content, "[1] video"); +}); + test("parseGeminiFunctionCalls reads search and fetch_context calls", () => { const calls = parseGeminiFunctionCalls({ candidates: [ diff --git a/export/app/lib/nativeTools/shared.ts b/export/app/lib/nativeTools/shared.ts @@ -24,17 +24,25 @@ export type NativeGatherContext = { // window text. Present only in expand mode with a pinned set (native only); // when undefined the fetch_context tool is not offered. runFetchContext?: (video: string, aroundSeconds?: number) => Promise<string>; + // Upserts a section of the running report document (report mode, native only). + // When undefined the update_report tool is not offered. + runUpdateReport?: (section: string, content: string) => Promise<string>; signal?: AbortSignal; }; export const SEARCH_TOOL_NAME = "search_transcripts"; export const FETCH_TOOL_NAME = "fetch_context"; +export const REPORT_TOOL_NAME = "update_report"; export const FINISH_TOOL_NAME = "finish"; // Bounds transcript reads per turn (independent of the search budget) so the // user's key isn't spent on unbounded fetching. export const FETCH_BUDGET = 4; +// Bounds report writes per turn (independent of the search budget). Report mode +// only needs a handful of section upserts per turn to keep the document current. +export const REPORT_BUDGET = 6; + export const SEARCH_TOOL_DESCRIPTION = "Search the video-transcript archive for excerpts relevant to a query. " + "Returns matching videos with timestamped snippet lines. Use focused keyword " + @@ -46,6 +54,11 @@ export const FETCH_TOOL_DESCRIPTION = "result) and optionally a timestamp in seconds to centre on. The returned " + "lines are added to your citable excerpts for that video."; +export const REPORT_TOOL_DESCRIPTION = + "Create or update a section of the running report document that persists " + + "across turns. Pass the section heading and its full new markdown content; " + + "the section is replaced (or appended if new)."; + export const FINISH_TOOL_DESCRIPTION = "Call this when you have gathered enough excerpts (or none are needed) and " + "are ready to answer."; @@ -81,16 +94,36 @@ export const FETCH_PARAMS = { required: ["video"], } as const; +// JSON-Schema for the update_report tool's input. +export const REPORT_PARAMS = { + type: "object", + properties: { + section: { + type: "string", + description: "The section heading (without the leading ##).", + }, + content: { + type: "string", + description: "The full new markdown content for the section.", + }, + }, + required: ["section", "content"], +} as const; + export const EMPTY_PARAMS = { type: "object", properties: {} } as const; // Shape returned by every provider's tool-call parser: `query` for search, -// `video`/`aroundSeconds` for fetch_context (all optional, filled per tool). +// `video`/`aroundSeconds` for fetch_context, `section`/`content` for +// update_report (all optional, filled per tool — omitted keys stay absent so +// deep-equality against a search/finish call still matches). export type ParsedToolCall = { id: string; name: string; query: string; video: string; aroundSeconds?: number; + section?: string; + content?: string; }; // Thrown when a native tool-calling request fails in a way that suggests the diff --git a/export/app/lib/searchAgent.ts b/export/app/lib/searchAgent.ts @@ -17,6 +17,7 @@ import { import { retrieve, type RetrievedVideo } from "./askRetrieval"; import { answerSystemPrompt, + applyReportPatch, buildApiMessages, buildGroundedContent, collectPriorPool, @@ -49,6 +50,8 @@ export type AgentEvent = // The model read more of a pinned video's transcript (fetch_context tool). | { type: "fetch_start"; ref: string; label: string } | { type: "fetch_done"; ref: string; count: number } + // The model upserted a section of the running report (update_report tool). + | { type: "report_update"; section: string } // Fired once gather is done, BEFORE the answer streams — carries the grounding // so the UI can persist it and a failed answer stream can be retried without // re-searching. @@ -225,6 +228,12 @@ export type RunAskTurnOptions = { // Human-edited context: replaces the history reconstructed from `prior` (used // when the user has pruned the context in the context panel). historyOverride?: ChatMessage[]; + // Report mode (native providers only): the running report document carried in, + // and a flag turning the mode on. When on, the prior Q&A text replay is dropped + // in favour of a compact "report so far" history (the excerpt pool still flows), + // and the update_report tool is offered so the model keeps the report current. + report?: string; + reportMode?: boolean; }; export type AskTurnResult = { @@ -232,6 +241,9 @@ export type AskTurnResult = { videos: RetrievedVideo[]; groundedContent: string; truncated: boolean; + // The report after this turn (possibly updated via update_report). Echoes the + // input report unchanged when report mode is off or nothing was written. + report: string; }; export async function runAskTurn( @@ -250,7 +262,28 @@ export async function runAskTurn( onEvent, } = opts; const budget = opts.budget ?? DEFAULT_BUDGET; - const history = opts.historyOverride ?? buildApiMessages(prior); + const reportMode = opts.reportMode === true; + // The running report — mutated in place by the update_report executor below and + // returned in the result so the caller can persist it. + let report = opts.report ?? ""; + + // Report mode COMPACTION: instead of replaying every prior Q&A turn as text, + // send a single "report so far" assistant message (empty report → no history at + // all). The cumulative excerpt pool still flows into the grounding unchanged — + // only the prior Q&A text replay is dropped, keeping long sessions bounded. + const compactHistory = (): ChatMessage[] => + report.trim() === "" + ? [] + : [ + { + role: "assistant", + content: + "Report so far (keep it current with update_report):\n\n" + report, + }, + ]; + const history = reportMode + ? compactHistory() + : opts.historyOverride ?? buildApiMessages(prior); // Retry path: grounding already gathered on a prior attempt — skip gather and // go straight to (re)streaming the answer over the same excerpts. @@ -272,6 +305,7 @@ export async function runAskTurn( videos: pv, groundedContent, truncated: opts.precomputedGrounding.truncated ?? false, + report, }; } @@ -359,6 +393,19 @@ export async function runAskTurn( ); }; + // Report mode: apply the model's section upsert to the running report, emit an + // event so the UI can show a live "updating report" indicator, and hand back a + // short confirmation as the tool result. Report writes don't consume the search + // budget (they have their own REPORT_BUDGET inside each provider loop). + const runUpdateReport = async ( + section: string, + content: string, + ): Promise<string> => { + report = applyReportPatch(report, section, content); + onEvent({ type: "report_update", section }); + return `Updated the "${section}" section.`; + }; + const gatherCtx: NativeGatherContext = { apiKey, model, @@ -372,12 +419,14 @@ export async function runAskTurn( const doGather = async (kind: "native" | "scripted"): Promise<void> => { const canFetch = kind === "native" && canFetchThisTurn; - let system = gatherSystemPrompt(aliases, kind, budget, canFetch); + const canReport = kind === "native" && reportMode; + let system = gatherSystemPrompt(aliases, kind, budget, canFetch, canReport); if (canFetchThisTurn) system += `\n\n${buildSeedDigest(seedVideos)}`; const ctx: NativeGatherContext = { ...gatherCtx, system, runFetchContext: canFetch ? runFetchContext : undefined, + runUpdateReport: canReport ? runUpdateReport : undefined, }; if (kind === "scripted") return scriptedGather(ctx, provider); switch (provider) { @@ -420,17 +469,21 @@ export async function runAskTurn( ].slice(0, MAX_CONTEXT_VIDEOS); const groundedContent = buildGroundedContent(question, finalVideos); + // In report mode, rebuild the compact history from the NOW-updated report so the + // answer phase sees the freshest report the gather loop wrote. + const answerHistory = reportMode ? compactHistory() : history; + onEvent({ type: "answer_start", videos: finalVideos, groundedContent }); const answer = await askStream({ provider, apiKey, model, system: answerSystemPrompt(aliases), - messages: [...history, { role: "user", content: groundedContent }], + messages: [...answerHistory, { role: "user", content: groundedContent }], maxTokens: 2048, signal, onDelta: (chunk) => onEvent({ type: "delta", text: chunk }), }); - return { answer, videos: finalVideos, groundedContent, truncated }; + return { answer, videos: finalVideos, groundedContent, truncated, report }; } diff --git a/export/e2e/ask-chat.spec.ts b/export/e2e/ask-chat.spec.ts @@ -610,6 +610,99 @@ test.describe("ask chat", () => { ).toBeVisible(); }); + test("report mode: the model maintains a report and the history is compacted", async ({ + page, + }) => { + await installRoutes(page); + let reportToolOffered = false; + const answerBodies: string[] = []; + await page.route("https://api.anthropic.com/**", async (route) => { + if (route.request().method() === "OPTIONS") { + await route.fulfill({ status: 204, headers: CORS }); + return; + } + const body = route.request().postDataJSON() as { + system?: string; + tools?: { name?: string }[]; + messages?: { role: string; content: unknown }[]; + }; + const system = body.system ?? ""; + // Answer phase → capture the messages so we can assert compaction. + if (system.includes("Markdown")) { + answerBodies.push(JSON.stringify(body.messages ?? [])); + await route.fulfill({ + status: 200, + headers: { ...CORS, "content-type": "text/event-stream" }, + body: sse("Report-mode answer [1]"), + }); + return; + } + // Native gather turn: write a report section once, then finish. + if (Array.isArray(body.tools)) { + if (body.tools.some((t) => t.name === "update_report")) { + reportToolOffered = true; + } + const msgs = body.messages ?? []; + const lastUser = [...msgs].reverse().find((m) => m.role === "user"); + const isToolResult = Array.isArray(lastUser?.content); + await route.fulfill({ + status: 200, + headers: { ...CORS, "content-type": "application/json" }, + body: isToolResult + ? toolUse("finish", {}) + : toolUse("update_report", { + section: "Findings", + content: "Alpha finding recorded from the archive.", + }), + }); + return; + } + await route.fulfill({ + status: 200, + headers: { ...CORS, "content-type": "text/event-stream" }, + body: sse("DONE"), + }); + }); + await page.goto("/ask/"); + await keyIn(page, "Native tools"); + + // Enable report mode (the toggle lives in the Report panel header). + await page.getByLabel("Report mode").check(); + + // Turn 1 establishes the report + some history. + await ask(page, "first question about the alpha topic"); + await expect(page.getByText(/Report-mode answer/)).toBeVisible(); + expect(reportToolOffered).toBe(true); + + // The Report panel shows the section the model wrote. + await page.locator("summary", { hasText: "Report" }).click(); + await expect( + page.getByText("Alpha finding recorded from the archive."), + ).toBeVisible(); + + // Turn 2: the answer request must be COMPACTED — it carries the "Report so + // far" summary and does NOT replay turn 1's full question text. + await ask(page, "second question about the beta topic"); + await expect(page.getByText(/Report-mode answer/).nth(1)).toBeVisible(); + + const lastAnswer = answerBodies[answerBodies.length - 1]; + expect(lastAnswer).toContain("Report so far"); + expect(lastAnswer).toContain("Alpha finding recorded from the archive."); + expect(lastAnswer).not.toContain("first question about the alpha topic"); + }); + + test("report mode is disabled on the scripted transport", async ({ page }) => { + await installRoutes(page); + await page.goto("/ask/"); + // keyIn fills the key first (proving hydration) before switching to Scripted, + // so the mode click isn't lost to a pre-hydration no-op. + await keyIn(page, "Scripted"); + // The toggle is present but disabled (no tool to maintain the report). + const toggle = page.getByLabel("Report mode"); + await expect(toggle).toBeVisible(); + await expect(toggle).toBeDisabled(); + }); + test("editing the context changes what the next turn sends", async ({ page, }) => {