Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 4782bea368f426a1f31f0133910023c725a6d0da
parent d0097cc6a45878a5e4c5bbb9cdc0aa6045aecf56
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Wed, 22 Jul 2026 11:00:27 -0400

Ask chat: clickable precise-moment report citations — [n @ mm:ss] + a persisted source registry

The Report panel now mirrors the chat: citation markers render as in-page
links that open the transcript modal seeked to the cited line (snapped to
the nearest real snippet second), with a numbered Sources list beneath.
Chat citations upgraded to the same [n @ mm:ss] form. New shared renderer
in citations.tsx; reportSources registry persisted with the convo + saved
chats; sweep batches numbered by global registry index.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

Diffstat:
Mexport/CHANGELOG.md | 1+
Mexport/app/ask/AskChat.tsx | 2++
Mexport/app/ask/MessageBubble.tsx | 139+++++--------------------------------------------------------------------------
Mexport/app/ask/ReportPanel.tsx | 34+++++++++++++++++++++++++++++++---
Mexport/app/ask/askChatStorage.test.ts | 48++++++++++++++++++++++++++++++++++++++++++++++++
Mexport/app/ask/askChatStorage.ts | 8++++++++
Aexport/app/ask/citations.test.ts | 94+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aexport/app/ask/citations.tsx | 222+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mexport/app/ask/useAskChat.ts | 50+++++++++++++++++++++++++++++++++++++++++++++++++-
Mexport/app/lib/askConversation.test.ts | 67++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-
Mexport/app/lib/askConversation.ts | 93+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++------------
Mexport/app/lib/askRetrieval.test.ts | 31+++++++++++++++++++++++++++++++
Mexport/app/lib/askRetrieval.ts | 13++++++++++---
Mexport/app/lib/searchAgent.ts | 62++++++++++++++++++++++++++++++++++++++++++++++++++++++++++----
Mexport/e2e/ask-chat.spec.ts | 65+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
15 files changed, 772 insertions(+), 157 deletions(-)

diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md @@ -1,6 +1,7 @@ # Changelog ## [Unreleased] +- **"Ask AI" — the Report panel's citations are now clickable, precise to the exact moment cited.** The chat's answer bubbles already linked their `[1]`, `[2]`… markers to the cited transcript; the **Report** panel — the persistent document Report mode and whole-corpus sweeps maintain — rendered as plain text, so the citations a user works from most were dead. Now the report mirrors the chat: it renders each citation as an in-page link and lists its **numbered sources** beneath the document, and clicking one **opens the transcript modal seeked to the cited line**. Citations are also upgraded — in the report *and* the chat — to a precise-moment **`[n @ mm:ss]`** form (the source's number plus the specific line's timestamp, e.g. `[3 @ 12:34]`), so a citation jumps to the exact moment instead of the video's first matched snippet. A slightly-off model timestamp is **snapped to the nearest real transcript line**, and it stays graceful: a bare `[n]` still resolves to the first snippet, and an out-of-range number or unparseable time stays plain text rather than becoming a broken link. Making this work in the report (a single accumulated string, unlike a chat message that carries its own sources) required two new pieces: a **persisted report source registry** — every video the report cites, deduped and stored light (no excerpt text) — and **global citation numbering** so a report's `[n]` is stable across every batch/turn folded in, not renumbered per write. The registry is saved with the conversation and each saved chat, so a report's clickable citations survive a reload; New chat clears it; and on the federated hub /ask (no player) citations fall back to scroll-to-source, exactly like the chat. See `export/app/ask/citations.tsx` (new shared renderer — `linkifyCitations` + `CitationLink` + `CitationSources`, with `[n @ mm:ss]` parsing and nearest-snippet snapping), `export/app/ask/{MessageBubble,ReportPanel,AskChat,useAskChat,askChatStorage}.tsx/ts`, `export/app/lib/{askRetrieval,askConversation,searchAgent}.ts` (optional/global-registry numbering + `mergeReportSources` + the `[n @ mm:ss]` prompt updates), and `export/e2e/ask-chat.spec.ts`. - **MCP server: switch which corpus you're reading on the fly — and it sticks across reconnects.** The MCP server was pinned to a single corpus at launch (`--hub` / `--remote` / `--local` or the `TRANSCRIPT_*` env), so re-aiming it at a different site or a local build meant editing the MCP config and restarting. Now, connected to (say) the **archilyzer hub**, you can retarget it from inside a session with three new **read-only** tools: **`list_sources`** shows the active corpus and, in a hub context, its member sites (`siteId · title · url`); **`use_source`** switches the active corpus — to a single hub member (`site:` — which becomes a plain single-site source and so regains **full channel-group + alias** support), a **federated subset** of the hub (`sites:[…]`, unknown tokens reported not dropped), or an arbitrary `remote:` URL / `local:` dir / `hub:` URL; and **`reset_source`** returns to the startup source. The selection is **persisted** to a small state file (under `$XDG_STATE_HOME/yt-dlp-transcript-mcp`, overridable with `TRANSCRIPT_MCP_STATE_DIR`) keyed by the **startup** source, so it survives a reconnect and two differently-configured servers keep independent selections. Everything the sweep/search tools do reads the current active source. Still strictly **read-only** — this only changes *which* already-published static shards are read, the same capability the startup flags already grant this locally-run tool; nothing is written to any corpus. See `mcp/src/{sources,sourceController,source,server,index}.ts`, `mcp/src/sourceController.test.ts`, and `mcp/README.md`. - **MCP server: run the corpus sweep through Claude Code itself — on plan usage, no API key.** The in-browser "corpus sweep" (fold every matching transcript into a running report) needs a BYO AI key, and a free key conks out fast on a real 30k-video corpus. The `mcp/` server now lets **Claude Code be the sweep engine** instead, via a first-class **`sweep` prompt** (a slash command, `/mcp__<name>__sweep query="k cups" channel="chrissie-mayr"`) plus the tools to drive it. `search_transcripts` gains **paging** — it reports the full `total` and `has_more`, so the whole match set can be enumerated with `offset` (and `include_snippets:false` for a cheap worklist) — and is now **alias-aware**: a plain query that matches a curated search alias also searches the alias regex (e.g. `k cups` → also `cake cup`), with the footer naming which aliases fired; reaching the scan cap is surfaced as **PARTIAL** coverage rather than hidden. A new **`get_transcripts`** tool batch-reads up to 20 videos in one call as bounded, timestamped **excerpt windows** around the matches (alias-correct) — or full transcripts without a query — so a sweep stays token-bounded. The `sweep` prompt walks Claude through search → enumerate → plan `ceil(N/batch)` batches → per-batch windowed read + cross-referenced upsert of cited findings (*title + [mm:ss]*) into a markdown report it maintains with its own Write/Edit tools. The sweep is now **group-aware and multi-channel**: `search_transcripts` scope is additive over one-or-more channels (`channel`/`channels`) and/or channel **groups** (`group`/`groups`, matched by group **id or display name** — `group="other"` ≡ `group="Extended Universe"` — and expanded to the group's channels), with the footer naming the resolved scope and flagging any channel/group token that matched nothing so a typo isn't silently a whole-corpus scan; `get_transcripts` takes matching `channels` lookup hints; and `list_channels` now organizes channels under their groups with a compact `id · name · N channels` cheat-sheet. Crucially the `sweep` prompt is now **pick-first**: invoked with **no** scope it makes Claude list the groups/channels and **ask which to sweep — or to confirm the whole corpus — before enumerating**, rather than silently scanning everything (an explicit `search_transcripts` call with no selector still means "all"). (Hub mode defers per-site group resolution — multi-channel scoping still works there.) The MCP stays strictly **read-only**; only the report file is written, in Claude's own working directory. See `mcp/src/{search,source,server}.ts`, `mcp/src/search.test.ts`, and `mcp/README.md`. - **"Ask AI" now paces itself to your key's rate limit instead of failing.** A whole-corpus sweep on a free-tier key used to fire provider calls as fast as the loop could produce them, blow straight through the per-minute request cap, and stop dead on the first HTTP 429 (and a rate-limit *mid-answer* surfaced as a hard error, because only the sweep path ever handled 429). Now a single client-side limiter sits behind **every** AI call: it **spaces requests** to a conservative, free-tier-safe **requests-per-minute** default chosen per model (e.g. Gemini `*-pro` → 5/min, `*flash` → 10/min; Claude/OpenAI higher), so a paid key runs fast and a free key just runs *slowly* rather than erroring. When a 429 does land, the limiter **honours the provider's own retry hint** (the `Retry-After` header, or Gemini's `RetryInfo` retry-delay) — or an exponential backoff — and **re-issues the request** (safe for streaming: the retry happens before any answer text is emitted), and it **self-lowers** the rate after a 429 so later calls pace slower. Only a 429 that outlasts the retries falls through to the existing **pause/checkpoint** (the right home for a daily-quota cap you resume tomorrow). Provider settings gain an editable **Requests per minute** field (with the effective spacing, e.g. *"≈12s between requests"*, and a reset-to-default), and a paced sweep shows a **"Rate limited — retrying in Ns"** line in its progress strip so it never looks frozen. See `export/app/lib/rateLimit.ts` (the whole mechanism), `export/app/lib/askProvider.ts` (`PausableError` + `acquire`/retry in `askStream`, 429 in `ensureOk`), `export/app/lib/nativeTools/shared.ts` (`acquire`/retry in `postJson`), `export/app/ask/{useAskChat,ProviderSettings,PinnedResultsPanel,AskChat}.tsx`, and `export/e2e/ask-workspace.spec.ts`. diff --git a/export/app/ask/AskChat.tsx b/export/app/ask/AskChat.tsx @@ -276,6 +276,8 @@ export default function AskChat() { <ReportPanel report={s.report} + sources={s.reportSources} + onOpenCitation={onOpenCitation} reportMode={s.reportMode} setReportMode={s.setReportMode} available={s.reportModeAvailable} diff --git a/export/app/ask/MessageBubble.tsx b/export/app/ask/MessageBubble.tsx @@ -1,90 +1,12 @@ "use client"; -import { useId, useState, type AnchorHTMLAttributes } from "react"; +import { useId, useState } from "react"; import { CheckIcon, CopyIcon, PencilIcon, RotateCwIcon } from "lucide-react"; import { Markdown } from "yt-dlp-transcript-common/components/Markdown"; import type { UiMessage } from "../lib/askConversation"; -import type { RetrievedVideo } from "../lib/askRetrieval"; +import { CitationSources, linkifyCitations, makeCitationLink } from "./citations"; import { PipelineStatus } from "./PipelineStatus"; -// Turn `[n]` citation markers in an answer into links to the matching source in -// the list below. Skips fenced/inline code, and only links numbers that have a -// source (1..count). The anchor id is scoped per message via `cid`. -function linkifyCitations(md: string, maxN: number, cid: string): string { - if (maxN <= 0) return md; - return md - .split(/(```[\s\S]*?```|`[^`]*`)/g) - .map((seg, i) => - i % 2 === 1 - ? seg - : seg.replace(/\[(\d+)\]/g, (m, d) => { - const n = parseInt(d, 10); - return n >= 1 && n <= maxN ? `[[${n}]](#cite-${cid}-${n})` : m; - }), - ) - .join(""); -} - -// Scroll to and briefly highlight the source `<li>` for citation `n` — the -// fallback when there's no player (hub /ask) or the source has no timestamp. -function scrollToSource(cid: string, n: number) { - const el = document.getElementById(`cite-${cid}-${n}`); - if (!el) return; - el.scrollIntoView({ behavior: "smooth", block: "center" }); - el.classList.add("bg-brand-soft"); - setTimeout(() => el.classList.remove("bg-brand-soft"), 900); -} - -// Build the link renderer for answer Markdown, closing over this message's -// sources + citation-scope id + the (optional) modal opener. A `#cite-<cid>-<n>` -// anchor opens the transcript modal at the cited video's first matched line when -// a player is available; otherwise it scrolls to the source below. Any other -// link opens externally. -function makeCitationLink( - cid: string, - sources: RetrievedVideo[], - onOpenCitation?: (slug: string, seconds?: number) => void, -) { - return function CitationLink({ - href, - children, - ...rest - }: AnchorHTMLAttributes<HTMLAnchorElement>) { - if (typeof href === "string" && href.startsWith("#cite-")) { - return ( - <a - href={href} - onClick={(e) => { - e.preventDefault(); - const n = parseInt(href.match(/-(\d+)$/)?.[1] ?? "", 10); - const src = Number.isFinite(n) ? sources[n - 1] : undefined; - const seconds = src?.snippets[0]?.seconds; - if (onOpenCitation && src && typeof seconds === "number") { - onOpenCitation(src.key, seconds); - } else { - scrollToSource(cid, n); - } - }} - className="font-mono text-brand no-underline hover:underline" - > - {children} - </a> - ); - } - return ( - <a - href={href} - target="_blank" - rel="noopener noreferrer" - className="text-brand underline decoration-brand/40 hover:decoration-brand" - {...rest} - > - {children} - </a> - ); - }; -} - function CopyButton({ text }: { text: string }) { const [state, setState] = useState<"idle" | "copied" | "failed">("idle"); const copy = async () => { @@ -265,57 +187,12 @@ export function MessageBubble({ )} {!isUser && message.sources && message.sources.length > 0 && ( - <ol className="ml-1 flex flex-col gap-1 text-xs text-muted-foreground"> - {message.sources.map((s, si) => ( - <li - key={s.key} - id={`cite-${cid}-${si + 1}`} - className="scroll-mt-20 rounded transition-colors animate-in fade-in motion-reduce:animate-none" - style={{ animationDelay: `${Math.min(si, 8) * 40}ms`, animationFillMode: "both" }} - > - <span className="font-mono text-brand">[{si + 1}]</span>{" "} - {s.url ? ( - <a - href={s.url} - target="_blank" - rel="noopener noreferrer" - className="hover:underline" - > - {s.title} - </a> - ) : ( - s.title - )}{" "} - <span className="text-muted-foreground/70"> - — {s.channel} - {s.siteTitle ? ` · ${s.siteTitle}` : ""} ·{" "} - {s.snippets.map((sn, sni) => ( - <span key={sn.seconds}> - {sni > 0 ? ", " : ""} - {onOpenCitation ? ( - <button - type="button" - onClick={() => onOpenCitation(s.key, sn.seconds)} - title="Open the transcript at this moment" - className="font-mono text-brand transition-colors hover:underline" - > - {sn.clock} - </button> - ) : ( - sn.clock - )} - </span> - ))} - </span> - </li> - ))} - {message.truncated && ( - <li className="text-muted-foreground/60"> - (some searches were truncated; ask more specifically for fuller - coverage) - </li> - )} - </ol> + <CitationSources + cid={cid} + sources={message.sources} + onOpenCitation={onOpenCitation} + truncated={message.truncated} + /> )} </div> ); diff --git a/export/app/ask/ReportPanel.tsx b/export/app/ask/ReportPanel.tsx @@ -1,8 +1,10 @@ "use client"; -import { useState } from "react"; +import { useId, useState } from "react"; import { FileTextIcon, Loader2Icon } from "lucide-react"; import { Markdown } from "yt-dlp-transcript-common/components/Markdown"; +import type { RetrievedVideo } from "../lib/askRetrieval"; +import { CitationSources, linkifyCitations, makeCitationLink } from "./citations"; // The running "report" surface for report mode: a persistent markdown document // the model maintains via the update_report tool while the user keeps chatting. @@ -12,6 +14,8 @@ import { Markdown } from "yt-dlp-transcript-common/components/Markdown"; // Collapsible like the Context panel so it stays out of the way until wanted. export function ReportPanel({ report, + sources, + onOpenCitation, reportMode, setReportMode, available, @@ -20,6 +24,12 @@ export function ReportPanel({ sweeping = false, }: { report: string; + // The report's source registry — what its `[n]` / `[n @ mm:ss]` citations index + // into. Drives the clickable citations + the numbered Sources list below. + sources: RetrievedVideo[]; + // Open the transcript modal seeked to a cited line. Undefined in hub /ask (no + // player), where citations fall back to scroll-to-source within the list below. + onOpenCitation?: (slug: string, seconds?: number) => void; reportMode: boolean; setReportMode: (on: boolean) => void; available: boolean; @@ -30,6 +40,10 @@ export function ReportPanel({ sweeping?: boolean; }) { const [open, setOpen] = useState(false); + // Stable per-panel id so the report's citation anchors don't collide with a + // chat message's. + const cid = useId().replace(/:/g, ""); + const CitationLink = makeCitationLink(cid, sources, onOpenCitation); const hasReport = report.trim() !== ""; // While a sweep runs, keep the panel expanded (so its document is visibly // growing) regardless of the user's toggle. @@ -84,11 +98,25 @@ export function ReportPanel({ <div aria-live="polite" className={ - "rounded-md border border-border bg-background px-3 py-2 text-sm text-foreground transition-opacity " + + "flex flex-col gap-3 rounded-md border border-border bg-background px-3 py-2 text-sm text-foreground transition-opacity " + (busyWriting ? "opacity-60" : "opacity-100") } > - <Markdown className="leading-relaxed">{report}</Markdown> + <Markdown className="leading-relaxed" linkComponent={CitationLink}> + {linkifyCitations(report, sources.length, cid)} + </Markdown> + {sources.length > 0 && ( + <div className="border-t border-border pt-2"> + <p className="mb-1 text-xs font-medium text-muted-foreground"> + Sources + </p> + <CitationSources + cid={cid} + sources={sources} + onOpenCitation={onOpenCitation} + /> + </div> + )} </div> ) : ( <p className="text-xs text-muted-foreground"> diff --git a/export/app/ask/askChatStorage.test.ts b/export/app/ask/askChatStorage.test.ts @@ -104,10 +104,58 @@ test("a sweep checkpoint round-trips (serialize → resume yields the remaining assert.equal(cp.reason, "usage limit"); }); +test("a report source registry round-trips (and old saves without one still load)", () => { + rawStore().removeItem(KEY); + const chat = baseChat({ + report: "# Findings\n\nClaim [1 @ 0:05].\n", + reportSources: [ + { + key: "ch/v1", + videoId: "v1", + title: "A video", + channel: "Chan", + uploadDate: "20200101", + snippets: [{ clock: "0:05", seconds: 5, text: "" }], + }, + ], + }); + saveSavedChats({ v: 1, items: { research: chat }, activeName: "research" }); + const back = loadSavedChats().items.research; + assert.equal(back.reportSources?.length, 1); + assert.equal(back.reportSources?.[0].key, "ch/v1"); + assert.equal(back.reportSources?.[0].snippets[0].seconds, 5); + + // An old save with no reportSources field loads with it simply absent. + rawStore().setItem( + KEY, + JSON.stringify({ v: 1, items: { legacy: baseChat({ report: "old" }) }, activeName: "legacy" }), + ); + const legacy = loadSavedChats().items.legacy; + assert.equal(legacy.report, "old"); + assert.equal(legacy.reportSources, undefined); +}); + test("savedChatsEqual: equal on the same restorable content", () => { assert.ok(savedChatsEqual(baseChat(), baseChat())); }); +test("savedChatsEqual: dirty when the report source registry differs", () => { + const withSources = baseChat({ + reportSources: [ + { + key: "ch/v1", + videoId: "v1", + title: "A", + channel: "C", + uploadDate: "20200101", + snippets: [{ clock: "0:05", seconds: 5, text: "" }], + }, + ], + }); + assert.ok(!savedChatsEqual(baseChat(), withSources)); + assert.ok(savedChatsEqual(withSources, withSources)); +}); + test("savedChatsEqual: dirty when messages / report / checkpoint differ", () => { assert.ok(!savedChatsEqual(baseChat(), baseChat({ messages: [msg("user", "x")] }))); assert.ok(!savedChatsEqual(baseChat(), baseChat({ report: "changed" }))); diff --git a/export/app/ask/askChatStorage.ts b/export/app/ask/askChatStorage.ts @@ -11,6 +11,7 @@ // conversation. import type { UiMessage } from "../lib/askConversation"; +import type { RetrievedVideo } from "../lib/askRetrieval"; const KEY = "ytdlp-tb:ai:chats"; const VERSION = 1; @@ -42,6 +43,9 @@ export type SavedChat = { detached?: boolean; enrichments?: Record<string, { clock: string; seconds: number; text: string }[]>; report?: string; + // The report's source registry (what its `[n]` citations index into), trimmed of + // snippet text. Absent on chats saved before this existed. + reportSources?: RetrievedVideo[]; reportMode?: boolean; // Sweep checkpoint (present only for a paused sweep). checkpoint?: SweepCheckpoint | null; @@ -94,6 +98,9 @@ function parseChat(raw: unknown): SavedChat | null { chat.enrichments = r.enrichments as SavedChat["enrichments"]; } if (typeof r.report === "string") chat.report = r.report; + if (Array.isArray(r.reportSources)) { + chat.reportSources = r.reportSources as RetrievedVideo[]; + } if (typeof r.reportMode === "boolean") chat.reportMode = r.reportMode; if (r.checkpoint !== undefined) chat.checkpoint = parseCheckpoint(r.checkpoint); if (typeof r.model === "string") chat.model = r.model; @@ -174,6 +181,7 @@ export function savedChatsEqual(a: SavedChat, b: SavedChat): boolean { detached: c.detached ?? false, enrichments: c.enrichments ?? {}, report: c.report ?? "", + reportSources: c.reportSources ?? [], reportMode: c.reportMode ?? false, checkpoint: c.checkpoint ?? null, }); diff --git a/export/app/ask/citations.test.ts b/export/app/ask/citations.test.ts @@ -0,0 +1,94 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + linkifyCitations, + parseClock, + resolveCitationSeconds, +} from "./citations"; +import type { RetrievedVideo } from "../lib/askRetrieval"; + +const CID = "r0"; + +function video(over: Partial<RetrievedVideo> = {}): RetrievedVideo { + return { + key: "ch/v1", + videoId: "v1", + title: "A video", + channel: "Chan", + uploadDate: "20200101", + snippets: [ + { clock: "0:05", seconds: 5, text: "" }, + { clock: "1:00", seconds: 60, text: "" }, + { clock: "2:30", seconds: 150, text: "" }, + ], + ...over, + }; +} + +test("parseClock: mm:ss and h:mm:ss, rejects malformed", () => { + assert.equal(parseClock("0:05"), 5); + assert.equal(parseClock("12:34"), 12 * 60 + 34); + assert.equal(parseClock("1:02:03"), 3723); + assert.equal(parseClock("5"), null); // no colon + assert.equal(parseClock("1:2:3:4"), null); // too many parts + assert.equal(parseClock("1:xx"), null); // non-numeric +}); + +test("linkifyCitations: bare [n] links with a plain anchor", () => { + const out = linkifyCitations("see point [1] here", 3, CID); + assert.equal(out, "see point [[1]](#cite-r0-1) here"); +}); + +test("linkifyCitations: [n @ mm:ss] encodes the parsed second into the anchor", () => { + const out = linkifyCitations("claim [2 @ 12:34]", 3, CID); + // Anchor tail is -<n>-<seconds>; display text keeps the full marker. + assert.equal(out, "claim [[2 @ 12:34]](#cite-r0-2-754)"); +}); + +test("linkifyCitations: tolerates h:mm:ss and stray spaces", () => { + assert.equal( + linkifyCitations("[3 @ 1:02:03]", 3, CID), + "[[3 @ 1:02:03]](#cite-r0-3-3723)", + ); + assert.equal( + linkifyCitations("[ 1 @ 0:05 ]", 3, CID), + "[[ 1 @ 0:05 ]](#cite-r0-1-5)", + ); +}); + +test("linkifyCitations: out-of-range n stays plain text", () => { + assert.equal(linkifyCitations("[5] and [2]", 3, CID), "[5] and [[2]](#cite-r0-2)"); + assert.equal(linkifyCitations("[9 @ 1:00]", 3, CID), "[9 @ 1:00]"); +}); + +test("linkifyCitations: maxN 0 is a no-op", () => { + assert.equal(linkifyCitations("[1] [2]", 0, CID), "[1] [2]"); +}); + +test("linkifyCitations: skips citations inside code spans", () => { + assert.equal( + linkifyCitations("real [1] but `code [2]` and ```\n[3]\n``` fenced", 3, CID), + "real [[1]](#cite-r0-1) but `code [2]` and ```\n[3]\n``` fenced", + ); +}); + +test("resolveCitationSeconds: no parsed time → first snippet", () => { + assert.equal(resolveCitationSeconds(video(), NaN), 5); +}); + +test("resolveCitationSeconds: snaps a parsed time to the nearest known snippet", () => { + // 58s is closest to the 60s snippet; 140s closest to 150s. + assert.equal(resolveCitationSeconds(video(), 58), 60); + assert.equal(resolveCitationSeconds(video(), 140), 150); + // An exact hit stays put. + assert.equal(resolveCitationSeconds(video(), 5), 5); +}); + +test("resolveCitationSeconds: passes a parsed time through when the source has no snippets", () => { + assert.equal(resolveCitationSeconds(video({ snippets: [] }), 42), 42); +}); + +test("resolveCitationSeconds: undefined for a missing source", () => { + assert.equal(resolveCitationSeconds(undefined, 5), undefined); + assert.equal(resolveCitationSeconds(undefined, NaN), undefined); +}); diff --git a/export/app/ask/citations.tsx b/export/app/ask/citations.tsx @@ -0,0 +1,222 @@ +"use client"; + +// Shared citation rendering for the /ask surfaces (chat answers AND the report). +// +// A model answer/report cites sources with bracketed markers: the classic `[n]` +// and the precise-moment `[n @ mm:ss]` (also tolerated: `[n @ h:mm:ss]` and stray +// spaces). `linkifyCitations` rewrites those into in-page Markdown links whose +// anchor encodes the source number and, when present, the cited second; +// `makeCitationLink` resolves a clicked anchor to `sources[n-1]` and opens the +// transcript modal at that moment (snapped to the nearest real line so a slightly +// off model timestamp still lands on a transcript cue), falling back to a +// scroll-to-source when there's no player (hub /ask). `CitationSources` is the +// numbered source list rendered beneath the document. Both the chat and the report +// reuse this module so their citations behave identically. + +import type { AnchorHTMLAttributes } from "react"; +import type { RetrievedVideo } from "../lib/askRetrieval"; + +// Parse a clock string (`mm:ss` or `h:mm:ss`) to whole seconds; null if malformed. +export function parseClock(s: string): number | null { + const parts = s.split(":"); + if (parts.length < 2 || parts.length > 3) return null; + let total = 0; + for (const p of parts) { + const n = parseInt(p, 10); + if (!Number.isFinite(n) || !/^\d+$/.test(p)) return null; + total = total * 60 + n; + } + return total; +} + +// One citation marker: a source number, optional `@ mm:ss` moment, tolerating +// stray spaces around the number, the `@`, and the time. +const CITE_RE = /\[\s*(\d+)\s*(?:@\s*(\d{1,2}(?::\d{2}){1,2})\s*)?\]/g; + +// Turn `[n]` / `[n @ mm:ss]` citation markers into links to the matching source in +// the list below. Skips fenced/inline code, and only links numbers that have a +// source (1..maxN). The anchor id is scoped per surface via `cid`; a parsed moment +// is encoded as a trailing `-<seconds>` segment so the click can seek to it. +export function linkifyCitations(md: string, maxN: number, cid: string): string { + if (maxN <= 0) return md; + return md + .split(/(```[\s\S]*?```|`[^`]*`)/g) + .map((seg, i) => + i % 2 === 1 + ? seg + : seg.replace(CITE_RE, (m, d, t) => { + const n = parseInt(d, 10); + if (!(n >= 1 && n <= maxN)) return m; + const secs = typeof t === "string" ? parseClock(t) : null; + const anchor = + secs != null + ? `#cite-${cid}-${n}-${secs}` + : `#cite-${cid}-${n}`; + return `[${m}](${anchor})`; + }), + ) + .join(""); +} + +// Scroll to and briefly highlight the source `<li>` for citation `n` — the +// fallback when there's no player (hub /ask) or the source has no timestamp. +export function scrollToSource(cid: string, n: number) { + const el = document.getElementById(`cite-${cid}-${n}`); + if (!el) return; + el.scrollIntoView({ behavior: "smooth", block: "center" }); + el.classList.add("bg-brand-soft"); + setTimeout(() => el.classList.remove("bg-brand-soft"), 900); +} + +// Resolve the second a citation should seek to. A parsed `mm:ss` is validated by +// snapping to the nearest known snippet second for that source (guards a +// hallucinated timestamp); passed through as-is if the source has no snippets. No +// parsed time → the source's first matched line (today's behavior). Undefined when +// there's nothing to seek to. +export function resolveCitationSeconds( + src: RetrievedVideo | undefined, + parsed: number, +): number | undefined { + if (!src) return undefined; + const snippets = src.snippets ?? []; + if (Number.isFinite(parsed)) { + if (snippets.length === 0) return parsed; + let best = snippets[0].seconds; + let bestDist = Math.abs(best - parsed); + for (const s of snippets) { + const dist = Math.abs(s.seconds - parsed); + if (dist < bestDist) { + bestDist = dist; + best = s.seconds; + } + } + return best; + } + return snippets[0]?.seconds; +} + +// Build the link renderer for citation Markdown, closing over this surface's +// sources + citation-scope id + the (optional) modal opener. A `#cite-<cid>-<n>` +// (optionally `-<seconds>`) anchor opens the transcript modal at the cited moment +// when a player is available; otherwise it scrolls to the source below. Any other +// link opens externally. +export function makeCitationLink( + cid: string, + sources: RetrievedVideo[], + onOpenCitation?: (slug: string, seconds?: number) => void, +) { + return function CitationLink({ + href, + children, + ...rest + }: AnchorHTMLAttributes<HTMLAnchorElement>) { + if (typeof href === "string" && href.startsWith("#cite-")) { + return ( + <a + href={href} + onClick={(e) => { + e.preventDefault(); + // The anchor tail is `-<n>` or `-<n>-<seconds>`; the cid always starts + // with a letter, so the dash before it never precedes a digit and this + // captures the number (+ optional second) unambiguously. + const mm = href.match(/-(\d+)(?:-(\d+))?$/); + const n = mm ? parseInt(mm[1], 10) : NaN; + const parsed = mm && mm[2] ? parseInt(mm[2], 10) : NaN; + const src = Number.isFinite(n) ? sources[n - 1] : undefined; + const seconds = resolveCitationSeconds(src, parsed); + if (onOpenCitation && src && typeof seconds === "number") { + onOpenCitation(src.key, seconds); + } else { + scrollToSource(cid, n); + } + }} + className="font-mono text-brand no-underline hover:underline" + > + {children} + </a> + ); + } + return ( + <a + href={href} + target="_blank" + rel="noopener noreferrer" + className="text-brand underline decoration-brand/40 hover:decoration-brand" + {...rest} + > + {children} + </a> + ); + }; +} + +// The numbered source list rendered beneath a cited document — one `<li>` per +// source (its `#cite-<cid>-<n>` anchor is the scroll-to-source target) with a +// per-timestamp button that opens the transcript modal at that moment (plain text +// when there's no player). Shared by the chat answer and the report so both lists +// look and behave identically. +export function CitationSources({ + cid, + sources, + onOpenCitation, + truncated, +}: { + cid: string; + sources: RetrievedVideo[]; + onOpenCitation?: (slug: string, seconds?: number) => void; + truncated?: boolean; +}) { + if (sources.length === 0) return null; + return ( + <ol className="ml-1 flex flex-col gap-1 text-xs text-muted-foreground"> + {sources.map((s, si) => ( + <li + key={s.key} + id={`cite-${cid}-${si + 1}`} + className="scroll-mt-20 rounded transition-colors animate-in fade-in motion-reduce:animate-none" + style={{ animationDelay: `${Math.min(si, 8) * 40}ms`, animationFillMode: "both" }} + > + <span className="font-mono text-brand">[{si + 1}]</span>{" "} + {s.url ? ( + <a + href={s.url} + target="_blank" + rel="noopener noreferrer" + className="hover:underline" + > + {s.title} + </a> + ) : ( + s.title + )}{" "} + <span className="text-muted-foreground/70"> + — {s.channel} + {s.siteTitle ? ` · ${s.siteTitle}` : ""} ·{" "} + {s.snippets.map((sn, sni) => ( + <span key={sn.seconds}> + {sni > 0 ? ", " : ""} + {onOpenCitation ? ( + <button + type="button" + onClick={() => onOpenCitation(s.key, sn.seconds)} + title="Open the transcript at this moment" + className="font-mono text-brand transition-colors hover:underline" + > + {sn.clock} + </button> + ) : ( + sn.clock + )} + </span> + ))} + </span> + </li> + ))} + {truncated && ( + <li className="text-muted-foreground/60"> + (some searches were truncated; ask more specifically for fuller coverage) + </li> + )} + </ol> + ); +} diff --git a/export/app/ask/useAskChat.ts b/export/app/ask/useAskChat.ts @@ -33,7 +33,9 @@ import { buildApiMessages, buildGroundedContent, chunk, + mergeReportSources, parseContext, + reportNumbering, serializeContext, type SearchStep, type UiMessage, @@ -90,6 +92,10 @@ type PersistedConvo = { // so a long research session's report survives reloads. report?: string; reportMode?: boolean; + // The report's SOURCE REGISTRY — every video the report's `[n]` citations index + // into, ordered + deduped, trimmed (no snippet text). Persisted alongside the + // report so its clickable citations survive a reload. + reportSources?: RetrievedVideo[]; // A paused sweep's checkpoint (remaining batches + directive), so a paused sweep // survives a reload and can be resumed — even under a different model. checkpoint?: SweepCheckpoint | null; @@ -177,6 +183,11 @@ export function useAskChat() { // model maintains a persistent markdown report via update_report, and the // conversation is compacted into that report instead of replaying every turn. const [report, setReport] = useState(""); + // The report's source registry: every video cited across all report writes + // (report-mode turns + sweep batches), ordered + deduped by key, trimmed. The + // report's `[n]`/`[n @ mm:ss]` citations resolve against this. Grows as writes + // land; cleared on New chat / a fresh sweep. + const [reportSources, setReportSources] = useState<RetrievedVideo[]>([]); const [reportMode, setReportModeState] = useState(false); // True while a turn is writing to the report (drives the panel shimmer); // cleared when the turn ends. @@ -273,6 +284,8 @@ export function useAskChat() { strictRef.current = strictGrounding; const reportRef = useRef(report); reportRef.current = report; + const reportSourcesRef = useRef(reportSources); + reportSourcesRef.current = reportSources; const reportModeRef = useRef(reportMode); reportModeRef.current = reportMode; const maxAnswerTokensRef = useRef(maxAnswerTokens); @@ -346,6 +359,9 @@ export function useAskChat() { setEnrichments(parsed.enrichments); } if (typeof parsed.report === "string") setReport(parsed.report); + if (Array.isArray(parsed.reportSources)) { + setReportSources(parsed.reportSources); + } if (typeof parsed.reportMode === "boolean") { setReportModeState(parsed.reportMode); } @@ -399,6 +415,7 @@ export function useAskChat() { detached, enrichments, report, + reportSources, reportMode, checkpoint: sweepCheckpoint, }) @@ -431,6 +448,7 @@ export function useAskChat() { detached, enrichments, report, + reportSources, reportMode, sweepCheckpoint, ]); @@ -729,6 +747,7 @@ export function useAskChat() { seedVideos: args.seedVideos, report: reportRef.current, reportMode: effectiveReportMode, + reportSources: reportSourcesRef.current, maxAnswerTokens: maxAnswerTokensRef.current, onDebug: pushDebug, precomputedGrounding: args.precomputedGrounding?.groundedContent @@ -743,6 +762,11 @@ export function useAskChat() { if (effectiveReportMode && result.report !== reportRef.current) { setReport(result.report); } + // Grow the report source registry with this turn's gathered videos (in the + // global-index order the report writes cited them by), trimmed for storage. + if (effectiveReportMode && result.reportSources) { + setReportSources(mergeReportSources([], result.reportSources)); + } patchAt(assistantIndex, (m) => ({ ...m, content: m.content || result.answer, @@ -899,6 +923,18 @@ export function useAskChat() { return; } setReportUpdating(true); + // Merge this batch into the report source registry BEFORE building its + // context, so its excerpts are numbered by their GLOBAL registry index — + // making the report's `[n]` stable across every batch and resolving + // against the persisted source list. Keep the ref current synchronously + // so the next batch numbers on top of this one. + const mergedSources = mergeReportSources( + reportSourcesRef.current, + batches[i], + ); + reportSourcesRef.current = mergedSources; + setReportSources(mergedSources); + const numberOf = reportNumbering(mergedSources); let r; try { r = await runReportChunk({ @@ -911,6 +947,7 @@ export function useAskChat() { count: total, report, aliases, + numberOf, signal: ac.signal, onEvent: (e) => { if (e.type === "report_update") { @@ -1025,7 +1062,13 @@ export function useAskChat() { ]); const startReport = mode === "fresh" ? "" : reportRef.current; - if (mode === "fresh") setReport(""); + if (mode === "fresh") { + setReport(""); + // A fresh report starts from an empty source registry too (a new document's + // `[n]` must not resolve against the previous report's sources). + setReportSources([]); + reportSourcesRef.current = []; + } await driveSweep({ remaining: videos, @@ -1189,6 +1232,7 @@ export function useAskChat() { setEnrichments({}); setStrictGroundingState(true); setReport(""); + setReportSources([]); setReportModeState(false); setSweepCheckpoint(null); // A new chat is a blank working conversation not derived from any saved chat — @@ -1307,6 +1351,7 @@ export function useAskChat() { detached, enrichments, report, + reportSources, reportMode, checkpoint: sweepCheckpoint, provider, @@ -1319,6 +1364,7 @@ export function useAskChat() { detached, enrichments, report, + reportSources, reportMode, sweepCheckpoint, provider, @@ -1418,6 +1464,7 @@ export function useAskChat() { setDetached(chat.detached ?? false); setEnrichments(chat.enrichments ?? {}); setReport(chat.report ?? ""); + setReportSources(chat.reportSources ?? []); setReportModeState(chat.reportMode ?? false); setSweepCheckpoint(chat.checkpoint ?? null); writeChats((s) => { @@ -1446,6 +1493,7 @@ export function useAskChat() { setMarkdownOn, // report mode report, + reportSources, reportMode, setReportMode, reportModeAvailable, diff --git a/export/app/lib/askConversation.test.ts b/export/app/lib/askConversation.test.ts @@ -9,8 +9,10 @@ import { buildGroundedContent, chunk, collectPriorPool, + mergeReportSources, renderAliasGlossary, renderUsedAliasContext, + reportNumbering, gatherSystemPrompt, answerSystemPrompt, serializeContext, @@ -255,6 +257,65 @@ test("buildAccumulationContent emits the directive, a batch label, and numbered assert.match(out, /\[1\] "First" — Rekieta/); assert.match(out, /\[0:05\] alpha here/); assert.match(out, /\[2\] "Second" — Rekieta/); + // The label steers the model to the precise-moment citation form. + assert.match(out, /\[n @ mm:ss\]/); +}); + +test("buildAccumulationContent numbers by a global registry index when given numberOf", () => { + const videos = [ + video({ key: "a", title: "First" }), + video({ key: "b", title: "Second" }), + ]; + const registry = mergeReportSources( + [video({ key: "x", title: "Earlier" }), video({ key: "y", title: "Also earlier" })], + videos, + ); + const out = buildAccumulationContent( + "d", + videos, + 2, + 3, + reportNumbering(registry), + ); + // The batch's videos are global entries 3 and 4 (after the two prior sources). + assert.match(out, /\[3\] "First" — Rekieta/); + assert.match(out, /\[4\] "Second" — Rekieta/); +}); + +test("mergeReportSources appends new videos, dedupes by key, and trims snippet text", () => { + const base = mergeReportSources([], [ + video({ key: "a", title: "A", snippets: [{ clock: "0:05", seconds: 5, text: "keep me?" }] }), + ]); + // Snippet text is trimmed for storage, but clock/seconds survive. + assert.equal(base[0].snippets[0].text, ""); + assert.equal(base[0].snippets[0].seconds, 5); + assert.equal(base[0].snippets[0].clock, "0:05"); + + const merged = mergeReportSources(base, [ + // "a" reappears with a new moment → merged into the existing entry. + video({ key: "a", title: "A", snippets: [{ clock: "1:00", seconds: 60, text: "x" }] }), + // "b" is new → appended after "a". + video({ key: "b", title: "B" }), + ]); + assert.deepEqual(merged.map((v) => v.key), ["a", "b"]); + // "a" now carries both moments (5s + 60s), deduped/merged by time. + assert.deepEqual( + merged[0].snippets.map((s) => s.seconds).sort((x, y) => x - y), + [5, 60], + ); +}); + +test("reportNumbering maps a video to its 1-based registry index by key", () => { + const registry = mergeReportSources([], [ + video({ key: "a" }), + video({ key: "b" }), + video({ key: "c" }), + ]); + const numberOf = reportNumbering(registry); + assert.equal(numberOf(video({ key: "a" }), 0), 1); + assert.equal(numberOf(video({ key: "c" }), 0), 3); + // A key not in the registry falls back to the local index + 1. + assert.equal(numberOf(video({ key: "zzz" }), 6), 7); }); test("accumulationSystemPrompt carries the update_report directive + focused alias block", () => { @@ -264,6 +325,9 @@ test("accumulationSystemPrompt carries the update_report directive + focused ali assert.match(p, /Do NOT search/); assert.match(p, /fetch_context tool/); assert.match(p, /finish tool/); + // Precise-moment, globally-stable citation form. + assert.match(p, /\[n @ mm:ss\]/); + assert.match(p, /GLOBAL and stable/); // The FOCUSED used-alias block (not the full "search the full term" glossary). assert.match(p, /TERMS USED IN THIS SEARCH/); assert.match(p, /Graham Platner: platner, plattner, platter/); @@ -292,7 +356,8 @@ test("applyReportPatch composes across sweep chunks: fold section A then B → b test("answerSystemPrompt asks for Markdown + citations and carries the FOCUSED used-alias block", () => { const p = answerSystemPrompt(ALIASES); assert.match(p, /Markdown/); - assert.match(p, /\[1\]/); + // Precise-moment citation form: [n @ mm:ss] (backward-tolerant of a bare [n]). + assert.match(p, /\[n @ mm:ss\]/); // The focused block (not the full glossary) — no "search the full term" directive. assert.match(p, /TERMS USED IN THIS SEARCH/); assert.match(p, /Graham Platner: platner, plattner, platter/); diff --git a/export/app/lib/askConversation.ts b/export/app/lib/askConversation.ts @@ -95,6 +95,61 @@ export function collectPriorPool( return order.map((k) => byKey.get(k)!); } +// ─── report source registry (what a report's global `[n]` indexes into) ─── + +// The report's persisted source registry: every video cited across all report +// writes (report-mode turns + sweep batches), ordered, deduped by `key`. A +// report's `[n]` is this list's n-th entry, and `[n @ mm:ss]` seeks to the cited +// moment. Stored TRIMMED — snippet text is dropped (only clock/seconds + the video +// meta are needed to resolve or label a citation) so a large sweep's registry +// stays storage-light. +const REPORT_SNIPPETS_CAP = 30; + +// Drop snippet text (keep clock/seconds) so the persisted registry is light. The +// RetrievedVideo shape is preserved (text becomes ""), so it still renders in the +// source list and resolves a citation's second. +function trimReportSource(v: RetrievedVideo): RetrievedVideo { + return { + ...v, + snippets: v.snippets.map((s) => ({ clock: s.clock, seconds: s.seconds, text: "" })), + }; +} + +// Fold a batch/turn's videos into the report registry: append videos not seen +// before (by `key`, first-seen order preserved) and merge snippet times into ones +// already present (a video cited again in a later batch keeps all its moments). +// Mirrors the cumulative-pool dedup in collectPriorPool. Pure + testable. +export function mergeReportSources( + existing: RetrievedVideo[], + batch: RetrievedVideo[], +): RetrievedVideo[] { + const order: string[] = existing.map((v) => v.key); + const byKey = new Map<string, RetrievedVideo>(existing.map((v) => [v.key, v])); + for (const raw of batch) { + const v = trimReportSource(raw); + const cur = byKey.get(v.key); + if (!cur) { + order.push(v.key); + byKey.set(v.key, v); + } else { + byKey.set(v.key, { + ...cur, + snippets: mergeSnippets(cur.snippets, v.snippets, REPORT_SNIPPETS_CAP), + }); + } + } + return order.map((k) => byKey.get(k)!); +} + +// A 1-based global-index lookup over a registry, keyed by video `key` — the +// numbering a report write uses so its `[n]` matches the registry's source list. +export function reportNumbering( + registry: RetrievedVideo[], +): (v: RetrievedVideo, index: number) => number { + const indexByKey = new Map(registry.map((v, i) => [v.key, i + 1])); + return (v, index) => indexByKey.get(v.key) ?? index + 1; +} + // ─── report mode (persistent markdown document, maintained via update_report) ─── // Upsert a `## <section>` block into the running report. If a section with that @@ -341,7 +396,9 @@ export function gatherSystemPrompt( ? " You are maintaining a running report document that persists across turns. " + "After gathering evidence, call the update_report tool to record findings in " + "well-titled sections (pass the section heading and its full markdown " + - "content); keep the report the source of truth. Then answer briefly." + "content), citing each source as [n @ mm:ss] — the excerpt's bracketed " + + "number plus the cited line's timestamp; keep the report the source of " + + "truth. Then answer briefly." : ""; const tail = mode === "native" @@ -367,8 +424,10 @@ export function answerSystemPrompt(usedAliases: SearchAlias[]): string { "answer on the transcript excerpts provided in this conversation. Excerpts " + "may appear in earlier turns, and a follow-up request (reformatting, " + "summarising, expanding) should reuse the relevant excerpts already " + - "provided. Cite the excerpts you use by their bracketed number, e.g. [1]. " + - "Each excerpt shows a video title, channel, and timestamped lines. If the " + + "provided. Cite each excerpt you use as [n @ mm:ss]: its bracketed number " + + "plus the specific line's timestamp (e.g. [1 @ 4:12]); a bare [n] is fine " + + "when no single line applies. Each excerpt shows a video title, channel, and " + + "timestamped lines. If the " + "excerpts do not contain enough to answer, say so plainly rather than " + "guessing. Format your answer in GitHub-flavored Markdown."; const context = renderUsedAliasContext(usedAliases); @@ -397,29 +456,35 @@ export function accumulationSystemPrompt(usedAliases: SearchAlias[]): string { "You are building a running report by reading transcript excerpts in " + "batches. Cross-reference this batch against the report so far. Use the " + "update_report tool to add or merge findings — claims, contradictions with " + - "earlier claims, and their sources (title + timestamp, cited by [n]) — into " + - "well-titled sections, keeping the report the single source of truth. Do NOT " + - "search; these excerpts are your evidence, but you may call the fetch_context " + - "tool to read more transcript around a specific line when one is too thin to " + - "judge. Call the finish tool when you are done with this batch."; + "earlier claims, and their sources — into well-titled sections, keeping the " + + "report the single source of truth. CITE each source as [n @ mm:ss]: the " + + "excerpt's bracketed number plus the specific line's timestamp (e.g. " + + "[3 @ 12:34]); the numbers are GLOBAL and stable across batches, so reuse a " + + "source's number if it reappears. Do NOT search; these excerpts are your " + + "evidence, but you may call the fetch_context tool to read more transcript " + + "around a specific line when one is too thin to judge. Call the finish tool " + + "when you are done with this batch."; const context = renderUsedAliasContext(usedAliases); return context ? `${base}\n\n${context}` : base; } // The per-chunk user message for a sweep: the report directive, a "Batch i of n" -// label, and this batch's numbered excerpts (same citation format as a grounded -// turn, reusing buildContext). Numbering restarts per batch — the raw excerpts -// are discarded once the batch folds into the report, so cross-batch [n] identity -// isn't needed. +// label, and this batch's excerpts (same format as a grounded turn, reusing +// buildContext). `numberOf` labels each excerpt by its GLOBAL registry index so a +// report's `[n]` is stable across every batch — the source registry the report's +// citations resolve against persists even though the raw excerpts are discarded +// once the batch folds in. Omitted → 1-based per-batch numbering (legacy). export function buildAccumulationContent( directive: string, videos: RetrievedVideo[], index: number, count: number, + numberOf?: (video: RetrievedVideo, i: number) => number, ): string { return ( `${directive}\n\n` + - `Batch ${index} of ${count} — transcript excerpts you may cite (by number):\n` + - buildContext(videos) + `Batch ${index} of ${count} — transcript excerpts (cite each as [n @ mm:ss]: ` + + `the bracketed number + the line's timestamp):\n` + + buildContext(videos, numberOf) ); } diff --git a/export/app/lib/askRetrieval.test.ts b/export/app/lib/askRetrieval.test.ts @@ -226,3 +226,34 @@ test("buildContext numbers videos and indents snippets", () => { assert.match(ctx, /\[1\] "Travel Bans" — Rekieta/); assert.match(ctx, /\[0:05\] about travel bans/); }); + +test("buildContext honours a numberOf override (global registry index)", () => { + const videos = [ + { + key: "a", + videoId: "a", + title: "First", + channel: "Rekieta", + uploadDate: "20200101", + snippets: [{ clock: "0:05", seconds: 5, text: "alpha" }], + }, + { + key: "b", + videoId: "b", + title: "Second", + channel: "Rekieta", + uploadDate: "20200101", + snippets: [{ clock: "1:00", seconds: 60, text: "beta" }], + }, + ]; + // Number the batch as if it were entries 41/42 of a large sweep's registry. + const index = new Map([ + ["a", 41], + ["b", 42], + ]); + const ctx = buildContext(videos, (v) => index.get(v.key)!); + assert.match(ctx, /\[41\] "First" — Rekieta/); + assert.match(ctx, /\[42\] "Second" — Rekieta/); + // The default 1-based numbering is NOT used. + assert.doesNotMatch(ctx, /\[1\] "First"/); +}); diff --git a/export/app/lib/askRetrieval.ts b/export/app/lib/askRetrieval.ts @@ -324,12 +324,19 @@ export function retrieve( } // Assemble numbered retrieved excerpts into the context block appended to the -// user's question (unchanged contract — the system prompt cites by [n]). -export function buildContext(videos: RetrievedVideo[]): string { +// user's question (the system prompt cites by [n]). By default each video is +// numbered 1-based by its position (the chat/grounded path). `numberOf` overrides +// this so a report write can number by GLOBAL registry index — making a report's +// `[n]` stable across every batch/turn folded into it, not just within one batch. +export function buildContext( + videos: RetrievedVideo[], + numberOf?: (video: RetrievedVideo, index: number) => number, +): string { if (videos.length === 0) return "(no matching transcript excerpts were found)"; return videos .map((v, i) => { - const head = `[${i + 1}] "${v.title}" — ${v.channel}${v.siteTitle ? ` (${v.siteTitle})` : ""}`; + const n = numberOf ? numberOf(v, i) : i + 1; + const head = `[${n}] "${v.title}" — ${v.channel}${v.siteTitle ? ` (${v.siteTitle})` : ""}`; const lines = v.snippets .map((s) => ` [${s.clock}] ${s.text}`) .join("\n"); diff --git a/export/app/lib/searchAgent.ts b/export/app/lib/searchAgent.ts @@ -168,13 +168,22 @@ async function scriptedGather( // ─── result formatting fed back to the model during gather ─── -export function formatResultsForModel(videos: RetrievedVideo[]): string { +// Feed a search's results back to the model during gather. `numberOf` overrides +// the per-result 1-based index (report mode uses it to show the GLOBAL registry +// number, so a report-mode turn's report write cites the same stable `[n]` the +// registry's source list uses); the "n." lead-in then becomes a bracketed "[n]" +// so the model reads it as the citation number. +export function formatResultsForModel( + videos: RetrievedVideo[], + numberOf?: (video: RetrievedVideo, index: number) => number, +): string { if (videos.length === 0) return "No transcript excerpts matched that query."; return videos .map((v, i) => { const site = v.siteTitle ? ` (${v.siteTitle})` : ""; const date = v.uploadDate ? ` [${v.uploadDate}]` : ""; - const head = `${i + 1}. "${v.title}" — ${v.channel}${site}${date}`; + const lead = numberOf ? `[${numberOf(v, i)}]` : `${i + 1}.`; + const head = `${lead} "${v.title}" — ${v.channel}${site}${date}`; const lines = v.snippets.map((s) => ` ${s.clock} ${s.text}`).join("\n"); return lines ? `${head}\n${lines}` : head; }) @@ -246,6 +255,12 @@ export type RunAskTurnOptions = { // and the update_report tool is offered so the model keeps the report current. report?: string; reportMode?: boolean; + // Report mode: the persisted report SOURCE REGISTRY carried in — every video + // cited across the report so far, ordered + deduped by key. The turn numbers its + // gathered videos by their GLOBAL index in this registry (so a report write cites + // the same stable `[n]` the registry's source list uses) and returns the registry + // grown with this turn's new videos (see AskTurnResult.reportSources). + reportSources?: RetrievedVideo[]; // Max output tokens for the answer phase (default DEFAULT_ANSWER_TOKENS). maxAnswerTokens?: number; // Optional capture of each provider call (finish reason, payloads) for the @@ -261,6 +276,10 @@ export type AskTurnResult = { // The report after this turn (possibly updated via update_report). Echoes the // input report unchanged when report mode is off or nothing was written. report: string; + // Report mode: the source registry grown with this turn's gathered videos, in + // the GLOBAL-index order the turn's report writes cited them by. Undefined when + // report mode is off (the caller then leaves its registry untouched). + reportSources?: RetrievedVideo[]; // The answer's normalized finish reason and (if blocked) the raw block reason, // so the UI can flag truncation / a safety block instead of a silent stop. finishReason?: FinishReason; @@ -289,6 +308,23 @@ export async function runAskTurn( // returned in the result so the caller can persist it. let report = opts.report ?? ""; + // Report mode: number this turn's gathered videos by their GLOBAL index in the + // carried-in report source registry, so a report write cites the same stable + // `[n]` the registry's source list uses (and `[n @ mm:ss]` seeks the moment). + // `reportOrder` accumulates the registry (base + this turn's new videos, in + // first-seen discovery order) and is returned so the caller can persist it. + const reportBase = opts.reportSources ?? []; + const reportOrder: RetrievedVideo[] = [...reportBase]; + const reportIndexByKey = new Map(reportBase.map((v, i) => [v.key, i + 1])); + const reportNumberOf = (v: RetrievedVideo): number => { + const existing = reportIndexByKey.get(v.key); + if (existing != null) return existing; + reportOrder.push(v); + const n = reportOrder.length; + reportIndexByKey.set(v.key, n); + return n; + }; + // Capture the answer phase's finish reason (+ any block reason) so the result // can flag truncation/blocking instead of a silent stop; also forward every // provider call to the debug sink for the export. @@ -358,6 +394,8 @@ export async function runAskTurn( truncated: (opts.precomputedGrounding.truncated ?? false) || answerFinish === "length", report, + // Strict/retry path skips gather → no report write → registry unchanged. + reportSources: reportMode ? reportBase : undefined, finishReason: answerFinish, blockReason: answerBlock, }; @@ -405,7 +443,9 @@ export async function runAskTurn( if (r.truncated) truncated = true; queries.push(query); onEvent({ type: "search_done", query, count: r.videos.length }); - return formatResultsForModel(r.videos); + // Report mode shows the GLOBAL registry number so the report write's `[n]` + // resolves against the persisted source list. + return formatResultsForModel(r.videos, reportMode ? reportNumberOf : undefined); }; // Expand mode with a pinned set: let the model read more of a grounding video's @@ -548,6 +588,9 @@ export async function runAskTurn( groundedContent, truncated: truncated || answerFinish === "length", report, + // The registry grown with this turn's gathered videos, in the global-index + // order the report writes cited them by. + reportSources: reportMode ? reportOrder : undefined, finishReason: answerFinish, blockReason: answerBlock, }; @@ -569,6 +612,11 @@ export type RunReportChunkOptions = { // The running report carried in; the returned report threads to the next chunk. report: string; aliases: SearchAlias[]; + // Number each batch video by its GLOBAL registry index (so the report's `[n]` is + // stable across every batch, resolving against the persisted source list). The + // caller merges the batch into its registry first, then passes this. Omitted → + // 1-based per-batch numbering. + numberOf?: (video: RetrievedVideo, index: number) => number; signal?: AbortSignal; onEvent: (e: AgentEvent) => void; onDebug?: (rec: DebugCall) => void; @@ -671,7 +719,13 @@ export async function runReportChunk( const system = `${accumulationSystemPrompt([...usedAliases.values()])}\n\n` + buildSeedDigest(opts.videos); - const question = buildAccumulationContent(directive, opts.videos, index, count); + const question = buildAccumulationContent( + directive, + opts.videos, + index, + count, + opts.numberOf, + ); const ctx: NativeGatherContext = { apiKey, diff --git a/export/e2e/ask-chat.spec.ts b/export/e2e/ask-chat.spec.ts @@ -890,6 +890,71 @@ test.describe("ask chat", () => { await expect(page.getByText("Recorded finding 4.")).toHaveCount(0); }); + test("report citations are clickable and open the transcript at the cited moment", async ({ + page, + }) => { + await installRoutes(page); + await page.route("https://api.anthropic.com/**", async (route) => { + if (route.request().method() === "OPTIONS") { + await route.fulfill({ status: 204, headers: CORS }); + return; + } + const body = route.request().postDataJSON() as { + system?: string; + tools?: { name?: string }[]; + messages?: { role: string; content: unknown }[]; + }; + const system = body.system ?? ""; + const msgs = body.messages ?? []; + const lastUser = [...msgs].reverse().find((m) => m.role === "user"); + const isToolResult = Array.isArray(lastUser?.content); + // Sweep gather turn: fold in ONE section citing a precise moment, then finish. + if (Array.isArray(body.tools) && /running report/.test(system)) { + await route.fulfill({ + status: 200, + headers: { ...CORS, "content-type": "application/json" }, + body: isToolResult + ? toolUse("finish", {}) + : toolUse("update_report", { + section: "Findings", + content: "The archive discusses this at length [1 @ 0:05].", + }), + }); + return; + } + await route.fulfill({ + status: 200, + headers: { ...CORS, "content-type": "text/event-stream" }, + body: sse("DONE"), + }); + }); + + // Sweep the whole search (3 fixtures) into a report in a single batch. + await searchThenChat(page, ALL_RESULT_TREE, 3); + await keyIn(page, "Native tools"); + await expect(page.getByText(/Grounded in 3 results/)).toBeVisible(); + await page + .getByRole("button", { name: /Build report from all 3 results/ }) + .click(); + await expect(page.getByText(/Built a report from 3 results/)).toBeVisible(); + + // The report renders the [n @ mm:ss] marker as an in-page citation anchor, + // with a numbered Sources list beneath the document. + const cite = page.getByRole("link", { name: "[1 @ 0:05]", exact: true }); + await expect(cite).toBeVisible(); + await expect(cite).toHaveAttribute("href", /^#cite-/); + await expect( + page.getByRole("listitem").filter({ hasText: /alpha|Transcript|Chat/ }).first(), + ).toBeVisible(); + + // Clicking it opens the transcript modal seeked to the cited second (5s) — + // the same modal a chat citation opens, via replaceState on the /ask route. + await cite.click(); + await expect(page).toHaveURL(/[?&]v=/); + await expect(page).toHaveURL(/[?&]t=5\b/); + await expect(page.getByRole("button", { name: "Close player" })).toBeVisible(); + }); + test("corpus sweep: Stop aborts mid-sweep and keeps the partial report", async ({ page, }) => {