commit 4782bea368f426a1f31f0133910023c725a6d0da
parent d0097cc6a45878a5e4c5bbb9cdc0aa6045aecf56
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Wed, 22 Jul 2026 11:00:27 -0400
Ask chat: clickable precise-moment report citations — [n @ mm:ss] + a persisted source registry
The Report panel now mirrors the chat: citation markers render as in-page
links that open the transcript modal seeked to the cited line (snapped to
the nearest real snippet second), with a numbered Sources list beneath.
Chat citations upgraded to the same [n @ mm:ss] form. New shared renderer
in citations.tsx; reportSources registry persisted with the convo + saved
chats; sweep batches numbered by global registry index.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Diffstat:
15 files changed, 772 insertions(+), 157 deletions(-)
diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md
@@ -1,6 +1,7 @@
# Changelog
## [Unreleased]
+- **"Ask AI" — the Report panel's citations are now clickable, precise to the exact moment cited.** The chat's answer bubbles already linked their `[1]`, `[2]`… markers to the cited transcript; the **Report** panel — the persistent document Report mode and whole-corpus sweeps maintain — rendered as plain text, so the citations a user works from most were dead. Now the report mirrors the chat: it renders each citation as an in-page link and lists its **numbered sources** beneath the document, and clicking one **opens the transcript modal seeked to the cited line**. Citations are also upgraded — in the report *and* the chat — to a precise-moment **`[n @ mm:ss]`** form (the source's number plus the specific line's timestamp, e.g. `[3 @ 12:34]`), so a citation jumps to the exact moment instead of the video's first matched snippet. A slightly-off model timestamp is **snapped to the nearest real transcript line**, and it stays graceful: a bare `[n]` still resolves to the first snippet, and an out-of-range number or unparseable time stays plain text rather than becoming a broken link. Making this work in the report (a single accumulated string, unlike a chat message that carries its own sources) required two new pieces: a **persisted report source registry** — every video the report cites, deduped and stored light (no excerpt text) — and **global citation numbering** so a report's `[n]` is stable across every batch/turn folded in, not renumbered per write. The registry is saved with the conversation and each saved chat, so a report's clickable citations survive a reload; New chat clears it; and on the federated hub /ask (no player) citations fall back to scroll-to-source, exactly like the chat. See `export/app/ask/citations.tsx` (new shared renderer — `linkifyCitations` + `CitationLink` + `CitationSources`, with `[n @ mm:ss]` parsing and nearest-snippet snapping), `export/app/ask/{MessageBubble,ReportPanel,AskChat,useAskChat,askChatStorage}.tsx/ts`, `export/app/lib/{askRetrieval,askConversation,searchAgent}.ts` (optional/global-registry numbering + `mergeReportSources` + the `[n @ mm:ss]` prompt updates), and `export/e2e/ask-chat.spec.ts`.
- **MCP server: switch which corpus you're reading on the fly — and it sticks across reconnects.** The MCP server was pinned to a single corpus at launch (`--hub` / `--remote` / `--local` or the `TRANSCRIPT_*` env), so re-aiming it at a different site or a local build meant editing the MCP config and restarting. Now, connected to (say) the **archilyzer hub**, you can retarget it from inside a session with three new **read-only** tools: **`list_sources`** shows the active corpus and, in a hub context, its member sites (`siteId · title · url`); **`use_source`** switches the active corpus — to a single hub member (`site:` — which becomes a plain single-site source and so regains **full channel-group + alias** support), a **federated subset** of the hub (`sites:[…]`, unknown tokens reported not dropped), or an arbitrary `remote:` URL / `local:` dir / `hub:` URL; and **`reset_source`** returns to the startup source. The selection is **persisted** to a small state file (under `$XDG_STATE_HOME/yt-dlp-transcript-mcp`, overridable with `TRANSCRIPT_MCP_STATE_DIR`) keyed by the **startup** source, so it survives a reconnect and two differently-configured servers keep independent selections. Everything the sweep/search tools do reads the current active source. Still strictly **read-only** — this only changes *which* already-published static shards are read, the same capability the startup flags already grant this locally-run tool; nothing is written to any corpus. See `mcp/src/{sources,sourceController,source,server,index}.ts`, `mcp/src/sourceController.test.ts`, and `mcp/README.md`.
- **MCP server: run the corpus sweep through Claude Code itself — on plan usage, no API key.** The in-browser "corpus sweep" (fold every matching transcript into a running report) needs a BYO AI key, and a free key conks out fast on a real 30k-video corpus. The `mcp/` server now lets **Claude Code be the sweep engine** instead, via a first-class **`sweep` prompt** (a slash command, `/mcp__<name>__sweep query="k cups" channel="chrissie-mayr"`) plus the tools to drive it. `search_transcripts` gains **paging** — it reports the full `total` and `has_more`, so the whole match set can be enumerated with `offset` (and `include_snippets:false` for a cheap worklist) — and is now **alias-aware**: a plain query that matches a curated search alias also searches the alias regex (e.g. `k cups` → also `cake cup`), with the footer naming which aliases fired; reaching the scan cap is surfaced as **PARTIAL** coverage rather than hidden. A new **`get_transcripts`** tool batch-reads up to 20 videos in one call as bounded, timestamped **excerpt windows** around the matches (alias-correct) — or full transcripts without a query — so a sweep stays token-bounded. The `sweep` prompt walks Claude through search → enumerate → plan `ceil(N/batch)` batches → per-batch windowed read + cross-referenced upsert of cited findings (*title + [mm:ss]*) into a markdown report it maintains with its own Write/Edit tools. The sweep is now **group-aware and multi-channel**: `search_transcripts` scope is additive over one-or-more channels (`channel`/`channels`) and/or channel **groups** (`group`/`groups`, matched by group **id or display name** — `group="other"` ≡ `group="Extended Universe"` — and expanded to the group's channels), with the footer naming the resolved scope and flagging any channel/group token that matched nothing so a typo isn't silently a whole-corpus scan; `get_transcripts` takes matching `channels` lookup hints; and `list_channels` now organizes channels under their groups with a compact `id · name · N channels` cheat-sheet. Crucially the `sweep` prompt is now **pick-first**: invoked with **no** scope it makes Claude list the groups/channels and **ask which to sweep — or to confirm the whole corpus — before enumerating**, rather than silently scanning everything (an explicit `search_transcripts` call with no selector still means "all"). (Hub mode defers per-site group resolution — multi-channel scoping still works there.) The MCP stays strictly **read-only**; only the report file is written, in Claude's own working directory. See `mcp/src/{search,source,server}.ts`, `mcp/src/search.test.ts`, and `mcp/README.md`.
- **"Ask AI" now paces itself to your key's rate limit instead of failing.** A whole-corpus sweep on a free-tier key used to fire provider calls as fast as the loop could produce them, blow straight through the per-minute request cap, and stop dead on the first HTTP 429 (and a rate-limit *mid-answer* surfaced as a hard error, because only the sweep path ever handled 429). Now a single client-side limiter sits behind **every** AI call: it **spaces requests** to a conservative, free-tier-safe **requests-per-minute** default chosen per model (e.g. Gemini `*-pro` → 5/min, `*flash` → 10/min; Claude/OpenAI higher), so a paid key runs fast and a free key just runs *slowly* rather than erroring. When a 429 does land, the limiter **honours the provider's own retry hint** (the `Retry-After` header, or Gemini's `RetryInfo` retry-delay) — or an exponential backoff — and **re-issues the request** (safe for streaming: the retry happens before any answer text is emitted), and it **self-lowers** the rate after a 429 so later calls pace slower. Only a 429 that outlasts the retries falls through to the existing **pause/checkpoint** (the right home for a daily-quota cap you resume tomorrow). Provider settings gain an editable **Requests per minute** field (with the effective spacing, e.g. *"≈12s between requests"*, and a reset-to-default), and a paced sweep shows a **"Rate limited — retrying in Ns"** line in its progress strip so it never looks frozen. See `export/app/lib/rateLimit.ts` (the whole mechanism), `export/app/lib/askProvider.ts` (`PausableError` + `acquire`/retry in `askStream`, 429 in `ensureOk`), `export/app/lib/nativeTools/shared.ts` (`acquire`/retry in `postJson`), `export/app/ask/{useAskChat,ProviderSettings,PinnedResultsPanel,AskChat}.tsx`, and `export/e2e/ask-workspace.spec.ts`.
diff --git a/export/app/ask/AskChat.tsx b/export/app/ask/AskChat.tsx
@@ -276,6 +276,8 @@ export default function AskChat() {
<ReportPanel
report={s.report}
+ sources={s.reportSources}
+ onOpenCitation={onOpenCitation}
reportMode={s.reportMode}
setReportMode={s.setReportMode}
available={s.reportModeAvailable}
diff --git a/export/app/ask/MessageBubble.tsx b/export/app/ask/MessageBubble.tsx
@@ -1,90 +1,12 @@
"use client";
-import { useId, useState, type AnchorHTMLAttributes } from "react";
+import { useId, useState } from "react";
import { CheckIcon, CopyIcon, PencilIcon, RotateCwIcon } from "lucide-react";
import { Markdown } from "yt-dlp-transcript-common/components/Markdown";
import type { UiMessage } from "../lib/askConversation";
-import type { RetrievedVideo } from "../lib/askRetrieval";
+import { CitationSources, linkifyCitations, makeCitationLink } from "./citations";
import { PipelineStatus } from "./PipelineStatus";
-// Turn `[n]` citation markers in an answer into links to the matching source in
-// the list below. Skips fenced/inline code, and only links numbers that have a
-// source (1..count). The anchor id is scoped per message via `cid`.
-function linkifyCitations(md: string, maxN: number, cid: string): string {
- if (maxN <= 0) return md;
- return md
- .split(/(```[\s\S]*?```|`[^`]*`)/g)
- .map((seg, i) =>
- i % 2 === 1
- ? seg
- : seg.replace(/\[(\d+)\]/g, (m, d) => {
- const n = parseInt(d, 10);
- return n >= 1 && n <= maxN ? `[[${n}]](#cite-${cid}-${n})` : m;
- }),
- )
- .join("");
-}
-
-// Scroll to and briefly highlight the source `<li>` for citation `n` — the
-// fallback when there's no player (hub /ask) or the source has no timestamp.
-function scrollToSource(cid: string, n: number) {
- const el = document.getElementById(`cite-${cid}-${n}`);
- if (!el) return;
- el.scrollIntoView({ behavior: "smooth", block: "center" });
- el.classList.add("bg-brand-soft");
- setTimeout(() => el.classList.remove("bg-brand-soft"), 900);
-}
-
-// Build the link renderer for answer Markdown, closing over this message's
-// sources + citation-scope id + the (optional) modal opener. A `#cite-<cid>-<n>`
-// anchor opens the transcript modal at the cited video's first matched line when
-// a player is available; otherwise it scrolls to the source below. Any other
-// link opens externally.
-function makeCitationLink(
- cid: string,
- sources: RetrievedVideo[],
- onOpenCitation?: (slug: string, seconds?: number) => void,
-) {
- return function CitationLink({
- href,
- children,
- ...rest
- }: AnchorHTMLAttributes<HTMLAnchorElement>) {
- if (typeof href === "string" && href.startsWith("#cite-")) {
- return (
- <a
- href={href}
- onClick={(e) => {
- e.preventDefault();
- const n = parseInt(href.match(/-(\d+)$/)?.[1] ?? "", 10);
- const src = Number.isFinite(n) ? sources[n - 1] : undefined;
- const seconds = src?.snippets[0]?.seconds;
- if (onOpenCitation && src && typeof seconds === "number") {
- onOpenCitation(src.key, seconds);
- } else {
- scrollToSource(cid, n);
- }
- }}
- className="font-mono text-brand no-underline hover:underline"
- >
- {children}
- </a>
- );
- }
- return (
- <a
- href={href}
- target="_blank"
- rel="noopener noreferrer"
- className="text-brand underline decoration-brand/40 hover:decoration-brand"
- {...rest}
- >
- {children}
- </a>
- );
- };
-}
-
function CopyButton({ text }: { text: string }) {
const [state, setState] = useState<"idle" | "copied" | "failed">("idle");
const copy = async () => {
@@ -265,57 +187,12 @@ export function MessageBubble({
)}
{!isUser && message.sources && message.sources.length > 0 && (
- <ol className="ml-1 flex flex-col gap-1 text-xs text-muted-foreground">
- {message.sources.map((s, si) => (
- <li
- key={s.key}
- id={`cite-${cid}-${si + 1}`}
- className="scroll-mt-20 rounded transition-colors animate-in fade-in motion-reduce:animate-none"
- style={{ animationDelay: `${Math.min(si, 8) * 40}ms`, animationFillMode: "both" }}
- >
- <span className="font-mono text-brand">[{si + 1}]</span>{" "}
- {s.url ? (
- <a
- href={s.url}
- target="_blank"
- rel="noopener noreferrer"
- className="hover:underline"
- >
- {s.title}
- </a>
- ) : (
- s.title
- )}{" "}
- <span className="text-muted-foreground/70">
- — {s.channel}
- {s.siteTitle ? ` · ${s.siteTitle}` : ""} ·{" "}
- {s.snippets.map((sn, sni) => (
- <span key={sn.seconds}>
- {sni > 0 ? ", " : ""}
- {onOpenCitation ? (
- <button
- type="button"
- onClick={() => onOpenCitation(s.key, sn.seconds)}
- title="Open the transcript at this moment"
- className="font-mono text-brand transition-colors hover:underline"
- >
- {sn.clock}
- </button>
- ) : (
- sn.clock
- )}
- </span>
- ))}
- </span>
- </li>
- ))}
- {message.truncated && (
- <li className="text-muted-foreground/60">
- (some searches were truncated; ask more specifically for fuller
- coverage)
- </li>
- )}
- </ol>
+ <CitationSources
+ cid={cid}
+ sources={message.sources}
+ onOpenCitation={onOpenCitation}
+ truncated={message.truncated}
+ />
)}
</div>
);
diff --git a/export/app/ask/ReportPanel.tsx b/export/app/ask/ReportPanel.tsx
@@ -1,8 +1,10 @@
"use client";
-import { useState } from "react";
+import { useId, useState } from "react";
import { FileTextIcon, Loader2Icon } from "lucide-react";
import { Markdown } from "yt-dlp-transcript-common/components/Markdown";
+import type { RetrievedVideo } from "../lib/askRetrieval";
+import { CitationSources, linkifyCitations, makeCitationLink } from "./citations";
// The running "report" surface for report mode: a persistent markdown document
// the model maintains via the update_report tool while the user keeps chatting.
@@ -12,6 +14,8 @@ import { Markdown } from "yt-dlp-transcript-common/components/Markdown";
// Collapsible like the Context panel so it stays out of the way until wanted.
export function ReportPanel({
report,
+ sources,
+ onOpenCitation,
reportMode,
setReportMode,
available,
@@ -20,6 +24,12 @@ export function ReportPanel({
sweeping = false,
}: {
report: string;
+ // The report's source registry — what its `[n]` / `[n @ mm:ss]` citations index
+ // into. Drives the clickable citations + the numbered Sources list below.
+ sources: RetrievedVideo[];
+ // Open the transcript modal seeked to a cited line. Undefined in hub /ask (no
+ // player), where citations fall back to scroll-to-source within the list below.
+ onOpenCitation?: (slug: string, seconds?: number) => void;
reportMode: boolean;
setReportMode: (on: boolean) => void;
available: boolean;
@@ -30,6 +40,10 @@ export function ReportPanel({
sweeping?: boolean;
}) {
const [open, setOpen] = useState(false);
+ // Stable per-panel id so the report's citation anchors don't collide with a
+ // chat message's.
+ const cid = useId().replace(/:/g, "");
+ const CitationLink = makeCitationLink(cid, sources, onOpenCitation);
const hasReport = report.trim() !== "";
// While a sweep runs, keep the panel expanded (so its document is visibly
// growing) regardless of the user's toggle.
@@ -84,11 +98,25 @@ export function ReportPanel({
<div
aria-live="polite"
className={
- "rounded-md border border-border bg-background px-3 py-2 text-sm text-foreground transition-opacity " +
+ "flex flex-col gap-3 rounded-md border border-border bg-background px-3 py-2 text-sm text-foreground transition-opacity " +
(busyWriting ? "opacity-60" : "opacity-100")
}
>
- <Markdown className="leading-relaxed">{report}</Markdown>
+ <Markdown className="leading-relaxed" linkComponent={CitationLink}>
+ {linkifyCitations(report, sources.length, cid)}
+ </Markdown>
+ {sources.length > 0 && (
+ <div className="border-t border-border pt-2">
+ <p className="mb-1 text-xs font-medium text-muted-foreground">
+ Sources
+ </p>
+ <CitationSources
+ cid={cid}
+ sources={sources}
+ onOpenCitation={onOpenCitation}
+ />
+ </div>
+ )}
</div>
) : (
<p className="text-xs text-muted-foreground">
diff --git a/export/app/ask/askChatStorage.test.ts b/export/app/ask/askChatStorage.test.ts
@@ -104,10 +104,58 @@ test("a sweep checkpoint round-trips (serialize → resume yields the remaining
assert.equal(cp.reason, "usage limit");
});
+test("a report source registry round-trips (and old saves without one still load)", () => {
+ rawStore().removeItem(KEY);
+ const chat = baseChat({
+ report: "# Findings\n\nClaim [1 @ 0:05].\n",
+ reportSources: [
+ {
+ key: "ch/v1",
+ videoId: "v1",
+ title: "A video",
+ channel: "Chan",
+ uploadDate: "20200101",
+ snippets: [{ clock: "0:05", seconds: 5, text: "" }],
+ },
+ ],
+ });
+ saveSavedChats({ v: 1, items: { research: chat }, activeName: "research" });
+ const back = loadSavedChats().items.research;
+ assert.equal(back.reportSources?.length, 1);
+ assert.equal(back.reportSources?.[0].key, "ch/v1");
+ assert.equal(back.reportSources?.[0].snippets[0].seconds, 5);
+
+ // An old save with no reportSources field loads with it simply absent.
+ rawStore().setItem(
+ KEY,
+ JSON.stringify({ v: 1, items: { legacy: baseChat({ report: "old" }) }, activeName: "legacy" }),
+ );
+ const legacy = loadSavedChats().items.legacy;
+ assert.equal(legacy.report, "old");
+ assert.equal(legacy.reportSources, undefined);
+});
+
test("savedChatsEqual: equal on the same restorable content", () => {
assert.ok(savedChatsEqual(baseChat(), baseChat()));
});
+test("savedChatsEqual: dirty when the report source registry differs", () => {
+ const withSources = baseChat({
+ reportSources: [
+ {
+ key: "ch/v1",
+ videoId: "v1",
+ title: "A",
+ channel: "C",
+ uploadDate: "20200101",
+ snippets: [{ clock: "0:05", seconds: 5, text: "" }],
+ },
+ ],
+ });
+ assert.ok(!savedChatsEqual(baseChat(), withSources));
+ assert.ok(savedChatsEqual(withSources, withSources));
+});
+
test("savedChatsEqual: dirty when messages / report / checkpoint differ", () => {
assert.ok(!savedChatsEqual(baseChat(), baseChat({ messages: [msg("user", "x")] })));
assert.ok(!savedChatsEqual(baseChat(), baseChat({ report: "changed" })));
diff --git a/export/app/ask/askChatStorage.ts b/export/app/ask/askChatStorage.ts
@@ -11,6 +11,7 @@
// conversation.
import type { UiMessage } from "../lib/askConversation";
+import type { RetrievedVideo } from "../lib/askRetrieval";
const KEY = "ytdlp-tb:ai:chats";
const VERSION = 1;
@@ -42,6 +43,9 @@ export type SavedChat = {
detached?: boolean;
enrichments?: Record<string, { clock: string; seconds: number; text: string }[]>;
report?: string;
+ // The report's source registry (what its `[n]` citations index into), trimmed of
+ // snippet text. Absent on chats saved before this existed.
+ reportSources?: RetrievedVideo[];
reportMode?: boolean;
// Sweep checkpoint (present only for a paused sweep).
checkpoint?: SweepCheckpoint | null;
@@ -94,6 +98,9 @@ function parseChat(raw: unknown): SavedChat | null {
chat.enrichments = r.enrichments as SavedChat["enrichments"];
}
if (typeof r.report === "string") chat.report = r.report;
+ if (Array.isArray(r.reportSources)) {
+ chat.reportSources = r.reportSources as RetrievedVideo[];
+ }
if (typeof r.reportMode === "boolean") chat.reportMode = r.reportMode;
if (r.checkpoint !== undefined) chat.checkpoint = parseCheckpoint(r.checkpoint);
if (typeof r.model === "string") chat.model = r.model;
@@ -174,6 +181,7 @@ export function savedChatsEqual(a: SavedChat, b: SavedChat): boolean {
detached: c.detached ?? false,
enrichments: c.enrichments ?? {},
report: c.report ?? "",
+ reportSources: c.reportSources ?? [],
reportMode: c.reportMode ?? false,
checkpoint: c.checkpoint ?? null,
});
diff --git a/export/app/ask/citations.test.ts b/export/app/ask/citations.test.ts
@@ -0,0 +1,94 @@
+import { test } from "node:test";
+import assert from "node:assert/strict";
+import {
+ linkifyCitations,
+ parseClock,
+ resolveCitationSeconds,
+} from "./citations";
+import type { RetrievedVideo } from "../lib/askRetrieval";
+
+const CID = "r0";
+
+function video(over: Partial<RetrievedVideo> = {}): RetrievedVideo {
+ return {
+ key: "ch/v1",
+ videoId: "v1",
+ title: "A video",
+ channel: "Chan",
+ uploadDate: "20200101",
+ snippets: [
+ { clock: "0:05", seconds: 5, text: "" },
+ { clock: "1:00", seconds: 60, text: "" },
+ { clock: "2:30", seconds: 150, text: "" },
+ ],
+ ...over,
+ };
+}
+
+test("parseClock: mm:ss and h:mm:ss, rejects malformed", () => {
+ assert.equal(parseClock("0:05"), 5);
+ assert.equal(parseClock("12:34"), 12 * 60 + 34);
+ assert.equal(parseClock("1:02:03"), 3723);
+ assert.equal(parseClock("5"), null); // no colon
+ assert.equal(parseClock("1:2:3:4"), null); // too many parts
+ assert.equal(parseClock("1:xx"), null); // non-numeric
+});
+
+test("linkifyCitations: bare [n] links with a plain anchor", () => {
+ const out = linkifyCitations("see point [1] here", 3, CID);
+ assert.equal(out, "see point [[1]](#cite-r0-1) here");
+});
+
+test("linkifyCitations: [n @ mm:ss] encodes the parsed second into the anchor", () => {
+ const out = linkifyCitations("claim [2 @ 12:34]", 3, CID);
+ // Anchor tail is -<n>-<seconds>; display text keeps the full marker.
+ assert.equal(out, "claim [[2 @ 12:34]](#cite-r0-2-754)");
+});
+
+test("linkifyCitations: tolerates h:mm:ss and stray spaces", () => {
+ assert.equal(
+ linkifyCitations("[3 @ 1:02:03]", 3, CID),
+ "[[3 @ 1:02:03]](#cite-r0-3-3723)",
+ );
+ assert.equal(
+ linkifyCitations("[ 1 @ 0:05 ]", 3, CID),
+ "[[ 1 @ 0:05 ]](#cite-r0-1-5)",
+ );
+});
+
+test("linkifyCitations: out-of-range n stays plain text", () => {
+ assert.equal(linkifyCitations("[5] and [2]", 3, CID), "[5] and [[2]](#cite-r0-2)");
+ assert.equal(linkifyCitations("[9 @ 1:00]", 3, CID), "[9 @ 1:00]");
+});
+
+test("linkifyCitations: maxN 0 is a no-op", () => {
+ assert.equal(linkifyCitations("[1] [2]", 0, CID), "[1] [2]");
+});
+
+test("linkifyCitations: skips citations inside code spans", () => {
+ assert.equal(
+ linkifyCitations("real [1] but `code [2]` and ```\n[3]\n``` fenced", 3, CID),
+ "real [[1]](#cite-r0-1) but `code [2]` and ```\n[3]\n``` fenced",
+ );
+});
+
+test("resolveCitationSeconds: no parsed time → first snippet", () => {
+ assert.equal(resolveCitationSeconds(video(), NaN), 5);
+});
+
+test("resolveCitationSeconds: snaps a parsed time to the nearest known snippet", () => {
+ // 58s is closest to the 60s snippet; 140s closest to 150s.
+ assert.equal(resolveCitationSeconds(video(), 58), 60);
+ assert.equal(resolveCitationSeconds(video(), 140), 150);
+ // An exact hit stays put.
+ assert.equal(resolveCitationSeconds(video(), 5), 5);
+});
+
+test("resolveCitationSeconds: passes a parsed time through when the source has no snippets", () => {
+ assert.equal(resolveCitationSeconds(video({ snippets: [] }), 42), 42);
+});
+
+test("resolveCitationSeconds: undefined for a missing source", () => {
+ assert.equal(resolveCitationSeconds(undefined, 5), undefined);
+ assert.equal(resolveCitationSeconds(undefined, NaN), undefined);
+});
diff --git a/export/app/ask/citations.tsx b/export/app/ask/citations.tsx
@@ -0,0 +1,222 @@
+"use client";
+
+// Shared citation rendering for the /ask surfaces (chat answers AND the report).
+//
+// A model answer/report cites sources with bracketed markers: the classic `[n]`
+// and the precise-moment `[n @ mm:ss]` (also tolerated: `[n @ h:mm:ss]` and stray
+// spaces). `linkifyCitations` rewrites those into in-page Markdown links whose
+// anchor encodes the source number and, when present, the cited second;
+// `makeCitationLink` resolves a clicked anchor to `sources[n-1]` and opens the
+// transcript modal at that moment (snapped to the nearest real line so a slightly
+// off model timestamp still lands on a transcript cue), falling back to a
+// scroll-to-source when there's no player (hub /ask). `CitationSources` is the
+// numbered source list rendered beneath the document. Both the chat and the report
+// reuse this module so their citations behave identically.
+
+import type { AnchorHTMLAttributes } from "react";
+import type { RetrievedVideo } from "../lib/askRetrieval";
+
+// Parse a clock string (`mm:ss` or `h:mm:ss`) to whole seconds; null if malformed.
+export function parseClock(s: string): number | null {
+ const parts = s.split(":");
+ if (parts.length < 2 || parts.length > 3) return null;
+ let total = 0;
+ for (const p of parts) {
+ const n = parseInt(p, 10);
+ if (!Number.isFinite(n) || !/^\d+$/.test(p)) return null;
+ total = total * 60 + n;
+ }
+ return total;
+}
+
+// One citation marker: a source number, optional `@ mm:ss` moment, tolerating
+// stray spaces around the number, the `@`, and the time.
+const CITE_RE = /\[\s*(\d+)\s*(?:@\s*(\d{1,2}(?::\d{2}){1,2})\s*)?\]/g;
+
+// Turn `[n]` / `[n @ mm:ss]` citation markers into links to the matching source in
+// the list below. Skips fenced/inline code, and only links numbers that have a
+// source (1..maxN). The anchor id is scoped per surface via `cid`; a parsed moment
+// is encoded as a trailing `-<seconds>` segment so the click can seek to it.
+export function linkifyCitations(md: string, maxN: number, cid: string): string {
+ if (maxN <= 0) return md;
+ return md
+ .split(/(```[\s\S]*?```|`[^`]*`)/g)
+ .map((seg, i) =>
+ i % 2 === 1
+ ? seg
+ : seg.replace(CITE_RE, (m, d, t) => {
+ const n = parseInt(d, 10);
+ if (!(n >= 1 && n <= maxN)) return m;
+ const secs = typeof t === "string" ? parseClock(t) : null;
+ const anchor =
+ secs != null
+ ? `#cite-${cid}-${n}-${secs}`
+ : `#cite-${cid}-${n}`;
+ return `[${m}](${anchor})`;
+ }),
+ )
+ .join("");
+}
+
+// Scroll to and briefly highlight the source `<li>` for citation `n` — the
+// fallback when there's no player (hub /ask) or the source has no timestamp.
+export function scrollToSource(cid: string, n: number) {
+ const el = document.getElementById(`cite-${cid}-${n}`);
+ if (!el) return;
+ el.scrollIntoView({ behavior: "smooth", block: "center" });
+ el.classList.add("bg-brand-soft");
+ setTimeout(() => el.classList.remove("bg-brand-soft"), 900);
+}
+
+// Resolve the second a citation should seek to. A parsed `mm:ss` is validated by
+// snapping to the nearest known snippet second for that source (guards a
+// hallucinated timestamp); passed through as-is if the source has no snippets. No
+// parsed time → the source's first matched line (today's behavior). Undefined when
+// there's nothing to seek to.
+export function resolveCitationSeconds(
+ src: RetrievedVideo | undefined,
+ parsed: number,
+): number | undefined {
+ if (!src) return undefined;
+ const snippets = src.snippets ?? [];
+ if (Number.isFinite(parsed)) {
+ if (snippets.length === 0) return parsed;
+ let best = snippets[0].seconds;
+ let bestDist = Math.abs(best - parsed);
+ for (const s of snippets) {
+ const dist = Math.abs(s.seconds - parsed);
+ if (dist < bestDist) {
+ bestDist = dist;
+ best = s.seconds;
+ }
+ }
+ return best;
+ }
+ return snippets[0]?.seconds;
+}
+
+// Build the link renderer for citation Markdown, closing over this surface's
+// sources + citation-scope id + the (optional) modal opener. A `#cite-<cid>-<n>`
+// (optionally `-<seconds>`) anchor opens the transcript modal at the cited moment
+// when a player is available; otherwise it scrolls to the source below. Any other
+// link opens externally.
+export function makeCitationLink(
+ cid: string,
+ sources: RetrievedVideo[],
+ onOpenCitation?: (slug: string, seconds?: number) => void,
+) {
+ return function CitationLink({
+ href,
+ children,
+ ...rest
+ }: AnchorHTMLAttributes<HTMLAnchorElement>) {
+ if (typeof href === "string" && href.startsWith("#cite-")) {
+ return (
+ <a
+ href={href}
+ onClick={(e) => {
+ e.preventDefault();
+ // The anchor tail is `-<n>` or `-<n>-<seconds>`; the cid always starts
+ // with a letter, so the dash before it never precedes a digit and this
+ // captures the number (+ optional second) unambiguously.
+ const mm = href.match(/-(\d+)(?:-(\d+))?$/);
+ const n = mm ? parseInt(mm[1], 10) : NaN;
+ const parsed = mm && mm[2] ? parseInt(mm[2], 10) : NaN;
+ const src = Number.isFinite(n) ? sources[n - 1] : undefined;
+ const seconds = resolveCitationSeconds(src, parsed);
+ if (onOpenCitation && src && typeof seconds === "number") {
+ onOpenCitation(src.key, seconds);
+ } else {
+ scrollToSource(cid, n);
+ }
+ }}
+ className="font-mono text-brand no-underline hover:underline"
+ >
+ {children}
+ </a>
+ );
+ }
+ return (
+ <a
+ href={href}
+ target="_blank"
+ rel="noopener noreferrer"
+ className="text-brand underline decoration-brand/40 hover:decoration-brand"
+ {...rest}
+ >
+ {children}
+ </a>
+ );
+ };
+}
+
+// The numbered source list rendered beneath a cited document — one `<li>` per
+// source (its `#cite-<cid>-<n>` anchor is the scroll-to-source target) with a
+// per-timestamp button that opens the transcript modal at that moment (plain text
+// when there's no player). Shared by the chat answer and the report so both lists
+// look and behave identically.
+export function CitationSources({
+ cid,
+ sources,
+ onOpenCitation,
+ truncated,
+}: {
+ cid: string;
+ sources: RetrievedVideo[];
+ onOpenCitation?: (slug: string, seconds?: number) => void;
+ truncated?: boolean;
+}) {
+ if (sources.length === 0) return null;
+ return (
+ <ol className="ml-1 flex flex-col gap-1 text-xs text-muted-foreground">
+ {sources.map((s, si) => (
+ <li
+ key={s.key}
+ id={`cite-${cid}-${si + 1}`}
+ className="scroll-mt-20 rounded transition-colors animate-in fade-in motion-reduce:animate-none"
+ style={{ animationDelay: `${Math.min(si, 8) * 40}ms`, animationFillMode: "both" }}
+ >
+ <span className="font-mono text-brand">[{si + 1}]</span>{" "}
+ {s.url ? (
+ <a
+ href={s.url}
+ target="_blank"
+ rel="noopener noreferrer"
+ className="hover:underline"
+ >
+ {s.title}
+ </a>
+ ) : (
+ s.title
+ )}{" "}
+ <span className="text-muted-foreground/70">
+ — {s.channel}
+ {s.siteTitle ? ` · ${s.siteTitle}` : ""} ·{" "}
+ {s.snippets.map((sn, sni) => (
+ <span key={sn.seconds}>
+ {sni > 0 ? ", " : ""}
+ {onOpenCitation ? (
+ <button
+ type="button"
+ onClick={() => onOpenCitation(s.key, sn.seconds)}
+ title="Open the transcript at this moment"
+ className="font-mono text-brand transition-colors hover:underline"
+ >
+ {sn.clock}
+ </button>
+ ) : (
+ sn.clock
+ )}
+ </span>
+ ))}
+ </span>
+ </li>
+ ))}
+ {truncated && (
+ <li className="text-muted-foreground/60">
+ (some searches were truncated; ask more specifically for fuller coverage)
+ </li>
+ )}
+ </ol>
+ );
+}
diff --git a/export/app/ask/useAskChat.ts b/export/app/ask/useAskChat.ts
@@ -33,7 +33,9 @@ import {
buildApiMessages,
buildGroundedContent,
chunk,
+ mergeReportSources,
parseContext,
+ reportNumbering,
serializeContext,
type SearchStep,
type UiMessage,
@@ -90,6 +92,10 @@ type PersistedConvo = {
// so a long research session's report survives reloads.
report?: string;
reportMode?: boolean;
+ // The report's SOURCE REGISTRY — every video the report's `[n]` citations index
+ // into, ordered + deduped, trimmed (no snippet text). Persisted alongside the
+ // report so its clickable citations survive a reload.
+ reportSources?: RetrievedVideo[];
// A paused sweep's checkpoint (remaining batches + directive), so a paused sweep
// survives a reload and can be resumed — even under a different model.
checkpoint?: SweepCheckpoint | null;
@@ -177,6 +183,11 @@ export function useAskChat() {
// model maintains a persistent markdown report via update_report, and the
// conversation is compacted into that report instead of replaying every turn.
const [report, setReport] = useState("");
+ // The report's source registry: every video cited across all report writes
+ // (report-mode turns + sweep batches), ordered + deduped by key, trimmed. The
+ // report's `[n]`/`[n @ mm:ss]` citations resolve against this. Grows as writes
+ // land; cleared on New chat / a fresh sweep.
+ const [reportSources, setReportSources] = useState<RetrievedVideo[]>([]);
const [reportMode, setReportModeState] = useState(false);
// True while a turn is writing to the report (drives the panel shimmer);
// cleared when the turn ends.
@@ -273,6 +284,8 @@ export function useAskChat() {
strictRef.current = strictGrounding;
const reportRef = useRef(report);
reportRef.current = report;
+ const reportSourcesRef = useRef(reportSources);
+ reportSourcesRef.current = reportSources;
const reportModeRef = useRef(reportMode);
reportModeRef.current = reportMode;
const maxAnswerTokensRef = useRef(maxAnswerTokens);
@@ -346,6 +359,9 @@ export function useAskChat() {
setEnrichments(parsed.enrichments);
}
if (typeof parsed.report === "string") setReport(parsed.report);
+ if (Array.isArray(parsed.reportSources)) {
+ setReportSources(parsed.reportSources);
+ }
if (typeof parsed.reportMode === "boolean") {
setReportModeState(parsed.reportMode);
}
@@ -399,6 +415,7 @@ export function useAskChat() {
detached,
enrichments,
report,
+ reportSources,
reportMode,
checkpoint: sweepCheckpoint,
})
@@ -431,6 +448,7 @@ export function useAskChat() {
detached,
enrichments,
report,
+ reportSources,
reportMode,
sweepCheckpoint,
]);
@@ -729,6 +747,7 @@ export function useAskChat() {
seedVideos: args.seedVideos,
report: reportRef.current,
reportMode: effectiveReportMode,
+ reportSources: reportSourcesRef.current,
maxAnswerTokens: maxAnswerTokensRef.current,
onDebug: pushDebug,
precomputedGrounding: args.precomputedGrounding?.groundedContent
@@ -743,6 +762,11 @@ export function useAskChat() {
if (effectiveReportMode && result.report !== reportRef.current) {
setReport(result.report);
}
+ // Grow the report source registry with this turn's gathered videos (in the
+ // global-index order the report writes cited them by), trimmed for storage.
+ if (effectiveReportMode && result.reportSources) {
+ setReportSources(mergeReportSources([], result.reportSources));
+ }
patchAt(assistantIndex, (m) => ({
...m,
content: m.content || result.answer,
@@ -899,6 +923,18 @@ export function useAskChat() {
return;
}
setReportUpdating(true);
+ // Merge this batch into the report source registry BEFORE building its
+ // context, so its excerpts are numbered by their GLOBAL registry index —
+ // making the report's `[n]` stable across every batch and resolving
+ // against the persisted source list. Keep the ref current synchronously
+ // so the next batch numbers on top of this one.
+ const mergedSources = mergeReportSources(
+ reportSourcesRef.current,
+ batches[i],
+ );
+ reportSourcesRef.current = mergedSources;
+ setReportSources(mergedSources);
+ const numberOf = reportNumbering(mergedSources);
let r;
try {
r = await runReportChunk({
@@ -911,6 +947,7 @@ export function useAskChat() {
count: total,
report,
aliases,
+ numberOf,
signal: ac.signal,
onEvent: (e) => {
if (e.type === "report_update") {
@@ -1025,7 +1062,13 @@ export function useAskChat() {
]);
const startReport = mode === "fresh" ? "" : reportRef.current;
- if (mode === "fresh") setReport("");
+ if (mode === "fresh") {
+ setReport("");
+ // A fresh report starts from an empty source registry too (a new document's
+ // `[n]` must not resolve against the previous report's sources).
+ setReportSources([]);
+ reportSourcesRef.current = [];
+ }
await driveSweep({
remaining: videos,
@@ -1189,6 +1232,7 @@ export function useAskChat() {
setEnrichments({});
setStrictGroundingState(true);
setReport("");
+ setReportSources([]);
setReportModeState(false);
setSweepCheckpoint(null);
// A new chat is a blank working conversation not derived from any saved chat —
@@ -1307,6 +1351,7 @@ export function useAskChat() {
detached,
enrichments,
report,
+ reportSources,
reportMode,
checkpoint: sweepCheckpoint,
provider,
@@ -1319,6 +1364,7 @@ export function useAskChat() {
detached,
enrichments,
report,
+ reportSources,
reportMode,
sweepCheckpoint,
provider,
@@ -1418,6 +1464,7 @@ export function useAskChat() {
setDetached(chat.detached ?? false);
setEnrichments(chat.enrichments ?? {});
setReport(chat.report ?? "");
+ setReportSources(chat.reportSources ?? []);
setReportModeState(chat.reportMode ?? false);
setSweepCheckpoint(chat.checkpoint ?? null);
writeChats((s) => {
@@ -1446,6 +1493,7 @@ export function useAskChat() {
setMarkdownOn,
// report mode
report,
+ reportSources,
reportMode,
setReportMode,
reportModeAvailable,
diff --git a/export/app/lib/askConversation.test.ts b/export/app/lib/askConversation.test.ts
@@ -9,8 +9,10 @@ import {
buildGroundedContent,
chunk,
collectPriorPool,
+ mergeReportSources,
renderAliasGlossary,
renderUsedAliasContext,
+ reportNumbering,
gatherSystemPrompt,
answerSystemPrompt,
serializeContext,
@@ -255,6 +257,65 @@ test("buildAccumulationContent emits the directive, a batch label, and numbered
assert.match(out, /\[1\] "First" — Rekieta/);
assert.match(out, /\[0:05\] alpha here/);
assert.match(out, /\[2\] "Second" — Rekieta/);
+ // The label steers the model to the precise-moment citation form.
+ assert.match(out, /\[n @ mm:ss\]/);
+});
+
+test("buildAccumulationContent numbers by a global registry index when given numberOf", () => {
+ const videos = [
+ video({ key: "a", title: "First" }),
+ video({ key: "b", title: "Second" }),
+ ];
+ const registry = mergeReportSources(
+ [video({ key: "x", title: "Earlier" }), video({ key: "y", title: "Also earlier" })],
+ videos,
+ );
+ const out = buildAccumulationContent(
+ "d",
+ videos,
+ 2,
+ 3,
+ reportNumbering(registry),
+ );
+ // The batch's videos are global entries 3 and 4 (after the two prior sources).
+ assert.match(out, /\[3\] "First" — Rekieta/);
+ assert.match(out, /\[4\] "Second" — Rekieta/);
+});
+
+test("mergeReportSources appends new videos, dedupes by key, and trims snippet text", () => {
+ const base = mergeReportSources([], [
+ video({ key: "a", title: "A", snippets: [{ clock: "0:05", seconds: 5, text: "keep me?" }] }),
+ ]);
+ // Snippet text is trimmed for storage, but clock/seconds survive.
+ assert.equal(base[0].snippets[0].text, "");
+ assert.equal(base[0].snippets[0].seconds, 5);
+ assert.equal(base[0].snippets[0].clock, "0:05");
+
+ const merged = mergeReportSources(base, [
+ // "a" reappears with a new moment → merged into the existing entry.
+ video({ key: "a", title: "A", snippets: [{ clock: "1:00", seconds: 60, text: "x" }] }),
+ // "b" is new → appended after "a".
+ video({ key: "b", title: "B" }),
+ ]);
+ assert.deepEqual(merged.map((v) => v.key), ["a", "b"]);
+ // "a" now carries both moments (5s + 60s), deduped/merged by time.
+ assert.deepEqual(
+ merged[0].snippets.map((s) => s.seconds).sort((x, y) => x - y),
+ [5, 60],
+ );
+});
+
+test("reportNumbering maps a video to its 1-based registry index by key", () => {
+ const registry = mergeReportSources([], [
+ video({ key: "a" }),
+ video({ key: "b" }),
+ video({ key: "c" }),
+ ]);
+ const numberOf = reportNumbering(registry);
+ assert.equal(numberOf(video({ key: "a" }), 0), 1);
+ assert.equal(numberOf(video({ key: "c" }), 0), 3);
+ // A key not in the registry falls back to the local index + 1.
+ assert.equal(numberOf(video({ key: "zzz" }), 6), 7);
});
test("accumulationSystemPrompt carries the update_report directive + focused alias block", () => {
@@ -264,6 +325,9 @@ test("accumulationSystemPrompt carries the update_report directive + focused ali
assert.match(p, /Do NOT search/);
assert.match(p, /fetch_context tool/);
assert.match(p, /finish tool/);
+ // Precise-moment, globally-stable citation form.
+ assert.match(p, /\[n @ mm:ss\]/);
+ assert.match(p, /GLOBAL and stable/);
// The FOCUSED used-alias block (not the full "search the full term" glossary).
assert.match(p, /TERMS USED IN THIS SEARCH/);
assert.match(p, /Graham Platner: platner, plattner, platter/);
@@ -292,7 +356,8 @@ test("applyReportPatch composes across sweep chunks: fold section A then B → b
test("answerSystemPrompt asks for Markdown + citations and carries the FOCUSED used-alias block", () => {
const p = answerSystemPrompt(ALIASES);
assert.match(p, /Markdown/);
- assert.match(p, /\[1\]/);
+ // Precise-moment citation form: [n @ mm:ss] (backward-tolerant of a bare [n]).
+ assert.match(p, /\[n @ mm:ss\]/);
// The focused block (not the full glossary) — no "search the full term" directive.
assert.match(p, /TERMS USED IN THIS SEARCH/);
assert.match(p, /Graham Platner: platner, plattner, platter/);
diff --git a/export/app/lib/askConversation.ts b/export/app/lib/askConversation.ts
@@ -95,6 +95,61 @@ export function collectPriorPool(
return order.map((k) => byKey.get(k)!);
}
+// ─── report source registry (what a report's global `[n]` indexes into) ───
+
+// The report's persisted source registry: every video cited across all report
+// writes (report-mode turns + sweep batches), ordered, deduped by `key`. A
+// report's `[n]` is this list's n-th entry, and `[n @ mm:ss]` seeks to the cited
+// moment. Stored TRIMMED — snippet text is dropped (only clock/seconds + the video
+// meta are needed to resolve or label a citation) so a large sweep's registry
+// stays storage-light.
+const REPORT_SNIPPETS_CAP = 30;
+
+// Drop snippet text (keep clock/seconds) so the persisted registry is light. The
+// RetrievedVideo shape is preserved (text becomes ""), so it still renders in the
+// source list and resolves a citation's second.
+function trimReportSource(v: RetrievedVideo): RetrievedVideo {
+ return {
+ ...v,
+ snippets: v.snippets.map((s) => ({ clock: s.clock, seconds: s.seconds, text: "" })),
+ };
+}
+
+// Fold a batch/turn's videos into the report registry: append videos not seen
+// before (by `key`, first-seen order preserved) and merge snippet times into ones
+// already present (a video cited again in a later batch keeps all its moments).
+// Mirrors the cumulative-pool dedup in collectPriorPool. Pure + testable.
+export function mergeReportSources(
+ existing: RetrievedVideo[],
+ batch: RetrievedVideo[],
+): RetrievedVideo[] {
+ const order: string[] = existing.map((v) => v.key);
+ const byKey = new Map<string, RetrievedVideo>(existing.map((v) => [v.key, v]));
+ for (const raw of batch) {
+ const v = trimReportSource(raw);
+ const cur = byKey.get(v.key);
+ if (!cur) {
+ order.push(v.key);
+ byKey.set(v.key, v);
+ } else {
+ byKey.set(v.key, {
+ ...cur,
+ snippets: mergeSnippets(cur.snippets, v.snippets, REPORT_SNIPPETS_CAP),
+ });
+ }
+ }
+ return order.map((k) => byKey.get(k)!);
+}
+
+// A 1-based global-index lookup over a registry, keyed by video `key` — the
+// numbering a report write uses so its `[n]` matches the registry's source list.
+export function reportNumbering(
+ registry: RetrievedVideo[],
+): (v: RetrievedVideo, index: number) => number {
+ const indexByKey = new Map(registry.map((v, i) => [v.key, i + 1]));
+ return (v, index) => indexByKey.get(v.key) ?? index + 1;
+}
+
// ─── report mode (persistent markdown document, maintained via update_report) ───
// Upsert a `## <section>` block into the running report. If a section with that
@@ -341,7 +396,9 @@ export function gatherSystemPrompt(
? " You are maintaining a running report document that persists across turns. " +
"After gathering evidence, call the update_report tool to record findings in " +
"well-titled sections (pass the section heading and its full markdown " +
- "content); keep the report the source of truth. Then answer briefly."
+ "content), citing each source as [n @ mm:ss] — the excerpt's bracketed " +
+ "number plus the cited line's timestamp; keep the report the source of " +
+ "truth. Then answer briefly."
: "";
const tail =
mode === "native"
@@ -367,8 +424,10 @@ export function answerSystemPrompt(usedAliases: SearchAlias[]): string {
"answer on the transcript excerpts provided in this conversation. Excerpts " +
"may appear in earlier turns, and a follow-up request (reformatting, " +
"summarising, expanding) should reuse the relevant excerpts already " +
- "provided. Cite the excerpts you use by their bracketed number, e.g. [1]. " +
- "Each excerpt shows a video title, channel, and timestamped lines. If the " +
+ "provided. Cite each excerpt you use as [n @ mm:ss]: its bracketed number " +
+ "plus the specific line's timestamp (e.g. [1 @ 4:12]); a bare [n] is fine " +
+ "when no single line applies. Each excerpt shows a video title, channel, and " +
+ "timestamped lines. If the " +
"excerpts do not contain enough to answer, say so plainly rather than " +
"guessing. Format your answer in GitHub-flavored Markdown.";
const context = renderUsedAliasContext(usedAliases);
@@ -397,29 +456,35 @@ export function accumulationSystemPrompt(usedAliases: SearchAlias[]): string {
"You are building a running report by reading transcript excerpts in " +
"batches. Cross-reference this batch against the report so far. Use the " +
"update_report tool to add or merge findings — claims, contradictions with " +
- "earlier claims, and their sources (title + timestamp, cited by [n]) — into " +
- "well-titled sections, keeping the report the single source of truth. Do NOT " +
- "search; these excerpts are your evidence, but you may call the fetch_context " +
- "tool to read more transcript around a specific line when one is too thin to " +
- "judge. Call the finish tool when you are done with this batch.";
+ "earlier claims, and their sources — into well-titled sections, keeping the " +
+ "report the single source of truth. CITE each source as [n @ mm:ss]: the " +
+ "excerpt's bracketed number plus the specific line's timestamp (e.g. " +
+ "[3 @ 12:34]); the numbers are GLOBAL and stable across batches, so reuse a " +
+ "source's number if it reappears. Do NOT search; these excerpts are your " +
+ "evidence, but you may call the fetch_context tool to read more transcript " +
+ "around a specific line when one is too thin to judge. Call the finish tool " +
+ "when you are done with this batch.";
const context = renderUsedAliasContext(usedAliases);
return context ? `${base}\n\n${context}` : base;
}
// The per-chunk user message for a sweep: the report directive, a "Batch i of n"
-// label, and this batch's numbered excerpts (same citation format as a grounded
-// turn, reusing buildContext). Numbering restarts per batch — the raw excerpts
-// are discarded once the batch folds into the report, so cross-batch [n] identity
-// isn't needed.
+// label, and this batch's excerpts (same format as a grounded turn, reusing
+// buildContext). `numberOf` labels each excerpt by its GLOBAL registry index so a
+// report's `[n]` is stable across every batch — the source registry the report's
+// citations resolve against persists even though the raw excerpts are discarded
+// once the batch folds in. Omitted → 1-based per-batch numbering (legacy).
export function buildAccumulationContent(
directive: string,
videos: RetrievedVideo[],
index: number,
count: number,
+ numberOf?: (video: RetrievedVideo, i: number) => number,
): string {
return (
`${directive}\n\n` +
- `Batch ${index} of ${count} — transcript excerpts you may cite (by number):\n` +
- buildContext(videos)
+ `Batch ${index} of ${count} — transcript excerpts (cite each as [n @ mm:ss]: ` +
+ `the bracketed number + the line's timestamp):\n` +
+ buildContext(videos, numberOf)
);
}
diff --git a/export/app/lib/askRetrieval.test.ts b/export/app/lib/askRetrieval.test.ts
@@ -226,3 +226,34 @@ test("buildContext numbers videos and indents snippets", () => {
assert.match(ctx, /\[1\] "Travel Bans" — Rekieta/);
assert.match(ctx, /\[0:05\] about travel bans/);
});
+
+test("buildContext honours a numberOf override (global registry index)", () => {
+ const videos = [
+ {
+ key: "a",
+ videoId: "a",
+ title: "First",
+ channel: "Rekieta",
+ uploadDate: "20200101",
+ snippets: [{ clock: "0:05", seconds: 5, text: "alpha" }],
+ },
+ {
+ key: "b",
+ videoId: "b",
+ title: "Second",
+ channel: "Rekieta",
+ uploadDate: "20200101",
+ snippets: [{ clock: "1:00", seconds: 60, text: "beta" }],
+ },
+ ];
+ // Number the batch as if it were entries 41/42 of a large sweep's registry.
+ const index = new Map([
+ ["a", 41],
+ ["b", 42],
+ ]);
+ const ctx = buildContext(videos, (v) => index.get(v.key)!);
+ assert.match(ctx, /\[41\] "First" — Rekieta/);
+ assert.match(ctx, /\[42\] "Second" — Rekieta/);
+ // The default 1-based numbering is NOT used.
+ assert.doesNotMatch(ctx, /\[1\] "First"/);
+});
diff --git a/export/app/lib/askRetrieval.ts b/export/app/lib/askRetrieval.ts
@@ -324,12 +324,19 @@ export function retrieve(
}
// Assemble numbered retrieved excerpts into the context block appended to the
-// user's question (unchanged contract — the system prompt cites by [n]).
-export function buildContext(videos: RetrievedVideo[]): string {
+// user's question (the system prompt cites by [n]). By default each video is
+// numbered 1-based by its position (the chat/grounded path). `numberOf` overrides
+// this so a report write can number by GLOBAL registry index — making a report's
+// `[n]` stable across every batch/turn folded into it, not just within one batch.
+export function buildContext(
+ videos: RetrievedVideo[],
+ numberOf?: (video: RetrievedVideo, index: number) => number,
+): string {
if (videos.length === 0) return "(no matching transcript excerpts were found)";
return videos
.map((v, i) => {
- const head = `[${i + 1}] "${v.title}" — ${v.channel}${v.siteTitle ? ` (${v.siteTitle})` : ""}`;
+ const n = numberOf ? numberOf(v, i) : i + 1;
+ const head = `[${n}] "${v.title}" — ${v.channel}${v.siteTitle ? ` (${v.siteTitle})` : ""}`;
const lines = v.snippets
.map((s) => ` [${s.clock}] ${s.text}`)
.join("\n");
diff --git a/export/app/lib/searchAgent.ts b/export/app/lib/searchAgent.ts
@@ -168,13 +168,22 @@ async function scriptedGather(
// ─── result formatting fed back to the model during gather ───
-export function formatResultsForModel(videos: RetrievedVideo[]): string {
+// Feed a search's results back to the model during gather. `numberOf` overrides
+// the per-result 1-based index (report mode uses it to show the GLOBAL registry
+// number, so a report-mode turn's report write cites the same stable `[n]` the
+// registry's source list uses); the "n." lead-in then becomes a bracketed "[n]"
+// so the model reads it as the citation number.
+export function formatResultsForModel(
+ videos: RetrievedVideo[],
+ numberOf?: (video: RetrievedVideo, index: number) => number,
+): string {
if (videos.length === 0) return "No transcript excerpts matched that query.";
return videos
.map((v, i) => {
const site = v.siteTitle ? ` (${v.siteTitle})` : "";
const date = v.uploadDate ? ` [${v.uploadDate}]` : "";
- const head = `${i + 1}. "${v.title}" — ${v.channel}${site}${date}`;
+ const lead = numberOf ? `[${numberOf(v, i)}]` : `${i + 1}.`;
+ const head = `${lead} "${v.title}" — ${v.channel}${site}${date}`;
const lines = v.snippets.map((s) => ` ${s.clock} ${s.text}`).join("\n");
return lines ? `${head}\n${lines}` : head;
})
@@ -246,6 +255,12 @@ export type RunAskTurnOptions = {
// and the update_report tool is offered so the model keeps the report current.
report?: string;
reportMode?: boolean;
+ // Report mode: the persisted report SOURCE REGISTRY carried in — every video
+ // cited across the report so far, ordered + deduped by key. The turn numbers its
+ // gathered videos by their GLOBAL index in this registry (so a report write cites
+ // the same stable `[n]` the registry's source list uses) and returns the registry
+ // grown with this turn's new videos (see AskTurnResult.reportSources).
+ reportSources?: RetrievedVideo[];
// Max output tokens for the answer phase (default DEFAULT_ANSWER_TOKENS).
maxAnswerTokens?: number;
// Optional capture of each provider call (finish reason, payloads) for the
@@ -261,6 +276,10 @@ export type AskTurnResult = {
// The report after this turn (possibly updated via update_report). Echoes the
// input report unchanged when report mode is off or nothing was written.
report: string;
+ // Report mode: the source registry grown with this turn's gathered videos, in
+ // the GLOBAL-index order the turn's report writes cited them by. Undefined when
+ // report mode is off (the caller then leaves its registry untouched).
+ reportSources?: RetrievedVideo[];
// The answer's normalized finish reason and (if blocked) the raw block reason,
// so the UI can flag truncation / a safety block instead of a silent stop.
finishReason?: FinishReason;
@@ -289,6 +308,23 @@ export async function runAskTurn(
// returned in the result so the caller can persist it.
let report = opts.report ?? "";
+ // Report mode: number this turn's gathered videos by their GLOBAL index in the
+ // carried-in report source registry, so a report write cites the same stable
+ // `[n]` the registry's source list uses (and `[n @ mm:ss]` seeks the moment).
+ // `reportOrder` accumulates the registry (base + this turn's new videos, in
+ // first-seen discovery order) and is returned so the caller can persist it.
+ const reportBase = opts.reportSources ?? [];
+ const reportOrder: RetrievedVideo[] = [...reportBase];
+ const reportIndexByKey = new Map(reportBase.map((v, i) => [v.key, i + 1]));
+ const reportNumberOf = (v: RetrievedVideo): number => {
+ const existing = reportIndexByKey.get(v.key);
+ if (existing != null) return existing;
+ reportOrder.push(v);
+ const n = reportOrder.length;
+ reportIndexByKey.set(v.key, n);
+ return n;
+ };
+
// Capture the answer phase's finish reason (+ any block reason) so the result
// can flag truncation/blocking instead of a silent stop; also forward every
// provider call to the debug sink for the export.
@@ -358,6 +394,8 @@ export async function runAskTurn(
truncated:
(opts.precomputedGrounding.truncated ?? false) || answerFinish === "length",
report,
+ // Strict/retry path skips gather → no report write → registry unchanged.
+ reportSources: reportMode ? reportBase : undefined,
finishReason: answerFinish,
blockReason: answerBlock,
};
@@ -405,7 +443,9 @@ export async function runAskTurn(
if (r.truncated) truncated = true;
queries.push(query);
onEvent({ type: "search_done", query, count: r.videos.length });
- return formatResultsForModel(r.videos);
+ // Report mode shows the GLOBAL registry number so the report write's `[n]`
+ // resolves against the persisted source list.
+ return formatResultsForModel(r.videos, reportMode ? reportNumberOf : undefined);
};
// Expand mode with a pinned set: let the model read more of a grounding video's
@@ -548,6 +588,9 @@ export async function runAskTurn(
groundedContent,
truncated: truncated || answerFinish === "length",
report,
+ // The registry grown with this turn's gathered videos, in the global-index
+ // order the report writes cited them by.
+ reportSources: reportMode ? reportOrder : undefined,
finishReason: answerFinish,
blockReason: answerBlock,
};
@@ -569,6 +612,11 @@ export type RunReportChunkOptions = {
// The running report carried in; the returned report threads to the next chunk.
report: string;
aliases: SearchAlias[];
+ // Number each batch video by its GLOBAL registry index (so the report's `[n]` is
+ // stable across every batch, resolving against the persisted source list). The
+ // caller merges the batch into its registry first, then passes this. Omitted →
+ // 1-based per-batch numbering.
+ numberOf?: (video: RetrievedVideo, index: number) => number;
signal?: AbortSignal;
onEvent: (e: AgentEvent) => void;
onDebug?: (rec: DebugCall) => void;
@@ -671,7 +719,13 @@ export async function runReportChunk(
const system =
`${accumulationSystemPrompt([...usedAliases.values()])}\n\n` +
buildSeedDigest(opts.videos);
- const question = buildAccumulationContent(directive, opts.videos, index, count);
+ const question = buildAccumulationContent(
+ directive,
+ opts.videos,
+ index,
+ count,
+ opts.numberOf,
+ );
const ctx: NativeGatherContext = {
apiKey,
diff --git a/export/e2e/ask-chat.spec.ts b/export/e2e/ask-chat.spec.ts
@@ -890,6 +890,71 @@ test.describe("ask chat", () => {
await expect(page.getByText("Recorded finding 4.")).toHaveCount(0);
});
+ test("report citations are clickable and open the transcript at the cited moment", async ({
+ page,
+ }) => {
+ await installRoutes(page);
+ await page.route("https://api.anthropic.com/**", async (route) => {
+ if (route.request().method() === "OPTIONS") {
+ await route.fulfill({ status: 204, headers: CORS });
+ return;
+ }
+ const body = route.request().postDataJSON() as {
+ system?: string;
+ tools?: { name?: string }[];
+ messages?: { role: string; content: unknown }[];
+ };
+ const system = body.system ?? "";
+ const msgs = body.messages ?? [];
+ const lastUser = [...msgs].reverse().find((m) => m.role === "user");
+ const isToolResult = Array.isArray(lastUser?.content);
+ // Sweep gather turn: fold in ONE section citing a precise moment, then finish.
+ if (Array.isArray(body.tools) && /running report/.test(system)) {
+ await route.fulfill({
+ status: 200,
+ headers: { ...CORS, "content-type": "application/json" },
+ body: isToolResult
+ ? toolUse("finish", {})
+ : toolUse("update_report", {
+ section: "Findings",
+ content: "The archive discusses this at length [1 @ 0:05].",
+ }),
+ });
+ return;
+ }
+ await route.fulfill({
+ status: 200,
+ headers: { ...CORS, "content-type": "text/event-stream" },
+ body: sse("DONE"),
+ });
+ });
+
+ // Sweep the whole search (3 fixtures) into a report in a single batch.
+ await searchThenChat(page, ALL_RESULT_TREE, 3);
+ await keyIn(page, "Native tools");
+ await expect(page.getByText(/Grounded in 3 results/)).toBeVisible();
+ await page
+ .getByRole("button", { name: /Build report from all 3 results/ })
+ .click();
+ await expect(page.getByText(/Built a report from 3 results/)).toBeVisible();
+
+ // The report renders the [n @ mm:ss] marker as an in-page citation anchor,
+ // with a numbered Sources list beneath the document.
+ const cite = page.getByRole("link", { name: "[1 @ 0:05]", exact: true });
+ await expect(cite).toBeVisible();
+ await expect(cite).toHaveAttribute("href", /^#cite-/);
+ await expect(
+ page.getByRole("listitem").filter({ hasText: /alpha|Transcript|Chat/ }).first(),
+ ).toBeVisible();
+
+ // Clicking it opens the transcript modal seeked to the cited second (5s) —
+ // the same modal a chat citation opens, via replaceState on the /ask route.
+ await cite.click();
+ await expect(page).toHaveURL(/[?&]v=/);
+ await expect(page).toHaveURL(/[?&]t=5\b/);
+ await expect(page.getByRole("button", { name: "Close player" })).toBeVisible();
+ });
+
test("corpus sweep: Stop aborts mid-sweep and keeps the partial report", async ({
page,
}) => {