commit 4a3188ad4d1d2eef8a78595d265aec8901d24c17
parent 2e302924bd7a9fd510bbd56efcd11d15e7a15afb
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Tue, 11 Aug 2026 02:40:15 -0400
Record the measured baseline table and the prebuilt-index decision
No remote here, so the plan's "table in the PR description" goes in the
README, which is the canonical design doc for this package.
The numbers say don't build a search index yet: scoped and filtered
questions now land between 0.04s and 2.6s, and only the unfiltered
whole-corpus scan is still slow — a sweep's one-off first step. If that
ever becomes the bottleneck the right artifact is a token -> (channel,
page) postings list, which prunes pages the same way the filter planner
already does.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Diffstat:
1 file changed, 34 insertions(+), 0 deletions(-)
diff --git a/mcp/README.md b/mcp/README.md
@@ -382,6 +382,40 @@ as if it weren't. Every run prints a **corpus fingerprint** first, because the
composed dir gets rebuilt and a before/after that silently spans two corpora is
worse than no measurement at all.
+**The run this work was built against** — corpus `local:export/public`,
+5 channels, 170 transcript pages, 1,288 MB, 3,330 videos in `summaries/`;
+3 repetitions, median, taken at 0.64 load/core (i.e. trusted — no `UNRELIABLE`
+stamp):
+
+| query | wall ms | page reads | MB parsed |
+|---|---|---|---|
+| `list_channels` (cold) | 1 | 0 | 0 |
+| rare term, whole corpus | 16,967 | 170 | 1,261 |
+| common term, whole corpus | 16,446 | 170 | 1,288 |
+| common term, one channel | 42 | 4 | 27 |
+| common term + one upload year | 2,647 | **32** | **202** |
+| common term + missing states only | **823** | **8** | **63** |
+| `enumerate_matches`, whole corpus | 23,422 | 170 | 1,288 |
+| `get_transcripts` × 20 ids, one channel | 24 | 4 | 27 |
+
+Read it as three facts. **Filter-first is worth ~21× on a selective filter** (8
+pages against 170) and ~5× on a one-year date range — and the comparison to make
+is within this same run: the removed-videos question costs 0.82 s where the only
+previously-available way to ask it, an unfiltered scan, costs 16.4 s. **The
+unfiltered rows are unchanged at 170 reads**, which is the point — planning must
+never make a whole-corpus scan slower. **The 20-id batch is 4 reads**, cold; it
+was up to 20 re-reads of the same page.
+
+*On the prebuilt-index question* (deferred until the speed win was a number):
+these numbers say **not yet**. Scoped and filtered questions — the ones people
+actually ask — now land between 0.04 s and 2.6 s. Only the unfiltered
+whole-corpus scan is still slow, and that is a sweep's one-off first step before
+it goes on to read hundreds of transcripts. If it ever does become the
+bottleneck, the right artifact is a narrow token → `(channel, page)` postings
+list emitted at export time: it would prune pages exactly the way the filter
+planner already does, which makes that machinery its prerequisite rather than
+its competitor.
+
## Data source (pick one)
Resolved from flags or env — precedence hub > remote > local: