Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 4a71c45cba4d69016b8dc9a5c05d455bf23e01c1
parent c16d156adacdc9eb5bcadf150c19b906f2adab74
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sat, 12 Sep 2026 11:44:17 -0400

plans: the S3 record — one search pipeline, and the inversion it paid for

Appended to §Record. Sha range, every gate's actual numbers, and the eight
places the slice plan and the code disagreed — chiefly that `searchEval.ts`
is NOT a copy of the mcp evaluator (slug sets vs one record), that `rank.ts`
had one implementation to collect rather than two, and that the viewer's
"duplicate-collapse" is a badge, not a collapse.

Also recorded, because the next slice will need them: the recipe for
re-composing the bench corpus read-only against the primary checkout (S1's
1.3 GB site was gone — /home is at 100 %), and the reason the one editor
spec failure is a subset-setup dependency rather than a regression.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Diffstat:
Mplans/one-core-phase-2.md | 218+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
1 file changed, 218 insertions(+), 0 deletions(-)

diff --git a/plans/one-core-phase-2.md b/plans/one-core-phase-2.md @@ -382,3 +382,221 @@ client chunk. The prerender is the only step that fails, and it fails at base. `plans/tools/jeralyzer-corpus-2026-09-12.json` was not re-fetched; the live diff is S3's gate, after the search pipeline lands. + +### S3 — shipped 2026-09-12 + +Branch `one-core/phase-2-s3`, off `7f86aef` (the `integrate/2026-09-storage-priority` +tip with S1 merged). Six commits, `eaa50b6` → `99fe37a`, plus this note, unmerged. +**No URL shape moved, no `corpus.json` byte moved, no CONTRACT version moved, no +architecture allow-list entry added — and `"components"` is now in +`FORBIDDEN.lib`, so the list of forbidden edges grew while the list of excused +ones did not.** + +| commit | what | +|---|---| +| `eaa50b6` | `lib/search/{policy,window,evalTree,rank,collapse,leafPipeline}.ts` + a test file each | +| `7252b57` | `mcp/src/search.ts`: 1,563 lines → 1,206, orchestration only; the caps are `MCP_POLICY` | +| `07075d6` | the inversion: four `lib → components` edges deleted, `"components"` added to `FORBIDDEN.lib` | +| `fb6a31d` | the unused `LeafMatcher` import the previous commit left | +| `99fe37a` | the `SearchMode` doc comment travels with the type | +| _(this commit)_ | this note | + +#### What the slice actually found, where the brief and the code disagreed + +**1. `searchEval.ts` is not a COPY of the mcp evaluator, and merging them +would have been wrong.** The plan says `evalTree.ts` should hold +"`search.ts:1189-1338` plus `lib/searchEval.ts`'s copy, once". Read side +by side they are two algorithms over two data shapes: mcp's `evalNode` +answers *does this record match*, given a record it already holds; +searchEval's `evaluateSlugs`/`orchestrate` answer *which slugs survive*, +streaming, over per-leaf network pipelines, because the browser cannot +hold 1.3 GB of transcripts. Unifying them means materialising the corpus +in the browser. What the slice does instead: the per-record evaluator +moves into `lib/search/evalTree.ts` once, and **each file names the other +as its sibling and states the rule** — the meaning of AND / OR / negate +changes in both together, or it has silently forked. + +**2. `rank.ts` had one implementation to collect, not two.** The stub says +the mcp ordering and the viewer ordering "are the same intent written +twice". mcp does not rank at all: `searchTranscripts` emits in corpus scan +order and `SearchResult` has no sort. The orderings that DO exist are the +viewer's newest-first over `uploadDate` +(`SearchSessionContext.tsx`, inline) and `getThread`'s oldest-first over +`createdAt` tie-broken by id (`search.ts`, inline). Both are now +comparators in `rank.ts` and both call sites call them. **No relevance +ranking was invented** — that would have changed what every caller returns +under cover of a refactor. + +**3. `collapse.ts` has one consumer, and the viewer's "duplicate-collapse" +is not one.** `SearchResults.tsx:502` is `duplicateSiblings(...)` from +`components/duplicatesCache.ts` — it decorates a card with a sibling +*count* and a jump menu. It does not remove a row, and the result list is +never collapsed. So there was no second implementation to fold in. What +moved is mcp's `collapseDuplicates`, now **pure and synchronous** and +generic over a `CollapsibleHit`: the caller supplies the already-fetched +duplicate index, so the rules are testable without a transport and the +viewer can adopt them when it wants them, rather than having them imposed +by this slice. + +**4. A sixth module was required: `lib/search/leafPipeline.ts`.** The +inversion is impossible without it. `searchEval.ts` does not merely CALL +`runLeafPipeline` — it is typed on `LayerHit` and `LeafController`, so +injecting the function alone leaves the import. Moving the browser's +streaming leaf scanner down (three worker-pool drivers + the leaf adapter, +~450 lines) is what makes `lib/` self-contained; `components/searchPipeline.ts` +is now 65 lines that answer one question — WHICH FETCH — and supplies +`searchRuntime`. The brief allows new files under `lib/search/`; no new +`components/*.ts` was created, so `common/package.json` is untouched. + +**5. Four back-edges, not one.** The brief named `lib/searchEval.ts:35-38`. +Re-grepping `lib/` found four, and all four had to go before +`"components"` could join `FORBIDDEN.lib`: + +| edge | what it was | how it inverted | +|---|---|---| +| `searchEval → components/searchPipeline` | `runLeafPipeline`, `LayerHit`, `LeafController` | pipeline moved down; the fetch is injected as `SearchRuntime.runLeaf` | +| `searchEval → components/searchLayerCache` | `cacheKey` / `getCached` / `getCachedSync` / `putCached` / `CachedResult` | the memo is injected as `SearchRuntime.cache`; `CachedResult` moved to `searchEval` (an IndexedDB store does not get to define what a leaf result IS) and `searchLayerCache` re-exports it | +| `aiHandoff → components/searchPipeline` | `LayerHit` (type) | import path only | +| `searchQuery → components/urlState` | `SearchMode` (type) | moved to `searchQuery`, which interprets it; `urlState` re-exports | + +`runQueryTree` gains one required `runtime` field. Its three callers pass +`searchRuntime`: `SearchSessionContext.tsx`, `charts/useSearchSeries.ts` +and `export/app/lib/askRetrieval.ts`. No other import site moved — every +name `components/searchPipeline` exported it still exports, including the +`LayerHit` that `askRetrieval.ts` imports from it. + +**6. `window.ts` keeps TWO excerpt shapes on purpose.** The stub hoped +"the viewer's excerpt and the MCP's excerpt are the same excerpt". They +are not the same function: the MCP takes a ±45 s window around every +matched cue and merges them (the paragraph a sweep cites); the viewer +emits one row per matched cue, widened to the neighbouring cues ONLY when +the match straddles a caption break (the line a reader scans). Collapsing +them changes both outputs, and the viewer's is pinned by the export e2e's +rendered hit text. One module owns both, one `truncate`, one `clock`, no +duplicated code — and the file says out loud why they stay two. + +**7. `policy.ts` also names the uncapped one.** The plan said "the viewer +passes its own, uncapped". `VIEWER_POLICY` exists so there is one place +that says what uncapped MEANS, and a test pins the load-bearing +consequence (`truncate(t, Infinity)` is the identity; +`0 >= Infinity` is false). + +**8. `buildScanPlan` stayed in `mcp/src/search.ts`.** It is reader-driven +orchestration, `scanPlan.test.ts` sits beside it, and it calls the single +`passesFilters` from `lib/search/evalTree` — so the pruner and the scanner +still cannot disagree about what matches, which was the one property that +mattered. + +#### Gates + +- `pnpm -r exec tsc --noEmit` — clean in all six packages, after every commit. +- `pnpm --filter yt-dlp-transcript-common test` — **1128 passed / 0 failed** + (baseline **1077** at `7f86aef`; + 2 `policy`, + 8 `window`, + 2 `rank`, + + 6 `collapse`, + 21 `evalTree`, + 12 `leafPipeline` = 51; none lost). + The architecture test passes with `"components"` in `FORBIDDEN.lib` and the + ALLOWED ledger **byte-identical** to the base. +- `pnpm --filter yt-dlp-transcript-mcp test` — **205 passed / 0 failed**, + unchanged. +- `pnpm test:scripts` — **71 passed / 1 skipped**, unchanged. +- `pnpm --filter export exec next build` — **compiled successfully**, 11 static + pages. The only real test that none of this dragged `reader-fs.ts` or a + react-query module into a client chunk. (Prerender needs a composed + `export/public`: the committed fixture at + `plans/tools/compose-fixture-one-youtube-channel/public/` copied in, which is + gitignored — a bare checkout fails on a missing `summaries/manifest.json` + before it reaches any code this slice touched.) +- **compose-site byte-identity** over the FACTS.md fixture recipe: + `IDENTICAL modulo the build clock` for both trees — 16 files under `public/`, + 7 under `index/`. Nothing composed moved. +- `pnpm --filter yt-dlp-transcript-mcp bench --repeat 1 --force --local <site>`, + before at `7f86aef` and after at the tip, both over the same composed + hasanalyzer site. **Identical fingerprint** (5 channels, 170 transcript pages, + 3,358 summaries videos, stats + duplicates + digests all present) and **every + structural counter identical, byte for byte** — including the per-case scan + notes and both `pruned` flags: + + | case | reads | bytes parsed | note | + |---|---|---|---| + | cold-channels | 0 → 0 | 0 → 0 | | + | rare, whole corpus | 170 → 170 | 1,335,885,512 → 1,335,885,512 | 170 pages | + | common, whole corpus | 170 → 170 | 1,364,679,296 → 1,364,679,296 | 170 pages | + | channel-scoped | 4 → 4 | 28,847,054 → 28,847,054 | 4 pages | + | date-scoped (filter-first) | 32 → 32 | 213,451,503 → 213,451,503 | 28 pages · pruned | + | state-scoped (filter-first) | 9 → 9 | 73,399,988 → 73,399,988 | 9 pages · pruned | + | enumerate, whole corpus | 170 → 170 | 1,364,679,296 → 1,364,679,296 | 170 pages | + | get_transcripts × 20 ids | 4 → 4 | 28,847,054 → 28,847,054 | | + + Wall ms is noise and is not quoted (the box was above the load threshold for + both runs, and the bench marked them UNRELIABLE, as designed). + + **The site S1 benched was gone and had to be rebuilt** — `/home` is at 100 %, + and nothing 1.3 GB survived. Recipe, for the next slice: compose `hasanalyzer` + READ-ONLY against the primary checkout, by pointing `EXPORT_INDEX_DIR` at a + directory holding two symlinks (`shared`, `sites` → the primary checkout's + `export/.export-index`), `SITES_DIR` at a **copy** of `sites/hasanalyzer` with + `"archives": false` added, `TRANSCRIPTS_DIR` at a throwaway dir holding + `settings.json` `{}` plus the corpus-wide `duplicates*.json` / + `search-aliases.json`, and `EXPORT_PUBLIC_DIR` at scratch. Nothing writes + inside `transcripts/`: the compose cache is keyed off the public dir, the + archive builder is off, and the channel signer opens its LMDB under the + throwaway root. The BEFORE run reproduced S1's recorded numbers to the byte, + which is the evidence that the rebuilt corpus is the same corpus. Delete the + 2.3 GB output afterwards. +- **e2e**, behind the queue lock from the worktree (port block + 4000/4001/4010/4011/4020): export `e2e` **172 passed / 0 failed**, + `e2e:hub` **5 passed / 0 failed**, editor `export-search` + + `export-player-platform-cache` **21 passed / 0 failed**. +- `e2e:2origin` was **not run**: it is red on the base for a reason outside this + slice (the hub `/ask` prerender — see S1's §Record, verified there against + `c7f7b90` itself). Because S3 touches two of the three files that failure + names (`askRetrieval.ts`, `SearchSessionContext.tsx`), `build:hub` was run + anyway to confirm the failure is the SAME one: `INSTANCE_MODE=hub next build` + compiles successfully and type-checks, then dies on the same page with the + same message from the same two chunks — + `Error: useSearchSession must be used within a SearchSessionProvider`, + `Export encountered an error on /(workspace)/ask/page`, + `SearchSessionContext` + `AskChat`. Unchanged, and the compile+typecheck + passing is itself the evidence that the injected `searchRuntime` resolves in + hub mode too. +- **The live jeralyzer contract** (S1 left this as S3's gate): + `curl https://jeralyzer.pages.dev/corpus.json` is **byte-identical** to + `plans/tools/jeralyzer-corpus-2026-09-12.json` — 12,380 bytes, zero + structural differences, `generatedAt` included. Be honest about what that + does and does not prove: the site has not been rebuilt since the snapshot, + so this confirms the snapshot is still current, not that a rebuild would + match. **A rebuilt jeralyzer was not attempted** — 30 channels / 30,886 + videos, and `/home` is at 100 %. The real evidence that the emitted contract + did not move is the compose-site fixture byte-identity above: jeralyzer's + `corpus.json` comes out of the same `buildSiteCorpus` the fixture exercises, + and S3 does not open `corpus.ts`, `compose-site.ts` or `compose-hub.ts` at + all. + +#### One e2e failure, and why it is not this slice + +The editor subset failed 1 of 21 on its first run: +`export-search.spec.ts:612` "a site's own social links win over the global +default", with `ENOENT: … editor/test-settings.json` at `helpers.ts:166`. + +It is a spec-SUBSET setup dependency, not a regression. That file is created +only by `resetData()` (`editor/e2e/helpers.ts:19`), and `export-search.spec.ts` +never calls `resetData` — it reads a settings file some earlier spec in the full +suite wrote. Running only these two spec files in a fresh worktree can never +create it. Seeding it the way the suite does +(`cp e2e/fixtures/test-settings.default.json editor/test-settings.json`) and +re-running the same subset: **21 passed / 0 failed**. Nothing in S3 touches +settings, social links or the footer. + +#### Notes for the slices still in flight + +- `components/searchPipeline.ts` is now 65 lines and exports `searchRuntime`. + **S2a: if a cache's `fetchX` signature changes, the one place to rebind it is + `CACHE_FETCHERS` there** — not inside `lib/`, which no longer knows the caches + exist. +- `lib/search/leafPipeline.ts` uses bare `setTimeout`/`clearTimeout` rather than + `window.setTimeout`. Identical in a browser, and it is what lets the module be + driven from a test process with no DOM — which is how + `leafPipeline.test.ts` counts page reads over an in-memory `ArchiveReader`. +- The two sibling comments (`lib/searchEval.ts` ↔ `lib/search/evalTree.ts`) are + load-bearing prose: they are the only thing stopping the two evaluations of + the same algebra from drifting, because no test can compare a streaming + slug-set walk against a per-record boolean.