Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 13b082635204b22620d759c5035e972c90f7529b
parent ba963c52531d3b9fcd242f5331bbde13c24f394d
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Fri, 25 Sep 2026 14:15:33 -0400

plans: release 8 slice V — review fixes, re-gate, live Rumble cost numbers

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
Mplans/release-8.md | 48++++++++++++++++++++++++++++++++++++------------
1 file changed, 36 insertions(+), 12 deletions(-)

diff --git a/plans/release-8.md b/plans/release-8.md @@ -34,7 +34,10 @@ Nothing creates `data/<id>/`, which keeps the scan store's invariant. `metadataS | `9ed6cce8` | `common/controller/videoTitles.ts`: `readChannelVideoTitles(paths, slug, ids) → Map<id, {title, source: "index"\|"scan"\|"metadata"}>` and `readVideoMetadataForDisplay(paths, slug, id) → {title?, description?, webpageUrl?, uploader?, uploadDate?, duration?, source: "metadata"\|"scan"\|"none"}`. `.test.ts` has 6 cases on a real compressed LMDB with a 4 KB description: the three-source merge with first hit winning, bare ids absent and another channel's id not leaking; missing index and scan store; the id-as-title fallback skipped; the head read with an escaped title, a title past 16 KB and no dir created; the display reader's three outcomes; and the 5,000-id cost case | | `cdd6c3dc` | The list. `VideoRow.title?`. `computeVideoRows` takes an optional `titles` map, which the page fills once from `readChannelVideoTitles` over data-dir ids ∪ `undownloadedIds`. `matchesVideoQuery` (id OR title, case-insensitive) is used by both the client filter and the server's `?q=` ordering for prev/next. A row shows the title with the id on a muted mono line under it, and shows the id alone when there is no title. **The row's accessible name stays `open <id>`**, and the checkbox stays `select <id>`. The placeholder is now "Search title or id…". The empty copy ("No videos match this filter.") names no ids and is unchanged. The embedded detail pane's `video title` line reads the same map, and the page's private `loadVideoTitle` parser is gone | | `16ef0758` | The video page. `loadMeta` is now `readVideoMetadataForDisplay`, and the page's private parser is gone. When the source is `scan`, the header shows the scan's title, upload date and duration, plus an italic note "from the listing scan — not downloaded" (`aria-label="metadata source"`). A collapsed `<details>` **Description** block appears when either source has a description. `generateMetadata` behaves as before (title, else id) and now also finds scan titles. New `editor/e2e/video-titles.spec.ts` (2 tests) on `youtube-with-playlist` plus a seeded `metadata-scan.json`: `fake00000001` shows its `metadata.info.json` title, `fake00000002` its scan title, and `fake00000003` its bare id. "harbor" leaves one row, and the id still matches. The undownloaded page shows the h2, the scan note, `2024-03-15`, `12:34`, and a Description that is collapsed and opens on click. The downloaded page has no scan note | -| *(this commit)* | this record and two `[Unreleased]` bullets in `editor/CHANGELOG.md` | +| `ff3944cd` | this record and two `[Unreleased]` bullets in `editor/CHANGELOG.md` | +| `4e8bb285` | **Review M1.** Source 3 is memoized per channel in `videoTitles.ts` (`globalThis.__yttVideoTitleMemo__`: `Map<dataDir, {dataDirMtimeMs, titles}>`). Each render does one `stat` of `data/`. A changed mtime (a video dir added or removed) drops the channel's entry. Only ids missing from the memo are read, and misses are not memoized, so a dir still being downloaded picks up its title once the file lands. The memo holds at most 64 channels, least recently used evicted first. `resetVideoTitleMemo()` is called by `api/test/invalidate-cache` alongside the other process memos. `videoTitleMetadataReadCount()` is a test hook. New unit case: with `data/`'s mtime unchanged, the second call reads 0 files and still serves a title whose file was deleted; a new dir makes it read all 3 again; a miss is re-read later | +| `41facfe0` | **Review L1 + L2.** `matchesVideoQuery` moves above `parseFilters`' doc comment. `editor/app/channels/[slug]/lib/videoRows.test.ts` (3 cases) covers id hits, title hits (case-insensitive), a row without a title, and a blank query. This pins the server `?q=` path, which shares the predicate with the search box | +| *(this commit)* | the re-gate below and the corrected cost paragraph | **Gates** (worktree root, on `16ef0758`). tsc (`pnpm -r --no-bail --workspace-concurrency=1 exec tsc --noEmit`) was clean before each commit. common **1801/1801** = 1795 + 6 (`videoTitles`). @@ -49,15 +52,34 @@ exec next build` ok. The export build was not run, because the slice touches no (`PORT=4011`, `EXPORT_PORT=4010`); the `PORT:3011` printed by the e2e script is a default for when the env var is unset. Every heavy step started with ≥ 3 GB available. +**Re-gate after the review fixes** (on `41facfe0`). tsc was clean. common **1802/1802** = 1801 + 1 +(memo case). Editor unit **75/75** = 72 + 3 (`videoRows.test.ts`). `pnpm --filter editor exec next +build` ok. The memo touches the e2e reset route, so the EDITOR e2e re-ran the full nine-spec list +(`v-e2e2.log`): **62 passed, 0 failed, 3.6 m**, with no queue wait. test:scripts and mcp were not +re-run, because neither imports anything that changed. + **Numbers: none.** No `settings.json`, `site.json` or `config.json` key changed, and nothing is written: the slice only reads. -**Cost of the title map** (the brief's bar: "does not change the page's order of magnitude"). In -the unit test, on a synthetic 5,000-id channel (3,000 index hits, 1,500 scan hits, 400 -`metadata.info.json` head reads of 50 KB files, 100 bare ids), `readChannelVideoTitles` takes -**82 ms** warm. That is one per page render, on a page that already does a `readdir` of `data/` -and reads the snapshot, so the order of magnitude is unchanged. The live corpus was not measured: -the rules forbid opening its index from a worktree. +**Cost of the title map** (the brief's bar: "does not change the page's order of magnitude"). +The first version held that bar only for a YouTube-shaped channel, where most titles come from the +index. The review found the Rumble case. On Rumble every directory name misses the index, so every +row fell through to a `metadata.info.json` head read, on **every** render. Each row click +re-renders, because rows are `?video=` links on a force-dynamic page. `the-quartering-rumble` has +**8,049** dirs. The reviewer replayed the read pattern read-only and measured **5,024 ms cold / +224 ms warm per render** before the memo. After the memo (`4e8bb285`), measured read-only with +`readChannelVideoTitles` itself over the live `data/` (no index opened, `$T/v-memo-measure.mts`): +the **first render took 5,331 ms** (cold page cache, all 8,049 titled) and **memoized renders took +5.5 ms and 6.5 ms**. The cold first render is still paid once per editor process, and again after a +video dir is added or removed on that channel. In the unit test's synthetic 5,000-id mix (3,000 +index, 1,500 scan, 400 `metadata.info.json`, 100 bare) it is 82–148 ms first and ~75 ms memoized. +That run is dominated by the index walk and the scan store, which are not memoized: one LMDB range +and one JSON read, not N file reads. + +**The memo's one blind spot, accepted.** A `metadata.info.json` rewritten inside an *existing* video +dir does not change `data/`'s mtime, so the list keeps the old title until the editor restarts or a +video dir is added or removed. A video's title does not change after download, so this was +accepted. **Found and left.** - **Coverage is partial on `paramount-tactical`.** A read-only `jq` over live `snapshot.json` / @@ -66,11 +88,13 @@ the rules forbid opening its index from a worktree. step slice K's `keep-videos` needs. Every other channel's list is fully downloaded, so it is titled from the index and `metadata.info.json`. - **The index is keyed by metadata id and the list by directory name.** On Rumble (URL-slug dirs, - embed-id metadata) and on legacy `YYYYMMDD_<id>` dirs, the index lookup misses. Those rows fall - through to a `metadata.info.json` head read: same title, one file read each. A channel with - thousands of such dirs pays thousands of 16 KB reads per render. That was not measured on the - live corpus. If it shows, add a process-level memo keyed by path and mtime, as `recencyIndex.ts` - does for its tail reads. + embed-id metadata) and on legacy `YYYYMMDD_<id>` dirs, the index lookup misses. Those rows are + titled from `metadata.info.json`, and the memo above now carries that cost. Matching dir names to + index ids (for example through the `mtimes` sub-DB's `[slug, videoDir]` keys) would make the cold + first render cheap too, but is left for later. +- **L3 (review, left).** When a previously scanned video has a `data/<id>/` dir but no readable + `metadata.info.json` (a failed or partial download), the page falls back to the scan entry and says + "not downloaded" beside a files panel. This is rare, and the wording was left as it is. - **Rows stay sorted by id**, as before. Sorting by title or date is a separate ask. - `VideoListPane.tsx` lives at `editor/app/channels/[slug]/components/`, not under `videos/**`. The brief names "the list pane" in ownership, so it was edited as in scope. `ROW_ESTIMATE_PX` (30) was