commit 7719bdcfb87fad3edb1c9d9412327ae3f457ca5b
parent 80680d21b0af285c763a30d869bf180fce426dd4
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Fri, 25 Sep 2026 14:00:51 -0400
plans: release 8 record — slice V (video titles) as shipped; changelog bullets
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Diffstat:
2 files changed, 86 insertions(+), 0 deletions(-)
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -1,6 +1,8 @@
# Changelog
## [Unreleased]
+- **A channel's video list shows titles, and you can search by them.** On `/channels/<slug>/videos`, each row now shows the video's title, with its id in smaller type underneath. Search matches the title or the id, ignoring case. The title comes from the transcript index for transcribed videos, from the channel's metadata scan for videos that were listed but never downloaded, and otherwise from the video's `metadata.info.json`. A video none of these name shows its id, as before. Each row's selection checkbox and link are named by the id as before, and the order is unchanged (by id).
+- **An undownloaded video's page shows its title and details.** The video page used to show a bare id for any video without a `metadata.info.json`. If the channel's metadata scan has read the video, the page now shows its title, upload date and duration from the scan, marked *from the listing scan — not downloaded*. Any video page with a description, from either source, has a **Description** section, collapsed by default.
- **New ops action: `pnpm ops keep-videos` marks every video of a channel whose title or description matches a pattern as "do not clean".** It sets the same marker as the video page's *Do not clean* toggle, so the clean sweep, extra-format cleanup, wrong-format removal, the superseded-subs purge and saved-video eviction all leave those videos alone. The body is `{"slug", "match", "fields"?, "note"?, "dryRun"?}`. `match` is matched the way a channel's download filter *include* is: a case-insensitive regex over title + description. `fields: ["title"]` or `["description"]` narrows it to one half, and `dryRun: true` reports without writing. Videos that already carry the marker are counted and left as they are. The marker lives in the video's folder, so a match that was never downloaded is listed under `notDownloaded` and no folder is created for it. Run `download-missing` on those ids, then run `keep-videos` again. On a new channel, run `metadata-scan` first: a video with no scanned title cannot match, and the reply counts those as `unscanned`.
- **Auto-download no longer retries the same rate-limited video over and over; it moves on to the next one.** When a download answered HTTP 429, the runner paused the whole platform for a while and then picked the same video again, because it was still first in the queue. Each retry doubled the pause, up to 30 minutes. On 2026-09-24 one YouTube Short was retried 12 times this way and kept YouTube paused all evening. A YouTube 429 comes from the subtitle fetch for one video, not from the whole site. Now a rate-limited video is also **deferred for 6 hours**: auto-download skips it, so when the pause ends the runner takes the next video. The pause still grows only when *different* videos keep hitting the limit. Deferrals are kept in `.auto-queue/state.json` beside the platform cooldowns, so a restart does not retry the video early. The log line reads `… (attempt 1). <id> deferred 6h; next video after cooldown.` A manual Sync or *download missing* ignores deferrals and still fetches the video. When every video left is deferred, the runner reports that it is idle for that reason: "every pending video was rate-limited recently and is deferred".
- **The cooldown strip on `/operations/download` also lists deferred videos.** It is now a region named *Rate-limit cooldown*, with a *Platforms in cooldown* list (unchanged) and a *Deferred videos* list. Each deferred video links to its page and shows how long it has left (`alpha/a1 — 5h 59m left`). The strip appears when either list has something in it. Times over an hour now read `5h 59m` instead of `359m 58s`.
diff --git a/plans/release-8.md b/plans/release-8.md
@@ -0,0 +1,84 @@
+# Release 8 — video titles (+ follow-ups)
+
+`main` at `bb3dbb4c`, release 7 live 2026-09-25; this release starts with the operator's video-titles
+ask. Rules: `plans/tools/implementer-rules.md`. Record file: this file.
+
+## Record
+
+### Slice V, as shipped — video titles (2026-09-25)
+
+Branch `one-core/r8-video-titles` off `main` `bb3dbb4c`. The operator's ask (2026-09-25): "In the
+editor, the individual video view should show video metadata like the title, and `/videos` should
+show titles in the list when available and allow searching by them." There is no global `/videos`;
+the list is the per-channel workspace `editor/app/channels/[slug]/videos/page.tsx`. Before this
+slice a list row was built from `ChannelSnapshot` id lists and showed only the id, the search box
+matched only the id, and the video page read `metadata.info.json` with its own parser, so a video
+that was never downloaded showed a bare id even when the metadata scan knew what it was called.
+
+Three places hold a title, and the slice reads them cheapest first. The first one found wins:
+- **`index`**: the LMDB `sums` sub-DB, reached by a key-only walk of this channel's `byChannel`
+ range. A summary is decoded only for an id the list asked about. It is opened `readOnly` **with
+ `compression: true`**, for the reason `curatedTagsPreview.ts`'s `openIndex` gives: without it,
+ any value over ~1 KB throws. A missing or unopenable index gives no titles and is not an error.
+ An index title equal to the id is `summarize()`'s fallback, not a title, and is skipped.
+- **`scan`**: `channels/<slug>/metadata-scan.json`, read once through `loadMetadataScan`.
+- **`metadata`**: `data/<id>/metadata.info.json`, only for ids the first two did not name. It is a
+ 16 KB head read plus a regex (yt-dlp writes `id` then `title` first). The full parse
+ (`loadRawMetadataFromDir`) runs only when the head has no title. 16 reads run concurrently.
+
+Nothing creates `data/<id>/`, which keeps the scan store's invariant. `metadataScanStore.ts` and
+`channelSnapshot.ts` are unchanged.
+
+| sha | what |
+|---|---|
+| `9ed6cce8` | `common/controller/videoTitles.ts`: `readChannelVideoTitles(paths, slug, ids) → Map<id, {title, source: "index"\|"scan"\|"metadata"}>` and `readVideoMetadataForDisplay(paths, slug, id) → {title?, description?, webpageUrl?, uploader?, uploadDate?, duration?, source: "metadata"\|"scan"\|"none"}`. `.test.ts` has 6 cases on a real compressed LMDB with a 4 KB description: the three-source merge with first hit winning, bare ids absent and another channel's id not leaking; missing index and scan store; the id-as-title fallback skipped; the head read with an escaped title, a title past 16 KB and no dir created; the display reader's three outcomes; and the 5,000-id cost case |
+| `cdd6c3dc` | The list. `VideoRow.title?`. `computeVideoRows` takes an optional `titles` map, which the page fills once from `readChannelVideoTitles` over data-dir ids ∪ `undownloadedIds`. `matchesVideoQuery` (id OR title, case-insensitive) is used by both the client filter and the server's `?q=` ordering for prev/next. A row shows the title with the id on a muted mono line under it, and shows the id alone when there is no title. **The row's accessible name stays `open <id>`**, and the checkbox stays `select <id>`. The placeholder is now "Search title or id…". The empty copy ("No videos match this filter.") names no ids and is unchanged. The embedded detail pane's `video title` line reads the same map, and the page's private `loadVideoTitle` parser is gone |
+| `16ef0758` | The video page. `loadMeta` is now `readVideoMetadataForDisplay`, and the page's private parser is gone. When the source is `scan`, the header shows the scan's title, upload date and duration, plus an italic note "from the listing scan — not downloaded" (`aria-label="metadata source"`). A collapsed `<details>` **Description** block appears when either source has a description. `generateMetadata` behaves as before (title, else id) and now also finds scan titles. New `editor/e2e/video-titles.spec.ts` (2 tests) on `youtube-with-playlist` plus a seeded `metadata-scan.json`: `fake00000001` shows its `metadata.info.json` title, `fake00000002` its scan title, and `fake00000003` its bare id. "harbor" leaves one row, and the id still matches. The undownloaded page shows the h2, the scan note, `2024-03-15`, `12:34`, and a Description that is collapsed and opens on click. The downloaded page has no scan note |
+| *(this commit)* | this record and two `[Unreleased]` bullets in `editor/CHANGELOG.md` |
+
+**Gates** (worktree root, on `16ef0758`). tsc (`pnpm -r --no-bail --workspace-concurrency=1 exec
+tsc --noEmit`) was clean before each commit. common **1801/1801** = 1795 + 6 (`videoTitles`).
+Editor unit **72/72**. test:scripts **161 pass + 1 skip**. mcp **219/219**. `pnpm --filter editor
+exec next build` ok. The export build was not run, because the slice touches no file under
+`export/` and no `common/` module it imports. EDITOR e2e used `$T/v-specs.txt`. `videos.spec` and
+`channel-page.spec` do not exist, so the list is `video-titles`, `video-page`,
+`video-filter-combine`, and the six specs that drive the list or the embedded pane
+(`channel-embedded-video`, `bulk-actions`, `incomplete-transcript`, `download-format-guard`,
+`channel-storage`, `whisper`). Result (`v-e2e.log`): **62 passed, 0 failed, 4.0 m**, after a
+2 m 38 s queue wait behind `one-core/r8-state-share`. The run used this worktree's block
+(`PORT=4011`, `EXPORT_PORT=4010`); the `PORT:3011` printed by the e2e script is a default for when
+the env var is unset. Every heavy step started with ≥ 3 GB available.
+
+**Numbers: none.** No `settings.json`, `site.json` or `config.json` key changed, and nothing is
+written: the slice only reads.
+
+**Cost of the title map** (the brief's bar: "does not change the page's order of magnitude"). In
+the unit test, on a synthetic 5,000-id channel (3,000 index hits, 1,500 scan hits, 400
+`metadata.info.json` head reads of 50 KB files, 100 bare ids), `readChannelVideoTitles` takes
+**82 ms** warm. That is one per page render, on a page that already does a `readdir` of `data/`
+and reads the snapshot, so the order of magnitude is unchanged. The live corpus was not measured:
+the rules forbid opening its index from a worktree.
+
+**Found and left.**
+- **Coverage is partial on `paramount-tactical`.** A read-only `jq` over live `snapshot.json` /
+ `metadata-scan.json` found one channel with undownloaded ids: `paramount-tactical`, with **1,439**
+ and no `metadata-scan.json`. Those rows show ids until a metadata scan runs, which is the same
+ step slice K's `keep-videos` needs. Every other channel's list is fully downloaded, so it is
+ titled from the index and `metadata.info.json`.
+- **The index is keyed by metadata id and the list by directory name.** On Rumble (URL-slug dirs,
+ embed-id metadata) and on legacy `YYYYMMDD_<id>` dirs, the index lookup misses. Those rows fall
+ through to a `metadata.info.json` head read: same title, one file read each. A channel with
+ thousands of such dirs pays thousands of 16 KB reads per render. That was not measured on the
+ live corpus. If it shows, add a process-level memo keyed by path and mtime, as `recencyIndex.ts`
+ does for its tail reads.
+- **Rows stay sorted by id**, as before. Sorting by title or date is a separate ask.
+- `VideoListPane.tsx` lives at `editor/app/channels/[slug]/components/`, not under `videos/**`. The
+ brief names "the list pane" in ownership, so it was edited as in scope. `ROW_ESTIMATE_PX` (30) was
+ left alone. Titled rows are about 40 px, and the virtualizer measures each row, so only the first
+ scrollbar estimate is off.
+- The stage lists (`VideoIdList.tsx`, on the download/transcribe stage panels) still show bare ids.
+ They are not the `/videos` list, and titling them is outside this slice.
+- **Commit trailers** name `Claude Opus 5.5 (1M context)`, as the release-6 and release-7
+ implementers did.
+
+## Rollout