commit c084840a709ae5cac07384acab33eba9cb09e978
parent 21c44a807157a99a4e56de7ccb37f6b92a2ae44f
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Sat, 10 Oct 2026 01:59:50 -0400
plans: release 20 D3, D1 and D2, as shipped (the caption-track bug closed, recorded dates, one Twitch id)
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Diffstat:
2 files changed, 88 insertions(+), 0 deletions(-)
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -1,6 +1,8 @@
# Changelog
## [Unreleased]
+- **A channel that mirrors another's streams can date its videos by the stream, not the upload.** A new channel setting, "Recorded date from the title (regex)" on the Configure form (`recordedDate.titlePattern` in config.json; `pnpm ops channel-config` with `recordedDateTitlePattern`), names where a title carries the recording's date — a regex with the groups `year`, `month` (a number or a month name) and `day`. The index then gives each of that channel's videos a `recordedDate`, which coverage reads before the upload date; a title without a date, or with one after the upload, keeps the upload date. Changing the pattern re-dates the channel's videos at the next index build.
+- **A Twitch video's id is the one in its address.** Twitch records carried yt-dlp's `v<number>` while their folders and URLs use `<number>`, so anything that went from a record to its files (a clip fetch above all) missed them. Records now carry `<number>`; the next index build re-keys the existing ones once, and their transcript page addresses change with them.
- **A clip window of a video whose source is saved is cut from it, not fetched.** `fetch_clip`, `POST /api/media/fetch-window` and `pnpm ops fetch-windows` now cut a window out of the video's saved container (a persisted source, a full-source fetch, or media attached from a local archive) when it covers the seconds asked for, and answer at once as a cached window — no request to the platform, so a deleted channel's held videos are clippable. The window's sidecar records `source: "saved-video"`, and the video page marks it "cut from the saved video". A batch runs such windows as their own job on `clips:saved-video`, outside every platform's queue, hold and cooldown. A saved video whose file cannot be read (its drive unplugged, the file gone) is refused with the media guard's sentence rather than fetched; one that ends before the window is fetched as before.
- **An X fetch with a `limit` stops at that many posts.** "Fetch posts" with `limit` (`pnpm ops fetch-posts {"limit": 400}`) on a gallery-dl X channel read the whole history instead — a new channel walked 3,803 posts under the rate limit and held the platform queue for hours — because the cap counted media files, which a metadata-only read has almost none of. It now caps the posts themselves.
- **A home seeder of last resort, behind a VPN.** `archilyzer seed` seeds the playable torrents of the sites named in the new `settings.seeder` (`sites`, `trackers`, `maxUploadKiBps`, `maxConnections`, `pollSeconds`, `standbyAfterSeconds`, `bindInterface`; SETTINGS.md) to desktop clients over TCP and to browsers over WebRTC — but each torrent only while no other seeder has it: other seeders seen on every poll for `standbyAfterSeconds` puts that torrent on standby (it stops announcing and closes its peers, keeping the data), and it comes back at once when a leecher is waiting with no other source, or after the same window with no other seeder. Every change is logged with its reason. No DHT, no local discovery, no UPnP. `archilyzer tracker` is a self-hosted HTTP + WebSocket tracker that tracks only those torrents. `docker-compose.seeder.yml` (profile `seeder`) runs both inside a WireGuard container's network namespace (gluetun, its firewall always on), so a tunnel that is down means no network, never the home connection; the WireGuard config is yours (`SEEDER_WG_CONF`, required, mounted read-only). `archilyzer doctor` compares the seeder's egress address with the host's and fails when they are the same; it says "seeder not configured" until `seeder.sites` names a site.
diff --git a/plans/release-20.md b/plans/release-20.md
@@ -35,4 +35,90 @@ spec(s) D3 names, then the full editor suite once if any editor code changed.
(Each slice adds a "### Slice <X>, as shipped" section here, before "## Rollout".)
+Track D of the overnight batch (2026-10-10), on `r20/d2-r20` off `r20/integration` (`585be292`), after release 21
+D2 on the same branch; D3 first, then D1 → D2, as the merge order allows. Every fixture is synthetic; nothing read the
+corpus beyond config.json key names and shapes, and a count of two Twitch channels' directories (below).
+
+### Slice D3, as shipped — the caption-track bug, closed
+
+**Verified against the tree.** The `en` → 0 cues bug (filed 2026-10-01) is closed by `200a3105`: one rule in
+`common/lib/videoStatus.ts` (`CAPTION_TRACK_RULE_VERSION` 1) — pin > `en-orig` > `en` > regional > `en-en-*` by name,
+then the first track that parses to at least one cue (`readEnglishVttCues`) — and `parseVtt` reads a document with no
+inline timing tag as plain cue blocks, the livestream `en` shape that used to parse to nothing. Every cue reader goes
+through the rule: the index, normalize, report compose, the editor's track reader, and umtool's `cues.mjs` copy (held
+equal by `captionTrack.test.ts`); `buildIndex.ts`'s other `parseVtt` call is the non-English and live-chat tracks. An
+index built under the old rule is re-read once for the records the rule reaches.
+
+**Added:** `controller/buildIndexCaptionTrack.test.ts` +1 through the real `buildIndex` — an empty `en-orig` before an
+`en` with text, and an empty `en` before an `en-US` with text (no `en-orig`), both read the track with text. FACTS
+gains "The caption-track rule" (and a superseded note on the old "a caption fixture must carry timing tags" line);
+STATE closes the bug.
+
+**Found and left (report only, as ruled): the stats cache does not see the caption-track pass.** `buildStats` keys a
+video's stat on its index `mtimes` record (`indexSignature`; `plans/stats-cache-key.md`, merged `10cefd15`, in the
+tree unchanged: schema 6, the whole record plus `indexKey`). The caption-track pass re-reads a record's cues and writes
+the SAME `mtimes` record back (no input file moved), so a record the pass took from no text to text keeps
+`hasTranscript: false` and its old `cueCount` in the stats until something else moves its record.
+
+### Slice D1, as shipped — recorded dates for VOD-mirror channels
+
+**What it does.** `common/lib/recordedDate.ts` (pure, client-safe): a channel's `recordedDate: {titlePattern}` is a
+case-insensitive regex source naming the groups `year` (four digits, or two read as 20YY), `month` (1–12, or an English
+month name, whole or cut to three or more letters) and `day`; it must compile, name all three, and pass the download
+filter's safety check (at most 200 characters, no nested quantifier), or the config reader drops the whole rule.
+`deriveRecordedDate` gives `YYYYMMDD` for a real day no later than the upload date (a later title date is one the
+title mentions, not the recording's), else nothing. `coverageDate(summary)` is `recordedDate ?? uploadDate`, the hook
+coverage reads.
+- **Config:** `ChannelConfig.recordedDate` through `CHANNEL_CONFIG_COERCIONS` / `channelConfigSchema`, with a nested
+ key table; CHANNEL.md regenerated (`archilyzer docs files`).
+- **The index** (`buildIndex.ts`): after a record's summary is settled (from metadata or a fresh `transcript.cues.json`),
+ `recordedDate` is set from the channel's compiled rule or deleted. `recordedDateRule:<slug>` in the index meta holds
+ the pattern a channel's records were derived under; a build that finds another re-processes that channel's records
+ once (`Recorded dates: <slug>: its rule changed; N record(s) re-derived.`), a held channel left for a later build.
+ No channel has a rule today, so the first build after the rollout re-processes nothing.
+- **`TranscriptSummary.recordedDate?`**, omitted when absent: it reaches a channel's published transcript pages only
+ once that channel has a rule, and no page changes until then.
+- **Setting it:** the Configure form's Advanced section gains "Recorded date from the title (regex)"
+ (`recordedDateTitlePattern`, refused with why when it could not be a rule; in `CHANNEL_FORM_FIELDS`, so clearing it
+ clears the key), and `pnpm ops channel-config {"slug":…,"patch":{"recordedDateTitlePattern":…}}` sets it through the
+ same parser.
+- **Coverage** (release 19 A9's `channel_coverage`) is not on this branch; `coverageDate` is what it reads at merge.
+
+**Which channels are mirrors** was read from config.json shapes only: no key marks one; the candidates are the
+channels whose name or URL says VODs, mirror or archive (listed in the track report). No channel's config was changed;
+the patterns are the operator's to write.
+
+### Slice D2, as shipped — one Twitch id
+
+**What was found.** A Twitch VOD's directory is the canonical id (the URL's `/videos/<n>`, `extractVideoId`), while
+`summarize` took yt-dlp's native `meta.id`, `v<n>`: the record's id and slug said `v<n>`, so a reader joining a record
+to `data/<id>/` (MCP `fetch_clip`, any sidecar lookup) missed, and the fallback page URL would have been
+`/videos/v<n>`. Two channels carry Twitch VODs; no curated tag or site report references a `v<n>` id (counts 0).
+
+**What it does.** `canonicalTwitchVideoId` (`lib/videoId.ts`) is the one normalization: a `v` followed by digits
+loses the `v`. `summarize` uses it for a Twitch record; `readNormalizedTranscript` applies it (`normalizeSummaryId`)
+to a `transcript.cues.json` written before, so the index and report compose read canonical ids from either source; the
+index re-processes, once, every record it holds under a v-id (`Twitch ids v1: N record(s) indexed under the native
+v-id, re-keyed.`, the platform-labels shape, recorded only with no channel held), and the existing re-key path removes
+the old key. The Twitch player is still handed `v<n>` (`TwitchPlayer.tsx`).
+
+**What it changes in public:** at the next index build and site build, every Twitch record's id and slug, and so its
+transcript page URL, move from `…/v<n>` to `…/<n>`.
+
+**Commits**
+
+| Commit | What |
+|---|---|
+| `04a075b3` | `common:` D3 — the fall-through index test; FACTS "The caption-track rule"; STATE closes the bug |
+| `96409e24` | `common:` D1 — the `recordedDate` rule, its derivation, the index pass, the form field; CHANNEL.md |
+| `433815d4` | `common:` D2 — `canonicalTwitchVideoId`, `normalizeSummaryId`, the one-shot re-key, the player |
+
+**Gates.** `pnpm -r --no-bail --workspace-concurrency=1 exec tsc --noEmit` clean at each commit. New tests:
+`lib/recordedDate.test.ts` 7, `controller/buildIndexRecordedDate.test.ts` 2 (derived, absent without a rule, a date
+after the upload dropped; a changed rule re-derives once, a removed one drops the field from the pages),
+`controller/twitchId.test.ts` 4, `controller/buildIndexTwitchIds.test.ts` 2 (indexed canonical; an old v-id record
+re-keyed once), `controller/buildIndexCaptionTrack.test.ts` +1, editor `app/channels/components/parseChannelForm.test.ts`
+3. The whole common suite 3,661/3,661 (release 21 D2's tests included); editor unit 228/228; export unit 118/118;
+mcp 293/293; `archilyzer docs files --check` and `docs env --check` clean. E2E, one run in the foreground (start of the queue after the heavy slot freed): `transcript-source.spec.ts` (D3's index spec, 4) and `ops-api.spec.ts` (13, the channel-config round trip among them) — 17 passed, 1.5 min. Release end per the plan: the full editor suite once, the orchestrator's.
+
## Rollout