Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 24ce786e21fd861d128338b8b9089ad15f6d685e
parent a0accf89f61b6f636e5c4b57ddd8f84f7c014be9
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sat, 10 Oct 2026 02:05:07 -0400

plans: release 20 D2 held for an operator ruling — the two options, the readers that miss the join; changelog and STATE corrected

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
Meditor/CHANGELOG.md | 1-
Mplans/STATE.md | 2++
Mplans/release-20.md | 43+++++++++++++++++++++++++------------------
3 files changed, 27 insertions(+), 19 deletions(-)

diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -2,7 +2,6 @@ ## [Unreleased] - **A channel that mirrors another's streams can date its videos by the stream, not the upload.** A new channel setting, "Recorded date from the title (regex)" on the Configure form (`recordedDate.titlePattern` in config.json; `pnpm ops channel-config` with `recordedDateTitlePattern`), names where a title carries the recording's date — a regex with the groups `year`, `month` (a number or a month name) and `day`. The index then gives each of that channel's videos a `recordedDate`, which coverage reads before the upload date; a title without a date, or with one after the upload, keeps the upload date. Changing the pattern re-dates the channel's videos at the next index build. -- **A Twitch video's id is the one in its address.** Twitch records carried yt-dlp's `v<number>` while their folders and URLs use `<number>`, so anything that went from a record to its files (a clip fetch above all) missed them. Records now carry `<number>`; the next index build re-keys the existing ones once, and their transcript page addresses change with them. - **A clip window of a video whose source is saved is cut from it, not fetched.** `fetch_clip`, `POST /api/media/fetch-window` and `pnpm ops fetch-windows` now cut a window out of the video's saved container (a persisted source, a full-source fetch, or media attached from a local archive) when it covers the seconds asked for, and answer at once as a cached window — no request to the platform, so a deleted channel's held videos are clippable. The window's sidecar records `source: "saved-video"`, and the video page marks it "cut from the saved video". A batch runs such windows as their own job on `clips:saved-video`, outside every platform's queue, hold and cooldown. A saved video whose file cannot be read (its drive unplugged, the file gone) is refused with the media guard's sentence rather than fetched; one that ends before the window is fetched as before. - **An X fetch with a `limit` stops at that many posts.** "Fetch posts" with `limit` (`pnpm ops fetch-posts {"limit": 400}`) on a gallery-dl X channel read the whole history instead — a new channel walked 3,803 posts under the rate limit and held the platform queue for hours — because the cap counted media files, which a metadata-only read has almost none of. It now caps the posts themselves. - **A home seeder of last resort, behind a VPN.** `archilyzer seed` seeds the playable torrents of the sites named in the new `settings.seeder` (`sites`, `trackers`, `maxUploadKiBps`, `maxConnections`, `pollSeconds`, `standbyAfterSeconds`, `bindInterface`; SETTINGS.md) to desktop clients over TCP and to browsers over WebRTC — but each torrent only while no other seeder has it: other seeders seen on every poll for `standbyAfterSeconds` puts that torrent on standby (it stops announcing and closes its peers, keeping the data), and it comes back at once when a leecher is waiting with no other source, or after the same window with no other seeder. Every change is logged with its reason. No DHT, no local discovery, no UPnP. `archilyzer tracker` is a self-hosted HTTP + WebSocket tracker that tracks only those torrents. `docker-compose.seeder.yml` (profile `seeder`) runs both inside a WireGuard container's network namespace (gluetun, its firewall always on), so a tunnel that is down means no network, never the home connection; the WireGuard config is yours (`SEEDER_WG_CONF`, required, mounted read-only). `archilyzer doctor` compares the seeder's egress address with the host's and fails when they are the same; it says "seeder not configured" until `seeder.sites` names a site. diff --git a/plans/STATE.md b/plans/STATE.md @@ -25,6 +25,8 @@ first. Nothing of release 18 is live. - **The `en` → 0 cues caption bug (filed 2026-10-01) is CLOSED** by `200a3105` (en-orig first, empty tracks fall through, cue-block VTTs parse, a one-shot index re-read); verified against the tree and an index fixture by release 20 D3 (2026-10-10; FACTS, "The caption-track rule"). Left open: the stats cache does not see that pass's re-reads. +- **Release 20 D2 (Twitch ids) is HELD for an operator ruling**: built and kept on `r20/twitch-ids`, reverted on + `r20/d2-r20`; the two options are in `release-20.md` ("Slice D2 — held"). **Previously (2026-10-06): release 18 — publishing as queueable stages — is complete on `r18/integration`** (record: [`release-18.md`](release-18.md): slices S1 the stage contract, stamps, lock, bundles and CLI; S2 deploy hardening; S3 diff --git a/plans/release-20.md b/plans/release-20.md @@ -36,7 +36,7 @@ spec(s) D3 names, then the full editor suite once if any editor code changed. (Each slice adds a "### Slice <X>, as shipped" section here, before "## Rollout".) Track D of the overnight batch (2026-10-10), on `r20/d2-r20` off `r20/integration` (`585be292`), after release 21 -D2 on the same branch; D3 first, then D1 → D2, as the merge order allows. Every fixture is synthetic; nothing read the +D2 on the same branch; D3 first, then D1; D2 built, then held (below). Every fixture is synthetic; nothing read the corpus beyond config.json key names and shapes, and a count of two Twitch channels' directories (below). ### Slice D3, as shipped — the caption-track bug, closed @@ -88,22 +88,28 @@ coverage reads. channels whose name or URL says VODs, mirror or archive (listed in the track report). No channel's config was changed; the patterns are the operator's to write. -### Slice D2, as shipped — one Twitch id +### Slice D2 — one Twitch id: held, operator ruling owed -**What was found.** A Twitch VOD's directory is the canonical id (the URL's `/videos/<n>`, `extractVideoId`), while -`summarize` took yt-dlp's native `meta.id`, `v<n>`: the record's id and slug said `v<n>`, so a reader joining a record -to `data/<id>/` (MCP `fetch_clip`, any sidecar lookup) missed, and the fallback page URL would have been -`/videos/v<n>`. Two channels carry Twitch VODs; no curated tag or site report references a `v<n>` id (counts 0). +Built as `433815d4` and reverted on this branch; the build is kept on `r20/twitch-ids` (orchestrator, 2026-10-10: +it changes published URLs, which is the operator's ruling). -**What it does.** `canonicalTwitchVideoId` (`lib/videoId.ts`) is the one normalization: a `v` followed by digits -loses the `v`. `summarize` uses it for a Twitch record; `readNormalizedTranscript` applies it (`normalizeSummaryId`) -to a `transcript.cues.json` written before, so the index and report compose read canonical ids from either source; the -index re-processes, once, every record it holds under a v-id (`Twitch ids v1: N record(s) indexed under the native -v-id, re-keyed.`, the platform-labels shape, recorded only with no channel held), and the existing re-key path removes -the old key. The Twitch player is still handed `v<n>` (`TwitchPlayer.tsx`). +**What is there.** A Twitch VOD's directory is the canonical id, the URL's `/videos/<n>` (`extractVideoId`); +`summarize` takes yt-dlp's native `meta.id`, `v<n>`, so the record's id and slug say `v<n>`. Two channels carry +Twitch VODs: `hasanabi` (180 directories) and `shondo-twitch` (37). No curated tag or site report references a `v<n>` +id (counts 0). -**What it changes in public:** at the next index build and site build, every Twitch record's id and slug, and so its -transcript page URL, move from `…/v<n>` to `…/<n>`. +**Readers that miss the join today** — each opens `data/<record id>/` with the record's `v<n>`: +- MCP `fetch_clip` → the editor's fetch-window (`data/v<n>/` holds no metadata, so no URL and no cache); +- the evidence-clip tiers (`common/lib/evidenceClip.mjs`, `videoDirOf`) behind reports prepare and report-to-video's + local sources, and `publish/reportMedia.ts`'s per-record availability read; +- umtool's local cue lookup (`report-to-video/cues.mjs`, `data/<videoId>/transcript.cues.json`). + +**The two options:** +- **(a) Records take `<n>`** (what `r20/twitch-ids` does: `canonicalTwitchVideoId` in `summarize` and in + `readNormalizedTranscript`, a one-shot index re-key, the player still given `v<n>`). 217 public transcript URLs move + from `…/v<n>` to `…/<n>` (hasanabi 180, shondo-twitch 37); the hub needs tombstones for the old ones. +- **(b) The published `v<n>` stays**, and the join normalizes on the directory side (a reader maps a Twitch record id + `v<n>` to `data/<n>/`). No URL moves. **Commits** @@ -111,14 +117,15 @@ transcript page URL, move from `…/v<n>` to `…/<n>`. |---|---| | `04a075b3` | `common:` D3 — the fall-through index test; FACTS "The caption-track rule"; STATE closes the bug | | `96409e24` | `common:` D1 — the `recordedDate` rule, its derivation, the index pass, the form field; CHANNEL.md | -| `433815d4` | `common:` D2 — `canonicalTwitchVideoId`, `normalizeSummaryId`, the one-shot re-key, the player | +| `433815d4` | `common:` D2 — `canonicalTwitchVideoId`, `normalizeSummaryId`, the one-shot re-key, the player (kept on `r20/twitch-ids`) | +| `949f7e41` | `common:` revert of D2 — held for an operator ruling | **Gates.** `pnpm -r --no-bail --workspace-concurrency=1 exec tsc --noEmit` clean at each commit. New tests: `lib/recordedDate.test.ts` 7, `controller/buildIndexRecordedDate.test.ts` 2 (derived, absent without a rule, a date after the upload dropped; a changed rule re-derives once, a removed one drops the field from the pages), -`controller/twitchId.test.ts` 4, `controller/buildIndexTwitchIds.test.ts` 2 (indexed canonical; an old v-id record -re-keyed once), `controller/buildIndexCaptionTrack.test.ts` +1, editor `app/channels/components/parseChannelForm.test.ts` -3. The whole common suite 3,661/3,661 (release 21 D2's tests included); editor unit 228/228; export unit 118/118; +`controller/buildIndexCaptionTrack.test.ts` +1, editor `app/channels/components/parseChannelForm.test.ts` +3 (D2's 6 left with its revert and live on `r20/twitch-ids`). The whole common suite 3,661/3,661 with D2 in +(release 21 D2's tests included), and after the revert the touched index, normalize, caption-track and architecture tests 54/54, tsc clean; editor unit 228/228; export unit 118/118; mcp 293/293; `archilyzer docs files --check` and `docs env --check` clean. E2E, one run in the foreground (start of the queue after the heavy slot freed): `transcript-source.spec.ts` (D3's index spec, 4) and `ops-api.spec.ts` (13, the channel-config round trip among them) — 17 passed, 1.5 min. Release end per the plan: the full editor suite once, the orchestrator's. ## Rollout