Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit cb56ff9d4982a9579592f6448592819c76ddc2a9
parent c968ddf6c36ff18bdc434d87e8bb69487a74d648
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Fri,  9 Oct 2026 12:09:47 -0400

Merge branch 'r19/integration' into worktree-agent-a103bef2ee4055362

Diffstat:
MPLAN.md | 16++++++++++++++++
MREADME.md | 6++++++
Mcommon/jobs/jobKinds.test.ts | 8+++++---
Mmcp/README.md | 1+
Mplans/README.md | 6++++--
Mplans/STATE.md | 35++++++++++++++++++++++++++++-------
Mplans/editor-operations-ia.md | 5++++-
Aplans/landed-2026-10.md | 91+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mplans/one-core.md | 9+++++++--
Aplans/release-19.md | 123+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aplans/release-20.md | 38++++++++++++++++++++++++++++++++++++++
Mplans/stats-cache-key.md | 2+-
12 files changed, 324 insertions(+), 16 deletions(-)

diff --git a/PLAN.md b/PLAN.md @@ -675,3 +675,19 @@ ollama pull qwen2.5:7b # ~4.7 GB Q4; llama3.1:8b is the alternative The Claude Code lane needs the `claude` CLI installed and authenticated on the host, plus `digest.remoteEnabled` turned on in settings. + +## Beyond the AI track: releases + +The phases above are the AI track. Later work is planned and recorded one release per file under `plans/` +(releases 5–16: `plans/release-5.md` … `plans/release-16.md`). From release 17: + +| release | what | file | status (2026-10-09) | +|---|---|---|---| +| 17 | the media tier: a channel's big files on another drive, its text always on the corpus disk | [`plans/release-17.md`](plans/release-17.md) | complete, LIVE 2026-10-02 | +| 18 | publishing as queueable stages: one index, per-site bundles, serial builds, deploys checked live, the publish lane | [`plans/release-18.md`](plans/release-18.md) | complete on `r18/integration` (`main` merged in, `4cffda3f`); fast-forward + rollout owed | +| 19 | agents run the archive: ops/CLI/MCP (Track A), machine safety and tooling (Track B), OPERATING.md and the docs (Track C) | [`plans/release-19.md`](plans/release-19.md) | in flight | +| 20 | the data model and the index: recorded dates, Twitch ids, the caption-track bug closed | [`plans/release-20.md`](plans/release-20.md) | planned, after 19 | +| 21 | — | — | not yet planned | +| 22 | — | — | not yet planned | + +Work that landed on `main` without a plan of its own: [`plans/landed-2026-10.md`](plans/landed-2026-10.md). diff --git a/README.md b/README.md @@ -644,6 +644,12 @@ See [CONTRIBUTING.md](CONTRIBUTING.md) to work on the code. | [PUBLISH.md](PUBLISH.md) | Building and deploying sites: Pages + R2, cost-abuse protection, parallel builds in containers. | | [SCHEDULED_SYNC.md](SCHEDULED_SYNC.md) | Unattended per-channel syncing. | | [WORKTREES.md](WORKTREES.md) | Parallel checkouts and the port scheme. | +| [SETTINGS.md](SETTINGS.md) | Every `settings.json` key (generated). | +| [SITE.md](SITE.md) | Every `site.json` key (generated). | +| [CHANNEL.md](CHANNEL.md) | Every key of a channel's `config.json` (generated). | +| [REPORT.md](REPORT.md) | The cited report document, `report.json` (generated). | +| [CITATIONS.md](CITATIONS.md) | The citation model every cited document shares (generated). | +| [AGENTS.md](AGENTS.md) | Instructions for coding agents working in this repo. | | [mcp/README.md](mcp/README.md) | The MCP server: tools, links, client setup. | ## License diff --git a/common/jobs/jobKinds.test.ts b/common/jobs/jobKinds.test.ts @@ -274,10 +274,12 @@ test("release 18: the kinds no longer created keep their labels (seven of them n assert.equal(jobKindLabel("build-deploy-homepage"), "Build & deploy homepage"); }); -test("release 18: every drainable kind but the lane runners is an ingest kind", () => { +test("release 18: every drainable kind but the lane runners and the clip-window batch is an ingest kind", () => { const runners = new Set(["auto-transcribe", "auto-download", "auto-digest", "auto-backfill", "auto-publish"]); + // fetch-windows writes only `data/<id>/clips/`, a cache the index never reads. + const notIngest = new Set(["fetch-windows"]); for (const kind of jobKindIds()) { - if (!isDrainableKind(kind) || runners.has(kind)) continue; + if (!isDrainableKind(kind) || runners.has(kind) || notIngest.has(kind)) continue; assert.equal(isIngestKind(kind), true, `${kind} is drainable and per channel: an ingest kind`); } for (const kind of runners) assert.equal(isIngestKind(kind), false, kind); @@ -286,7 +288,7 @@ test("release 18: every drainable kind but the lane runners is an ingest kind", assert.equal(isIngestKind(kind), true, kind); } // Nothing that publishes, moves or only reads. - for (const kind of ["publish-build-site", "build-export", "relocate-channel-media", "scan-media", "fetch-window"]) { + for (const kind of ["publish-build-site", "build-export", "relocate-channel-media", "scan-media", "fetch-window", "fetch-windows"]) { assert.equal(isIngestKind(kind), false, kind); } assert.equal(isIngestKind("no-such-kind"), false); diff --git a/mcp/README.md b/mcp/README.md @@ -15,6 +15,7 @@ clip window; the MCP itself still writes nothing. | Tool | What it does | |------|--------------| | `list_channels` | List channels **organized under their channel groups** (name, slug, video count; site in hub mode), with a compact group cheat-sheet (`id · name · N channels`) for scoping. | +| `list_tags` | The **curated tags** the archive publishes (the operator's cross-channel vocabulary, what `tags` on `search_transcripts` / `enumerate_matches` filters by): each tag's id, label, group, how many videos carry it on this source, and the per-channel breakdown. Not the yt-dlp keywords `scopes:["tags"]` searches. A source that publishes none says so. | | `list_reports` / `get_report` | The cited reports a site publishes (corpus spec 5): the index (title, kind, claim and citation counts, a fact-check's verdict tally), then one report's sections and claims with every citation's **verbatim quote, original URL and the site's moment page**. See [Cited reports](#cited-reports). | | `search_transcripts` | Search captions for a term/phrase (or regex); returns matching videos with timestamped snippets — **each `[mm:ss]` is a clickable link to that exact moment** (or a compact `[mm:ss\|sec]` with `link_style:"base"`). Alias-aware, pageable, and **filterable** (`states`, `date_from`/`date_to`, `media_type`, `age`, `exclude`, `scopes`). A page that isn't the whole match set is flagged **above** the hits. | | `enumerate_matches` | A query's **complete** match set as a worklist (id/title/channel/date + batch count) in **one scan**. Takes the **same filters** as `search_transcripts`, so the two can never disagree about coverage. The tool to use whenever you need to count or cover everything. | diff --git a/plans/README.md b/plans/README.md @@ -2,7 +2,7 @@ The local-AI derived-corpus work spans many phases and months, and **context is cleared between phases** to keep token cost down. That only works if a cold agent can resume without -re-deriving anything. Five kinds of artifact, with deliberately different lifetimes: +re-deriving anything. Seven kinds of artifact, with deliberately different lifetimes: | File | Lifetime | Contents | | --- | --- | --- | @@ -10,7 +10,9 @@ re-deriving anything. Five kinds of artifact, with deliberately different lifeti | [`FACTS.md`](FACTS.md) | Append-only | Verified codebase facts with `file:line` anchors. | | [`STATE.md`](STATE.md) | Rewritten each session | Phase status, decisions log, open questions, commands known to pass. | | `phase-N-<slug>.md` | Written at phase start, archived at merge | File-level detail for the phase in flight. | -| `<topic>.md` design docs | Live until the design is fully shipped | A model that spans phases and outlives any one of them: [`unified-operations-model.md`](unified-operations-model.md) (how media-derived work is dispatched — the backend half) and [`editor-operations-ia.md`](editor-operations-ia.md) (what the editor's screens are about — the UI half, nine slices, slice 1 shipped). Read the matching one before touching a dispatcher or a page. | +| `<topic>.md` design docs | Live until the design is fully shipped | A model that spans phases and outlives any one of them: [`unified-operations-model.md`](unified-operations-model.md) (how media-derived work is dispatched — the backend half) and [`editor-operations-ia.md`](editor-operations-ia.md) (what the editor's screens are about — the UI half, nine slices, complete). Read the matching one before touching a dispatcher or a page. | +| `release-N.md` | Written at a release's start; its record grows per slice | One release: rulings, slices, the "as shipped" records, the rollout. Indexed at the end of `../PLAN.md`. | +| `landed-<yyyy-mm>.md` | Written once | Work that reached `main` without a plan of its own: merges and what they landed. | `FACTS.md` is the most valuable file here. It holds exactly the material that is expensive to establish and cheap to get wrong — the original draft of this plan contained several diff --git a/plans/STATE.md b/plans/STATE.md @@ -3,7 +3,27 @@ The working memory for the local-AI derived-corpus work. Rewritten at the end of every session, before context is cleared. See [`README.md`](README.md) for the protocol. -**Now (2026-10-06): release 18 — publishing as queueable stages — is complete on `r18/integration`** (record: +**Now (2026-10-09): release 18 is complete on `r18/integration`, and `main` is merged into it (`4cffda3f`, 94 +commits ahead of `main` `e921f82f`, 0 behind: a fast-forward).** The merge brought main's ~70 commits since 2026-10-06 +(umtool articles + notes, `ops transcribe` with word timings, fetch-windows, metadata refresh, channel +create/rename/delete over ops; record: [`landed-2026-10.md`](landed-2026-10.md)); the changelogs keep both sides, r18 +first. Nothing of release 18 is live. +- **Owed, in order:** fast-forward `main` to `r18/integration` from a session that is not worktree-isolated (the + primary clean first); the rollout per `~/reports/release-18/RUNBOOK.html` (editor rebuild + restart, Publish now, + jeralyzer preview, hub tombstones, step 5b `r18-eva-sites.sh` + production deploys of hasanalyzer and bonnellyzer, + policies, the lane on; the script guard sha updated to the merge); then pruning — the five `r18-*` slice worktrees + and their branches, `render-local-sources`; the ~170 merged branches and the stale worktrees (`diet-series`, + `export-gumroad-tip`, `r13-phase-5`) are listed for the operator to approve, never deleted unasked. +- **The five working sessions** (Super Hasanalyzer, Super Jeralyzer, Super Jasolyzer, candace owens corpus analysis, + unified media aggregation) are doing corpus and content work and hold no unmerged code. The tooling gaps they + reported are releases 19–20. +- **Releases 19–20 are in flight on three tracks** (index: `PLAN.md`, "Beyond the AI track: releases"): Track A + [`release-19.md`](release-19.md) A1–A9 (agents run the archive: ops, CLI, MCP) then + [`release-20.md`](release-20.md) (data model); Track B release 19 B1–B6 (machine safety and tooling); Track C docs + and plans (OPERATING.md and `archilyzer docs cli`, doc fixes, homepage docs, FACTS). Integration branch + `r19/integration`, branched from `4cffda3f`. + +**Previously (2026-10-06): release 18 — publishing as queueable stages — is complete on `r18/integration`** (record: [`release-18.md`](release-18.md): slices S1 the stage contract, stamps, lock, bundles and CLI; S2 deploy hardening; S3 the status, the queue and the publish lane; S4 the surfaces; S5 the image half; S6 the records — each reviewed, merged `--no-ff`; `main` `edadc712` merged in first). **Not on `main` yet, and nothing of it is live.** @@ -102,8 +122,9 @@ written), [`release-14.md`](release-14.md), [`release-15.md`](release-15.md). Th - The JSX entity/whitespace hazard sweep: 22 texts in 19 files run a word into the element before them (FACTS, "Use with AI is the homepage's AI and MCP doc"; release 16 DX). -**Now (2026-09-28, night): the stats cache key fix — built, reviewed (SHIP AFTER FIXES, then SHIP -on re-review; every touch-up done), not merged.** The branch is `fix/stats-cache-key`, and [`stats-cache-key.md`](stats-cache-key.md) +**Previously (2026-09-28, night): the stats cache key fix — built, reviewed (SHIP AFTER FIXES, then SHIP +on re-review; every touch-up done), not merged.** (Merged `10cefd15` the same evening; the recount rolled out with +release 15, the homepage at 77,547 transcripts on 2026-09-30.) The branch is `fix/stats-cache-key`, and [`stats-cache-key.md`](stats-cache-key.md) holds the record, the review and the rollout. FACTS has "The stats cache key". - **What it fixes:** the homepage showed Jasolyzer as 0 transcripts, 0 channels, 0 hours while it served 1,889 videos. Instance-wide it showed 49,798 transcripts of about 77,000. @@ -127,7 +148,7 @@ holds the record, the review and the rollout. FACTS has "The stats cache key". and rolled out, the rollout's step 3 `Diff:` check stays the safeguard: thousands removed means a drive was missing. -**Now (2026-09-28, evening): release 12 — the source mirror — is merged to `main` and NOT rolled +**Previously (2026-09-28, evening): release 12 — the source mirror — is merged to `main` and NOT rolled out.** [`release-12.md`](release-12.md) holds Q's and R's records, their reviews, "Merged" and "Rollout". The operator's runbook is `~/reports/release-12/RUNBOOK.html`, with its scripts in `~/reports/release-12/scripts/`. The plan is [`source-mirror.md`](source-mirror.md). @@ -164,7 +185,7 @@ out.** [`release-12.md`](release-12.md) holds Q's and R's records, their reviews - **Baselines now:** common **2,149**, homepage unit **7**, homepage e2e **36**; editor unit 85, `test:scripts` 185 + 1 and mcp 269 are unchanged. -**Now (2026-09-28, afternoon): release 11 is LIVE as 0.10.0, and Jasolyzer is launched.** The +**Previously (2026-09-28, afternoon): release 11 is LIVE as 0.10.0, and Jasolyzer is launched.** The rollout ran 12:11–14:05 ([`release-11.md`](release-11.md), "Rollout, as done", every deploy and job id): the cut (editor `24c8352e`, export `bf6904e8`), ONE :3001 restart (`BUILD_ID` `vWCJb87ktCy5akih_pM9X`, on `main` `e6c5d2e3`), a Settings save (two ran; the second byte-identical — `buildPipeline.mode` dropped, the @@ -187,7 +208,7 @@ hub and the homepage. **The LM chat-only tier is closed as moot** (Legal Mindset Phase 5 (P0, the plan doc `one-core-phase-5.md`, awaits the operator's OK and five rulings). - **Worktrees:** the seven `r11-*` stay until release 12's are gone (index-based ports). -**Now (2026-09-28, morning): release 11 is merged to `main` in full and NOT rolled out.** The +**Previously (2026-09-28, morning): release 11 is merged to `main` in full and NOT rolled out.** The overnight of 2026-09-28 ([`release-11.md`](release-11.md): every slice's record, then "Integration, as merged", then "Rollout"). `main` = `eb28a341` + three integration `plans:` commits (FACTS `eeff74c3`, `brand-and-themes.md` `3643c974`, then this STATE and the record); no code @@ -334,7 +355,7 @@ the history, newest first from "Accent swap and mark contrast rolled out". - re-uploading the YouTube picture, watermark and banner (`~/reports/archilyzer-media/brand/`). - **Next:** release 11, the overnight of 2026-09-28 ([`release-11.md`](release-11.md)). -**Now (2026-09-26, afternoon): release 10 is merged to `main` in full and NOT rolled out.** It is +**Previously (2026-09-26, afternoon): release 10 is merged to `main` in full and NOT rolled out.** It is the brand (S0–S3: the Found-line mark and base × accent reader themes; its own final commit `bb6f378f`, then cut as editor + export 0.9.0 at 01:37, `0e0d8f59` + `31798769`), plus brand S4 and the lows L1 + L2. Those three merged 2026-09-26 12:52 on the operator's word, before the rollout, diff --git a/plans/editor-operations-ia.md b/plans/editor-operations-ia.md @@ -1,6 +1,9 @@ # The editor: one noun, not a pile -**Status: the model is decided; slices 1, 2 and 3 have shipped.** This is the UI half of +**Status: COMPLETE (verified against the tree 2026-10-09).** Slices 1–8 shipped (records below). Slice 9 shipped +as one-core Phase 1 slice 1.3 (`253826b` → `451c454`, 2026-09-07; record `one-core-phase-1.md`): `digestSweep.ts`, +`backfillSweep.ts`, `arbiter.ts`, `SweepLane.tsx` and `SweepScope.tsx` are gone, the sweep resume hooks left +`instrumentation.ts`, and each lane has one console. This is the UI half of [`unified-operations-model.md`](unified-operations-model.md), which states the backend half ("four schedulers for what is really one kind of work"; steps 1, 5, 6 open). Read that one first if you are touching a dispatcher; read this one if you are touching a page. diff --git a/plans/landed-2026-10.md b/plans/landed-2026-10.md @@ -0,0 +1,91 @@ +# Landed without a plan — `main`, 2026-10-05 → 2026-10-09 + +`main`'s first-parent history `39abec18..e921f82f`: 70 entries, 246 commits. Merged into `r18/integration` at `f2fd11a7` +(through `edadc712`), `ff9c6b7d` (through `f6137a04`) and `4cffda3f` (the rest). The report-site slices have their own record, +[`report-sites.md`](report-sites.md); everything else here had none. Each row is the merge and what it landed. +FACTS has the verified details for the starred rows ("Verified 2026-10-09"). + +## Report sites (recorded in `report-sites.md`) + +| merge | what | +|---|---| +| `6c140bc1` … `b94742c7` | R0–R8 + R4b: the citation model and report document, site keys, evidence media, compose, pages, consumers, converters, the editor's Reports tab, the cited-site e2e | +| `89d65ae0` | `site.json` `search` switch replaces `publish`; search off = report-only; `export start` serves moment dirs | +| `30447593` | per-report revision history (a bare git repo per report, history page + `history.json`, dumb-HTTP clone) | +| `a1a98c2c` | per-report HTML, PDF, Markdown and evidence-pack exports; `reports-export` job; shared publish size limit | + +## Reports and citations + +| merge | what | +|---|---| +| `32c49f45` | a site's report ids use the report document's id rule | +| `bf5050b7` | a report claim may carry a short title; a source may carry a note | +| `7473e3bb` | a report quote is checked against every transcript of the record; its moment shows the one that matched | +| `a01b9f63` | export: one report = the home page; the index without a duplicate title | +| `4f6ee08a` | report series tag; no index of one; lean index cards | +| `53014ad1` | report polish: claim flag, series line, byline + attribution, source rail, citation origin, three reading tiers | +| `b26eaff0`, `f986152f` | report header: our heading, one date/revision line, the reviewed document as a citation card | +| `4fa5c46b`, `694217b9` | evidence clips stop at the recording's end; prepare and compose share the clamped cited span | +| `f6137a04` | a `#<id>` link in report markdown stays in the tab | +| `125163ab` | a report can carry a video | +| `e99f0e4f` | a cite label may hold an editorial `[insertion]` | +| `7f688cf9` | export: `serve-out` binds `HOST` when set | + +## Sources and posts + +| merge | what | +|---|---| +| `9cb6bf99` | a metadata-scan members-only or private verdict keeps a video out of the auto-download queue | +| `96812e72` | `archilyzer feeds backfill-metadata`: a podcast feed's titles, dates and descriptions onto feed-imported records through a history-recorded writer; job kind + ops route | +| `21cc5741` | XenForo threads as a posts source: paced fetch, saved-page import, forum session Connect, post capture | +| `ea3bc5cc` | forum mirror hosts | +| `c85b7155` | archive.org items and files as a video/audio source: polite import, provenance, file player, archive.org + torrent links on citations | +| `dc6a3c3c` | archive.org files over BitTorrent via aria2c, seeded; verified direct-download fallback; no yt-dlp for archive.org | +| `5f88bfee`, `c9fcdc80` | archive.org file records: clean titles, dates from the file name or folder; `archive-org refresh` | +| `b8466d24` | BitChute as a platform: detection, polite import and pacing, native player, relabel pass | +| `33db0dd4` | Wayback captures named by what they copy; `wayback.json` provenance; `wayback refresh`; archived-copy links | +| `edadc712` | a one-off Odysee or BitChute import waits a floor after the last one there | +| `0adfcb23` | a renamed social channel's posts read as their new slug | +| `dafdd904` | an older X walk can start at a date; Kiwi Farms player uploads are video media | +| `283dce07` | forum capture downloads the media a post's page holds now | + +## Transcripts + +| merge | what | +|---|---| +| `200a3105` | one caption-track rule: en-orig first, empty tracks fall through; cue-block VTTs parse; a one-shot index re-read | +| `e506f738` | every English track searchable; the reader and MCP choose a track; alternates only where words differ | +| `4f1d647f` ★ | `pnpm ops transcribe` / `POST /api/ops/transcribe`: one local file through the editor's workers (`3524be34`, `6814ae6f`) | +| `772ae618` ★ | `"words": true` returns the engine's word timings (`6d94995f`) | + +## Editor and ops + +| merge | what | +|---|---| +| `41029f40` ★ | refresh one video's metadata from the video page and over ops; refuses a video with no `data/<id>/` (`26dad6a0`, `9c1df123`) | +| `9e1e8cc2` ★ | rename and delete a social channel from its page; channel create/rename/delete/sites over ops (`b16653af`) | +| `0c9fbe9d` | Rumble clip windows retry with `-extension_picky 0` | +| `de11b507` ★ | fetch-windows: a paced batch of clip windows, one job per platform; the /jobs row names who asked, for what, how many (`414a7bb2`, `4e404ff9`, `103d81af`) | +| `f011109d`, `1a17d841` | a window's own URL picks its queue and its platform args | +| `a2448b11` | one Rumble page load per window | +| `903a4d0a` | Rumble windows 120 s apart, across runs | +| `20ad060a` | page titles lose their lead paragraphs; /channels' header wraps | +| `19b03898` | e2e: export-search's filter checkboxes located by exact name | + +## umtool and report-to-video + +| merge | what | +|---|---| +| `e2c6eb14` | report-to-video assembles a long cut in crossfaded batches, then joins them | +| `ddaef3f0` | deck: posts ride the whole clip, `shotMaxHeight`, web platform, per-post variant; `snapLead`/`snapTail` | +| `8cf1201f` | post cards: accent, logo, flag | +| `de8eebab` | umtool takes: alternative renders of part of a cut, side by side, judged | +| `fac257cf`, `3d09ee71`, `526b2b22` | teaser rows that change in place (per-line role and hold); a second tier pops with its line; a small tier above | +| `f294d107` | thread rail, flips panel, panning screenshots, the date always shows | +| `38c11e41` | the renderer runs without a display | +| `dc35dfd7` | dim dots ahead; outcomes wait for the rail | +| `e921f82f` ★ | umtool articles + notes: `/sites` (every site's articles, site pages, media, workspace files); the notes store (`notes.json` beside a report, never published, exported or committed to history), `/api/notes`, `umtool notes` CLI + agent digest; timed, take and edit notes; the timeline and structure editors; kinds declared on the registry (`d319982f` … `847b8188`, 20 commits) | + +## Plans commits + +`1af693a4`, `86e49763`, `53c8d776`, `771a55be` — the report-sites plan and its record. diff --git a/plans/one-core.md b/plans/one-core.md @@ -317,7 +317,7 @@ https://jeralyzer.pages.dev/corpus.json` before and after, byte-identical output > `config.json`, sidecars) shipped 2026-09-24 (`3241fed2`, `ef88ac4d`). Release 4 shipped the > same evening: slice P (/channels rack polish, `77a32de2`), slice W (one write idiom, > `ddad13f4`) and slice 3b (the `VideoPanel` split, `e172749b`). Release 4 is not yet live -> on :3001. Phase 4 is next. +> on :3001. Phase 4 is next (done 2026-09-28, below). > Records, gates and what is still open: [`one-core-phase-3.md`](one-core-phase-3.md). 1. **`common/views/`**: move `buildActiveJobs`, `buildWorkers`, `lanes.ts`, @@ -459,7 +459,12 @@ counters and the private tmp + rename code at 26 write sites.* ### Phase 4 — CLI, entry points, config, docs (3 slices) -> **Next** (after the exports-off and Rumble releases, see +> **DONE 2026-09-28** (release 11): slice 3 checkpoints A (`cb9d02b2`: `archilyzer doctor`, `run`, `mcp`, every bin a +> subcommand, generated `ENVIRONMENT.md`, `PUBLISH.md` absorbing the deploy docs) and B (`eb28a341`: test-only env vars +> `E2E_`-prefixed, the harnesses read `ports.mjs`); the build-mode stub dropped (O6c). The optional SETUP/PLAN.md +> consolidations were not done. Record: `release-11.md`, STATE "release 11". +> +> *Was:* **Next** (after the exports-off and Rumble releases, see > [`one-core-phase-3.md`](one-core-phase-3.md#next--phase-4)). Starting points, checked on > `e172749b`: > - `editor/app/sites/lib/buildDeployCore.ts` is 578 lines and imports nothing from diff --git a/plans/release-19.md b/plans/release-19.md @@ -0,0 +1,123 @@ +# Release 19 — agents run the archive (ops, CLI and MCP; machine safety; the operator's guide) + +Written 2026-10-09 from the "next batch" plan (releases 19–22), against `r18/integration` `4cffda3f` (release 18 +complete, `main` `e921f82f` merged in). Rules: `plans/tools/implementer-rules.md` — three parallel Opus implementers, +one per track, each in its own worktree branched from `4cffda3f`; the parent merges each track into `r19/integration` +`--no-ff` on a clean tree, then fast-forwards `main`; records state rulings never reasons; counts-only privacy greps +before every merge; implementers' commits carry their OWN model's `Co-Authored-By` + the session line; never boot a +second editor against the real corpus; the homepage build gate is `build:nodata`; memory gate (`free -m` available +≥ 6 GB) before any build, render or e2e, every `next build` under a 5 GB `systemd-run` scope. Record file: this file +(shape of `plans/release-18.md`). Release 20 (the data model) is [`release-20.md`](release-20.md). + +## Context + +- Five working sessions (Super Hasanalyzer, Super Jeralyzer, Super Jasolyzer, candace owens corpus analysis, unified + media aggregation) do corpus and content work and hold no unmerged code. Every gap they report is repo-wide + tooling: wrappers that find `WORKER_TOKEN`, `joblog.sh`/`watchjobs.sh` job watchers, hand-run archive.org imports, + hard-linking another site's `export/out` aside, two OOMs from concurrent heavy work. +- Precedents to reuse: `editor/app/api/ops/_lib.ts` (`ops`, `opsAuth`, field readers, `jobResponse`), route unit + tests (`editor/app/api/ops/{fetch-windows,transcribe,rename-channel}/route.test.ts`), `scripts/archilyzer-ops.mjs` + (`ACTIONS`, `GETTERS`, `--wait`), the views (`common/views/names.ts`, `/api/view/*`), `mcp/src/fetchClip.ts`. +- Tests: the editor e2e is 130 specs / ~709 tests / ~40 min, serial, behind a machine-wide lock; the editor has 29 + unit-test files. No CI, no root `test` / `typecheck`; the editor's `lint` has no config. +- **Precondition (Wave 0, release 18's rollout):** `main` fast-forwarded to `r18/integration`, the release 18 rollout + run, the `r18-*` worktrees pruned. Every `editor/app/api/ops` or `archilyzer-ops.mjs` change merges after it. + +## Rulings (operator, 2026-10-09; do not re-open) + +- **The MCP gains archival writes through the editor** — the `fetch_clip` pattern: the MCP server asks the editor + (editor URL + `WORKER_TOKEN`), the editor writes. Settings, storage and deletes stay CLI/ops-only. +- **Three parallel tracks**: A owns the e2e slot (code over ops/CLI/MCP, sequential), B is machine safety and tooling + (no editor e2e; one umtool run), C is docs and plans (no e2e). +- **Release 18 lands first**, by a session that may touch the primary checkout. +- **e2e budget:** one editor/umtool/export e2e run at a time machine-wide; the full editor suite only at a release's + end. Every slice carries its class: `[none]` docs/plans only, `[unit]` unit tests are its gate, `[spec]` one or + two named specs, `[suite]` the full suite. +- Polite scraping applies to every import and listing: one stream, real gaps, back off on 429. + +## Slices + +### Track A — release 19 (owner: the Track A implementer; sequential; owns the e2e slot) + +| slice | class | what | after | +|---|---|---|---| +| **A1** `pnpm ops` finds its token | `[unit]` | read `WORKER_TOKEN` / `ARCHILYZER_EDITOR_URL` from `editor/.env` when unset; `--wait` polls a real authed route instead of the `/api/jobs/active` rewrite | Wave 0 | +| **A2** jobs over ops | `[unit]` + `[spec ops-api]` | `get jobs [--active\|--failed\|--kind\|--slug]`, `get job <id> [--tail]`, `job cancel\|retry\|retry-failed\|drain\|promote\|force-release <id…>`, `--wait` on many ids with queue position; wraps `editor/app/jobs/actions.ts` | A1 | +| **A3** read side | `[unit]` | `get settings\|storage\|sites\|workers\|auto-queue\|scheduler\|cleanup <slug>` over the existing view builders/readers; `/api/auto-queue/control` gated behind `opsAuth` (unauthenticated today) | A2 | +| **A4** archival writes | `[unit]` | `settings` patch (through `saveSettings` + the schema); `lane` start/stop/drain for every lane incl. `publish`; `clear-platform-hold`; `workers enable\|disable`; per-video `transcribe-one\|delete-file\|do-not-clean`; cleanup buckets; `relocate {dryRun}`; `archilyzer storage report` (tierable bytes per channel) | A3 | +| **A5** fetch queue hygiene | `[unit]` | the editor dedupes identical fetch-window requests (same slug/id/window → the existing job); `fetch-via-editor.mjs` waits without a timeout, printing queue position; small fetch-window jobs get their own queue key / priority instead of waiting behind a multi-hour `persist-videos` | A4 | +| **A6** imports | `[unit]` | `import-archive-org {items:[…] \| query}` with the polite inter-item gap, a `dryRun` that flags restricted (401) items, skip-if-held; `get remote-listing <slug>` for Odysee/BitChute (paced, one stream) diffed against held videos | A5 | +| **A7** cues without an index | `[unit]` | `pnpm ops build-cues --slug [--ids]` (normalize → `transcript.cues.json`); MCP `get_transcript` falls back to the VTT | A6 | +| **A8** publish edges | `[unit]` | `POST publish {verb:"build"}` takes `allowMissingMedia`; `pnpm ops lane` covers `publish`; `archilyzer publish build <id> --out <dir>` for a private build served elsewhere (cheap on release 18's bundles) | A7 | +| **A9** MCP archival writes | `[unit]` | on the `fetch_clip` pattern: `get_job` (status + log tail), `enqueue` for sync / download-missing ids / transcribe-bucket / fetch-posts / import-video (each → the existing ops route), `channel_coverage` (held videos by date, gaps), notes `list\|read\|reply` (umtool `/api/notes`, `UMTOOL_URL`); instructions in `mcp/src/instructions.ts`; the README tool table | A8 | +| release end | `[spec ops-api]` + `[suite]` | one `ops-api.spec` run, then the full editor suite once | A1–A9 | + +### Track B — release 19 (owner: the Track B implementer; no editor e2e) + +| slice | class | what | after | +|---|---|---|---| +| **B1** heavy-work gate | `[unit]` | `scripts/queue-lock.mjs` generalized into `pnpm heavy -- <cmd>`: one heavy slot machine-wide + a ≥ 6 GB free-memory floor; used by e2e, `next build` (publish stages, `build-site.sh`), `build-video.mjs` renders; a render may hold the transcription lane for its duration (ops `lane`); documented in WORKTREES.md + AGENTS.md | — | +| **B2** report-to-video robustness | `[unit]` | prune `*-frames` after the final mux (`--keep-frames` keeps them); `--chrome-only` builds missing segments; a manifest lint before render (teaser > 34 chars, image src relative to the manifest, a QR legible at 720p). `umtool/report-to-video/` is the Candace session's: coordinated with it | B1 | +| **B3** report CLI | `[unit]` | `archilyzer reports check <site> [--reports …]` (compose without a build); `reports verify-quotes <report.json>` (quote vs cue span, the en-orig check); `reports attach-video` encoding to fit the 24 MiB compose limit | — | +| **B4** ops transcribe | `[unit]` | `durationMs` excludes queue wait; a priority for short one-file jobs over long auto jobs | Wave 0 | +| **B5** umtool small debts | `[unit]` → `[spec umtool]` | per-worktree umtool e2e ports (`portFor`); `mix.spec` order dependence; `umtool window` writes edit notes; a selection spanning two blocks gets a Note button; the CLI finds `SITES_DIR` from the repo, not the cwd; umtool's /sites heading becomes "Articles" (naming hazard vs the editor's /sites). ONE umtool e2e run | B1 | +| **B6** test economy | `[unit]` | pure/API e2e moves to unit: `audio-check-classifier.spec` → common; `ops-api.spec` routes → `route.test.ts`; `view-route`, `worker-unit`, `media-file-abort`. Root `pnpm test` and `pnpm typecheck`; the editor gets an eslint config or loses its broken `lint` script | — | + +### Track C — docs and plans (owner: the Track C implementer; no e2e) + +| slice | class | what | +|---|---|---| +| **C1** plans hygiene | `[none]` | STATE "Now"; PLAN.md's release index; this file and `release-20.md`; the "landed without a plan" record ([`landed-2026-10.md`](landed-2026-10.md)); one-core Phase 4, `editor-operations-ia.md`, `stats-cache-key.md` statuses | +| **C2** OPERATING.md + `archilyzer docs cli [--check]` | `[unit]` | recipes (add a channel; sync and download; import from archive.org / Odysee / BitChute / Wayback; fetch and capture posts; transcribe; tag; prepare and export reports; publish) over `pnpm ops`, `archilyzer` and the MCP; a generated command reference from `common/bin/_cli.ts` `usage()` and `scripts/archilyzer-ops.mjs` `usage()` on the `env-docs.ts` / `file-schemas-docs.ts` pattern; linked from AGENTS.md, README, CONTRIBUTING (its hand-picked commands replaced by the link). Regenerated after Track A's actions land | +| **C3** doc fixes | `[none]` | `mcp/README.md` lists every tool; README's docs table adds AGENTS, SETTINGS, SITE, CHANNEL, REPORT, CITATIONS; RUNNING_IN_DOCKER's ops list points at the generated reference | +| **C4** homepage docs | `[unit]` | `homepage/content/docs/operate.md` gains the CLI/ops/MCP overview; gates `build:nodata` + `pnpm --filter homepage test` | +| **C5** FACTS refresh | `[none]` | verified facts for main's last ~70 commits (notes store, articles kind registry, ops transcribe word timings, fetch-windows pacing, metadata refresh) | + +### File ownership + +| track | owns | +|---|---| +| A | `editor/app/api/ops/**` (except B4's transcribe route), `scripts/archilyzer-ops.mjs` + its test, `mcp/src/**`, `editor/app/jobs/actions.ts`, `common/controller/fetchWindows.ts`, `fetch-via-editor.mjs` | +| B | `scripts/queue-lock.mjs` and the heavy gate, `umtool/**` (except `umtool/report-to-video/*`, the Candace session's), the e2e specs B6 moves, the report CLI, `common/controller/transcribeFile.ts`, `editor/app/api/ops/transcribe/**` | +| C | `plans/**` (except the records others append), root `*.md` docs, `homepage/content/docs/**`, `mcp/README.md`, the `archilyzer docs cli` generator | +| shared, append-only | `editor/CHANGELOG.md`, `AGENTS.md`, `common/bin/_cli.ts` (rows only) | + +### Graph and merge order + +``` +Wave 0 (release 18 lands) ──► A1 ──► A2 ──► … ──► A9 ──► release-end suite (Track A is sequential: A1–A5, then A6–A9) + └─► B4 +now: B1 ──► {B2, B5}; B3; B6 (B6 touches no release-18 file) +now: C1, C3 ──► C2 ──► C4, C5; C2 regenerated after A's actions land +``` + +- **Merge order:** Wave 0 before any `editor/app/api/ops` or `archilyzer-ops.mjs` change (they conflict with release + 18). C1, C3 and B6 touch no release-18 file and may merge first. Track A merges slice by slice in its own order; + B and C merge whenever a slice is green. The parent regenerates the command reference (C2) after each Track A merge + that adds an ops action or CLI row. +- **Do not touch:** `umtool/report-to-video/*` without the Candace session; `transcripts/**` except through the + writers; the working sessions' `~/reports/*` workspaces. + +## Verification + +- Every Track A slice: route unit tests (the `route.test.ts` pattern) + `scripts/archilyzer-ops.test.mjs`; MCP tools + via `pnpm --filter yt-dlp-transcript-mcp test` and one live `/ask`-style session against the editor; one + `ops-api.spec` run per release; the full editor suite once at release end. +- B1: a unit test of the lock + a manual two-build collision showing the second waits. B6: the moved tests pass and + the editor suite shrinks by the moved count. B5: one umtool e2e run. +- C2: `archilyzer docs cli --check` in the gates beside `docs env --check` and `docs files --check`; every command a + recipe names exists (grepped against the generated reference). C4: homepage `build:nodata` + unit. +- Every track: `pnpm -r --no-bail --workspace-concurrency=1 exec tsc --noEmit` clean per commit; counts-only privacy + greps (home paths in added lines 0, no `transcripts/` content). + +## Record + +(Each slice adds a "#### Slice <X>, as shipped" section under its track, in merge order.) + +### Track A + +### Track B + +### Track C + +## Rollout diff --git a/plans/release-20.md b/plans/release-20.md @@ -0,0 +1,38 @@ +# Release 20 — the data model and the index (recorded dates, Twitch ids, the caption-track bug closed) + +Written 2026-10-09 from the "next batch" plan (releases 19–22), against `r18/integration` `4cffda3f`. Track A's, after +release 19 ([`release-19.md`](release-19.md)); small. Rules: `plans/tools/implementer-rules.md` and release 19's +rules line, unchanged. Record file: this file (shape of `plans/release-18.md`). + +## Context + +- A VOD-mirror channel's videos are dated by their upload to the mirror, not by the stream they copy; coverage + ("held videos by date, gaps", release 19 A9 `channel_coverage`) reads the wrong date for them. +- A Twitch video's directory is `<n>` while its record carries `v<n>`; readers that join the two miss. +- The `en` track → 0 cues bug (an `en` VTT that parses to 0 cues preferred over `en-orig`) was filed 2026-10-01. + Main's caption-track merges since (`200a3105` transcripts/en-track-fallback: en-orig first, empty tracks fall + through, cue-block VTTs parse, a one-shot index re-read; `e506f738` transcripts/multi-track) are the candidate fix. +- `plans/stats-cache-key.md`: merged `10cefd15` (2026-09-28); the recount rolled out with release 15 (verified by + release 19 C1). + +## Slices + +| slice | class | what | after | +|---|---|---|---| +| **D1** recorded dates | `[unit]` | `recordedDate` for VOD-mirror channels, derived from the title by a per-channel rule; coverage reads it before the upload date | release 19 | +| **D2** Twitch ids | `[unit]` | one normalization of the Twitch id (`<n>` directory vs `v<n>` record) at the index boundary | D1 | +| **D3** the caption-track bug | `[unit]` + `[spec]` for the index | verify against the tree and a fixture that the `en` → 0 cues case is closed by `200a3105`; close it in STATE and FACTS, or fix what remains | — | + +**Owner:** the Track A implementer. **Merge order:** D3 may go first (verification); D1 → D2. Release end: the index +spec(s) D3 names, then the full editor suite once if any editor code changed. + +## Verification + +- Unit tests per slice; the index spec for D3; tsc clean per commit; `docs files --check` if a schema changes. +- Counts-only privacy greps before each merge. + +## Record + +(Each slice adds a "### Slice <X>, as shipped" section here, before "## Rollout".) + +## Rollout diff --git a/plans/stats-cache-key.md b/plans/stats-cache-key.md @@ -1,6 +1,6 @@ # The stats cache key (fix, 2026-09-28) -Branch `fix/stats-cache-key` from `main` `ac438bbc`. Merged: not yet. Rolled out: not yet. +Branch `fix/stats-cache-key` from `main` `ac438bbc`. **Merged `10cefd15` (2026-09-28). Rolled out with release 15** (the recount; the homepage at 77,547 transcripts on 2026-09-30, `release-15.md` "Rollout"). Verified in the tree 2026-10-09: the cache keys on `indexSignature` (`common/controller/buildStats.ts:119`), and `STATS_DOWNGRADE_ENV` (`:134`) guards a downgrade. ## What was wrong