Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 166a1ac691b412199b8b4c387f0ef1b7a56293ac
parent 93e59abdd569b113a1728de0c31b50f4681da298
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Fri,  9 Oct 2026 13:13:38 -0400

Merge track C (C2 OPERATING.md + docs cli, C4 homepage docs, C5 FACTS, the track's record) into r19/integration

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
MAGENTS.md | 6++++++
ACOMMANDS.md | 134+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
MCONTRIBUTING.md | 14++++----------
AOPERATING.md | 160+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
MREADME.md | 2++
MRUNNING_IN_DOCKER.md | 3+++
Mcommon/bin/archilyzer.ts | 7+++++++
Acommon/bin/cli-docs.test.ts | 140+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/bin/cli-docs.ts | 274+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mhomepage/CHANGELOG.md | 1+
Mhomepage/content/docs/operate.md | 28++++++++++++++++++++++++++++
Mplans/FACTS.md | 110+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mplans/release-19.md | 52+++++++++++++++++++++++++++++++++++++++++++++-------
13 files changed, 914 insertions(+), 17 deletions(-)

diff --git a/AGENTS.md b/AGENTS.md @@ -85,6 +85,12 @@ another markdown file that will drift from it. Use `/ask` for a question answered in the conversation and `/sweep` for a cited report written to a file. +**Running an archive — adding channels, syncing, importing, transcribing, tagging, +reports, publishing — is [OPERATING.md](OPERATING.md)**: recipes over `pnpm ops` (the +running editor's actions), `pnpm archilyzer` and the MCP. Every command and action is in +[COMMANDS.md](COMMANDS.md), generated by `pnpm archilyzer docs cli` from their help; after +adding an ops action or a CLI row, regenerate it (`--check` is a gate). + ## Clips and report-to-video The high-value loop: point the MCP at a public instance, ask about a subject, then pull diff --git a/COMMANDS.md b/COMMANDS.md @@ -0,0 +1,134 @@ +# Command reference + +<!-- GENERATED by common/bin/cli-docs.ts from the `archilyzer` command table (common/bin/archilyzer.ts) and `pnpm ops --help` (scripts/archilyzer-ops.mjs) — do not edit by hand. Regenerate: `pnpm archilyzer docs cli`. --> + +Every `archilyzer` command and every `pnpm ops` action, from their own help. Recipes that chain them: [OPERATING.md](OPERATING.md). + +## `pnpm archilyzer` + +Run from the repo root (in the container: `docker compose exec editor pnpm archilyzer …`). `--help` after a command prints its line. A command marked *passthrough* parses its own flags. + +| command | arguments | what it does | +|---|---|---| +| `archilyzer index` | | rebuild the LMDB transcript index | +| `archilyzer build stats` | | rebuild the stats datasets (reads the index; the data phase's second step) | +| `archilyzer build templates` | | bake each site's chart templates into its export staging dir (the data phase's third step) | +| `archilyzer build archives` | | warm the shared archive-zip cache for every enabled site's channels, once | +| `archilyzer compose site` | `<id> [--allow-missing-media]` | compose one site's export/public (default: SITE\_ID); --allow-missing-media lets a report citation with no prepared media through | +| `archilyzer compose hub` | | compose the hub's export/public (hub-sites.json, corpus.json, …) | +| `archilyzer compose homepage` | | compose homepage/public (whole-pool stats + landing summary) | +| `archilyzer publish index` | | update the index: the LMDB index, the stats datasets and the chart templates in one child (8 GB heap), then the index stamp every build reads | +| `archilyzer publish build` | `<id\|all> [--runner local\|docker\|auto] [--force] [--skip-archives]` | build a site (or every stale one) into its bundle &lt;exportBuildsDir&gt;/&lt;id&gt;/out from the current index; --runner docker builds every site in containers (host only); a fresh site is a no-op without --force | +| `archilyzer publish deploy` | `<id\|all> [--preview <branch>] [--to local] [--force]` | ship a site's bundle to its Pages project (a preview with --preview), or with --to local into ARCHILYZER\_SITE\_OUT; a bundle already deployed there is a no-op without --force | +| `archilyzer publish hub` | `[--deploy \| --deploy-only] [--preview <branch>] [--force]` | build the hub into its bundle &lt;exportBuildsDir&gt;/\_hub/out, then (--deploy) ship it; --deploy-only ships the bundle as built | +| `archilyzer publish homepage` | `[--deploy \| --deploy-only] [--preview <branch>] [--to local] [--force]` | build homepage/out (source mirror included), then (--deploy) ship it; --deploy-only ships it as built | +| `archilyzer publish status` | `[--json]` | the publish status: the index, the lane, and per site / hub / homepage its policy and its built, deployed and live chips, then the plan Publish now would run | +| `archilyzer publish now` | | run what Publish now runs — the index update when it is stale, then each policy target's build and deploy (site.json publish.auto; settings.json publish.hub / publish.homepage), never forced — one stage at a time IN THIS PROCESS under the publish lock, as the editor's queue would (the CLI has no queue); exit 0 when every stage ran or was a no-op | +| `archilyzer stage` | `<kind> <target> --run-id <id> [--preview <b>] [--to local] [--runner docker] [--force] [--skip-archives] [--allow-missing-media] [--index-after <ms>] [--built-after <ms>]` | INTERNAL: one publish stage, as the editor's job runs it (exit 0 ran/no-op, 1 failed, 2 usage, 3 precondition not met, 130 cancelled) | +| `archilyzer build site` | `<id> [--nodata] [--skip-archives] [--allow-missing-media]` | alias: publish index (not with --nodata) + publish build &lt;id&gt; --force (default id: SITE\_ID) | +| `archilyzer build all` | `[--skip-archives]` | alias: publish index + publish build all --runner auto (containers when an engine answers, else serially on the host) | +| `archilyzer build hub` | | compose:hub + INSTANCE\_MODE=hub next build into export/out — a raw build, unstamped; to deploy the hub build it with `publish hub` | +| `archilyzer build homepage` | `[--no-source]` | compose + source publish + next build in homepage/ (reads the index as it stands); --no-source removes the published source instead | +| `archilyzer reports prepare` | `<id>` | cut every clip and copy every post capture the site's published reports cite into its report-media cache, before its build, then export the reports as files (reports export) when nothing is missing (exit 1 when a citation lacks media or an export fails; default id: SITE\_ID) | +| `archilyzer reports export` | `<id> [--report <reportId>] [--formats html,pdf,md,zip] [--allow-missing-media]` | write the site's published reports as files (report.html, report.pdf, report.md, evidence-pack.zip) into its report-exports staging, where compose publishes them from (exit 1 on a problem; a PDF skipped for want of a browser is a note; default id: SITE\_ID) | +| `archilyzer reports convert` | `<sweep\|ask\|manifest> <in> --out <report.json> [--channels-dir <dir>] [--id <id>] [--title <title>]` | a /sweep report (markdown), an /ask answer or a report-to-video manifest as a report.json, written only when it validates (--channels-dir: widen spans from the cues, find posts' channels) | +| `archilyzer reports to-manifest` | `<report.json> --out <manifest.json> [--channels-dir <dir>] [--site-origin <url>]` | a starter report-to-video manifest from a report (a chapter card per section, a claim's still and clips stamped with its verdict, its posts) | +| `archilyzer source publish` | `[--force] [--check] [--keep-scratch]` | the scrubbed git mirror, raw tree, history pages (stagit, when installed) and tarball into homepage/public, behind the denied-literal gate (--check: audit and count, write nothing) | +| `archilyzer source audit` | `[<git dir>]` | the denied-literal gate (+ gitleaks) over a git dir; default the published homepage/public/source/archilyzer.git | +| `archilyzer deploy site` | `<id> [--preview <branch>]` | alias: publish deploy &lt;id&gt; \[--preview &lt;branch&gt;\] — ship the site's bundle to its Pages project (default id: SITE\_ID) | +| `archilyzer deploy hub` | `[--preview <branch>]` | alias: publish hub --deploy-only \[--preview &lt;branch&gt;\] — ship the hub's bundle to homepage.json's Pages project | +| `archilyzer deploy homepage` | `[--preview <branch>]` | alias: publish homepage --deploy-only \[--preview &lt;branch&gt;\] — ship homepage/out to the Pages project archilyzer | +| `archilyzer run` | `<operation> <channel> [ids…] [--lane local\|remote]` | run one catalogued operation over a channel offline, as the editor's job does (sync, downloads and transcription are refused: they run in the editor; it does not see the editor's lanes, so not beside one on the same channel) | +| `archilyzer archive-org refresh` | `<slug> [--dry-run]` | bring a channel's archive.org file records up to their provenance: a name-only mirror's title (where the item gives the file none) and upload date from its file name; offline, through the metadata history; prints old → new | +| `archilyzer wayback refresh` | `<slug> [--titles <json>] [--dry-run]` | bring a channel's Wayback Machine copies up to the Wayback rules: wayback.json (original URL, capture time), the dir renamed to its canonical id through the snapshot's reconcile (roster moved with it), and with --titles (a file of id → {title, upload\_date}) a raw file's title and date; offline, skips a record a live job holds; prints old → new | +| `archilyzer feeds backfill-metadata` | `<slug> [--feed <url>] [--dry-run]` | complete a podcast channel's records (title, date, description, duration) from its RSS feed: one fetch of the feed (default: the channel's url), no media; --dry-run counts matched / unmatched / already complete and writes nothing | +| `archilyzer duplicates` | `[--threshold N] [--all-durations] [--blocking title\|duration\|both] [--near F] [--tolerance N] …` | on-demand duplicate detection (after index + stats) *(passthrough)* | +| `archilyzer posts fetch` | `--slug <channel> [--full \| --older [--floor YYYY-MM-DD] [--from YYYY-MM-DD] [--force]] [--limit N] [--pages N]` | fetch a social channel's posts into its posts corpus (--older: walk back below the oldest archived post; --pages: a forum thread's latest N pages) *(passthrough)* | +| `archilyzer posts import-html` | `<slug> <file-or-dir>… [--dry-run]` | import forum thread pages saved from a browser ("Save page as", .html) into a forum-thread channel: new posts appended, edited ones updated, nothing fetched | +| `archilyzer posts check` | `--slug <channel> [--mode stale\|unchecked\|all] [--limit N]` | which archived posts were deleted at the source *(passthrough)* | +| `archilyzer diarize backfill` | `[--dry-run] [--scope transcribed\|channel:<slug>\|video:<slug>/<id>] [--limit N] [--force] …` | diarize videos whose audio is still on disk *(passthrough)* | +| `archilyzer digest plan` | `[--lane local\|remote] [--channels a,b] [--top N] [--json] [--census] …` | price the digest backfill; writes nothing *(passthrough)* | +| `archilyzer digest validate` | `<channel> [<channel> …]` | score digests already on disk *(passthrough)* | +| `archilyzer reconcile video-dirs` | `[--channel <slug>] [--dry-run] [--verbose]` | rename video dirs to the canonical id layout *(passthrough)* | +| `archilyzer verify transcripts` | `--channel <slug>` | list duplicate and missing transcripts *(passthrough)* | +| `archilyzer migrate channel-priority` | `[--dry-run]` | the one-shot channel-priority migration (plans/channel-priority.md, S5) *(passthrough)* | +| `archilyzer storage migrate-tier` | `<slug>…\|--all [--order smallest] [--include-large] [--dry-run] [--reclaim]` | bring a channel off the retired whole-directory layout onto the media tier, its text home to the corpus disk (editor stopped; --all stops before the three big-text channels) *(passthrough)* | +| `archilyzer brand media` | `[--out <dir>] [--video-kit]` | render the Archilyzer Media channel's assets *(passthrough)* | +| `archilyzer mcp` | `[--local <dir>\|--remote <url>\|--hub <url>]` | start the MCP server on stdio (as `pnpm --filter yt-dlp-transcript-mcp exec tsx src/index.ts`) *(passthrough)* | +| `archilyzer sync tick` | | POST one scheduler tick to the editor (SYNC\_TICK\_URL, SYNC\_TICK\_TOKEN) | +| `archilyzer docs env` | `[--check]` | write ENVIRONMENT.md from the declared env-var list (lib/envVars.ts) | +| `archilyzer docs files` | `[--check]` | write SITE.md + CHANNEL.md + REPORT.md + CITATIONS.md from the file schemas | +| `archilyzer docs cli` | `[--check]` | write COMMANDS.md from this command table and `pnpm ops --help` | +| `archilyzer settings example` | `[--check]` | write settings.json.example + SETTINGS.md from the schema | +| `archilyzer doctor` | `[--json]` | read-only report: node, the checkout, the corpus, settings, every tool, the port block; exit 1 on a failure | +| `archilyzer release show` | `[editor\|export]` | latest release, its date, how many bullets wait under \[Unreleased\] | +| `archilyzer release cut` | `<editor\|export\|all> <X.Y.Z\|next\|next-minor> [--commit] [--date YYYY-MM-DD]` | \[Unreleased\] -&gt; a dated heading; all = both, one version, two commits | + +## `pnpm ops` + +Drives a running editor over HTTP (`/api/ops/*`, the same actions its pages run), gated by `WORKER_TOKEN`. A body is `--json '<object>'` or `--file <path>`; `--wait` follows a job to its end. + +| action | what it does | example | +|---|---|---| +| `channel-priority` | — | `pnpm ops channel-priority --json '{"slugs":["x"],"operation":"download","tier":"paused"}'` | +| `channel-config` | channel-config changes a channel as its Configure form does: {"slug"} and any of "patch" (form field names; "" clears one), "sites" (the WHOLE membership set: \[{"siteId", "groupId"? \| "newGroupName"?}\], \[\] = on no site; an unknown site id is refused), "excludeFromBuild" and "excludeFromCleanup" (set to the value given, not toggled). | `pnpm ops channel-config --json '{"slug":"x","patch":{"downloadFilterExclude":"rerun"}}'`<br>`pnpm ops channel-config --json '{"slug":"x","sites":[{"siteId":"anilyzer"}]}'`<br>`pnpm ops channel-config --json '{"slug":"x","sites":[],"excludeFromBuild":true}'` | +| `create-channel` | create-channel is the New channel form: {"fields": {"name", "handling": "youtube"\|"transcribe", "url"?, "platform"?, "sourceKind"?, "postFetcher"?, "socialHandle"?, …}} with channel-config's patch keys; "slug"? (else derived from the name), "sites"? (absent = on no site). "fetchPlaylist", "fetchPostsNow" and "prioritizeDownload" are the form's checkboxes, OFF unless true; a job they start comes back as jobId(s), so --wait follows it. | `pnpm ops create-channel --json '{"fields":{"name":"Example (X)","handling":"transcribe","url":"https://x.com/example"}}'` | +| `rename-channel` | rename-channel moves a channel to a new slug, as Danger → Rename does: {"slug", "newSlug"}. Refused while the channel is busy (a job, a lane unit, media in transition) or when the new slug is taken. Old links break. | `pnpm ops rename-channel --json '{"slug":"old-slug","newSlug":"new-slug"}'` | +| `delete-channel` | delete-channel removes a channel's whole directory, as Danger → Delete does: {"slug", "confirm"} — "confirm" must repeat the slug. No undo outside the transcripts/ repo's own history. | `pnpm ops delete-channel --json '{"slug":"x","confirm":"x"}'` | +| `metadata-scan` | — | `pnpm ops metadata-scan --json '{"slug":"the-quartering"}'` | +| `refresh-metadata` | refresh-metadata re-reads ONE video's metadata.info.json from its source (no subtitles, no media) on the platform's queue: {"slug", "id"}. The job's log ends with what the source now says — live\_status, formats, audio-only formats and whether any is non-fragmented, English captions, the keys that changed. An id with no data/&lt;id&gt;/ is refused (a refresh re-reads a video already archived), as are archive.org and Wayback records. | `pnpm ops refresh-metadata --json '{"slug":"the-quartering","id":"<videoId>"}' --wait` | +| `import-video` | — | `pnpm ops import-video --json '{"slug":"demo-archive","url":"https://archive.org/details/example-item"}'` | +| `import-archive-org` | — | `pnpm ops import-archive-org --json '{"slug":"demo-archive","item":"example-item","match":"\\.mp4$"}' --wait` | +| `feed-metadata` | — | `pnpm ops feed-metadata --json '{"slug":"demo-podcast","dryRun":true}' --wait` | +| `refresh-report` | — | `pnpm ops refresh-report --json '{"all":true}'` | +| `sync` | — | `pnpm ops sync --json '{"slug":"the-quartering"}' --wait` | +| `download-missing` | — | | +| `retry-bucket` | retry-bucket runs one bucket of a channel's report as one job, past any lane hold: {"slug", "bucket"}. "ids": \[...\] runs only those videos, and every one must be in the bucket (a stray id is refused, named); a job run with ids is not replayable, as a checkbox selection in the UI is not. | | +| `transcribe-bucket` | transcribe-bucket transcribes a channel's "downloaded, not transcribed" bucket on the transcription queue, as the channel page's Transcribe button does: {"slug"}. "ids": \[...\] narrows it the same way as on retry-bucket. | | +| `fetch-posts` | fetch-posts fetches a social channel's new posts: {"slug"}. "full": true re-walks the whole timeline; "older": true walks back from the oldest archived post through search (X; needs a login), saving its place for the next run, down to "floor": "YYYY-MM-DD" when given; "from": "YYYY-MM-DD" starts the walk afresh there, replacing its saved place (and a "complete") — for a gap above one surviving old post. "limit": N caps the posts one run reads; "pages": N caps the pages (a forum thread: its latest N pages). "full" and "older" together are refused. An older walk over an account that shows no posts (nothing archived, and the last timeline fetch read none) is refused unless "force": true. A drained fetch stops at its next resume point and the next run resumes. | | +| `capture-posts` | capture-posts captures archived posts of a social channel (X, forum): a screenshot of each through the connected X profile (a forum thread: its host's forum profile), and its attached media through gallery-dl (a forum thread: the same profile), into the channel's posts-media/&lt;id&gt;/: {"slug", "ids": \[...\]}. Every id must be in the channel's posts archive. "shots": false or "media": false skips that half; posts already captured are skipped unless "force": true. A post that links to an X Article also gets the article (article.json, .md, .png and its images) unless "articles": false; both halves off with "articles": true reads only the articles. Paced like a post fetch, on its queue. | | +| `publish` | publish runs publish stages on the editor's publish queue, one at a time, under one run id: {"verb": …}. "index" updates the index; "build" builds "siteId"/"siteIds" (forced; the index first when stale; "runner": "docker" builds every site in containers); "deploy" ships their built bundles (production, "preview": "&lt;branch&gt;", or "to": "local"; "force" redeploys a bundle already shipped there); "hub" / "homepage" build them, {"deploy": true} deploys after; "now" is Publish now (the stale index, then each policy target); "stale" builds every stale site. The answer lists every job ({target, kind, jobId}) and --wait follows them all. A site never built is refused: "no build of &lt;id&gt; in &lt;dir&gt; — archilyzer publish build &lt;id&gt;". `get publish` is the status. | | +| `build-index`, `build-site`, `build-deploy`, `deploy-site`, `build-hub`, `deploy-hub`, `build-homepage`, `deploy-homepage` | build-index, build-site, build-deploy, deploy-site, build-hub, deploy-hub, build-homepage and deploy-homepage are publish's aliases, with their old bodies and answers ("skipData" is accepted and ignored). | `pnpm ops build-deploy --json '{"siteIds":["anilyzer","jeralyzer"]}' --wait`<br>`pnpm ops build-site --json '{"siteId":"anilyzer"}' --wait`<br>`pnpm ops deploy-site --json '{"siteId":"anilyzer","preview":"tags-exclude"}' --wait`<br>`pnpm ops build-hub --wait`<br>`pnpm ops build-hub --json '{"deploy":true}' --wait`<br>`pnpm ops deploy-hub --wait`<br>`pnpm ops build-homepage --json '{"deploy":true}' --wait`<br>`pnpm ops deploy-homepage --json '{"preview":"refresh"}' --wait` | +| `build-site`, `build-deploy`, `deploy-site` | build-site, build-deploy and deploy-site all take "siteId" (one) or "siteIds" (a list). | | +| `build-hub` | build-hub builds the hub into its bundle; {"deploy": true} deploys it after, and deploy-hub ships the one already built. Both deploy to the Pages project set on /sites under Hub, and take "preview" too. | | +| `build-homepage` | build-homepage builds the homepage package into homepage/out; {"deploy": true} deploys it after (only if the build succeeded), and deploy-homepage ships the one already built. Both deploy to the Pages project archilyzer (https://archilyzer.pages.dev), production unless "preview" is given. | | +| `relocate` | — | `pnpm ops relocate --json '{"slugs":["x"],"locationId":"platter"}'` | +| `relocate-back` | — | | +| `evict-clips` | — | | +| `reports-prepare` | reports-prepare cuts every clip and copies every post capture a site's published reports cite into its report-media cache, before its build: {"siteId"}. The job fails, naming each one, when a citation lacks media. When nothing is missing it then exports the reports, as reports-export. | `pnpm ops reports-prepare --json '{"siteId":"demo-site"}' --wait` | +| `reports-export` | reports-export writes each published report as report.html, report.pdf, report.md and evidence-pack.zip for the site's build to publish: {"siteId", "reportId"?, "formats"?: \["html","pdf","md","zip"\]}. | `pnpm ops reports-export --json '{"siteId":"demo-site","formats":["html","md"]}' --wait` | +| `lane` | — | `pnpm ops lane --json '{"lane":"download","held":true}'` | +| `tags` | — | `pnpm ops tags --json '{"op":"define","tag":{"id":"eva-collab","label":"Collab"}}'` | +| `tag-videos` | — | `pnpm ops tag-videos --file ids.json` | +| `keep-videos` | — | | +| `persist-videos` | persist-videos saves specific videos, across channels, to the saved-video store: {"items": \[{"slug", "id"}, ...\]}. "format": "original" \| "video\_720" (default: each channel's own). "replace": "above-height" also re-fetches a saved one whose height is unknown or above that quality (default "never"). "gapMs" pauses between downloads (default the batch gap), "minFreeMemMb" waits for that much free memory before each. "dryRun": true answers with the buckets (saved, wrongHeight, toFetch, noUrl, unknown) and starts nothing. One job per channel, on its download queue; a low disk or a rate limit stops it, and running the same body again resumes — saved videos are skipped. | `pnpm ops persist-videos --file list.json --wait` | +| `fetch-windows` | fetch-windows fetches clip windows, one paced job per platform queue (YouTube and Rumble side by side): {"siteId"} fetches every window the site's published reports cite and the disk does not hold; {"items": \[{"slug", "id", "from", "to", "clipId"?, "reason"?, "pad"?, "webpageUrl"?}, ...\], "requestedBy", "manifest"?} fetches a list. "maxHeight" caps the source height (default 720). "dryRun": true lists the windows per platform, the ones already on disk ("cached") and the ones no fetch can fill ("unfetchable": deleted, off the site) and starts nothing. A platform cooling down or held is refused for its group; a 429, or two 403s in a row, backs the platform off and stops its job. Running the same body again resumes — fetched windows are cached. | `pnpm ops fetch-windows --json '{"siteId":"demo-site","dryRun":true}'`<br>`pnpm ops fetch-windows --file windows.json --wait` | +| `cut-release` | cut-release turns a changelog's \[Unreleased\] into "## \[&lt;version&gt;\] - &lt;date&gt;": {"workspace": "editor" \| "export" \| "all", "version": "X.Y.Z" \| "next" \| "next-minor", "commit": boolean (default false), "date": "YYYY-MM-DD" (default today)}. "all" cuts both with ONE version and commits each ("Release &lt;workspace&gt; &lt;version&gt;") — or neither: every check runs before either file is written. Only an editor built from release 10 or later has the route (an older one answers 404); with no editor running, `archilyzer release cut` does the same locally. | `pnpm ops cut-release --json '{"workspace":"all","version":"next","commit":true}'` | +| `transcribe` | transcribe runs ONE local file through a local transcription worker, as a job: {"path": "/abs/file"} (audio or video), "start"/"end" (seconds) for a window, "workerId" (a settings worker id; default: the one auto-transcribe would get), "out" (an absolute path for the result JSON, never inside the corpus). The result is {path, window, worker: {id, appId, model, device}, cues: \[{start, end, text}\], text, ...}, cue times on the file's own clock. "words": true adds words: \[{w, start, end, conf?}\] on the same clock -- from parakeet, which keeps its word timestamps; \[\] from an engine that does not. With --wait it is printed on stdout (the response and the log go to stderr), so `pnpm ops transcribe ... --wait \| jq -r .text` works. | `pnpm ops transcribe --json '{"path":"/abs/clip.mp4","start":120,"end":150}' --wait`<br>`pnpm ops transcribe --json '{"path":"/abs/a.wav","workerId":"parakeet-cpu","out":"/tmp/a.json"}' --wait` | + +### Usage, flags and environment + +```text +Usage: pnpm ops <action> [--json '<body>' | --file <path>] [--wait] + [--wait-timeout <seconds>] [--quiet] + pnpm ops get channel <slug> [--counts] + pnpm ops get channels + pnpm ops get tags [<tagId>] + pnpm ops get publish + pnpm ops list + +--wait follows the job's log and survives a poll that fails (a busy + in-process build starves the server): it backs off and, after three + failures, asks /api/jobs/active whether the job is still there. +--wait-timeout <seconds> gives up and exits 1 instead of waiting forever. + Default: no timeout — the queue may legitimately hold a job for hours. + +"preview": "<branch>" on deploy-site or build-deploy makes it a Cloudflare + Pages PREVIEW instead of production: the same bundle goes to a branch + alias, https://<branch>.<project>.pages.dev, and the live site is left + alone. The alias is printed after the response. Lowercase letters, + digits and dashes, up to 28 characters; "main" is refused. + +Env: ARCHILYZER_EDITOR_URL (default http://localhost:3001), WORKER_TOKEN, + ARCHILYZER_AGENT (provenance of a tag write; default "cli") +``` diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md @@ -146,16 +146,10 @@ The same controllers the editor uses are one command line, `common/bin/archilyze which is how you drive the pipeline headlessly or from cron. `pnpm archilyzer <command>` from the repo root is the short form of `pnpm --filter yt-dlp-transcript-common exec tsx bin/archilyzer.ts <command>`; `pnpm archilyzer ---help` lists every command. - -```bash -pnpm archilyzer doctor # read-only: can this machine do what it is configured to? -pnpm archilyzer index # the LMDB index -pnpm archilyzer run diarization <channel> [ids…] # one catalogued operation, offline, as the editor's job -pnpm archilyzer build site <id> # publish: see PUBLISH.md -pnpm archilyzer verify transcripts --channel <slug> -pnpm archilyzer mcp # the MCP server on stdio -``` +--help` lists every command. Every command and every `pnpm ops` action is in +[COMMANDS.md](COMMANDS.md), generated by `pnpm archilyzer docs cli` from the two help +texts (`--check` fails when it is stale, or when [OPERATING.md](OPERATING.md) names a +command that does not exist); recipes that chain them are in OPERATING.md. The table is `archilyzer.ts`; the machinery (parser, lookup, usage) is `_cli.ts`. Every file in `common/bin/` is reachable from a row — a test fails otherwise. A bin that diff --git a/OPERATING.md b/OPERATING.md @@ -0,0 +1,160 @@ +# Operating an archive + +Recipes for running an archive from a shell or an agent. Every command named here is in +[COMMANDS.md](COMMANDS.md), generated from the commands' own help. + +## The three surfaces + +| surface | what it is | needs | +|---|---|---| +| `pnpm ops <action>` | the running editor's actions over HTTP (`/api/ops/*`): the same code a click runs, as jobs on the editor's queues | a running editor; `ARCHILYZER_EDITOR_URL` (default `http://localhost:3001`) and `WORKER_TOKEN` (the editor's own, from `editor/.env`) | +| `pnpm archilyzer <command>` | the core's CLI: index, compose, publish stages, reports, offline refreshes, doctor | the checkout and its `transcripts/`; in the container, `docker compose exec editor pnpm archilyzer …` | +| the MCP server | reads a published archive (search, transcripts, reports); `fetch_clip` asks the editor for clip media | registered as `archilyzer` ([AGENTS.md](AGENTS.md), [mcp/README.md](mcp/README.md)) | + +- A job-starting `pnpm ops` action answers with a `jobId` once the job is queued. `--wait` follows its log to the + end and exits with its status; a platform queue may hold a job for hours. +- What is there: `pnpm ops get channels`, `pnpm ops get channel <slug> --counts`, `pnpm ops get publish`, + `pnpm archilyzer doctor`. +- Every fetch goes through the editor (paced per platform, cookie-aware, provenanced). Never run yt-dlp, whisper or + parakeet by hand, and never edit `transcripts/**` by hand: each file has one writer, named below. + +## Add a channel + +```sh +pnpm ops create-channel --json '{"fields":{"name":"Example","handling":"youtube","url":"https://www.youtube.com/@example"}}' +pnpm ops channel-config --json '{"slug":"example","patch":{"downloadFilterExclude":"#shorts"}}' +pnpm ops channel-config --json '{"slug":"example","sites":[{"siteId":"jeralyzer"}]}' +pnpm ops get channel example +``` + +- `handling`: `"youtube"` fetches the platform's captions; `"transcribe"` downloads audio and transcribes it here. + `fields` takes the New channel form's field names, the same as `channel-config`'s `patch` + ([RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md), "Driving the editor without a browser"). Every `config.json` key: + [CHANNEL.md](CHANNEL.md). +- `sites` is the whole membership set; absent or `[]` keeps the channel on no site. +- `"fetchPlaylist": true` on `create-channel` stores the playlist at once (the form's "Fetch playlist now"); a social channel's is `"fetchPostsNow": true`. +- Rename or delete: `pnpm ops rename-channel --json '{"slug":"old","newSlug":"new"}'`, + `pnpm ops delete-channel --json '{"slug":"x","confirm":"x"}'` — both refused while the channel is busy. + +## Sync and download + +```sh +pnpm ops sync --json '{"slug":"example"}' --wait +pnpm ops metadata-scan --json '{"slug":"example"}' --wait +pnpm ops download-missing --json '{"slug":"example"}' --wait +pnpm ops retry-bucket --json '{"slug":"example","bucket":"<bucket>"}' --wait +pnpm ops refresh-metadata --json '{"slug":"example","id":"<videoId>"}' --wait +pnpm ops persist-videos --json '{"items":[{"slug":"example","id":"<videoId>"}],"dryRun":true}' +``` + +- The lanes do this unattended: `pnpm ops lane --json '{"lane":"download","enabled":true,"action":"start"}'`; + `"held": true` pauses dispatch and keeps the runner's place. Lanes: `transcription`, `download`, `digest`, + `backfill`. +- A channel's priority: `pnpm ops channel-priority --json '{"slugs":["example"],"tier":"low"}'` + (`"operation"` pins one operation's tier). +- Bucket names and sizes: `pnpm ops get channel example`. +- Keep videos out of cleanup by title or description: `pnpm ops keep-videos --json '{"slug":"example","match":"interview","dryRun":true}'`. + +## Import from archive.org, Odysee, BitChute and the Wayback Machine + +```sh +pnpm ops import-archive-org --json '{"slug":"example-archive","item":"<item>","match":"\\.mp4$","dryRun":true}' --wait +pnpm ops import-archive-org --json '{"slug":"example-archive","item":"<item>","match":"\\.mp4$"}' --wait +pnpm ops import-video --json '{"slug":"example","url":"https://www.bitchute.com/video/<id>/"}' --wait +pnpm ops import-video --json '{"slug":"example","url":"https://odysee.com/@example:0/<video>:0"}' --wait +pnpm ops import-video --json '{"slug":"example","url":"https://web.archive.org/web/<timestamp>/<original-url>"}' --wait +``` + +- `import-archive-org` takes one item's files (`files` exact, or `match` a regex), one at a time on archive.org's + queue; files already held are skipped. archive.org is fetched over BitTorrent when it can be, never by yt-dlp. +- A one-off Odysee or BitChute import is paced like that platform's own downloads. A whole Odysee or BitChute + channel is a channel with that URL, then `sync`. +- A Wayback capture is named by what it copies; `wayback.json` beside it records the capture. +- Bring records up to the current rules, offline (dry run first): + `pnpm archilyzer archive-org refresh <slug> --dry-run`, `pnpm archilyzer wayback refresh <slug> --dry-run`, + `pnpm archilyzer feeds backfill-metadata <slug> --dry-run` (or `pnpm ops feed-metadata` on the editor). + +## Fetch and capture posts + +```sh +pnpm ops create-channel --json '{"fields":{"name":"Example (X)","handling":"transcribe","url":"https://x.com/example"}}' +pnpm ops fetch-posts --json '{"slug":"example-x"}' --wait +pnpm ops fetch-posts --json '{"slug":"example-x","older":true}' --wait +pnpm ops capture-posts --json '{"slug":"example-x","ids":["<postId>"]}' --wait +``` + +- An X, Bluesky or XenForo URL makes a social channel (its handle derived from the URL). `"older": true` walks back below the oldest + archived post and resumes from its saved place; `"full": true` re-walks the timeline. +- `capture-posts` saves a screenshot and the attached media of named posts (and a linked X Article). +- Forum pages saved from a browser: `pnpm archilyzer posts import-html <slug> <file-or-dir> --dry-run`. +- Which archived posts were deleted at the source: `pnpm archilyzer posts check --slug <slug>`. + +## Transcribe + +```sh +pnpm ops lane --json '{"lane":"transcription","enabled":true,"action":"start"}' +pnpm ops transcribe-bucket --json '{"slug":"example"}' --wait +pnpm ops transcribe --json '{"path":"/abs/clip.mp4","start":120,"end":150,"words":true}' --wait +``` + +- The transcription lane takes every channel's downloaded, untranscribed audio; `transcribe-bucket` runs one + channel's now. +- `transcribe` runs one local file (or a window of it) through the corpus's own engine and model; the result JSON + (cues, text, and with `"words": true` the word timings parakeet keeps) is printed on stdout with `--wait`. + `"out"` writes it to a file outside the corpus. + +## Tag + +```sh +pnpm ops get tags +pnpm ops tags --json '{"op":"define","tag":{"id":"example-collab","label":"Collab"}}' +pnpm ops tag-videos --json '{"tag":"example-collab","op":"add","videos":[{"slug":"example","id":"<videoId>"}]}' +pnpm ops tag-videos --file ids.json +``` + +- `transcripts/tags.json` has one writer, and `tags` and `tag-videos` go through it, recording who asked + (`ARCHILYZER_AGENT`). `remove` unpins; `suppress` rejects a rule's hit. +- A tag's `sites` limits the sites it exists on. +- Read side: the MCP's `list_tags`, and `tags` on `search_transcripts` / `enumerate_matches`. + +## Prepare and export reports + +```sh +pnpm archilyzer reports convert sweep <sweep.md> --out <report.json> +pnpm ops fetch-windows --json '{"siteId":"example-site","dryRun":true}' +pnpm ops fetch-windows --json '{"siteId":"example-site"}' --wait +pnpm ops reports-prepare --json '{"siteId":"example-site"}' --wait +pnpm ops reports-export --json '{"siteId":"example-site","formats":["html","md"]}' --wait +``` + +- A report is `transcripts/sites/<site>/reports/<id>/report.json` ([REPORT.md](REPORT.md), + [CITATIONS.md](CITATIONS.md)); `reports convert` makes one from a `/sweep` report, an `/ask` answer or a + report-to-video manifest. +- `fetch-windows` with `siteId` fetches every clip window the site's published reports cite that the disk does not + hold, one paced job per platform; a re-run resumes. +- `reports-prepare` cuts the cited clips and copies the cited post captures, then exports; it fails naming each + citation that lacks media. Offline: `pnpm archilyzer reports prepare <siteId>`, + `pnpm archilyzer reports export <siteId>`. +- One cited moment's media, from an agent: the MCP's `fetch_clip`. + +## Publish + +```sh +pnpm ops get publish +pnpm ops publish --json '{"verb":"now"}' --wait +pnpm ops publish --json '{"verb":"build","siteId":"example-site"}' --wait +pnpm ops publish --json '{"verb":"deploy","siteId":"example-site","preview":"check"}' --wait +pnpm archilyzer publish status +``` + +- Publishing is stages on one queue: update the index, build a site into its bundle, deploy the bundle, each + checked live. `now` runs what Publish now runs (the stale index, then each site's policy). Offline, under the same + lock: `pnpm archilyzer publish index`, `publish build <id>`, `publish deploy <id> --preview <branch>`. +- Details, policies and the hub and homepage: [PUBLISH.md](PUBLISH.md). + +## Research with the MCP + +- `/ask` answers a question with citations in the conversation; `/sweep` writes a cited report to a file. +- Search first, then pull only the cited seconds with `fetch_clip`. The editor fetches only for a channel it + already archives. +- Tools and their arguments: [mcp/README.md](mcp/README.md). diff --git a/README.md b/README.md @@ -640,6 +640,8 @@ See [CONTRIBUTING.md](CONTRIBUTING.md) to work on the code. | [SETUP.md](SETUP.md) | Full per-OS install, transcription backends. | | [ENVIRONMENT.md](ENVIRONMENT.md) | Every environment variable, by audience (generated). | | [CONTRIBUTING.md](CONTRIBUTING.md) | Workspace layout, tests, the `archilyzer` CLI, internals. | +| [OPERATING.md](OPERATING.md) | Running an archive from a shell or an agent: recipes over `pnpm ops`, `archilyzer` and the MCP. | +| [COMMANDS.md](COMMANDS.md) | Every `archilyzer` command and `pnpm ops` action (generated). | | [RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md) | `docker compose up` for the whole stack: exposure model, auth, GPU. | | [PUBLISH.md](PUBLISH.md) | Building and deploying sites: Pages + R2, cost-abuse protection, parallel builds in containers. | | [SCHEDULED_SYNC.md](SCHEDULED_SYNC.md) | Unattended per-channel syncing. | diff --git a/RUNNING_IN_DOCKER.md b/RUNNING_IN_DOCKER.md @@ -413,6 +413,9 @@ pnpm ops get channels # every channel, its kind and its sites pnpm ops list # every action name ``` +These are examples. Every action and getter, with its body, is in [COMMANDS.md](COMMANDS.md) (generated from +`pnpm ops --help` and `pnpm archilyzer --help`); recipes that chain them are in [OPERATING.md](OPERATING.md). + Four things to know before you script against it: - **A job-starting action returns a `jobId` and does not stream.** The job may diff --git a/common/bin/archilyzer.ts b/common/bin/archilyzer.ts @@ -542,6 +542,13 @@ export const COMMANDS: Command[] = [ (await import("./file-schemas-docs")).main({ check: flags.check === true }), }, { + path: ["docs", "cli"], + usage: "[--check] write COMMANDS.md from this command table and `pnpm ops --help`", + flags: { check: "boolean" }, + run: async ({ flags }) => + (await import("./cli-docs")).main({ check: flags.check === true }), + }, + { path: ["settings", "example"], usage: "[--check] write settings.json.example + SETTINGS.md from the schema", flags: { check: "boolean" }, diff --git a/common/bin/cli-docs.test.ts b/common/bin/cli-docs.test.ts @@ -0,0 +1,140 @@ +import assert from "node:assert/strict"; +import { readFileSync } from "node:fs"; +import path from "node:path"; +import { test } from "node:test"; +import { fileURLToPath } from "node:url"; +import { + COMMANDS_FILE, + generateCommandsMarkdown, + mdCell, + opsActions, + opsExamples, + opsParagraphs, + RECIPES_FILE, + renderCommandsMarkdown, + splitUsage, + unknownCommandsIn, + unknownRecipeCommands, +} from "./cli-docs"; + +const REPO = path.resolve(path.dirname(fileURLToPath(import.meta.url)), "..", ".."); + +const OPS_USAGE = [ + "Usage: pnpm ops <action> [--json '<body>'] [--wait]", + " pnpm ops list", + "", + "Actions: sync, build-site, deploy-site, lane", + "", + "--wait follows the job's log.", + "", + 'build-site and deploy-site take "siteId" or', + ' "siteIds" (a list).', + "", + 'deploy-site ships a built bundle: {"siteId"}. See `a | b`.', + "", + "Env: WORKER_TOKEN", +].join("\n"); + +test("splitUsage splits at the first double space; a usage with none is all text", () => { + assert.deepEqual(splitUsage("<id> [--force] build it now"), { + args: "<id> [--force]", + text: "build it now", + }); + assert.deepEqual(splitUsage("rebuild the index"), { args: "", text: "rebuild the index" }); +}); + +test("mdCell escapes markup outside code and only the pipe inside it", () => { + assert.equal(mdCell("a <id> | *b* _c_ [x]"), "a &lt;id&gt; \\| \\*b\\* \\_c\\_ \\[x\\]"); + assert.equal(mdCell("run `x | y <z>` then"), "run `x \\| y <z>` then"); +}); + +test("opsActions reads the Actions line, and refuses a usage without one", () => { + assert.deepEqual(opsActions(OPS_USAGE), ["sync", "build-site", "deploy-site", "lane"]); + assert.throws(() => opsActions("Usage: nothing"), /no `Actions:` line/); +}); + +test("opsParagraphs attributes a paragraph to every action it opens with", () => { + const ps = opsParagraphs(OPS_USAGE, opsActions(OPS_USAGE)); + assert.deepEqual( + ps.map((p) => p.actions), + [[], [], ["build-site", "deploy-site"], ["deploy-site"], []], + ); + assert.ok(!ps.some((p) => p.text.startsWith("Actions: "))); +}); + +test("opsExamples takes header lines of known actions only, padding collapsed", () => { + const source = [ + "// pnpm ops sync --json '{\"slug\":\"x\"}' --wait", + "// pnpm ops get channel x", + "// pnpm ops lane", + "// pnpm ops sync --json '{\"slug\":\"y\"}'", + "const x = 1; // pnpm ops sync not a header line", + ].join("\n"); + const ex = opsExamples(source, ["sync", "lane"]); + assert.deepEqual(ex.get("sync"), [ + "pnpm ops sync --json '{\"slug\":\"x\"}' --wait", + "pnpm ops sync --json '{\"slug\":\"y\"}'", + ]); + assert.deepEqual(ex.get("lane"), ["pnpm ops lane"]); + assert.equal(ex.has("get"), false); +}); + +test("the rendered reference lists every command and every action exactly once as a row head", () => { + const md = renderCommandsMarkdown( + [ + { path: ["index"], usage: "rebuild the LMDB transcript index" }, + { path: ["publish", "build"], usage: "<id|all> [--force] build a site" }, + { path: ["mcp"], usage: "[--local <dir>] start the MCP server", passthrough: true }, + ], + OPS_USAGE, + new Map([["deploy-site", ["pnpm ops deploy-site --json '{}'"]]]), + ); + assert.match(md, /^# Command reference\n\n<!-- GENERATED by common\/bin\/cli-docs\.ts/); + assert.match(md, /\| `archilyzer index` \| \| rebuild the LMDB transcript index \|/); + assert.match(md, /\| `archilyzer publish build` \| `<id\\\|all> \[--force\]` \| build a site \|/); + assert.match(md, /\| `archilyzer mcp` \| .* \*\(passthrough\)\* \|/); + // An action with no paragraph still has its row. + assert.match(md, /\| `sync` \| — \| \|/); + assert.match(md, /\| `lane` \| — \| \|/); + // The shared paragraph is one row naming both; deploy-site's example sits on + // that first row, not on its own paragraph's. + assert.match( + md, + /\| `build-site`, `deploy-site` \| build-site and deploy-site take "siteId" or "siteIds" \(a list\)\. \| `pnpm ops deploy-site --json '\{\}'` \|/, + ); + assert.match(md, /\| `deploy-site` \| deploy-site ships a built bundle: \{"siteId"\}\. See `a \\\| b`\. \| \|/); + // The general paragraphs are printed as they stand, in a text block. + assert.match(md, /```text\nUsage: pnpm ops <action>[^]*--wait follows the job's log\.\n\nEnv: WORKER_TOKEN\n```\n$/); + assert.doesNotMatch(md, /Actions: sync/); +}); + +test("unknownCommandsIn names what no action, getter or command answers to", () => { + const cli = [{ path: ["publish", "build"], usage: "" }, { path: ["doctor"], usage: "" }]; + const md = [ + "`pnpm ops <action>` and `pnpm archilyzer <command>` are placeholders.", + "pnpm ops sync --json '{}' --wait", + "pnpm ops list", + "pnpm ops get channel x and pnpm ops get bogus", + "pnpm ops frobnicate --wait", + "`pnpm archilyzer publish build <id>`, `pnpm archilyzer doctor`.", + "pnpm archilyzer publish destroy x", + ].join("\n"); + const usage = OPS_USAGE.replace("pnpm ops list", "pnpm ops list\n pnpm ops get channel <slug>"); + assert.deepEqual(unknownCommandsIn(md, cli, usage), [ + "pnpm ops get bogus", + "pnpm ops frobnicate", + "pnpm archilyzer publish destroy x", + ]); +}); + +test(`every command ${RECIPES_FILE} names exists`, async () => { + assert.deepEqual(await unknownRecipeCommands(), []); +}); + +test(`${COMMANDS_FILE} is what the command table and pnpm ops --help generate`, async () => { + assert.equal( + readFileSync(path.join(REPO, COMMANDS_FILE), "utf8"), + await generateCommandsMarkdown(), + `${COMMANDS_FILE} is stale: run \`archilyzer docs cli\``, + ); +}); diff --git a/common/bin/cli-docs.ts b/common/bin/cli-docs.ts @@ -0,0 +1,274 @@ +// WRITE COMMANDS.md FROM THE TWO COMMAND SURFACES' OWN HELP. +// +// archilyzer docs cli [--check] +// +// The `archilyzer` rows come from the command table (`COMMANDS`, archilyzer.ts): +// each row's path and its one-line usage. The `pnpm ops` rows come from +// `usage()` in scripts/archilyzer-ops.mjs — the text `pnpm ops --help` prints — +// split into paragraphs: a paragraph that opens with action names documents +// those actions; every other paragraph (the usage lines, --wait, preview, +// the env) is printed as it stands. Each action also gets the examples the +// script's header comment gives it (`// pnpm ops <action> …` lines). An +// action with neither is still listed, so the reference never omits one. +// +// `--check` writes nothing and returns 1 when the committed file differs from +// what the two sources generate (the same claim cli-docs.test.ts makes). Both +// modes then check OPERATING.md: every `pnpm ops …` / `pnpm archilyzer …` its +// recipes name must exist. The sibling of env-docs.ts (ENVIRONMENT.md) and +// file-schemas-docs.ts. + +import { readFile, writeFile } from "node:fs/promises"; +import path from "node:path"; +import { fileURLToPath, pathToFileURL } from "node:url"; +import { runIfEntryPoint } from "./_cli"; + +const REPO = path.resolve(path.dirname(fileURLToPath(import.meta.url)), "..", ".."); + +export const COMMANDS_FILE = "COMMANDS.md"; +const OPS_SCRIPT = path.join("scripts", "archilyzer-ops.mjs"); + +// What the renderer needs of a CLI row — structural, so the tests never import +// the table. +export type CliRow = { path: readonly string[]; usage: string; passthrough?: boolean }; + +/** A usage line split at its first double space: the arguments, then the text. */ +export function splitUsage(usage: string): { args: string; text: string } { + const at = usage.indexOf(" "); + if (at < 0) return { args: "", text: usage.trim() }; + return { args: usage.slice(0, at).trim(), text: usage.slice(at).trim() }; +} + +/** + * Plain help text as one Markdown table cell: backtick spans are kept as code, + * and everything else is escaped so `<id>`, `*`, `_` and `[x]` print as typed. + * A `|` is escaped everywhere, code included (GFM reads it as a cell edge). + */ +export function mdCell(text: string): string { + return text + .split(/(`[^`]*`)/) + .map((part, i) => + i % 2 === 1 + ? part.replace(/\|/g, "\\|") + : part + .replace(/\\/g, "\\\\") + .replace(/\|/g, "\\|") + .replace(/</g, "&lt;") + .replace(/>/g, "&gt;") + .replace(/([*_[\]])/g, "\\$1"), + ) + .join(""); +} + +function codeCell(text: string): string { + return text ? `\`${text.replace(/\|/g, "\\|")}\`` : ""; +} + +/** The action list `usage()` prints on its `Actions: a, b, c` line. */ +export function opsActions(opsUsage: string): string[] { + const line = opsUsage.split("\n").find((l) => l.startsWith("Actions: ")); + if (!line) throw new Error("pnpm ops usage: no `Actions:` line"); + return line + .slice("Actions: ".length) + .split(",") + .map((a) => a.trim()) + .filter(Boolean); +} + +export type OpsParagraph = { actions: string[]; text: string }; + +/** + * The usage text's paragraphs, each with the actions it opens with + * ("build-site, build-deploy and deploy-site all take …" → three). A paragraph + * that opens with no action name has `actions: []`. The `Actions:` line itself + * is left out: the table replaces it. + */ +export function opsParagraphs(opsUsage: string, actions: readonly string[]): OpsParagraph[] { + const known = new Set(actions); + const out: OpsParagraph[] = []; + for (const raw of opsUsage.split(/\n\s*\n/)) { + const text = raw.replace(/\s+$/, ""); + if (!text.trim() || text.startsWith("Actions: ")) continue; + const opened: string[] = []; + for (const word of text.trim().split(/[\s,]+/)) { + if (known.has(word)) opened.push(word); + else if (word === "and" && opened.length > 0) continue; + else break; + } + out.push({ actions: opened, text }); + } + return out; +} + +/** + * The examples in the ops script's header comment, per action: every + * `// pnpm ops <action> …` line whose action is a known one, its padding + * collapsed. (`get` and `list` lines are not actions; the usage block has them.) + */ +export function opsExamples(source: string, actions: readonly string[]): Map<string, string[]> { + const known = new Set(actions); + const out = new Map<string, string[]>(); + for (const line of source.split("\n")) { + const m = /^\/\/\s+pnpm ops ([a-z][a-z-]*)(\s+.*)?$/.exec(line); + if (!m || !known.has(m[1])) continue; + const rest = m[2] ? ` ${m[2].trim().replace(/\s{2,}/g, " ")}` : ""; + out.set(m[1], [...(out.get(m[1]) ?? []), `pnpm ops ${m[1]}${rest}`]); + } + return out; +} + +const oneLine = (s: string) => s.replace(/\s+/g, " ").trim(); + +export function renderCommandsMarkdown( + cli: readonly CliRow[], + opsUsage: string, + examples: ReadonlyMap<string, readonly string[]> = new Map(), +): string { + const actions = opsActions(opsUsage); + const paragraphs = opsParagraphs(opsUsage, actions); + const lines: string[] = [ + "# Command reference", + "", + "<!-- GENERATED by common/bin/cli-docs.ts from the `archilyzer` command table (common/bin/archilyzer.ts) and `pnpm ops --help` (scripts/archilyzer-ops.mjs) — do not edit by hand. Regenerate: `pnpm archilyzer docs cli`. -->", + "", + "Every `archilyzer` command and every `pnpm ops` action, from their own help. Recipes that chain them: [OPERATING.md](OPERATING.md).", + "", + "## `pnpm archilyzer`", + "", + "Run from the repo root (in the container: `docker compose exec editor pnpm archilyzer …`). `--help` after a command prints its line. A command marked *passthrough* parses its own flags.", + "", + "| command | arguments | what it does |", + "|---|---|---|", + ]; + for (const row of cli) { + const { args, text } = splitUsage(row.usage); + const note = row.passthrough ? " *(passthrough)*" : ""; + lines.push( + `| \`archilyzer ${row.path.join(" ")}\` | ${codeCell(args)} | ${mdCell(oneLine(text))}${note} |`, + ); + } + lines.push( + "", + "## `pnpm ops`", + "", + "Drives a running editor over HTTP (`/api/ops/*`, the same actions its pages run), gated by `WORKER_TOKEN`. A body is `--json '<object>'` or `--file <path>`; `--wait` follows a job to its end.", + "", + "| action | what it does | example |", + "|---|---|---|", + ); + // One row per help paragraph that opens with actions, in the order of their + // first action; an action no paragraph opens with gets a row of its own. An + // action's examples go on the first row that names it. + type Row = { actions: string[]; text: string; examples: string[] }; + const rows: Row[] = []; + const emitted = new Set<OpsParagraph>(); + for (const action of actions) { + const mine = paragraphs.filter((p) => p.actions.includes(action)); + if (mine.length === 0) rows.push({ actions: [action], text: "", examples: [] }); + for (const p of mine) { + if (emitted.has(p)) continue; + emitted.add(p); + rows.push({ actions: p.actions, text: oneLine(p.text), examples: [] }); + } + } + for (const action of actions) { + const row = rows.find((r) => r.actions.includes(action)); + row?.examples.push(...(examples.get(action) ?? [])); + } + for (const r of rows) { + const names = r.actions.map((a) => `\`${a}\``).join(", "); + const ex = r.examples.map(codeCell).join("<br>"); + lines.push(`| ${names} | ${r.text ? mdCell(r.text) : "—"} | ${ex} |`); + } + lines.push("", "### Usage, flags and environment", "", "```text"); + for (const p of paragraphs) { + if (p.actions.length === 0) lines.push(p.text, ""); + } + if (lines[lines.length - 1] === "") lines.pop(); + lines.push("```", ""); + return lines.join("\n"); +} + +/** + * Every `pnpm ops …` and `pnpm archilyzer …` a document names that is not a + * real action, getter or command — what keeps OPERATING.md's recipes runnable. + * A placeholder (`pnpm ops <action>`) is not a name. + */ +export function unknownCommandsIn( + markdown: string, + cli: readonly CliRow[], + opsUsage: string, +): string[] { + const actions = new Set(opsActions(opsUsage)); + const nouns = new Set([...opsUsage.matchAll(/pnpm ops get ([a-z][a-z-]*)/g)].map((m) => m[1])); + const bad = new Set<string>(); + for (const m of markdown.matchAll(/pnpm ops ([^\s`'"]+)(?:[ \t]+([^\s`'"]+))?/g)) { + const [, word, next] = m; + if (word.startsWith("<") || word === "list") continue; + if (word === "get") { + if (!next || !nouns.has(next)) bad.add(`pnpm ops get ${next ?? ""}`.trim()); + continue; + } + if (!actions.has(word)) bad.add(`pnpm ops ${word}`); + } + for (const m of markdown.matchAll(/pnpm archilyzer((?:[ \t]+[a-z][a-z-]*)+)/g)) { + const words = m[1].trim().split(/\s+/); + if (!cli.some((r) => r.path.every((w, i) => words[i] === w))) { + bad.add(`pnpm archilyzer ${words.join(" ")}`); + } + } + return [...bad]; +} + +// The recipes the reference backs. +export const RECIPES_FILE = "OPERATING.md"; + +/** `usage()` of scripts/archilyzer-ops.mjs — what `pnpm ops --help` prints. */ +export async function loadOpsUsage(repo = REPO): Promise<string> { + const mod = (await import(pathToFileURL(path.join(repo, OPS_SCRIPT)).href)) as { + usage?: () => string; + }; + if (typeof mod.usage !== "function") { + throw new Error(`${OPS_SCRIPT} exports no usage()`); + } + return mod.usage(); +} + +export async function generateCommandsMarkdown(repo = REPO): Promise<string> { + const { COMMANDS } = await import("./archilyzer"); + const opsUsage = await loadOpsUsage(repo); + const source = await readFile(path.join(repo, OPS_SCRIPT), "utf8"); + return renderCommandsMarkdown(COMMANDS, opsUsage, opsExamples(source, opsActions(opsUsage))); +} + +/** OPERATING.md's names that no command answers to (see unknownCommandsIn). */ +export async function unknownRecipeCommands(repo = REPO): Promise<string[]> { + const { COMMANDS } = await import("./archilyzer"); + const recipes = await readFile(path.join(repo, RECIPES_FILE), "utf8"); + return unknownCommandsIn(recipes, COMMANDS, await loadOpsUsage(repo)); +} + +// Writes COMMANDS.md (or, with `check`, compares it), then checks that every +// command OPERATING.md names exists. 1 when either is wrong. +export async function main(opts: { check?: boolean } = {}): Promise<number> { + const file = path.join(REPO, COMMANDS_FILE); + const want = await generateCommandsMarkdown(); + let code = 0; + if (opts.check) { + const have = await readFile(file, "utf8").catch(() => ""); + if (have !== want) { + console.error(`${COMMANDS_FILE} is stale — regenerate it with \`archilyzer docs cli\``); + code = 1; + } + } else { + await writeFile(file, want); + console.log(`wrote ${COMMANDS_FILE}`); + } + const unknown = await unknownRecipeCommands(); + if (unknown.length > 0) { + console.error(`${RECIPES_FILE} names commands that do not exist: ${unknown.join("; ")}`); + code = 1; + } + return code; +} + +runIfEntryPoint(import.meta.url, () => main({ check: process.argv.includes("--check") })); diff --git a/homepage/CHANGELOG.md b/homepage/CHANGELOG.md @@ -1,6 +1,7 @@ # Homepage Changelog ## [Unreleased] +- **The *Running an archive* doc says how to run one from a shell or an agent.** A new section, *From a shell or an agent*, names the three surfaces — `pnpm ops` (the running editor's actions over HTTP, with `ARCHILYZER_EDITOR_URL` and `WORKER_TOKEN`), `pnpm archilyzer` (the core's command line) and the MCP server — with three example commands, and links the source's OPERATING.md (recipes) and COMMANDS.md (every command and action). - **The homepage builds and deploys from the Docker image too.** It is a publish stage like a site's: `archilyzer publish homepage [--deploy]` (or its row on /sites → Publish) builds it — the `/source` mirror included, from the repository `docker-compose.source.yml` mounts read-only — into `homepage/out`, stamps it, and deploys that build to its Pages project, or with `--to local` into the volume the container's `homepage` service serves. The image has the pinned git-filter-repo; it has no gitleaks or stagit, so a container build skips the secret scan with a warning and ships no history pages. - **A site that publishes only its reports is not on the homepage.** A site with `publish: "cited"` has no Official Instances card, chart series, `/stats` entry or recent item, is not in `channel-sites.json`, and the channels only it carries count in no total — as an unlisted site, whatever its listing setting says. A site with reports that publishes its full corpus is listed as before. - **The AI and MCP doc has a Ten-minute setup.** Right after the MCP server's introduction, one block runs Claude Code against a published archive, the Jeralyzer as the example: clone the source (or unpack the tarball on Downloads), `pnpm install`, `claude mcp add archilyzer`, start `claude` and try `/ask`; then what it needs, why the server must be registered as `archilyzer` (the shipped `/ask` and `/sweep` call `mcp__archilyzer__…`), the two optional editor lines for `fetch_clip`, `TRANSCRIPT_HUB_URL`, where the `mcp.json` form for other clients is, and WSL2 on Windows. "What it can do" is a heading of its own after it. Every archive's **Use with AI** link now lands on this page. diff --git a/homepage/content/docs/operate.md b/homepage/content/docs/operate.md @@ -107,4 +107,32 @@ Audio for recordings that have disappeared upstream is protected from cleanup, o the reasoning that a local copy of something no longer available anywhere is the one thing you cannot re-fetch. +## From a shell or an agent + +Everything above can be run without a browser, three ways: + +- **`pnpm ops <action>`** sends the editor the same action a click does — add a + channel, sync, download, import, transcribe, tag, fetch posts, prepare reports, + publish — over HTTP. It needs the running editor's address + (`ARCHILYZER_EDITOR_URL`) and its `WORKER_TOKEN`. A job-starting action returns + once the job is queued; `--wait` follows it to the end. +- **`pnpm archilyzer <command>`** is the core's own command line: the index, + publish stages (`publish now`, `publish build <id>`, `publish deploy <id>`), + reports, offline refreshes, and `doctor`, which says what the machine can do. +- **The MCP server** lets an AI assistant read a published archive — search, + transcripts, reports — and ask the editor for the media behind a cited moment. + See [Use with AI](/docs/ai-and-mcp/). + +```sh +pnpm ops create-channel --json '{"fields":{"name":"Example","handling":"youtube","url":"https://www.youtube.com/@example"}}' +pnpm ops sync --json '{"slug":"example"}' --wait +pnpm ops publish --json '{"verb":"now"}' --wait +``` + +Every fetch still goes through the editor's paced, per-platform queues; nothing +here downloads around them. Recipes for each task are in +[OPERATING.md](https://archilyzer.pages.dev/source/tree/OPERATING.md), and every +command and action in +[COMMANDS.md](https://archilyzer.pages.dev/source/tree/COMMANDS.md). + Next: [Deploy to Cloudflare](/docs/deploy-cloudflare/). diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -8974,3 +8974,113 @@ Build stats jobs carry a "Superseded by release 18" line pointing here. `Authentication error [code: 10000]` and exits 1 with no argv sidecar. `--version` answers `4.147.0`. `publish.spec.ts` reads and writes both under `test-transcripts/.export-builds` (`:32`, `:65`, `:174`). + +## Main's October work, outside release 18 (verified 2026-10-09, `4cffda3f`) + +The record is [`landed-2026-10.md`](landed-2026-10.md). Every anchor below was read at `4cffda3f`. + +### `pnpm ops transcribe` — one local file through the editor's workers + +- **The body is `path`, `start`, `end`, `workerId`, `out`, `words`** (`TRANSCRIBE_FILE_BODY_KEYS`, + `common/controller/transcribeFile.ts:71-78`); any other key is a 400. `path` is absolute and may be anywhere, + the corpus included (`:15-17`); `out` is refused inside any corpus root (`:307`); `"words"` must be a boolean + (`:171`). +- **Job kind `transcribe-file`** (`:64`), `queueKey: ""` (`:547`): parallel at the queue; the worker pool + serialises it (`common/jobs/jobKinds.ts:884-896`, not drainable, not replayable). Only `kind: "local"` workers; + default = the one auto-transcribe would get. +- **The audio is always a 16 kHz mono WAV cut by ffmpeg** into scratch (`:21-25`, `:374`), so parakeet's wrapper + never writes beside the source. +- **The result** (`TranscribeFileResult`, `:107`) carries `cues` (seconds, shifted back onto the SOURCE file's + clock), `text`, `worker`, `durationMs`, and with `"words": true` a `words` array of + `TranscribedWord = {w, start, end, conf?}` (`:97`). Only parakeet emits word timings (`--words`, + `common/lib/transcriptionApps.ts:252-256`); any other engine gives `words: []`. +- **`durationMs` includes the wait for a worker**: `started` is taken at `:444`, before the ffmpeg cut and the + "Waiting for a free local worker…" lease (`:472`); it is read at `:511`. (Release 19 B4 changes this.) +- **The job's last log line is `@@transcribe-result <json>`** (`TRANSCRIBE_RESULT_MARKER`, `:68`, the same + string at `scripts/archilyzer-ops.mjs:205`); `pnpm ops transcribe --wait` prints that JSON alone on stdout + (`resultMarker`, `:331`). + +### Refresh one video's metadata + +- **Job kind `refresh-metadata`** (`REFRESH_METADATA_JOB_KIND`, `common/controller/refreshVideoMetadataJob.ts:47`; + `common/jobs/jobKinds.ts:648-656`: platform-queued, replayable, `needsText`). It runs on the channel's + download queue (`resolveQueueKey(downloadQueueKey(channelConfig), …)`, `refreshVideoMetadataJob.ts:95`); a + rate limit records the platform's backoff, a clean pass settles it (`:115-116`). +- **One yt-dlp spawn**: `--skip-download --write-info-json`, no subtitles, `--sleep-requests 1`, the channel's + extra args before the negations (`buildRefreshArgs`, `common/controller/refreshVideoMetadata.ts:181`). +- **The rewrite goes through `withMetadataHistory`** (`:424`) as writer `"refresh"` (`:418`; + `common/lib/metadataHistory.ts:80`), so what moved lands in `metadata.history.json`. +- **Refused:** a video with no `data/<id>/` (`notFetchedRefusal`, `:92` — the directory is never created); a + record a non-yt-dlp writer owns — archive.org, Wayback, feed backfill (`NON_YTDLP_WRITERS`, `:154`); a held + or cooling platform. The same server action backs the video page's button and `POST /api/ops/refresh-metadata` + `{slug, id, queueKey?}`. + +### fetch-windows — a paced batch of clip windows + +- **One job per queue key, and the WINDOW'S URL picks it**, not the channel's (`planFetchWindows`, + `editor/app/channels/[slug]/videos/fetchWindowsAction.ts:165-169`, `queueKeyForUrl`). Job kind + `fetch-windows`: platform-queued, drainable, replayable, `needsText` (`common/jobs/jobKinds.ts:318-326`). +- **Pacing:** at least `CLIP_WINDOW_MIN_GAP_SECONDS` = 30 between two network fetches + (`common/controller/fetchWindows.ts:82`), Rumble 120 (`CLIP_WINDOW_PLATFORM_MIN_GAP_SECONDS`, `:90-91`). The + gap ACROSS runs is keyed `clip-window:<platform>` (`:387`) in `common/jobs/platformGap.ts`, which is + process-wide and IN MEMORY (`globalThis.__yttPlatformGap__`, `platformGap.ts:13-20`): an editor restart + forgets it. +- **Cache:** a window is `data/<id>/clips/<from>-<to>.mp4`; `findContainingClipWindow` + (`common/lib/clipWindow-server.ts:79`) answers with the tightest held window that covers the ask, checked + before any pacing (`fetchWindows.ts:345`) and costing no pause. +- **Dedupe is within ONE request only** (`dedupeFetchWindowsItems`, `fetchWindows.ts:213-218`): nothing looks + for an existing job, so two requests for one window make two jobs (release 19 A5). +- **Rumble gets `-extension_picky 0` on the first try** (`common/ytdlp/fetchWindowManaged.ts:261`): every + attempt reloads the Rumble page, and Cloudflare refuses a share of loads. Elsewhere it is a retry on + `HLS_EXTENSION_REFUSED`, and a picky run that finds no HLS retries without it (`:65-67`, `:288-296`). +- **The /jobs row** says who asked, for what, how many: the first log line + ``Fetch windows: N window(s) for <requestedBy>[ · <manifest>]`` (`fetchWindows.ts:269`); progress metric + `clips`. + +### Channel create, rename and delete over ops + +- `create-channel`, `rename-channel`, `delete-channel` and `channel-config`'s `sites` are the New channel form + and the channel page's Danger zone over HTTP; there is no separate sites route. +- **A social URL makes a social channel**: `parseChannelForm` infers `sourceKind: "social"` from the platform + (X, Bluesky, XenForo; `isSocialPlatform`, `common/lib/platform.ts:42`) and refuses one with no derivable + handle (`editor/app/channels/components/parseChannelForm.ts:85-107`). +- **Confirmation:** delete needs `confirm` = the slug, rename the current slug + (`editor/app/channels/actions.ts:635`, `:695`). Both are refused while `channelMediaBusyReason` names a job or a + lane unit on the channel (`:644`, `:723`; `editor/app/channels/lib/mediaBusy.ts:48`), and delete while + `.relocating.json` exists (`common/controller/channels.ts:581-592`). +- A social channel's page now carries the same Danger zone (`data-social-danger`, + `editor/app/channels/[slug]/page.tsx:214`). + +### umtool's operator notes + +- **Where:** an article's notes are `SITES_DIR/<site>/reports/<report>/notes.json`, beside `report.json`; a + report-video project's are `<project>/notes.json` beside its manifest (`umtool/lib/annotations/targets.mjs:1-12`). + The article file is the ONE corpus file umtool writes, only through `isCorpusNotesFile` + (`umtool/lib/paths.mjs:252`: exactly `<site>/reports/<report>/notes.json`, realpath-checked). + `SITES_DIR` = `$SITES_DIR`, else `$TRANSCRIPTS_DIR/sites`, else `<repo>/transcripts/sites` (`paths.mjs:217-223`). +- **Shape:** `{format: "umtool-notes", version: 1, subject, source?, notes}` (`umtool/lib/annotations/shape.mjs:13-14`); + statuses `open|resolved|wontfix`, authors `operator|agent` (`:16-17`); anchors + `text|cite|section|whole|moment|entry|take|edit` (`:18`) — a `text` anchor is a quote with prefix/suffix, + re-located on read, never an offset (`umtool/lib/annotations/anchor.mjs`). +- **One writer, `writeOp`** (`umtool/lib/annotations/store.mjs:139`): a `<file>.lock` taken exclusively (stale + after `LOCK_STALE_MS` 30 s, `:28`), an mtime+size token (`notesToken`, `:50`) — a stale token is `StaleNotes` + (409, `:32`), an unparseable file is never overwritten (`NotesUnreadable`, `:41`) — then tmp + rename (`:152`); + the last note deleted removes the file. +- **Who:** `/api/notes` stamps every write `operator` (`umtool/app/api/notes/route.ts:46`); `umtool notes` stamps + `agent` (`umtool/lib/annotations/cli.mjs:67`, `:80`). The agent digest (`digest`, + `umtool/lib/annotations/digest.mjs:160`) is what both `umtool notes <target>` and + `GET /api/notes/context` print. +- **Never published:** compose and the report history read named files only; the guard is the test + `common/publish/composeReports.test.ts:794` (a sentinel in `notes.json` appears in no built file and no + history revision). + +### umtool kinds declare capabilities; nothing branches on a kind id + +- `PROJECT_KINDS` (`umtool/lib/projects/kinds.mjs`): `report-video` declares `notes: true` and + `linksArticles: true` (`:109`, `:112`); `song` and `sweep-report` neither. Callers ask + `kindTakesNotes` / `kindLinksArticles` (`:192`, `:195`) — `umtool/lib/annotations/targets.mjs:80`, `:164`; + `umtool/lib/articles/links.mjs:47` — and an e2e spec fails on a kind id written as a literal outside + `lib/projects/` and `components/projects/`. +- **/sites** lists every site's articles (`umtool/lib/articles/sites.ts`): published ids unioned with draft + report dirs, each with its open notes, its source and its linked video. An article row's `kind` is the REPORT's + kind (`factcheck|sweep`), not a project kind. diff --git a/plans/release-19.md b/plans/release-19.md @@ -43,6 +43,7 @@ second editor against the real corpus; the homepage build gate is `build:nodata` |---|---|---|---| | **A1** `pnpm ops` finds its token | `[unit]` | read `WORKER_TOKEN` / `ARCHILYZER_EDITOR_URL` from `editor/.env` when unset; `--wait` polls a real authed route instead of the `/api/jobs/active` rewrite | Wave 0 | | **A2** jobs over ops | `[unit]` + `[spec ops-api]` | `get jobs [--active\|--failed\|--kind\|--slug]`, `get job <id> [--tail]`, `job cancel\|retry\|retry-failed\|drain\|promote\|force-release <id…>`, `--wait` on many ids with queue position; wraps `editor/app/jobs/actions.ts` | A1 | +| **A2b** ops routes to unit tests | `[unit]` | the `ops-api.spec` route cases move to `route.test.ts` beside each route (moved from B6 on 2026-10-09: Track A owns the ops routes) | A2 | | **A3** read side | `[unit]` | `get settings\|storage\|sites\|workers\|auto-queue\|scheduler\|cleanup <slug>` over the existing view builders/readers; `/api/auto-queue/control` gated behind `opsAuth` (unauthenticated today) | A2 | | **A4** archival writes | `[unit]` | `settings` patch (through `saveSettings` + the schema); `lane` start/stop/drain for every lane incl. `publish`; `clear-platform-hold`; `workers enable\|disable`; per-video `transcribe-one\|delete-file\|do-not-clean`; cleanup buckets; `relocate {dryRun}`; `archilyzer storage report` (tierable bytes per channel) | A3 | | **A5** fetch queue hygiene | `[unit]` | the editor dedupes identical fetch-window requests (same slug/id/window → the existing job); `fetch-via-editor.mjs` waits without a timeout, printing queue position; small fetch-window jobs get their own queue key / priority instead of waiting behind a multi-hour `persist-videos` | A4 | @@ -57,11 +58,11 @@ second editor against the real corpus; the homepage build gate is `build:nodata` | slice | class | what | after | |---|---|---|---| | **B1** heavy-work gate | `[unit]` | `scripts/queue-lock.mjs` generalized into `pnpm heavy -- <cmd>`: one heavy slot machine-wide + a ≥ 6 GB free-memory floor; used by e2e, `next build` (publish stages, `build-site.sh`), `build-video.mjs` renders; a render may hold the transcription lane for its duration (ops `lane`); documented in WORKTREES.md + AGENTS.md | — | -| **B2** report-to-video robustness | `[unit]` | prune `*-frames` after the final mux (`--keep-frames` keeps them); `--chrome-only` builds missing segments; a manifest lint before render (teaser > 34 chars, image src relative to the manifest, a QR legible at 720p). `umtool/report-to-video/` is the Candace session's: coordinated with it | B1 | +| **B2** report-to-video robustness | `[unit]` | prune `*-frames` after the final mux (`--keep-frames` keeps them); `--chrome-only` builds missing segments; a manifest lint before render (teaser > 34 chars, image src relative to the manifest, a QR legible at 720p). Track B's, after B5; the Candace session reviews it before merge (ruled 2026-10-09) | B5 | | **B3** report CLI | `[unit]` | `archilyzer reports check <site> [--reports …]` (compose without a build); `reports verify-quotes <report.json>` (quote vs cue span, the en-orig check); `reports attach-video` encoding to fit the 24 MiB compose limit | — | | **B4** ops transcribe | `[unit]` | `durationMs` excludes queue wait; a priority for short one-file jobs over long auto jobs | Wave 0 | | **B5** umtool small debts | `[unit]` → `[spec umtool]` | per-worktree umtool e2e ports (`portFor`); `mix.spec` order dependence; `umtool window` writes edit notes; a selection spanning two blocks gets a Note button; the CLI finds `SITES_DIR` from the repo, not the cwd; umtool's /sites heading becomes "Articles" (naming hazard vs the editor's /sites). ONE umtool e2e run | B1 | -| **B6** test economy | `[unit]` | pure/API e2e moves to unit: `audio-check-classifier.spec` → common; `ops-api.spec` routes → `route.test.ts`; `view-route`, `worker-unit`, `media-file-abort`. Root `pnpm test` and `pnpm typecheck`; the editor gets an eslint config or loses its broken `lint` script | — | +| **B6** test economy | `[unit]` | pure/API e2e moves to unit: `audio-check-classifier.spec` → common; `view-route`, `worker-unit`, `media-file-abort`. Root `pnpm test` and `pnpm typecheck`; the editor gets an eslint config or loses its broken `lint` script | — | ### Track C — docs and plans (owner: the Track C implementer; no e2e) @@ -77,17 +78,17 @@ second editor against the real corpus; the homepage build gate is `build:nodata` | track | owns | |---|---| -| A | `editor/app/api/ops/**` (except B4's transcribe route), `scripts/archilyzer-ops.mjs` + its test, `mcp/src/**`, `editor/app/jobs/actions.ts`, `common/controller/fetchWindows.ts`, `fetch-via-editor.mjs` | -| B | `scripts/queue-lock.mjs` and the heavy gate, `umtool/**` (except `umtool/report-to-video/*`, the Candace session's), the e2e specs B6 moves, the report CLI, `common/controller/transcribeFile.ts`, `editor/app/api/ops/transcribe/**` | +| A | `editor/app/api/ops/**` (except B4's transcribe route), `scripts/archilyzer-ops.mjs` + its test, `mcp/src/**`, `editor/app/jobs/actions.ts`, `common/controller/fetchWindows.ts`, `fetch-via-editor.mjs`, the `ops-api.spec` cases A2b moves | +| B | `scripts/queue-lock.mjs` and the heavy gate, `umtool/**` (`umtool/report-to-video/*` for B2 only, the Candace session reviewing before merge), the e2e specs B6 moves, the report CLI, `common/controller/transcribeFile.ts`, `editor/app/api/ops/transcribe/**` | | C | `plans/**` (except the records others append), root `*.md` docs, `homepage/content/docs/**`, `mcp/README.md`, the `archilyzer docs cli` generator | | shared, append-only | `editor/CHANGELOG.md`, `AGENTS.md`, `common/bin/_cli.ts` (rows only) | ### Graph and merge order ``` -Wave 0 (release 18 lands) ──► A1 ──► A2 ──► … ──► A9 ──► release-end suite (Track A is sequential: A1–A5, then A6–A9) +Wave 0 (release 18 lands) ──► A1 ──► A2 ──► A2b ──► … ──► A9 ──► release-end suite (Track A is sequential: A1–A5, then A6–A9) └─► B4 -now: B1 ──► {B2, B5}; B3; B6 (B6 touches no release-18 file) +now: B1 ──► B5 ──► B2 (Candace session review); B3; B6 (B6 touches no release-18 file) now: C1, C3 ──► C2 ──► C4, C5; C2 regenerated after A's actions land ``` @@ -95,7 +96,7 @@ now: C1, C3 ──► C2 ──► C4, C5; C2 regenerated after A's actions la 18). C1, C3 and B6 touch no release-18 file and may merge first. Track A merges slice by slice in its own order; B and C merge whenever a slice is green. The parent regenerates the command reference (C2) after each Track A merge that adds an ops action or CLI row. -- **Do not touch:** `umtool/report-to-video/*` without the Candace session; `transcripts/**` except through the +- **Do not touch:** `umtool/report-to-video/*` except B2, which the Candace session reviews before merge; `transcripts/**` except through the writers; the working sessions' `~/reports/*` workspaces. ## Verification @@ -120,4 +121,41 @@ now: C1, C3 ──► C2 ──► C4, C5; C2 regenerated after A's actions la ### Track C +#### Slices C1–C5, as shipped (2026-10-09) + +Branch `worktree-agent-ad0ba0d0eda014c20` from `4cffda3f`; C1 and C3 merged into `r19/integration` at `e5e7c55e`. + +| commit | what | +|---|---| +| `11cb1b09` | C1: STATE "Now" (2026-10-09; older "Now" entries read "Previously"); this file and `release-20.md`; PLAN.md "Beyond the AI track: releases"; [`landed-2026-10.md`](landed-2026-10.md) (main's first-parent `39abec18..e921f82f`, 70 entries); one-core Phase 4 DONE 2026-09-28; `editor-operations-ia.md` COMPLETE (slice 9 = one-core Phase 1 slice 1.3); `stats-cache-key.md` merged `10cefd15`, rolled out with release 15; `plans/README.md` names the release and landed files | +| `2010600d` | C3: `mcp/README.md` lists `list_tags` (now every tool in `TOOLS`); README's docs table adds SETTINGS, SITE, CHANNEL, REPORT, CITATIONS, AGENTS | +| `86e4e436` | C2: `archilyzer docs cli [--check]` (`common/bin/cli-docs.ts` + test) → `COMMANDS.md`; `OPERATING.md`; links from AGENTS.md, README, CONTRIBUTING, RUNNING_IN_DOCKER | +| `d9564923` | C5: FACTS "Main's October work, outside release 18"; this file's A2b / B2 scope change | +| `e35523a4` | C4: `homepage/content/docs/operate.md` "From a shell or an agent"; homepage changelog | +| `b064c835` | merge `r19/integration` (release 21's plan, the jobKinds test fix) | + +**How the reference is made.** `COMMANDS.md` is a whole generated file, as ENVIRONMENT.md and SETTINGS.md are; +OPERATING.md links it. The `archilyzer` rows come from `COMMANDS` (`common/bin/archilyzer.ts`, where the table +lives — the `docs cli` row is there, not in `_cli.ts`). The `pnpm ops` rows come from the exported `usage()` of +`scripts/archilyzer-ops.mjs`, read by dynamic import, and from its header comment's `// pnpm ops <action> …` +example lines; the script is not changed. A help paragraph that opens with action names is that row's text; the +other paragraphs print as they stand. `--check` (and a plain run) also fails when OPERATING.md names a +`pnpm ops …` or `pnpm archilyzer …` that does not exist. + +**Gates** (logs `$T/c-gates-*.log`, `$T/c-homepage-build.log`): tsc clean (before C2's commit and after the +merge); `docs env|files|cli --check` 0; `bin/cli-docs.test.ts` 9/9; after the merge `bin/cli-docs.test.ts`, +`bin/_cli.test.ts`, `architecture.test.ts` and `jobs/jobKinds.test.ts` 49/49; homepage unit 23/23; homepage `build:nodata` ok (45 s, 5 GB scope, after the e2e run +ended). The whole common suite was started once and stopped on the parent's throttle (the editor suite was +running); only the changed files' tests ran. No e2e (class `[none]`/`[unit]`). + +**Found and left:** +- 15 ops actions have no help paragraph in `usage()`: `channel-priority`, `metadata-scan`, `import-video`, + `import-archive-org`, `feed-metadata`, `refresh-report`, `sync`, `download-missing`, `relocate`, + `relocate-back`, `evict-clips`, `lane`, `tags`, `tag-videos`, `keep-videos`. COMMANDS.md lists them with their + header examples (none for `download-missing`, `relocate-back`, `evict-clips`, `keep-videos`); OPERATING.md's + recipes cover the common ones. A paragraph each in `usage()` is Track A's file. +- `pnpm ops transcribe`'s `durationMs` includes the wait for a worker (FACTS; B4). +- fetch-windows: the cross-run platform gap is in memory, and dedupe is within one request (FACTS; A5). +- C2 is regenerated after each Track A merge that adds an action or a CLI row (`pnpm archilyzer docs cli`). + ## Rollout