Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit ab142494a9039c18a3afe6a62830be8f203f5a51
parent 1afdd499994315a3d607c93cf94458878dbe841a
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Fri,  9 Oct 2026 11:41:22 -0400

docs: OPERATING.md, the agent's guide to running an archive, over a generated COMMANDS.md (archilyzer docs cli)

- `archilyzer docs cli [--check]` (common/bin/cli-docs.ts) writes COMMANDS.md from the two help texts: every
  `archilyzer` row from the command table (path, arguments, what it does) and every `pnpm ops` action from
  `usage()` in scripts/archilyzer-ops.mjs (each help paragraph on the actions it opens with, the script's header
  examples beside them, the general paragraphs as they stand). An action with no help text still has its row.
  Both modes then check OPERATING.md: a `pnpm ops …` or `pnpm archilyzer …` that does not exist is a failure.
  Read-only over the ops script (its exported `usage()` and its header comment); no change to it.
- OPERATING.md: recipes for adding a channel, sync and download, imports (archive.org, Odysee, BitChute,
  Wayback), posts, transcription, tags, reports and publishing, over `pnpm ops`, `archilyzer` and the MCP.
- Linked from AGENTS.md (one paragraph), README's docs table, CONTRIBUTING (its six hand-picked CLI commands
  replaced by the link) and RUNNING_IN_DOCKER's ops examples.
- Tests: cli-docs.test.ts (9: splitting, escaping, attribution, examples, the unknown-command check, OPERATING.md's
  commands exist, COMMANDS.md is current).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
MAGENTS.md | 6++++++
ACOMMANDS.md | 134+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
MCONTRIBUTING.md | 14++++----------
AOPERATING.md | 160+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
MREADME.md | 2++
MRUNNING_IN_DOCKER.md | 3+++
Mcommon/bin/archilyzer.ts | 7+++++++
Acommon/bin/cli-docs.test.ts | 140+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/bin/cli-docs.ts | 274+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
9 files changed, 730 insertions(+), 10 deletions(-)

diff --git a/AGENTS.md b/AGENTS.md @@ -85,6 +85,12 @@ another markdown file that will drift from it. Use `/ask` for a question answered in the conversation and `/sweep` for a cited report written to a file. +**Running an archive — adding channels, syncing, importing, transcribing, tagging, +reports, publishing — is [OPERATING.md](OPERATING.md)**: recipes over `pnpm ops` (the +running editor's actions), `pnpm archilyzer` and the MCP. Every command and action is in +[COMMANDS.md](COMMANDS.md), generated by `pnpm archilyzer docs cli` from their help; after +adding an ops action or a CLI row, regenerate it (`--check` is a gate). + ## Clips and report-to-video The high-value loop: point the MCP at a public instance, ask about a subject, then pull diff --git a/COMMANDS.md b/COMMANDS.md @@ -0,0 +1,134 @@ +# Command reference + +<!-- GENERATED by common/bin/cli-docs.ts from the `archilyzer` command table (common/bin/archilyzer.ts) and `pnpm ops --help` (scripts/archilyzer-ops.mjs) — do not edit by hand. Regenerate: `pnpm archilyzer docs cli`. --> + +Every `archilyzer` command and every `pnpm ops` action, from their own help. Recipes that chain them: [OPERATING.md](OPERATING.md). + +## `pnpm archilyzer` + +Run from the repo root (in the container: `docker compose exec editor pnpm archilyzer …`). `--help` after a command prints its line. A command marked *passthrough* parses its own flags. + +| command | arguments | what it does | +|---|---|---| +| `archilyzer index` | | rebuild the LMDB transcript index | +| `archilyzer build stats` | | rebuild the stats datasets (reads the index; the data phase's second step) | +| `archilyzer build templates` | | bake each site's chart templates into its export staging dir (the data phase's third step) | +| `archilyzer build archives` | | warm the shared archive-zip cache for every enabled site's channels, once | +| `archilyzer compose site` | `<id> [--allow-missing-media]` | compose one site's export/public (default: SITE\_ID); --allow-missing-media lets a report citation with no prepared media through | +| `archilyzer compose hub` | | compose the hub's export/public (hub-sites.json, corpus.json, …) | +| `archilyzer compose homepage` | | compose homepage/public (whole-pool stats + landing summary) | +| `archilyzer publish index` | | update the index: the LMDB index, the stats datasets and the chart templates in one child (8 GB heap), then the index stamp every build reads | +| `archilyzer publish build` | `<id\|all> [--runner local\|docker\|auto] [--force] [--skip-archives]` | build a site (or every stale one) into its bundle &lt;exportBuildsDir&gt;/&lt;id&gt;/out from the current index; --runner docker builds every site in containers (host only); a fresh site is a no-op without --force | +| `archilyzer publish deploy` | `<id\|all> [--preview <branch>] [--to local] [--force]` | ship a site's bundle to its Pages project (a preview with --preview), or with --to local into ARCHILYZER\_SITE\_OUT; a bundle already deployed there is a no-op without --force | +| `archilyzer publish hub` | `[--deploy \| --deploy-only] [--preview <branch>] [--force]` | build the hub into its bundle &lt;exportBuildsDir&gt;/\_hub/out, then (--deploy) ship it; --deploy-only ships the bundle as built | +| `archilyzer publish homepage` | `[--deploy \| --deploy-only] [--preview <branch>] [--to local] [--force]` | build homepage/out (source mirror included), then (--deploy) ship it; --deploy-only ships it as built | +| `archilyzer publish status` | `[--json]` | the publish status: the index, the lane, and per site / hub / homepage its policy and its built, deployed and live chips, then the plan Publish now would run | +| `archilyzer publish now` | | run what Publish now runs — the index update when it is stale, then each policy target's build and deploy (site.json publish.auto; settings.json publish.hub / publish.homepage), never forced — one stage at a time IN THIS PROCESS under the publish lock, as the editor's queue would (the CLI has no queue); exit 0 when every stage ran or was a no-op | +| `archilyzer stage` | `<kind> <target> --run-id <id> [--preview <b>] [--to local] [--runner docker] [--force] [--skip-archives] [--allow-missing-media] [--index-after <ms>] [--built-after <ms>]` | INTERNAL: one publish stage, as the editor's job runs it (exit 0 ran/no-op, 1 failed, 2 usage, 3 precondition not met, 130 cancelled) | +| `archilyzer build site` | `<id> [--nodata] [--skip-archives] [--allow-missing-media]` | alias: publish index (not with --nodata) + publish build &lt;id&gt; --force (default id: SITE\_ID) | +| `archilyzer build all` | `[--skip-archives]` | alias: publish index + publish build all --runner auto (containers when an engine answers, else serially on the host) | +| `archilyzer build hub` | | compose:hub + INSTANCE\_MODE=hub next build into export/out — a raw build, unstamped; to deploy the hub build it with `publish hub` | +| `archilyzer build homepage` | `[--no-source]` | compose + source publish + next build in homepage/ (reads the index as it stands); --no-source removes the published source instead | +| `archilyzer reports prepare` | `<id>` | cut every clip and copy every post capture the site's published reports cite into its report-media cache, before its build, then export the reports as files (reports export) when nothing is missing (exit 1 when a citation lacks media or an export fails; default id: SITE\_ID) | +| `archilyzer reports export` | `<id> [--report <reportId>] [--formats html,pdf,md,zip] [--allow-missing-media]` | write the site's published reports as files (report.html, report.pdf, report.md, evidence-pack.zip) into its report-exports staging, where compose publishes them from (exit 1 on a problem; a PDF skipped for want of a browser is a note; default id: SITE\_ID) | +| `archilyzer reports convert` | `<sweep\|ask\|manifest> <in> --out <report.json> [--channels-dir <dir>] [--id <id>] [--title <title>]` | a /sweep report (markdown), an /ask answer or a report-to-video manifest as a report.json, written only when it validates (--channels-dir: widen spans from the cues, find posts' channels) | +| `archilyzer reports to-manifest` | `<report.json> --out <manifest.json> [--channels-dir <dir>] [--site-origin <url>]` | a starter report-to-video manifest from a report (a chapter card per section, a claim's still and clips stamped with its verdict, its posts) | +| `archilyzer source publish` | `[--force] [--check] [--keep-scratch]` | the scrubbed git mirror, raw tree, history pages (stagit, when installed) and tarball into homepage/public, behind the denied-literal gate (--check: audit and count, write nothing) | +| `archilyzer source audit` | `[<git dir>]` | the denied-literal gate (+ gitleaks) over a git dir; default the published homepage/public/source/archilyzer.git | +| `archilyzer deploy site` | `<id> [--preview <branch>]` | alias: publish deploy &lt;id&gt; \[--preview &lt;branch&gt;\] — ship the site's bundle to its Pages project (default id: SITE\_ID) | +| `archilyzer deploy hub` | `[--preview <branch>]` | alias: publish hub --deploy-only \[--preview &lt;branch&gt;\] — ship the hub's bundle to homepage.json's Pages project | +| `archilyzer deploy homepage` | `[--preview <branch>]` | alias: publish homepage --deploy-only \[--preview &lt;branch&gt;\] — ship homepage/out to the Pages project archilyzer | +| `archilyzer run` | `<operation> <channel> [ids…] [--lane local\|remote]` | run one catalogued operation over a channel offline, as the editor's job does (sync, downloads and transcription are refused: they run in the editor; it does not see the editor's lanes, so not beside one on the same channel) | +| `archilyzer archive-org refresh` | `<slug> [--dry-run]` | bring a channel's archive.org file records up to their provenance: a name-only mirror's title (where the item gives the file none) and upload date from its file name; offline, through the metadata history; prints old → new | +| `archilyzer wayback refresh` | `<slug> [--titles <json>] [--dry-run]` | bring a channel's Wayback Machine copies up to the Wayback rules: wayback.json (original URL, capture time), the dir renamed to its canonical id through the snapshot's reconcile (roster moved with it), and with --titles (a file of id → {title, upload\_date}) a raw file's title and date; offline, skips a record a live job holds; prints old → new | +| `archilyzer feeds backfill-metadata` | `<slug> [--feed <url>] [--dry-run]` | complete a podcast channel's records (title, date, description, duration) from its RSS feed: one fetch of the feed (default: the channel's url), no media; --dry-run counts matched / unmatched / already complete and writes nothing | +| `archilyzer duplicates` | `[--threshold N] [--all-durations] [--blocking title\|duration\|both] [--near F] [--tolerance N] …` | on-demand duplicate detection (after index + stats) *(passthrough)* | +| `archilyzer posts fetch` | `--slug <channel> [--full \| --older [--floor YYYY-MM-DD] [--from YYYY-MM-DD] [--force]] [--limit N] [--pages N]` | fetch a social channel's posts into its posts corpus (--older: walk back below the oldest archived post; --pages: a forum thread's latest N pages) *(passthrough)* | +| `archilyzer posts import-html` | `<slug> <file-or-dir>… [--dry-run]` | import forum thread pages saved from a browser ("Save page as", .html) into a forum-thread channel: new posts appended, edited ones updated, nothing fetched | +| `archilyzer posts check` | `--slug <channel> [--mode stale\|unchecked\|all] [--limit N]` | which archived posts were deleted at the source *(passthrough)* | +| `archilyzer diarize backfill` | `[--dry-run] [--scope transcribed\|channel:<slug>\|video:<slug>/<id>] [--limit N] [--force] …` | diarize videos whose audio is still on disk *(passthrough)* | +| `archilyzer digest plan` | `[--lane local\|remote] [--channels a,b] [--top N] [--json] [--census] …` | price the digest backfill; writes nothing *(passthrough)* | +| `archilyzer digest validate` | `<channel> [<channel> …]` | score digests already on disk *(passthrough)* | +| `archilyzer reconcile video-dirs` | `[--channel <slug>] [--dry-run] [--verbose]` | rename video dirs to the canonical id layout *(passthrough)* | +| `archilyzer verify transcripts` | `--channel <slug>` | list duplicate and missing transcripts *(passthrough)* | +| `archilyzer migrate channel-priority` | `[--dry-run]` | the one-shot channel-priority migration (plans/channel-priority.md, S5) *(passthrough)* | +| `archilyzer storage migrate-tier` | `<slug>…\|--all [--order smallest] [--include-large] [--dry-run] [--reclaim]` | bring a channel off the retired whole-directory layout onto the media tier, its text home to the corpus disk (editor stopped; --all stops before the three big-text channels) *(passthrough)* | +| `archilyzer brand media` | `[--out <dir>] [--video-kit]` | render the Archilyzer Media channel's assets *(passthrough)* | +| `archilyzer mcp` | `[--local <dir>\|--remote <url>\|--hub <url>]` | start the MCP server on stdio (as `pnpm --filter yt-dlp-transcript-mcp exec tsx src/index.ts`) *(passthrough)* | +| `archilyzer sync tick` | | POST one scheduler tick to the editor (SYNC\_TICK\_URL, SYNC\_TICK\_TOKEN) | +| `archilyzer docs env` | `[--check]` | write ENVIRONMENT.md from the declared env-var list (lib/envVars.ts) | +| `archilyzer docs files` | `[--check]` | write SITE.md + CHANNEL.md + REPORT.md + CITATIONS.md from the file schemas | +| `archilyzer docs cli` | `[--check]` | write COMMANDS.md from this command table and `pnpm ops --help` | +| `archilyzer settings example` | `[--check]` | write settings.json.example + SETTINGS.md from the schema | +| `archilyzer doctor` | `[--json]` | read-only report: node, the checkout, the corpus, settings, every tool, the port block; exit 1 on a failure | +| `archilyzer release show` | `[editor\|export]` | latest release, its date, how many bullets wait under \[Unreleased\] | +| `archilyzer release cut` | `<editor\|export\|all> <X.Y.Z\|next\|next-minor> [--commit] [--date YYYY-MM-DD]` | \[Unreleased\] -&gt; a dated heading; all = both, one version, two commits | + +## `pnpm ops` + +Drives a running editor over HTTP (`/api/ops/*`, the same actions its pages run), gated by `WORKER_TOKEN`. A body is `--json '<object>'` or `--file <path>`; `--wait` follows a job to its end. + +| action | what it does | example | +|---|---|---| +| `channel-priority` | — | `pnpm ops channel-priority --json '{"slugs":["x"],"operation":"download","tier":"paused"}'` | +| `channel-config` | channel-config changes a channel as its Configure form does: {"slug"} and any of "patch" (form field names; "" clears one), "sites" (the WHOLE membership set: \[{"siteId", "groupId"? \| "newGroupName"?}\], \[\] = on no site; an unknown site id is refused), "excludeFromBuild" and "excludeFromCleanup" (set to the value given, not toggled). | `pnpm ops channel-config --json '{"slug":"x","patch":{"downloadFilterExclude":"rerun"}}'`<br>`pnpm ops channel-config --json '{"slug":"x","sites":[{"siteId":"anilyzer"}]}'`<br>`pnpm ops channel-config --json '{"slug":"x","sites":[],"excludeFromBuild":true}'` | +| `create-channel` | create-channel is the New channel form: {"fields": {"name", "handling": "youtube"\|"transcribe", "url"?, "platform"?, "sourceKind"?, "postFetcher"?, "socialHandle"?, …}} with channel-config's patch keys; "slug"? (else derived from the name), "sites"? (absent = on no site). "fetchPlaylist", "fetchPostsNow" and "prioritizeDownload" are the form's checkboxes, OFF unless true; a job they start comes back as jobId(s), so --wait follows it. | `pnpm ops create-channel --json '{"fields":{"name":"Example (X)","handling":"transcribe","url":"https://x.com/example"}}'` | +| `rename-channel` | rename-channel moves a channel to a new slug, as Danger → Rename does: {"slug", "newSlug"}. Refused while the channel is busy (a job, a lane unit, media in transition) or when the new slug is taken. Old links break. | `pnpm ops rename-channel --json '{"slug":"old-slug","newSlug":"new-slug"}'` | +| `delete-channel` | delete-channel removes a channel's whole directory, as Danger → Delete does: {"slug", "confirm"} — "confirm" must repeat the slug. No undo outside the transcripts/ repo's own history. | `pnpm ops delete-channel --json '{"slug":"x","confirm":"x"}'` | +| `metadata-scan` | — | `pnpm ops metadata-scan --json '{"slug":"the-quartering"}'` | +| `refresh-metadata` | refresh-metadata re-reads ONE video's metadata.info.json from its source (no subtitles, no media) on the platform's queue: {"slug", "id"}. The job's log ends with what the source now says — live\_status, formats, audio-only formats and whether any is non-fragmented, English captions, the keys that changed. An id with no data/&lt;id&gt;/ is refused (a refresh re-reads a video already archived), as are archive.org and Wayback records. | `pnpm ops refresh-metadata --json '{"slug":"the-quartering","id":"<videoId>"}' --wait` | +| `import-video` | — | `pnpm ops import-video --json '{"slug":"demo-archive","url":"https://archive.org/details/example-item"}'` | +| `import-archive-org` | — | `pnpm ops import-archive-org --json '{"slug":"demo-archive","item":"example-item","match":"\\.mp4$"}' --wait` | +| `feed-metadata` | — | `pnpm ops feed-metadata --json '{"slug":"demo-podcast","dryRun":true}' --wait` | +| `refresh-report` | — | `pnpm ops refresh-report --json '{"all":true}'` | +| `sync` | — | `pnpm ops sync --json '{"slug":"the-quartering"}' --wait` | +| `download-missing` | — | | +| `retry-bucket` | retry-bucket runs one bucket of a channel's report as one job, past any lane hold: {"slug", "bucket"}. "ids": \[...\] runs only those videos, and every one must be in the bucket (a stray id is refused, named); a job run with ids is not replayable, as a checkbox selection in the UI is not. | | +| `transcribe-bucket` | transcribe-bucket transcribes a channel's "downloaded, not transcribed" bucket on the transcription queue, as the channel page's Transcribe button does: {"slug"}. "ids": \[...\] narrows it the same way as on retry-bucket. | | +| `fetch-posts` | fetch-posts fetches a social channel's new posts: {"slug"}. "full": true re-walks the whole timeline; "older": true walks back from the oldest archived post through search (X; needs a login), saving its place for the next run, down to "floor": "YYYY-MM-DD" when given; "from": "YYYY-MM-DD" starts the walk afresh there, replacing its saved place (and a "complete") — for a gap above one surviving old post. "limit": N caps the posts one run reads; "pages": N caps the pages (a forum thread: its latest N pages). "full" and "older" together are refused. An older walk over an account that shows no posts (nothing archived, and the last timeline fetch read none) is refused unless "force": true. A drained fetch stops at its next resume point and the next run resumes. | | +| `capture-posts` | capture-posts captures archived posts of a social channel (X, forum): a screenshot of each through the connected X profile (a forum thread: its host's forum profile), and its attached media through gallery-dl (a forum thread: the same profile), into the channel's posts-media/&lt;id&gt;/: {"slug", "ids": \[...\]}. Every id must be in the channel's posts archive. "shots": false or "media": false skips that half; posts already captured are skipped unless "force": true. A post that links to an X Article also gets the article (article.json, .md, .png and its images) unless "articles": false; both halves off with "articles": true reads only the articles. Paced like a post fetch, on its queue. | | +| `publish` | publish runs publish stages on the editor's publish queue, one at a time, under one run id: {"verb": …}. "index" updates the index; "build" builds "siteId"/"siteIds" (forced; the index first when stale; "runner": "docker" builds every site in containers); "deploy" ships their built bundles (production, "preview": "&lt;branch&gt;", or "to": "local"; "force" redeploys a bundle already shipped there); "hub" / "homepage" build them, {"deploy": true} deploys after; "now" is Publish now (the stale index, then each policy target); "stale" builds every stale site. The answer lists every job ({target, kind, jobId}) and --wait follows them all. A site never built is refused: "no build of &lt;id&gt; in &lt;dir&gt; — archilyzer publish build &lt;id&gt;". `get publish` is the status. | | +| `build-index`, `build-site`, `build-deploy`, `deploy-site`, `build-hub`, `deploy-hub`, `build-homepage`, `deploy-homepage` | build-index, build-site, build-deploy, deploy-site, build-hub, deploy-hub, build-homepage and deploy-homepage are publish's aliases, with their old bodies and answers ("skipData" is accepted and ignored). | `pnpm ops build-deploy --json '{"siteIds":["anilyzer","jeralyzer"]}' --wait`<br>`pnpm ops build-site --json '{"siteId":"anilyzer"}' --wait`<br>`pnpm ops deploy-site --json '{"siteId":"anilyzer","preview":"tags-exclude"}' --wait`<br>`pnpm ops build-hub --wait`<br>`pnpm ops build-hub --json '{"deploy":true}' --wait`<br>`pnpm ops deploy-hub --wait`<br>`pnpm ops build-homepage --json '{"deploy":true}' --wait`<br>`pnpm ops deploy-homepage --json '{"preview":"refresh"}' --wait` | +| `build-site`, `build-deploy`, `deploy-site` | build-site, build-deploy and deploy-site all take "siteId" (one) or "siteIds" (a list). | | +| `build-hub` | build-hub builds the hub into its bundle; {"deploy": true} deploys it after, and deploy-hub ships the one already built. Both deploy to the Pages project set on /sites under Hub, and take "preview" too. | | +| `build-homepage` | build-homepage builds the homepage package into homepage/out; {"deploy": true} deploys it after (only if the build succeeded), and deploy-homepage ships the one already built. Both deploy to the Pages project archilyzer (https://archilyzer.pages.dev), production unless "preview" is given. | | +| `relocate` | — | `pnpm ops relocate --json '{"slugs":["x"],"locationId":"platter"}'` | +| `relocate-back` | — | | +| `evict-clips` | — | | +| `reports-prepare` | reports-prepare cuts every clip and copies every post capture a site's published reports cite into its report-media cache, before its build: {"siteId"}. The job fails, naming each one, when a citation lacks media. When nothing is missing it then exports the reports, as reports-export. | `pnpm ops reports-prepare --json '{"siteId":"demo-site"}' --wait` | +| `reports-export` | reports-export writes each published report as report.html, report.pdf, report.md and evidence-pack.zip for the site's build to publish: {"siteId", "reportId"?, "formats"?: \["html","pdf","md","zip"\]}. | `pnpm ops reports-export --json '{"siteId":"demo-site","formats":["html","md"]}' --wait` | +| `lane` | — | `pnpm ops lane --json '{"lane":"download","held":true}'` | +| `tags` | — | `pnpm ops tags --json '{"op":"define","tag":{"id":"eva-collab","label":"Collab"}}'` | +| `tag-videos` | — | `pnpm ops tag-videos --file ids.json` | +| `keep-videos` | — | | +| `persist-videos` | persist-videos saves specific videos, across channels, to the saved-video store: {"items": \[{"slug", "id"}, ...\]}. "format": "original" \| "video\_720" (default: each channel's own). "replace": "above-height" also re-fetches a saved one whose height is unknown or above that quality (default "never"). "gapMs" pauses between downloads (default the batch gap), "minFreeMemMb" waits for that much free memory before each. "dryRun": true answers with the buckets (saved, wrongHeight, toFetch, noUrl, unknown) and starts nothing. One job per channel, on its download queue; a low disk or a rate limit stops it, and running the same body again resumes — saved videos are skipped. | `pnpm ops persist-videos --file list.json --wait` | +| `fetch-windows` | fetch-windows fetches clip windows, one paced job per platform queue (YouTube and Rumble side by side): {"siteId"} fetches every window the site's published reports cite and the disk does not hold; {"items": \[{"slug", "id", "from", "to", "clipId"?, "reason"?, "pad"?, "webpageUrl"?}, ...\], "requestedBy", "manifest"?} fetches a list. "maxHeight" caps the source height (default 720). "dryRun": true lists the windows per platform, the ones already on disk ("cached") and the ones no fetch can fill ("unfetchable": deleted, off the site) and starts nothing. A platform cooling down or held is refused for its group; a 429, or two 403s in a row, backs the platform off and stops its job. Running the same body again resumes — fetched windows are cached. | `pnpm ops fetch-windows --json '{"siteId":"demo-site","dryRun":true}'`<br>`pnpm ops fetch-windows --file windows.json --wait` | +| `cut-release` | cut-release turns a changelog's \[Unreleased\] into "## \[&lt;version&gt;\] - &lt;date&gt;": {"workspace": "editor" \| "export" \| "all", "version": "X.Y.Z" \| "next" \| "next-minor", "commit": boolean (default false), "date": "YYYY-MM-DD" (default today)}. "all" cuts both with ONE version and commits each ("Release &lt;workspace&gt; &lt;version&gt;") — or neither: every check runs before either file is written. Only an editor built from release 10 or later has the route (an older one answers 404); with no editor running, `archilyzer release cut` does the same locally. | `pnpm ops cut-release --json '{"workspace":"all","version":"next","commit":true}'` | +| `transcribe` | transcribe runs ONE local file through a local transcription worker, as a job: {"path": "/abs/file"} (audio or video), "start"/"end" (seconds) for a window, "workerId" (a settings worker id; default: the one auto-transcribe would get), "out" (an absolute path for the result JSON, never inside the corpus). The result is {path, window, worker: {id, appId, model, device}, cues: \[{start, end, text}\], text, ...}, cue times on the file's own clock. "words": true adds words: \[{w, start, end, conf?}\] on the same clock -- from parakeet, which keeps its word timestamps; \[\] from an engine that does not. With --wait it is printed on stdout (the response and the log go to stderr), so `pnpm ops transcribe ... --wait \| jq -r .text` works. | `pnpm ops transcribe --json '{"path":"/abs/clip.mp4","start":120,"end":150}' --wait`<br>`pnpm ops transcribe --json '{"path":"/abs/a.wav","workerId":"parakeet-cpu","out":"/tmp/a.json"}' --wait` | + +### Usage, flags and environment + +```text +Usage: pnpm ops <action> [--json '<body>' | --file <path>] [--wait] + [--wait-timeout <seconds>] [--quiet] + pnpm ops get channel <slug> [--counts] + pnpm ops get channels + pnpm ops get tags [<tagId>] + pnpm ops get publish + pnpm ops list + +--wait follows the job's log and survives a poll that fails (a busy + in-process build starves the server): it backs off and, after three + failures, asks /api/jobs/active whether the job is still there. +--wait-timeout <seconds> gives up and exits 1 instead of waiting forever. + Default: no timeout — the queue may legitimately hold a job for hours. + +"preview": "<branch>" on deploy-site or build-deploy makes it a Cloudflare + Pages PREVIEW instead of production: the same bundle goes to a branch + alias, https://<branch>.<project>.pages.dev, and the live site is left + alone. The alias is printed after the response. Lowercase letters, + digits and dashes, up to 28 characters; "main" is refused. + +Env: ARCHILYZER_EDITOR_URL (default http://localhost:3001), WORKER_TOKEN, + ARCHILYZER_AGENT (provenance of a tag write; default "cli") +``` diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md @@ -146,16 +146,10 @@ The same controllers the editor uses are one command line, `common/bin/archilyze which is how you drive the pipeline headlessly or from cron. `pnpm archilyzer <command>` from the repo root is the short form of `pnpm --filter yt-dlp-transcript-common exec tsx bin/archilyzer.ts <command>`; `pnpm archilyzer ---help` lists every command. - -```bash -pnpm archilyzer doctor # read-only: can this machine do what it is configured to? -pnpm archilyzer index # the LMDB index -pnpm archilyzer run diarization <channel> [ids…] # one catalogued operation, offline, as the editor's job -pnpm archilyzer build site <id> # publish: see PUBLISH.md -pnpm archilyzer verify transcripts --channel <slug> -pnpm archilyzer mcp # the MCP server on stdio -``` +--help` lists every command. Every command and every `pnpm ops` action is in +[COMMANDS.md](COMMANDS.md), generated by `pnpm archilyzer docs cli` from the two help +texts (`--check` fails when it is stale, or when [OPERATING.md](OPERATING.md) names a +command that does not exist); recipes that chain them are in OPERATING.md. The table is `archilyzer.ts`; the machinery (parser, lookup, usage) is `_cli.ts`. Every file in `common/bin/` is reachable from a row — a test fails otherwise. A bin that diff --git a/OPERATING.md b/OPERATING.md @@ -0,0 +1,160 @@ +# Operating an archive + +Recipes for running an archive from a shell or an agent. Every command named here is in +[COMMANDS.md](COMMANDS.md), generated from the commands' own help. + +## The three surfaces + +| surface | what it is | needs | +|---|---|---| +| `pnpm ops <action>` | the running editor's actions over HTTP (`/api/ops/*`): the same code a click runs, as jobs on the editor's queues | a running editor; `ARCHILYZER_EDITOR_URL` (default `http://localhost:3001`) and `WORKER_TOKEN` (the editor's own, from `editor/.env`) | +| `pnpm archilyzer <command>` | the core's CLI: index, compose, publish stages, reports, offline refreshes, doctor | the checkout and its `transcripts/`; in the container, `docker compose exec editor pnpm archilyzer …` | +| the MCP server | reads a published archive (search, transcripts, reports); `fetch_clip` asks the editor for clip media | registered as `archilyzer` ([AGENTS.md](AGENTS.md), [mcp/README.md](mcp/README.md)) | + +- A job-starting `pnpm ops` action answers with a `jobId` once the job is queued. `--wait` follows its log to the + end and exits with its status; a platform queue may hold a job for hours. +- What is there: `pnpm ops get channels`, `pnpm ops get channel <slug> --counts`, `pnpm ops get publish`, + `pnpm archilyzer doctor`. +- Every fetch goes through the editor (paced per platform, cookie-aware, provenanced). Never run yt-dlp, whisper or + parakeet by hand, and never edit `transcripts/**` by hand: each file has one writer, named below. + +## Add a channel + +```sh +pnpm ops create-channel --json '{"fields":{"name":"Example","handling":"youtube","url":"https://www.youtube.com/@example"}}' +pnpm ops channel-config --json '{"slug":"example","patch":{"downloadFilterExclude":"#shorts"}}' +pnpm ops channel-config --json '{"slug":"example","sites":[{"siteId":"jeralyzer"}]}' +pnpm ops get channel example +``` + +- `handling`: `"youtube"` fetches the platform's captions; `"transcribe"` downloads audio and transcribes it here. + `fields` takes the New channel form's field names, the same as `channel-config`'s `patch` + ([RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md), "Driving the editor without a browser"). Every `config.json` key: + [CHANNEL.md](CHANNEL.md). +- `sites` is the whole membership set; absent or `[]` keeps the channel on no site. +- `"fetchPlaylist": true` on `create-channel` stores the playlist at once (the form's "Fetch playlist now"); a social channel's is `"fetchPostsNow": true`. +- Rename or delete: `pnpm ops rename-channel --json '{"slug":"old","newSlug":"new"}'`, + `pnpm ops delete-channel --json '{"slug":"x","confirm":"x"}'` — both refused while the channel is busy. + +## Sync and download + +```sh +pnpm ops sync --json '{"slug":"example"}' --wait +pnpm ops metadata-scan --json '{"slug":"example"}' --wait +pnpm ops download-missing --json '{"slug":"example"}' --wait +pnpm ops retry-bucket --json '{"slug":"example","bucket":"<bucket>"}' --wait +pnpm ops refresh-metadata --json '{"slug":"example","id":"<videoId>"}' --wait +pnpm ops persist-videos --json '{"items":[{"slug":"example","id":"<videoId>"}],"dryRun":true}' +``` + +- The lanes do this unattended: `pnpm ops lane --json '{"lane":"download","enabled":true,"action":"start"}'`; + `"held": true` pauses dispatch and keeps the runner's place. Lanes: `transcription`, `download`, `digest`, + `backfill`. +- A channel's priority: `pnpm ops channel-priority --json '{"slugs":["example"],"tier":"low"}'` + (`"operation"` pins one operation's tier). +- Bucket names and sizes: `pnpm ops get channel example`. +- Keep videos out of cleanup by title or description: `pnpm ops keep-videos --json '{"slug":"example","match":"interview","dryRun":true}'`. + +## Import from archive.org, Odysee, BitChute and the Wayback Machine + +```sh +pnpm ops import-archive-org --json '{"slug":"example-archive","item":"<item>","match":"\\.mp4$","dryRun":true}' --wait +pnpm ops import-archive-org --json '{"slug":"example-archive","item":"<item>","match":"\\.mp4$"}' --wait +pnpm ops import-video --json '{"slug":"example","url":"https://www.bitchute.com/video/<id>/"}' --wait +pnpm ops import-video --json '{"slug":"example","url":"https://odysee.com/@example:0/<video>:0"}' --wait +pnpm ops import-video --json '{"slug":"example","url":"https://web.archive.org/web/<timestamp>/<original-url>"}' --wait +``` + +- `import-archive-org` takes one item's files (`files` exact, or `match` a regex), one at a time on archive.org's + queue; files already held are skipped. archive.org is fetched over BitTorrent when it can be, never by yt-dlp. +- A one-off Odysee or BitChute import is paced like that platform's own downloads. A whole Odysee or BitChute + channel is a channel with that URL, then `sync`. +- A Wayback capture is named by what it copies; `wayback.json` beside it records the capture. +- Bring records up to the current rules, offline (dry run first): + `pnpm archilyzer archive-org refresh <slug> --dry-run`, `pnpm archilyzer wayback refresh <slug> --dry-run`, + `pnpm archilyzer feeds backfill-metadata <slug> --dry-run` (or `pnpm ops feed-metadata` on the editor). + +## Fetch and capture posts + +```sh +pnpm ops create-channel --json '{"fields":{"name":"Example (X)","handling":"transcribe","url":"https://x.com/example"}}' +pnpm ops fetch-posts --json '{"slug":"example-x"}' --wait +pnpm ops fetch-posts --json '{"slug":"example-x","older":true}' --wait +pnpm ops capture-posts --json '{"slug":"example-x","ids":["<postId>"]}' --wait +``` + +- An X, Bluesky or XenForo URL makes a social channel (its handle derived from the URL). `"older": true` walks back below the oldest + archived post and resumes from its saved place; `"full": true` re-walks the timeline. +- `capture-posts` saves a screenshot and the attached media of named posts (and a linked X Article). +- Forum pages saved from a browser: `pnpm archilyzer posts import-html <slug> <file-or-dir> --dry-run`. +- Which archived posts were deleted at the source: `pnpm archilyzer posts check --slug <slug>`. + +## Transcribe + +```sh +pnpm ops lane --json '{"lane":"transcription","enabled":true,"action":"start"}' +pnpm ops transcribe-bucket --json '{"slug":"example"}' --wait +pnpm ops transcribe --json '{"path":"/abs/clip.mp4","start":120,"end":150,"words":true}' --wait +``` + +- The transcription lane takes every channel's downloaded, untranscribed audio; `transcribe-bucket` runs one + channel's now. +- `transcribe` runs one local file (or a window of it) through the corpus's own engine and model; the result JSON + (cues, text, and with `"words": true` the word timings parakeet keeps) is printed on stdout with `--wait`. + `"out"` writes it to a file outside the corpus. + +## Tag + +```sh +pnpm ops get tags +pnpm ops tags --json '{"op":"define","tag":{"id":"example-collab","label":"Collab"}}' +pnpm ops tag-videos --json '{"tag":"example-collab","op":"add","videos":[{"slug":"example","id":"<videoId>"}]}' +pnpm ops tag-videos --file ids.json +``` + +- `transcripts/tags.json` has one writer, and `tags` and `tag-videos` go through it, recording who asked + (`ARCHILYZER_AGENT`). `remove` unpins; `suppress` rejects a rule's hit. +- A tag's `sites` limits the sites it exists on. +- Read side: the MCP's `list_tags`, and `tags` on `search_transcripts` / `enumerate_matches`. + +## Prepare and export reports + +```sh +pnpm archilyzer reports convert sweep <sweep.md> --out <report.json> +pnpm ops fetch-windows --json '{"siteId":"example-site","dryRun":true}' +pnpm ops fetch-windows --json '{"siteId":"example-site"}' --wait +pnpm ops reports-prepare --json '{"siteId":"example-site"}' --wait +pnpm ops reports-export --json '{"siteId":"example-site","formats":["html","md"]}' --wait +``` + +- A report is `transcripts/sites/<site>/reports/<id>/report.json` ([REPORT.md](REPORT.md), + [CITATIONS.md](CITATIONS.md)); `reports convert` makes one from a `/sweep` report, an `/ask` answer or a + report-to-video manifest. +- `fetch-windows` with `siteId` fetches every clip window the site's published reports cite that the disk does not + hold, one paced job per platform; a re-run resumes. +- `reports-prepare` cuts the cited clips and copies the cited post captures, then exports; it fails naming each + citation that lacks media. Offline: `pnpm archilyzer reports prepare <siteId>`, + `pnpm archilyzer reports export <siteId>`. +- One cited moment's media, from an agent: the MCP's `fetch_clip`. + +## Publish + +```sh +pnpm ops get publish +pnpm ops publish --json '{"verb":"now"}' --wait +pnpm ops publish --json '{"verb":"build","siteId":"example-site"}' --wait +pnpm ops publish --json '{"verb":"deploy","siteId":"example-site","preview":"check"}' --wait +pnpm archilyzer publish status +``` + +- Publishing is stages on one queue: update the index, build a site into its bundle, deploy the bundle, each + checked live. `now` runs what Publish now runs (the stale index, then each site's policy). Offline, under the same + lock: `pnpm archilyzer publish index`, `publish build <id>`, `publish deploy <id> --preview <branch>`. +- Details, policies and the hub and homepage: [PUBLISH.md](PUBLISH.md). + +## Research with the MCP + +- `/ask` answers a question with citations in the conversation; `/sweep` writes a cited report to a file. +- Search first, then pull only the cited seconds with `fetch_clip`. The editor fetches only for a channel it + already archives. +- Tools and their arguments: [mcp/README.md](mcp/README.md). diff --git a/README.md b/README.md @@ -640,6 +640,8 @@ See [CONTRIBUTING.md](CONTRIBUTING.md) to work on the code. | [SETUP.md](SETUP.md) | Full per-OS install, transcription backends. | | [ENVIRONMENT.md](ENVIRONMENT.md) | Every environment variable, by audience (generated). | | [CONTRIBUTING.md](CONTRIBUTING.md) | Workspace layout, tests, the `archilyzer` CLI, internals. | +| [OPERATING.md](OPERATING.md) | Running an archive from a shell or an agent: recipes over `pnpm ops`, `archilyzer` and the MCP. | +| [COMMANDS.md](COMMANDS.md) | Every `archilyzer` command and `pnpm ops` action (generated). | | [RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md) | `docker compose up` for the whole stack: exposure model, auth, GPU. | | [PUBLISH.md](PUBLISH.md) | Building and deploying sites: Pages + R2, cost-abuse protection, parallel builds in containers. | | [SCHEDULED_SYNC.md](SCHEDULED_SYNC.md) | Unattended per-channel syncing. | diff --git a/RUNNING_IN_DOCKER.md b/RUNNING_IN_DOCKER.md @@ -413,6 +413,9 @@ pnpm ops get channels # every channel, its kind and its sites pnpm ops list # every action name ``` +These are examples. Every action and getter, with its body, is in [COMMANDS.md](COMMANDS.md) (generated from +`pnpm ops --help` and `pnpm archilyzer --help`); recipes that chain them are in [OPERATING.md](OPERATING.md). + Four things to know before you script against it: - **A job-starting action returns a `jobId` and does not stream.** The job may diff --git a/common/bin/archilyzer.ts b/common/bin/archilyzer.ts @@ -542,6 +542,13 @@ export const COMMANDS: Command[] = [ (await import("./file-schemas-docs")).main({ check: flags.check === true }), }, { + path: ["docs", "cli"], + usage: "[--check] write COMMANDS.md from this command table and `pnpm ops --help`", + flags: { check: "boolean" }, + run: async ({ flags }) => + (await import("./cli-docs")).main({ check: flags.check === true }), + }, + { path: ["settings", "example"], usage: "[--check] write settings.json.example + SETTINGS.md from the schema", flags: { check: "boolean" }, diff --git a/common/bin/cli-docs.test.ts b/common/bin/cli-docs.test.ts @@ -0,0 +1,140 @@ +import assert from "node:assert/strict"; +import { readFileSync } from "node:fs"; +import path from "node:path"; +import { test } from "node:test"; +import { fileURLToPath } from "node:url"; +import { + COMMANDS_FILE, + generateCommandsMarkdown, + mdCell, + opsActions, + opsExamples, + opsParagraphs, + RECIPES_FILE, + renderCommandsMarkdown, + splitUsage, + unknownCommandsIn, + unknownRecipeCommands, +} from "./cli-docs"; + +const REPO = path.resolve(path.dirname(fileURLToPath(import.meta.url)), "..", ".."); + +const OPS_USAGE = [ + "Usage: pnpm ops <action> [--json '<body>'] [--wait]", + " pnpm ops list", + "", + "Actions: sync, build-site, deploy-site, lane", + "", + "--wait follows the job's log.", + "", + 'build-site and deploy-site take "siteId" or', + ' "siteIds" (a list).', + "", + 'deploy-site ships a built bundle: {"siteId"}. See `a | b`.', + "", + "Env: WORKER_TOKEN", +].join("\n"); + +test("splitUsage splits at the first double space; a usage with none is all text", () => { + assert.deepEqual(splitUsage("<id> [--force] build it now"), { + args: "<id> [--force]", + text: "build it now", + }); + assert.deepEqual(splitUsage("rebuild the index"), { args: "", text: "rebuild the index" }); +}); + +test("mdCell escapes markup outside code and only the pipe inside it", () => { + assert.equal(mdCell("a <id> | *b* _c_ [x]"), "a &lt;id&gt; \\| \\*b\\* \\_c\\_ \\[x\\]"); + assert.equal(mdCell("run `x | y <z>` then"), "run `x \\| y <z>` then"); +}); + +test("opsActions reads the Actions line, and refuses a usage without one", () => { + assert.deepEqual(opsActions(OPS_USAGE), ["sync", "build-site", "deploy-site", "lane"]); + assert.throws(() => opsActions("Usage: nothing"), /no `Actions:` line/); +}); + +test("opsParagraphs attributes a paragraph to every action it opens with", () => { + const ps = opsParagraphs(OPS_USAGE, opsActions(OPS_USAGE)); + assert.deepEqual( + ps.map((p) => p.actions), + [[], [], ["build-site", "deploy-site"], ["deploy-site"], []], + ); + assert.ok(!ps.some((p) => p.text.startsWith("Actions: "))); +}); + +test("opsExamples takes header lines of known actions only, padding collapsed", () => { + const source = [ + "// pnpm ops sync --json '{\"slug\":\"x\"}' --wait", + "// pnpm ops get channel x", + "// pnpm ops lane", + "// pnpm ops sync --json '{\"slug\":\"y\"}'", + "const x = 1; // pnpm ops sync not a header line", + ].join("\n"); + const ex = opsExamples(source, ["sync", "lane"]); + assert.deepEqual(ex.get("sync"), [ + "pnpm ops sync --json '{\"slug\":\"x\"}' --wait", + "pnpm ops sync --json '{\"slug\":\"y\"}'", + ]); + assert.deepEqual(ex.get("lane"), ["pnpm ops lane"]); + assert.equal(ex.has("get"), false); +}); + +test("the rendered reference lists every command and every action exactly once as a row head", () => { + const md = renderCommandsMarkdown( + [ + { path: ["index"], usage: "rebuild the LMDB transcript index" }, + { path: ["publish", "build"], usage: "<id|all> [--force] build a site" }, + { path: ["mcp"], usage: "[--local <dir>] start the MCP server", passthrough: true }, + ], + OPS_USAGE, + new Map([["deploy-site", ["pnpm ops deploy-site --json '{}'"]]]), + ); + assert.match(md, /^# Command reference\n\n<!-- GENERATED by common\/bin\/cli-docs\.ts/); + assert.match(md, /\| `archilyzer index` \| \| rebuild the LMDB transcript index \|/); + assert.match(md, /\| `archilyzer publish build` \| `<id\\\|all> \[--force\]` \| build a site \|/); + assert.match(md, /\| `archilyzer mcp` \| .* \*\(passthrough\)\* \|/); + // An action with no paragraph still has its row. + assert.match(md, /\| `sync` \| — \| \|/); + assert.match(md, /\| `lane` \| — \| \|/); + // The shared paragraph is one row naming both; deploy-site's example sits on + // that first row, not on its own paragraph's. + assert.match( + md, + /\| `build-site`, `deploy-site` \| build-site and deploy-site take "siteId" or "siteIds" \(a list\)\. \| `pnpm ops deploy-site --json '\{\}'` \|/, + ); + assert.match(md, /\| `deploy-site` \| deploy-site ships a built bundle: \{"siteId"\}\. See `a \\\| b`\. \| \|/); + // The general paragraphs are printed as they stand, in a text block. + assert.match(md, /```text\nUsage: pnpm ops <action>[^]*--wait follows the job's log\.\n\nEnv: WORKER_TOKEN\n```\n$/); + assert.doesNotMatch(md, /Actions: sync/); +}); + +test("unknownCommandsIn names what no action, getter or command answers to", () => { + const cli = [{ path: ["publish", "build"], usage: "" }, { path: ["doctor"], usage: "" }]; + const md = [ + "`pnpm ops <action>` and `pnpm archilyzer <command>` are placeholders.", + "pnpm ops sync --json '{}' --wait", + "pnpm ops list", + "pnpm ops get channel x and pnpm ops get bogus", + "pnpm ops frobnicate --wait", + "`pnpm archilyzer publish build <id>`, `pnpm archilyzer doctor`.", + "pnpm archilyzer publish destroy x", + ].join("\n"); + const usage = OPS_USAGE.replace("pnpm ops list", "pnpm ops list\n pnpm ops get channel <slug>"); + assert.deepEqual(unknownCommandsIn(md, cli, usage), [ + "pnpm ops get bogus", + "pnpm ops frobnicate", + "pnpm archilyzer publish destroy x", + ]); +}); + +test(`every command ${RECIPES_FILE} names exists`, async () => { + assert.deepEqual(await unknownRecipeCommands(), []); +}); + +test(`${COMMANDS_FILE} is what the command table and pnpm ops --help generate`, async () => { + assert.equal( + readFileSync(path.join(REPO, COMMANDS_FILE), "utf8"), + await generateCommandsMarkdown(), + `${COMMANDS_FILE} is stale: run \`archilyzer docs cli\``, + ); +}); diff --git a/common/bin/cli-docs.ts b/common/bin/cli-docs.ts @@ -0,0 +1,274 @@ +// WRITE COMMANDS.md FROM THE TWO COMMAND SURFACES' OWN HELP. +// +// archilyzer docs cli [--check] +// +// The `archilyzer` rows come from the command table (`COMMANDS`, archilyzer.ts): +// each row's path and its one-line usage. The `pnpm ops` rows come from +// `usage()` in scripts/archilyzer-ops.mjs — the text `pnpm ops --help` prints — +// split into paragraphs: a paragraph that opens with action names documents +// those actions; every other paragraph (the usage lines, --wait, preview, +// the env) is printed as it stands. Each action also gets the examples the +// script's header comment gives it (`// pnpm ops <action> …` lines). An +// action with neither is still listed, so the reference never omits one. +// +// `--check` writes nothing and returns 1 when the committed file differs from +// what the two sources generate (the same claim cli-docs.test.ts makes). Both +// modes then check OPERATING.md: every `pnpm ops …` / `pnpm archilyzer …` its +// recipes name must exist. The sibling of env-docs.ts (ENVIRONMENT.md) and +// file-schemas-docs.ts. + +import { readFile, writeFile } from "node:fs/promises"; +import path from "node:path"; +import { fileURLToPath, pathToFileURL } from "node:url"; +import { runIfEntryPoint } from "./_cli"; + +const REPO = path.resolve(path.dirname(fileURLToPath(import.meta.url)), "..", ".."); + +export const COMMANDS_FILE = "COMMANDS.md"; +const OPS_SCRIPT = path.join("scripts", "archilyzer-ops.mjs"); + +// What the renderer needs of a CLI row — structural, so the tests never import +// the table. +export type CliRow = { path: readonly string[]; usage: string; passthrough?: boolean }; + +/** A usage line split at its first double space: the arguments, then the text. */ +export function splitUsage(usage: string): { args: string; text: string } { + const at = usage.indexOf(" "); + if (at < 0) return { args: "", text: usage.trim() }; + return { args: usage.slice(0, at).trim(), text: usage.slice(at).trim() }; +} + +/** + * Plain help text as one Markdown table cell: backtick spans are kept as code, + * and everything else is escaped so `<id>`, `*`, `_` and `[x]` print as typed. + * A `|` is escaped everywhere, code included (GFM reads it as a cell edge). + */ +export function mdCell(text: string): string { + return text + .split(/(`[^`]*`)/) + .map((part, i) => + i % 2 === 1 + ? part.replace(/\|/g, "\\|") + : part + .replace(/\\/g, "\\\\") + .replace(/\|/g, "\\|") + .replace(/</g, "&lt;") + .replace(/>/g, "&gt;") + .replace(/([*_[\]])/g, "\\$1"), + ) + .join(""); +} + +function codeCell(text: string): string { + return text ? `\`${text.replace(/\|/g, "\\|")}\`` : ""; +} + +/** The action list `usage()` prints on its `Actions: a, b, c` line. */ +export function opsActions(opsUsage: string): string[] { + const line = opsUsage.split("\n").find((l) => l.startsWith("Actions: ")); + if (!line) throw new Error("pnpm ops usage: no `Actions:` line"); + return line + .slice("Actions: ".length) + .split(",") + .map((a) => a.trim()) + .filter(Boolean); +} + +export type OpsParagraph = { actions: string[]; text: string }; + +/** + * The usage text's paragraphs, each with the actions it opens with + * ("build-site, build-deploy and deploy-site all take …" → three). A paragraph + * that opens with no action name has `actions: []`. The `Actions:` line itself + * is left out: the table replaces it. + */ +export function opsParagraphs(opsUsage: string, actions: readonly string[]): OpsParagraph[] { + const known = new Set(actions); + const out: OpsParagraph[] = []; + for (const raw of opsUsage.split(/\n\s*\n/)) { + const text = raw.replace(/\s+$/, ""); + if (!text.trim() || text.startsWith("Actions: ")) continue; + const opened: string[] = []; + for (const word of text.trim().split(/[\s,]+/)) { + if (known.has(word)) opened.push(word); + else if (word === "and" && opened.length > 0) continue; + else break; + } + out.push({ actions: opened, text }); + } + return out; +} + +/** + * The examples in the ops script's header comment, per action: every + * `// pnpm ops <action> …` line whose action is a known one, its padding + * collapsed. (`get` and `list` lines are not actions; the usage block has them.) + */ +export function opsExamples(source: string, actions: readonly string[]): Map<string, string[]> { + const known = new Set(actions); + const out = new Map<string, string[]>(); + for (const line of source.split("\n")) { + const m = /^\/\/\s+pnpm ops ([a-z][a-z-]*)(\s+.*)?$/.exec(line); + if (!m || !known.has(m[1])) continue; + const rest = m[2] ? ` ${m[2].trim().replace(/\s{2,}/g, " ")}` : ""; + out.set(m[1], [...(out.get(m[1]) ?? []), `pnpm ops ${m[1]}${rest}`]); + } + return out; +} + +const oneLine = (s: string) => s.replace(/\s+/g, " ").trim(); + +export function renderCommandsMarkdown( + cli: readonly CliRow[], + opsUsage: string, + examples: ReadonlyMap<string, readonly string[]> = new Map(), +): string { + const actions = opsActions(opsUsage); + const paragraphs = opsParagraphs(opsUsage, actions); + const lines: string[] = [ + "# Command reference", + "", + "<!-- GENERATED by common/bin/cli-docs.ts from the `archilyzer` command table (common/bin/archilyzer.ts) and `pnpm ops --help` (scripts/archilyzer-ops.mjs) — do not edit by hand. Regenerate: `pnpm archilyzer docs cli`. -->", + "", + "Every `archilyzer` command and every `pnpm ops` action, from their own help. Recipes that chain them: [OPERATING.md](OPERATING.md).", + "", + "## `pnpm archilyzer`", + "", + "Run from the repo root (in the container: `docker compose exec editor pnpm archilyzer …`). `--help` after a command prints its line. A command marked *passthrough* parses its own flags.", + "", + "| command | arguments | what it does |", + "|---|---|---|", + ]; + for (const row of cli) { + const { args, text } = splitUsage(row.usage); + const note = row.passthrough ? " *(passthrough)*" : ""; + lines.push( + `| \`archilyzer ${row.path.join(" ")}\` | ${codeCell(args)} | ${mdCell(oneLine(text))}${note} |`, + ); + } + lines.push( + "", + "## `pnpm ops`", + "", + "Drives a running editor over HTTP (`/api/ops/*`, the same actions its pages run), gated by `WORKER_TOKEN`. A body is `--json '<object>'` or `--file <path>`; `--wait` follows a job to its end.", + "", + "| action | what it does | example |", + "|---|---|---|", + ); + // One row per help paragraph that opens with actions, in the order of their + // first action; an action no paragraph opens with gets a row of its own. An + // action's examples go on the first row that names it. + type Row = { actions: string[]; text: string; examples: string[] }; + const rows: Row[] = []; + const emitted = new Set<OpsParagraph>(); + for (const action of actions) { + const mine = paragraphs.filter((p) => p.actions.includes(action)); + if (mine.length === 0) rows.push({ actions: [action], text: "", examples: [] }); + for (const p of mine) { + if (emitted.has(p)) continue; + emitted.add(p); + rows.push({ actions: p.actions, text: oneLine(p.text), examples: [] }); + } + } + for (const action of actions) { + const row = rows.find((r) => r.actions.includes(action)); + row?.examples.push(...(examples.get(action) ?? [])); + } + for (const r of rows) { + const names = r.actions.map((a) => `\`${a}\``).join(", "); + const ex = r.examples.map(codeCell).join("<br>"); + lines.push(`| ${names} | ${r.text ? mdCell(r.text) : "—"} | ${ex} |`); + } + lines.push("", "### Usage, flags and environment", "", "```text"); + for (const p of paragraphs) { + if (p.actions.length === 0) lines.push(p.text, ""); + } + if (lines[lines.length - 1] === "") lines.pop(); + lines.push("```", ""); + return lines.join("\n"); +} + +/** + * Every `pnpm ops …` and `pnpm archilyzer …` a document names that is not a + * real action, getter or command — what keeps OPERATING.md's recipes runnable. + * A placeholder (`pnpm ops <action>`) is not a name. + */ +export function unknownCommandsIn( + markdown: string, + cli: readonly CliRow[], + opsUsage: string, +): string[] { + const actions = new Set(opsActions(opsUsage)); + const nouns = new Set([...opsUsage.matchAll(/pnpm ops get ([a-z][a-z-]*)/g)].map((m) => m[1])); + const bad = new Set<string>(); + for (const m of markdown.matchAll(/pnpm ops ([^\s`'"]+)(?:[ \t]+([^\s`'"]+))?/g)) { + const [, word, next] = m; + if (word.startsWith("<") || word === "list") continue; + if (word === "get") { + if (!next || !nouns.has(next)) bad.add(`pnpm ops get ${next ?? ""}`.trim()); + continue; + } + if (!actions.has(word)) bad.add(`pnpm ops ${word}`); + } + for (const m of markdown.matchAll(/pnpm archilyzer((?:[ \t]+[a-z][a-z-]*)+)/g)) { + const words = m[1].trim().split(/\s+/); + if (!cli.some((r) => r.path.every((w, i) => words[i] === w))) { + bad.add(`pnpm archilyzer ${words.join(" ")}`); + } + } + return [...bad]; +} + +// The recipes the reference backs. +export const RECIPES_FILE = "OPERATING.md"; + +/** `usage()` of scripts/archilyzer-ops.mjs — what `pnpm ops --help` prints. */ +export async function loadOpsUsage(repo = REPO): Promise<string> { + const mod = (await import(pathToFileURL(path.join(repo, OPS_SCRIPT)).href)) as { + usage?: () => string; + }; + if (typeof mod.usage !== "function") { + throw new Error(`${OPS_SCRIPT} exports no usage()`); + } + return mod.usage(); +} + +export async function generateCommandsMarkdown(repo = REPO): Promise<string> { + const { COMMANDS } = await import("./archilyzer"); + const opsUsage = await loadOpsUsage(repo); + const source = await readFile(path.join(repo, OPS_SCRIPT), "utf8"); + return renderCommandsMarkdown(COMMANDS, opsUsage, opsExamples(source, opsActions(opsUsage))); +} + +/** OPERATING.md's names that no command answers to (see unknownCommandsIn). */ +export async function unknownRecipeCommands(repo = REPO): Promise<string[]> { + const { COMMANDS } = await import("./archilyzer"); + const recipes = await readFile(path.join(repo, RECIPES_FILE), "utf8"); + return unknownCommandsIn(recipes, COMMANDS, await loadOpsUsage(repo)); +} + +// Writes COMMANDS.md (or, with `check`, compares it), then checks that every +// command OPERATING.md names exists. 1 when either is wrong. +export async function main(opts: { check?: boolean } = {}): Promise<number> { + const file = path.join(REPO, COMMANDS_FILE); + const want = await generateCommandsMarkdown(); + let code = 0; + if (opts.check) { + const have = await readFile(file, "utf8").catch(() => ""); + if (have !== want) { + console.error(`${COMMANDS_FILE} is stale — regenerate it with \`archilyzer docs cli\``); + code = 1; + } + } else { + await writeFile(file, want); + console.log(`wrote ${COMMANDS_FILE}`); + } + const unknown = await unknownRecipeCommands(); + if (unknown.length > 0) { + console.error(`${RECIPES_FILE} names commands that do not exist: ${unknown.join("; ")}`); + code = 1; + } + return code; +} + +runIfEntryPoint(import.meta.url, () => main({ check: process.argv.includes("--check") }));