Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 05148db6b454255b77d31e8116a804106f6b35da
parent 320b844af53698185305ad5a1218670672cc499f
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Fri,  9 Oct 2026 14:53:05 -0400

docs: COMMANDS.md regenerated for tracks A, D and E

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
MCOMMANDS.md | 58+++++++++++++++++++++++++++++++++++++++++++++++++++++-----
1 file changed, 53 insertions(+), 5 deletions(-)

diff --git a/COMMANDS.md b/COMMANDS.md @@ -59,6 +59,7 @@ Run from the repo root (in the container: `docker compose exec editor pnpm archi | `archilyzer reconcile video-dirs` | `[--channel <slug>] [--dry-run] [--verbose]` | rename video dirs to the canonical id layout *(passthrough)* | | `archilyzer verify transcripts` | `--channel <slug>` | list duplicate and missing transcripts *(passthrough)* | | `archilyzer migrate channel-priority` | `[--dry-run]` | the one-shot channel-priority migration (plans/channel-priority.md, S5) *(passthrough)* | +| `archilyzer storage report` | `[<slug>…] [--json]` | each channel's media (tierable), text and clip bytes and where its media is — off the reports, no drive touched; totals per place (the corpus disk's is what a move would free) | | `archilyzer storage migrate-tier` | `<slug>…\|--all [--order smallest] [--include-large] [--dry-run] [--reclaim]` | bring a channel off the retired whole-directory layout onto the media tier, its text home to the corpus disk (editor stopped; --all stops before the three big-text channels) *(passthrough)* | | `archilyzer brand media` | `[--out <dir>] [--video-kit]` | render the Archilyzer Media channel's assets *(passthrough)* | | `archilyzer mcp` | `[--local <dir>\|--remote <url>\|--hub <url>]` | start the MCP server on stdio (as `pnpm --filter yt-dlp-transcript-mcp exec tsx src/index.ts`) *(passthrough)* | @@ -101,19 +102,26 @@ Drives a running editor over HTTP (`/api/ops/*`, the same actions its pages run) | `build-site`, `build-deploy`, `deploy-site` | build-site, build-deploy and deploy-site all take "siteId" (one) or "siteIds" (a list). | | | `build-hub` | build-hub builds the hub into its bundle; {"deploy": true} deploys it after, and deploy-hub ships the one already built. Both deploy to the Pages project set on /sites under Hub, and take "preview" too. | | | `build-homepage` | build-homepage builds the homepage package into homepage/out; {"deploy": true} deploys it after (only if the build succeeded), and deploy-homepage ships the one already built. Both deploy to the Pages project archilyzer (https://archilyzer.pages.dev), production unless "preview" is given. | | -| `relocate` | — | `pnpm ops relocate --json '{"slugs":["x"],"locationId":"platter"}'` | +| `relocate` | relocate moves channels' media to a location: {"slugs", "locationId" \| "root"}. "dryRun": true answers each channel's preview (bytes to copy, free space both sides) and moves nothing — though, as on the Storage panel, a channel never tiered is first tiered in place on the corpus disk. | `pnpm ops relocate --json '{"slugs":["x"],"locationId":"platter"}'`<br>`pnpm ops relocate --json '{"slugs":["x"],"locationId":"platter","dryRun":true}'` | | `relocate-back` | — | | | `evict-clips` | — | | | `reports-prepare` | reports-prepare cuts every clip and copies every post capture a site's published reports cite into its report-media cache, before its build: {"siteId"}. The job fails, naming each one, when a citation lacks media. When nothing is missing it then exports the reports, as reports-export. | `pnpm ops reports-prepare --json '{"siteId":"demo-site"}' --wait` | | `reports-export` | reports-export writes each published report as report.html, report.pdf, report.md, slides.html, slides.pdf and evidence-pack.zip for the site's build to publish: {"siteId", "reportId"?, "formats"?: \["html","pdf","md", "slides","slides-pdf","zip"\]}. | `pnpm ops reports-export --json '{"siteId":"demo-site","formats":["html","md"]}' --wait` | -| `lane` | — | `pnpm ops lane --json '{"lane":"download","held":true}'` | +| `lane` | lane takes {"lane": "transcription"\|"download"\|"digest"\|"backfill"\| "publish", "enabled"?, "held"?, "action"?: "start"\|"stop"\|"drain"}. | `pnpm ops lane --json '{"lane":"download","held":true}'`<br>`pnpm ops lane --json '{"lane":"publish","action":"drain"}'` | | `tags` | — | `pnpm ops tags --json '{"op":"define","tag":{"id":"eva-collab","label":"Collab"}}'` | | `tag-videos` | — | `pnpm ops tag-videos --file ids.json` | | `keep-videos` | — | | | `persist-videos` | persist-videos saves specific videos, across channels, to the saved-video store: {"items": \[{"slug", "id"}, ...\]}. "format": "original" \| "video\_720" (default: each channel's own). "replace": "above-height" also re-fetches a saved one whose height is unknown or above that quality (default "never"). "gapMs" pauses between downloads (default the batch gap), "minFreeMemMb" waits for that much free memory before each. "dryRun": true answers with the buckets (saved, wrongHeight, toFetch, noUrl, unknown) and starts nothing. One job per channel, on its download queue; a low disk or a rate limit stops it, and running the same body again resumes — saved videos are skipped. | `pnpm ops persist-videos --file list.json --wait` | -| `fetch-windows` | fetch-windows fetches clip windows, one paced job per platform queue (YouTube and Rumble side by side): {"siteId"} fetches every window the site's published reports cite and the disk does not hold; {"items": \[{"slug", "id", "from", "to", "clipId"?, "reason"?, "pad"?, "webpageUrl"?}, ...\], "requestedBy", "manifest"?} fetches a list. "maxHeight" caps the source height (default 720). "dryRun": true lists the windows per platform, the ones already on disk ("cached") and the ones no fetch can fill ("unfetchable": deleted, off the site) and starts nothing. A platform cooling down or held is refused for its group; a 429, or two 403s in a row, backs the platform off and stops its job. Running the same body again resumes — fetched windows are cached. | `pnpm ops fetch-windows --json '{"siteId":"demo-site","dryRun":true}'`<br>`pnpm ops fetch-windows --file windows.json --wait` | +| `fetch-windows` | fetch-windows fetches clip windows, one paced job per platform queue (YouTube and Rumble side by side): {"siteId"} fetches every window the site's published reports cite and the disk does not hold; {"items": \[{"slug", "id", "from", "to", "clipId"?, "reason"?, "pad"?, "webpageUrl"?}, ...\], "requestedBy", "manifest"?} fetches a list. "maxHeight" caps the source height (default 720). "dryRun": true lists the windows per platform, the ones already on disk ("cached") and the ones no fetch can fill ("unfetchable": deleted, off the site) and starts nothing. A platform cooling down or held is refused for its group; a 429, or two 403s in a row, backs the platform off and stops its job. Running the same body again resumes — fetched windows are cached, and a window a queued or running job will already write is answered in "inFlight" with that job, whose id joins "jobIds" (so --wait follows it) and no second job is queued. Windows run on the platform's clip queue (clips:youtube), never behind its long downloads. | `pnpm ops fetch-windows --json '{"siteId":"demo-site","dryRun":true}'`<br>`pnpm ops fetch-windows --file windows.json --wait` | | `cut-release` | cut-release turns a changelog's \[Unreleased\] into "## \[&lt;version&gt;\] - &lt;date&gt;": {"workspace": "editor" \| "export" \| "all", "version": "X.Y.Z" \| "next" \| "next-minor", "commit": boolean (default false), "date": "YYYY-MM-DD" (default today)}. "all" cuts both with ONE version and commits each ("Release &lt;workspace&gt; &lt;version&gt;") — or neither: every check runs before either file is written. Only an editor built from release 10 or later has the route (an older one answers 404); with no editor running, `archilyzer release cut` does the same locally. | `pnpm ops cut-release --json '{"workspace":"all","version":"next","commit":true}'` | | `transcribe` | transcribe runs ONE local file through a local transcription worker, as a job: {"path": "/abs/file"} (audio or video), "start"/"end" (seconds) for a window, "workerId" (a settings worker id; default: the one auto-transcribe would get), "out" (an absolute path for the result JSON, never inside the corpus). The result is {path, window, worker: {id, appId, model, device}, cues: \[{start, end, text}\], text, ...}, cue times on the file's own clock. "words": true adds words: \[{w, start, end, conf?}\] on the same clock -- from parakeet, which keeps its word timestamps; \[\] from an engine that does not. With --wait it is printed on stdout (the response and the log go to stderr), so `pnpm ops transcribe ... --wait \| jq -r .text` works. | `pnpm ops transcribe --json '{"path":"/abs/clip.mp4","start":120,"end":150}' --wait`<br>`pnpm ops transcribe --json '{"path":"/abs/a.wav","workerId":"parakeet-cpu","out":"/tmp/a.json"}' --wait` | +| `settings` | settings patches settings.json through the editor's one writer: {"patch": {"&lt;top-level key&gt;": value, …}} (SETTINGS.md). A block's keys merge one level; an array, a scalar or an object nested in a block replaces. Refused before anything is written: an unknown key; a key another writer owns (channelPriority, autoQueue, workers, storage — the refusal names it); a value the schema would not keep as sent (clamped, dropped), named with what would be saved. Answers the saved values. | `pnpm ops settings --json '{"patch":{"minFreeDiskGB":20}}'` | +| `clear-platform-hold` | clear-platform-hold clears a platform's rate-limit hold, backoff and raised pace, as the lane strip's Clear hold: {"platform"}. Answers the sentence. | | +| `workers` | workers switches transcription workers as /workers does: {"op": "enable" \| "disable" \| "drain", "ids": \[...\]} — the live pool; a restart restores the launch default. An unknown id is named; the others still apply. | `pnpm ops workers --json '{"op":"disable","ids":["parakeet-cpu"]}'` | +| `transcribe-one` | transcribe-one transcribes one video: {"slug", "id", "file"?}. With "file" (an audio file in its dir) that file; without, the video page's Transcribe — the audio on disk, or downloaded first through the channel's managed path. A job; --wait follows it. | | +| `delete-file` | delete-file deletes one file of one video, as its Files list does: {"slug", "id", "file"}. A tiered media file goes with its bytes on the media drive (never a bare link); refused while that drive is not answering. No undo. | | +| `do-not-clean` | do-not-clean sets one video's "Do not clean" marker: {"slug", "id", "keep"?: true}; "keep": false removes it. (keep-videos is the bulk form.) | | +| `cleanup` | cleanup runs one of a channel's cleanup sweeps as a job: {"slug", "sweep": "transcribed" (audio of transcribed videos — the /cleanup total) \| "extra-formats" \| "wrong-format" \| "failed-transcriptions" (clears the list; deletes nothing)}. Do-not-clean videos are skipped; `get cleanup <slug>` says what each would reclaim. | `pnpm ops cleanup --json '{"slug":"x","sweep":"transcribed"}' --wait` | ### Usage, flags and environment @@ -124,14 +132,51 @@ Usage: pnpm ops <action> [--json '<body>' | --file <path>] [--wait] pnpm ops get channels pnpm ops get tags [<tagId>] pnpm ops get publish + pnpm ops get job <id> [--tail [<lines>]] + pnpm ops get jobs [--active | --failed] [--kind <kind>] [--slug <slug>] + [--limit <n>] + pnpm ops job cancel|drain|promote|force-release|retry <id>... [--wait] + pnpm ops job retry-failed [--wait] + pnpm ops job wait <id>... [--wait-timeout <seconds>] + pnpm ops get settings [<key>] + pnpm ops get storage | sites | workers | auto-queue | scheduler + pnpm ops get cleanup <slug> pnpm ops list --wait follows the job's log and survives a poll that fails (a busy in-process build starves the server): it backs off and, after three - failures, asks /api/jobs/active whether the job is still there. + failures, asks /api/ops/job/<id> whether the job is still there and how + it ended. While the job waits it prints its queue position on stderr. --wait-timeout <seconds> gives up and exits 1 instead of waiting forever. Default: no timeout — the queue may legitimately hold a job for hours. +get job <id> is one job: kind, channel, status, times, exit code, and while + it waits its queue (key, position, how many wait, the job at the head); + --tail [<lines>] adds the last lines of its log (40, at most 500). +get jobs lists jobs newest first (50; --limit, at most 500): --active is + every queued and running job in queue order, --failed the failed ones; + --kind and --slug narrow either. A filtered list looks through the + newest 2000 jobs. +job <verb> <id>... is a /jobs row's button: cancel (a running job is + stopped), drain (finish the sub-operations in flight, start no more), + promote (a queued job to the front of its queue), force-release (free a + wedged queue slot), retry (re-run from its replay descriptor, ahead of + the queue). retry-failed retries every failed job the editor still holds. + Each id is answered in "results"; one that could not be acted on makes + the answer ok: false without stopping the rest. --wait follows the jobs a + retry started. job wait <id>... follows jobs already running. + +get settings [<key>] is settings.json as the editor reads it (migrated and + defaulted; a secret is "<redacted>"), or one top-level key of it. +get storage is /storage: each location, mounted or not, free space, tiers. +get sites lists every site: id, title, siteUrl, Pages project, audience, + listed, search, publish policy, channels. (`get publish` is the status.) +get workers, get auto-queue and get scheduler are /workers, the four lanes + of /operations, and /operations/sync — the payloads those pages poll. +get cleanup <slug> is one channel's /cleanup row: what each sweep would + reclaim (they overlap — never add them), what holds the rest, and the + failed-transcriptions count. measured: false means unknown, not zero. + "preview": "<branch>" on deploy-site or build-deploy makes it a Cloudflare Pages PREVIEW instead of production: the same bundle goes to a branch alias, https://<branch>.<project>.pages.dev, and the live site is left @@ -139,5 +184,8 @@ Usage: pnpm ops <action> [--json '<body>' | --file <path>] [--wait] digits and dashes, up to 28 characters; "main" is refused. Env: ARCHILYZER_EDITOR_URL (default http://localhost:3001), WORKER_TOKEN, - ARCHILYZER_AGENT (provenance of a tag write; default "cli") + ARCHILYZER_AGENT (provenance of a tag write; default "cli"). + WORKER_TOKEN and ARCHILYZER_EDITOR_URL, when unset, are read from + editor/.env.local and editor/.env of this checkout (found from the + script, not the cwd), then of the main worktree. ```