Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 6bd5722378cd6cf08623a8962f7a9582865c9269
parent 0baa0fa1e03f3442c4e848ae3e20c2b63f84ea09
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sat, 10 Oct 2026 02:24:03 -0400

ops: a help paragraph for each of the twelve actions that had none

channel-priority, sync, download-missing, metadata-scan, import-video,
feed-metadata, refresh-report, relocate-back, evict-clips, tags, tag-videos
and keep-videos now open a paragraph in usage(), so COMMANDS.md's rows carry
their bodies (release 19 C2's found-and-left; relocate and lane gained theirs
in A4). COMMANDS.md regenerated.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
MCOMMANDS.md | 24++++++++++++------------
Mscripts/archilyzer-ops.mjs | 66++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
2 files changed, 78 insertions(+), 12 deletions(-)

diff --git a/COMMANDS.md b/COMMANDS.md @@ -78,22 +78,22 @@ Drives a running editor over HTTP (`/api/ops/*`, the same actions its pages run) | action | what it does | example | |---|---|---| -| `channel-priority` | — | `pnpm ops channel-priority --json '{"slugs":["x"],"operation":"download","tier":"paused"}'` | +| `channel-priority` | channel-priority sets channels' priority, as the /channels deck's tier control does: {"slugs", "tier": "normal" \| "low" \| "paused"} sets the base tier; with "operation" (sync, transcription, download, digest, backfill) it pins that operation's tier, "tier": null clearing it back to the base; "preset": "sync-only" \| "clear" is the deck's two shortcuts and takes no tier. A manual change clears an automatic pause. | `pnpm ops channel-priority --json '{"slugs":["x"],"operation":"download","tier":"paused"}'` | | `channel-config` | channel-config changes a channel as its Configure form does: {"slug"} and any of "patch" (form field names; "" clears one), "sites" (the WHOLE membership set: \[{"siteId", "groupId"? \| "newGroupName"?}\], \[\] = on no site; an unknown site id is refused), "excludeFromBuild" and "excludeFromCleanup" (set to the value given, not toggled). | `pnpm ops channel-config --json '{"slug":"x","patch":{"downloadFilterExclude":"rerun"}}'`<br>`pnpm ops channel-config --json '{"slug":"x","sites":[{"siteId":"anilyzer"}]}'`<br>`pnpm ops channel-config --json '{"slug":"x","sites":[],"excludeFromBuild":true}'` | | `create-channel` | create-channel is the New channel form: {"fields": {"name", "handling": "youtube"\|"transcribe", "url"?, "platform"?, "sourceKind"?, "postFetcher"?, "socialHandle"?, …}} with channel-config's patch keys; "slug"? (else derived from the name), "sites"? (absent = on no site). "fetchPlaylist", "fetchPostsNow" and "prioritizeDownload" are the form's checkboxes, OFF unless true; a job they start comes back as jobId(s), so --wait follows it. | `pnpm ops create-channel --json '{"fields":{"name":"Example (X)","handling":"transcribe","url":"https://x.com/example"}}'` | | `rename-channel` | rename-channel moves a channel to a new slug, as Danger → Rename does: {"slug", "newSlug"}. Refused while the channel is busy (a job, a lane unit, media in transition) or when the new slug is taken. Old links break. | `pnpm ops rename-channel --json '{"slug":"old-slug","newSlug":"new-slug"}'` | | `delete-channel` | delete-channel removes a channel's whole directory, as Danger → Delete does: {"slug", "confirm"} — "confirm" must repeat the slug. No undo outside the transcripts/ repo's own history. | `pnpm ops delete-channel --json '{"slug":"x","confirm":"x"}'` | -| `metadata-scan` | — | `pnpm ops metadata-scan --json '{"slug":"the-quartering"}'` | +| `metadata-scan` | metadata-scan reads every listed video's title, date and description into the channel's scan file, fetching no media: {"slug"}. Not held by the download pause and needs no disk floor; one request per video, so it respects the platform's cooldown and records one when pushed back. | `pnpm ops metadata-scan --json '{"slug":"the-quartering"}'` | | `refresh-metadata` | refresh-metadata re-reads ONE video's metadata.info.json from its source (no subtitles, no media) on the platform's queue: {"slug", "id"}. The job's log ends with what the source now says — live\_status, formats, audio-only formats and whether any is non-fragmented, English captions, the keys that changed. An id with no data/&lt;id&gt;/ is refused (a refresh re-reads a video already archived), as are archive.org and Wayback records. | `pnpm ops refresh-metadata --json '{"slug":"the-quartering","id":"<videoId>"}' --wait` | -| `import-video` | — | `pnpm ops import-video --json '{"slug":"demo-archive","url":"https://archive.org/details/example-item"}'` | +| `import-video` | import-video imports one video by URL into a channel: {"slug", "url"}. An archive.org URL becomes its canonical item or file (an item of several media files is refused — see import-archive-org); a BitChute or Odysee URL runs on that platform's queue at its pace, refused while it is held or cooling down, and refused when already on disk; a Wayback capture is named by what it copies. | `pnpm ops import-video --json '{"slug":"demo-archive","url":"https://archive.org/details/example-item"}'` | | `import-archive-org` | import-archive-org imports archive.org media into a channel, as one job on archive.org's queue: {"slug", "item", "files": \[...\] \| "match": "&lt;regex&gt;"} for one item; {"slug", "items": \["&lt;id&gt;" \| {"item", "files"? \| "match"?}, ...\]} for many (a bare id takes "match", else every media original); or {"slug", "query": "&lt;archive.org search&gt;", "limit"?: 100} for the first items a search finds (at most 500). One file at a time, a jittered pause between files and between items; a record already held (on disk, or in the saved-video store) is skipped. "dryRun": true lists each file as held, RESTRICTED (archive.org marks it not for download: a fetch would answer 401/403) or would get, and fetches nothing. A 401/403 skips the rest of its item; three such items in a row, a 429 or three failures in a row stop the job. The log ends with a summary: line; a re-run resumes. | `pnpm ops import-archive-org --json '{"slug":"demo-archive","item":"example-item","match":"\\.mp4$"}' --wait`<br>`pnpm ops import-archive-org --json '{"slug":"demo-archive","query":"collection:example-collection","dryRun":true}' --wait` | | `remote-listing` | remote-listing lists an Odysee or BitChute channel upstream and diffs it against what it holds: {"slug"}. One flat-playlist read on the platform's queue (refused while it is held or cooling down; a 429 backs it off), nothing written. `get remote-listing <slug>` waits for the job and prints {listed, held, notHeld: \[{id, url}\], heldNotListed: \[id\], ...} on stdout. | | | `attach-media` | attach-media copies each held video's file out of a LOCAL archive into the saved-video store, as its source container (nothing fetched): {"slug", "source"} — an absolute path to a directory, a .zip (read in place) or a .7z. The id is the folder's trailing "(&lt;id&gt;)", else the file's yt-dlp suffix; "items": \[{"id", "path"}\] names exact files (path inside the source), "match" narrows the folders by regex. "createRecords": true writes a record for a video the channel does not hold; "replace": true re-attaches over a saved container. "dryRun": true logs each folder's class (attach, not-held, already-attached, lost, no-media, unmatched, ambiguous) and the held videos with no media in it, and writes nothing. The log ends with a summary: line; a re-run resumes. | | | `prepare-playable` | prepare-playable remuxes each of a channel's saved containers, losslessly (-c copy), into a browser-playable mp4 (+faststart) or webm, and makes one single-file torrent per copy, under playable/ beside the saved-video store: {"slug"}. "ids" narrows it; "trackers": \[...\] is each .torrent's announce list (none by default; the infohash does not depend on it); "root" names another playable root. A video prepared from the same source (by sha256) is skipped; "dryRun": true logs each decision. | | -| `feed-metadata` | — | `pnpm ops feed-metadata --json '{"slug":"demo-podcast","dryRun":true}' --wait` | -| `refresh-report` | — | `pnpm ops refresh-report --json '{"all":true}'` | -| `sync` | — | `pnpm ops sync --json '{"slug":"the-quartering"}' --wait` | -| `download-missing` | — | | +| `feed-metadata` | feed-metadata completes a podcast channel's records from its RSS feed: one fetch of the channel's url, then title, date, description and duration into each record that lacks them: {"slug"}. "dryRun": true counts matched, unmatched and already complete, and writes nothing. | `pnpm ops feed-metadata --json '{"slug":"demo-podcast","dryRun":true}' --wait` | +| `refresh-report` | refresh-report regenerates a channel's report (the buckets and counts its page shows) on the one serial report queue: {"slug"} answers the job queued, or the one already waiting ("started": false); {"all": true} queues every channel and answers {queued, jobIds, skipped}. | `pnpm ops refresh-report --json '{"all":true}'` | +| `sync` | sync lists a channel and downloads what is new, as its Sync button does: {"slug"}. "full": true runs the periodic whole-listing sweep now (else the configured cadence decides); "queueKey" picks another queue. On the channel's platform queue, paced like its downloads. | `pnpm ops sync --json '{"slug":"the-quartering"}' --wait` | +| `download-missing` | download-missing downloads every listed video the channel does not hold: {"slug"}. "ignoreArchive": true drops the download archive so listed videos it names are fetched again (recovery from a stale archive); "abortOnError": false carries on past a failure (default: stop at the first that is not one video's own). Paced on the platform's queue. | | | `retry-bucket` | retry-bucket runs one bucket of a channel's report as one job, past any lane hold: {"slug", "bucket"}. "ids": \[...\] runs only those videos, and every one must be in the bucket (a stray id is refused, named); a job run with ids is not replayable, as a checkbox selection in the UI is not. | | | `transcribe-bucket` | transcribe-bucket transcribes a channel's "downloaded, not transcribed" bucket on the transcription queue, as the channel page's Transcribe button does: {"slug"}. "ids": \[...\] narrows it the same way as on retry-bucket. | | | `fetch-posts` | fetch-posts fetches a social channel's new posts: {"slug"}. "full": true re-walks the whole timeline; "older": true walks back from the oldest archived post through search (X; needs a login), saving its place for the next run, down to "floor": "YYYY-MM-DD" when given; "from": "YYYY-MM-DD" starts the walk afresh there, replacing its saved place (and a "complete") — for a gap above one surviving old post. "limit": N caps the posts one run reads; "pages": N caps the pages (a forum thread: its latest N pages). "full" and "older" together are refused. An older walk over an account that shows no posts (nothing archived, and the last timeline fetch read none) is refused unless "force": true. A drained fetch stops at its next resume point and the next run resumes. | | @@ -104,14 +104,14 @@ Drives a running editor over HTTP (`/api/ops/*`, the same actions its pages run) | `build-hub` | build-hub builds the hub into its bundle; {"deploy": true} deploys it after, and deploy-hub ships the one already built. Both deploy to the Pages project set on /sites under Hub, and take "preview" too. | | | `build-homepage` | build-homepage builds the homepage package into homepage/out; {"deploy": true} deploys it after (only if the build succeeded), and deploy-homepage ships the one already built. Both deploy to the Pages project archilyzer (https://archilyzer.pages.dev), production unless "preview" is given. | | | `relocate` | relocate moves channels' media to a location: {"slugs", "locationId" \| "root"}. "dryRun": true answers each channel's preview (bytes to copy, free space both sides) and moves nothing — though, as on the Storage panel, a channel never tiered is first tiered in place on the corpus disk. | `pnpm ops relocate --json '{"slugs":["x"],"locationId":"platter"}'`<br>`pnpm ops relocate --json '{"slugs":["x"],"locationId":"platter","dryRun":true}'` | -| `relocate-back` | — | | -| `evict-clips` | — | | +| `relocate-back` | relocate-back moves channels' media back onto the corpus disk, one job per channel on the relocation queue: {"slugs"}. A busy channel is answered in "skipped"; --wait follows the jobs started. | | +| `evict-clips` | evict-clips deletes cached clip windows (data/&lt;id&gt;/clips/) older than a number of days, as /storage's button does: {"olderThanDays", "slug"?, "dryRun"?}. BY AGE: nothing knows whether a umtool report still cites a window; an evicted window is fetched again when asked for. | | | `reports-prepare` | reports-prepare cuts every clip and copies every post capture a site's published reports cite into its report-media cache, before its build: {"siteId"}. The job fails, naming each one, when a citation lacks media. When nothing is missing it then exports the reports, as reports-export. | `pnpm ops reports-prepare --json '{"siteId":"demo-site"}' --wait` | | `reports-export` | reports-export writes each published report as report.html, report.pdf, report.md, slides.html, slides.pdf and evidence-pack.zip for the site's build to publish: {"siteId", "reportId"?, "formats"?: \["html","pdf","md", "slides","slides-pdf","zip"\]}. | `pnpm ops reports-export --json '{"siteId":"demo-site","formats":["html","md"]}' --wait` | | `lane` | lane takes {"lane": "transcription"\|"download"\|"digest"\|"backfill"\| "publish", "enabled"?, "held"?, "action"?: "start"\|"stop"\|"drain"}. | `pnpm ops lane --json '{"lane":"download","held":true}'`<br>`pnpm ops lane --json '{"lane":"publish","action":"drain"}'` | -| `tags` | — | `pnpm ops tags --json '{"op":"define","tag":{"id":"eva-collab","label":"Collab"}}'` | -| `tag-videos` | — | `pnpm ops tag-videos --file ids.json` | -| `keep-videos` | — | | +| `tags` | tags edits the curated-tag vocabulary through its one writer: {"op": "define", "tag": {...}} (the WHOLE definition: id, label, rules and "sites", the sites it exists on, absent = every site) or {"op": "remove", "tag": "&lt;id&gt;"}. `get tags [<id>]` reads it. | `pnpm ops tags --json '{"op":"define","tag":{"id":"eva-collab","label":"Collab"}}'` | +| `tag-videos` | tag-videos pins, unpins, suppresses or unsuppresses one tag on many videos in ONE write: {"tag", "op": "add" \| "remove" \| "suppress" \| "unsuppress", "videos": \[{"slug", "id"}, ...\]}. "remove" unpins only — a rule's hit survives; "suppress" rejects it. The write records who asked (ARCHILYZER\_AGENT); a big list goes in --file. | `pnpm ops tag-videos --file ids.json` | +| `keep-videos` | keep-videos sets the "Do not clean" marker on every held video of a channel whose title or description matches a download-filter pattern: {"slug", "match", "fields"?: \["title" \| "description"\], "note"?, "dryRun"?}. A match not held is reported in "notDownloaded", never created; a fresh channel wants metadata-scan first. | | | `persist-videos` | persist-videos saves specific videos, across channels, to the saved-video store: {"items": \[{"slug", "id"}, ...\]}. "format": "original" \| "video\_720" (default: each channel's own). "replace": "above-height" also re-fetches a saved one whose height is unknown or above that quality (default "never"). "gapMs" pauses between downloads (default the batch gap), "minFreeMemMb" waits for that much free memory before each. "dryRun": true answers with the buckets (saved, wrongHeight, toFetch, noUrl, unknown) and starts nothing. One job per channel, on its download queue; a low disk or a rate limit stops it, and running the same body again resumes — saved videos are skipped. | `pnpm ops persist-videos --file list.json --wait` | | `fetch-windows` | fetch-windows fetches clip windows, one paced job per platform queue (YouTube and Rumble side by side): {"siteId"} fetches every window the site's published reports cite and the disk does not hold; {"items": \[{"slug", "id", "from", "to", "clipId"?, "reason"?, "pad"?, "webpageUrl"?}, ...\], "requestedBy", "manifest"?} fetches a list. "maxHeight" caps the source height (default 720). "dryRun": true lists the windows per platform, the ones already on disk ("cached") and the ones no fetch can fill ("unfetchable": deleted, off the site) and starts nothing. A platform cooling down or held is refused for its group; a 429, or two 403s in a row, backs the platform off and stops its job. Running the same body again resumes — fetched windows are cached, and a window a queued or running job will already write is answered in "inFlight" with that job, whose id joins "jobIds" (so --wait follows it) and no second job is queued. Windows run on the platform's clip queue (clips:youtube), never behind its long downloads. | `pnpm ops fetch-windows --json '{"siteId":"demo-site","dryRun":true}'`<br>`pnpm ops fetch-windows --file windows.json --wait` | | `cut-release` | cut-release turns a changelog's \[Unreleased\] into "## \[&lt;version&gt;\] - &lt;date&gt;": {"workspace": "editor" \| "export" \| "all", "version": "X.Y.Z" \| "next" \| "next-minor", "commit": boolean (default false), "date": "YYYY-MM-DD" (default today)}. "all" cuts both with ONE version and commits each ("Release &lt;workspace&gt; &lt;version&gt;") — or neither: every check runs before either file is written. Only an editor built from release 10 or later has the route (an older one answers 404); with no editor running, `archilyzer release cut` does the same locally. | `pnpm ops cut-release --json '{"workspace":"all","version":"next","commit":true}'` | diff --git a/scripts/archilyzer-ops.mjs b/scripts/archilyzer-ops.mjs @@ -751,6 +751,72 @@ export function usage() { ' free space both sides) and moves nothing — though, as on the Storage', ' panel, a channel never tiered is first tiered in place on the corpus disk.', "", + 'channel-priority sets channels\' priority, as the /channels deck\'s tier', + ' control does: {"slugs", "tier": "normal" | "low" | "paused"} sets the', + ' base tier; with "operation" (sync, transcription, download, digest,', + ' backfill) it pins that operation\'s tier, "tier": null clearing it back', + ' to the base; "preset": "sync-only" | "clear" is the deck\'s two shortcuts', + ' and takes no tier. A manual change clears an automatic pause.', + "", + 'sync lists a channel and downloads what is new, as its Sync button does:', + ' {"slug"}. "full": true runs the periodic whole-listing sweep now (else', + ' the configured cadence decides); "queueKey" picks another queue. On the', + " channel's platform queue, paced like its downloads.", + "", + 'download-missing downloads every listed video the channel does not hold:', + ' {"slug"}. "ignoreArchive": true drops the download archive so listed', + ' videos it names are fetched again (recovery from a stale archive);', + ' "abortOnError": false carries on past a failure (default: stop at the', + ' first that is not one video\'s own). Paced on the platform\'s queue.', + "", + 'metadata-scan reads every listed video\'s title, date and description', + ' into the channel\'s scan file, fetching no media: {"slug"}. Not held by', + " the download pause and needs no disk floor; one request per video, so", + " it respects the platform's cooldown and records one when pushed back.", + "", + 'import-video imports one video by URL into a channel: {"slug", "url"}. An', + " archive.org URL becomes its canonical item or file (an item of several", + " media files is refused — see import-archive-org); a BitChute or Odysee", + " URL runs on that platform's queue at its pace, refused while it is held", + " or cooling down, and refused when already on disk; a Wayback capture is", + " named by what it copies.", + "", + 'feed-metadata completes a podcast channel\'s records from its RSS feed: one', + ' fetch of the channel\'s url, then title, date, description and duration', + ' into each record that lacks them: {"slug"}. "dryRun": true counts', + " matched, unmatched and already complete, and writes nothing.", + "", + 'refresh-report regenerates a channel\'s report (the buckets and counts its', + ' page shows) on the one serial report queue: {"slug"} answers the job', + ' queued, or the one already waiting ("started": false); {"all": true}', + " queues every channel and answers {queued, jobIds, skipped}.", + "", + 'relocate-back moves channels\' media back onto the corpus disk, one job', + ' per channel on the relocation queue: {"slugs"}. A busy channel is', + ' answered in "skipped"; --wait follows the jobs started.', + "", + 'evict-clips deletes cached clip windows (data/<id>/clips/) older than a', + ' number of days, as /storage\'s button does: {"olderThanDays", "slug"?,', + ' "dryRun"?}. BY AGE: nothing knows whether a umtool report still cites a', + " window; an evicted window is fetched again when asked for.", + "", + 'tags edits the curated-tag vocabulary through its one writer: {"op":', + ' "define", "tag": {...}} (the WHOLE definition: id, label, rules and', + ' "sites", the sites it exists on, absent = every site) or {"op":', + ' "remove", "tag": "<id>"}. `get tags [<id>]` reads it.', + "", + 'tag-videos pins, unpins, suppresses or unsuppresses one tag on many videos', + ' in ONE write: {"tag", "op": "add" | "remove" | "suppress" | "unsuppress",', + ' "videos": [{"slug", "id"}, ...]}. "remove" unpins only — a rule\'s hit', + ' survives; "suppress" rejects it. The write records who asked', + ' (ARCHILYZER_AGENT); a big list goes in --file.', + "", + 'keep-videos sets the "Do not clean" marker on every held video of a', + ' channel whose title or description matches a download-filter pattern:', + ' {"slug", "match", "fields"?: ["title" | "description"], "note"?,', + ' "dryRun"?}. A match not held is reported in "notDownloaded", never', + " created; a fresh channel wants metadata-scan first.", + "", 'channel-config changes a channel as its Configure form does: {"slug"} and', ' any of "patch" (form field names; "" clears one), "sites" (the WHOLE', ' membership set: [{"siteId", "groupId"? | "newGroupName"?}], [] = on no',