# Running an archive Day-to-day operation: adding channels, keeping them current, and getting from a pile of audio to a published site. ## Adding a channel Everything starts on the editor's **Channels** page. A channel needs a URL and a decision about how it is handled: - **Captions** — for platforms that already publish machine captions, those are fetched directly and no media is downloaded. This is fast and costs almost nothing. - **Transcribe** — the audio is downloaded and transcribed on your machine. This is the slow, expensive path, and the one that produces a transcript for material that never had captions at all. You can also set per-channel download arguments, an audio format, and whether to pass cookies from a browser profile for material that needs an account. ## The three fetch modes The pipeline offers three operations per channel, and the distinction matters once a catalogue gets large: - **Store playlist** — ask the platform for the channel's full list of videos and write it down. Nothing is downloaded. - **Download from playlist** — compare that list against what has already been fetched, and download only the difference. Filtering on our side rather than the downloader's avoids re-fetching metadata for thousands of entries you already have, which on some platforms is the difference between minutes and hours. - **Sync** — a quick incremental pass for new uploads. This is what runs on a schedule. ## Keeping it current without watching it Each channel can carry its own sync cadence — hourly, daily, every N minutes. A heartbeat runs frequently and the server decides which channels are actually due, so nothing is scheduled twice and a missed tick simply runs at the next one. There are no catch-up storms after the machine has been asleep. There are two ways to drive the heartbeat, and they are interchangeable: - **Internal timer** — the editor arms it at startup. Set a cadence in settings and there is nothing else to install. This is the recommended setup. - **External cron** — leave the internal timer off and have the system's cron hit the tick endpoint instead. Useful when schedules are managed centrally. Scheduled syncs run through the same queue as a manual click, so they cannot collide with one, and they appear live in the editor exactly like manual work. ## Transcribing Transcription is the bottleneck, and it is bounded by hardware rather than patience. A few things worth knowing: - Jobs run in parallel up to a configured limit. More is not always faster — when the machine runs out of memory, a GPU backend silently falls back to the CPU and everything gets slower at once. - Failures are retryable per channel, and there is a verification pass that re-checks transcripts against their audio and flags the ones that came out truncated. - Transcript coverage — how much of a recording's runtime the transcript actually spans — is tracked per video, because a transcript that covers 40% of a video will otherwise look exactly like a complete one. ## Building and publishing A build has two halves. First the index: transcripts are read out of the corpus and written as paginated JSON. Then the compose-and-build step turns that into a static site for each configured site. Builds are incremental. Channels whose contents haven't changed are skipped, so a rebuild after one new video does not re-process the whole archive. Publishing is a separate action from building, and both can be triggered from the editor. See [Deploy to Cloudflare](/docs/deploy-cloudflare/) for hosting, and [Building several sites at once](/docs/deploy-docker/) if you run more than one. ## Sites, groups, and one corpus Channels live in a single shared pool. A **site** is a selection of them with its own title, description, accent colour and domain. The accent is one of seven named colours or a hex of your own, and every page of the site wears it; a reader picks a light or dark ground (or the system's) with the header's theme toggle. One corpus can therefore publish several public archives without any data being duplicated — and a channel can appear on more than one. Within a site, channels can be arranged into named groups, which is what drives the navigation on the published pages. ## When a video disappears The archive periodically re-checks whether recordings still exist where they came from, and records the answer per video: available, unlisted, private, members-only, or deleted. That state is visible in the published site and can be filtered on. One caveat that matters, and that the software is careful about: a recording counts as *available* until something re-checks it and finds otherwise. Nobody re-checks tens of thousands of videos continuously. "Available" means **not known to be gone**, never "confirmed still there". Only the recordings marked as gone are evidence of anything. Audio for recordings that have disappeared upstream is protected from cleanup, on the reasoning that a local copy of something no longer available anywhere is the one thing you cannot re-fetch. ## From a shell or an agent Everything above can be run without a browser, three ways: - **`pnpm ops `** sends the editor the same action a click does — add a channel, sync, download, import, transcribe, tag, fetch posts, prepare reports, publish — over HTTP. It needs the running editor's address (`ARCHILYZER_EDITOR_URL`) and its `WORKER_TOKEN`. A job-starting action returns once the job is queued; `--wait` follows it to the end. - **`pnpm archilyzer `** is the core's own command line: the index, publish stages (`publish now`, `publish build `, `publish deploy `), reports, offline refreshes, and `doctor`, which says what the machine can do. - **The MCP server** lets an AI assistant read a published archive — search, transcripts, reports — and ask the editor for the media behind a cited moment. See [Use with AI](/docs/ai-and-mcp/). ```sh pnpm ops create-channel --json '{"fields":{"name":"Example","handling":"youtube","url":"https://www.youtube.com/@example"}}' pnpm ops sync --json '{"slug":"example"}' --wait pnpm ops publish --json '{"verb":"now"}' --wait ``` Every fetch still goes through the editor's paced, per-platform queues; nothing here downloads around them. Recipes for each task are in [OPERATING.md](https://archilyzer.pages.dev/source/tree/OPERATING.md), and every command and action in [COMMANDS.md](https://archilyzer.pages.dev/source/tree/COMMANDS.md). Next: [Deploy to Cloudflare](/docs/deploy-cloudflare/).