Archilyzer
← All docs

Running an archive

Adding channels, downloading, transcribing, and keeping an archive current without babysitting it.

Day-to-day operation: adding channels, keeping them current, and getting from a pile of audio to a published site.

Adding a channel

Everything starts on the editor's Channels page. A channel needs a URL and a decision about how it is handled:

  • Captions — for platforms that already publish machine captions, those are fetched directly and no media is downloaded. This is fast and costs almost nothing.
  • Transcribe — the audio is downloaded and transcribed on your machine. This is the slow, expensive path, and the one that produces a transcript for material that never had captions at all.

You can also set per-channel download arguments, an audio format, and whether to pass cookies from a browser profile for material that needs an account.

The three fetch modes

The pipeline offers three operations per channel, and the distinction matters once a catalogue gets large:

  • Store playlist — ask the platform for the channel's full list of videos and write it down. Nothing is downloaded.
  • Download from playlist — compare that list against what has already been fetched, and download only the difference. Filtering on our side rather than the downloader's avoids re-fetching metadata for thousands of entries you already have, which on some platforms is the difference between minutes and hours.
  • Sync — a quick incremental pass for new uploads. This is what runs on a schedule.

Keeping it current without watching it

Each channel can carry its own sync cadence — hourly, daily, every N minutes. A heartbeat runs frequently and the server decides which channels are actually due, so nothing is scheduled twice and a missed tick simply runs at the next one. There are no catch-up storms after the machine has been asleep.

There are two ways to drive the heartbeat, and they are interchangeable:

  • Internal timer — the editor arms it at startup. Set a cadence in settings and there is nothing else to install. This is the recommended setup.
  • External cron — leave the internal timer off and have the system's cron hit the tick endpoint instead. Useful when schedules are managed centrally.

Scheduled syncs run through the same queue as a manual click, so they cannot collide with one, and they appear live in the editor exactly like manual work.

Transcribing

Transcription is the bottleneck, and it is bounded by hardware rather than patience. A few things worth knowing:

  • Jobs run in parallel up to a configured limit. More is not always faster — when the machine runs out of memory, a GPU backend silently falls back to the CPU and everything gets slower at once.
  • Failures are retryable per channel, and there is a verification pass that re-checks transcripts against their audio and flags the ones that came out truncated.
  • Transcript coverage — how much of a recording's runtime the transcript actually spans — is tracked per video, because a transcript that covers 40% of a video will otherwise look exactly like a complete one.

Building and publishing

A build has two halves. First the index: transcripts are read out of the corpus and written as paginated JSON. Then the compose-and-build step turns that into a static site for each configured site.

Builds are incremental. Channels whose contents haven't changed are skipped, so a rebuild after one new video does not re-process the whole archive.

Publishing is a separate action from building, and both can be triggered from the editor. See Deploy to Cloudflare for hosting, and Building several sites at once if you run more than one.

Sites, groups, and one corpus

Channels live in a single shared pool. A site is a selection of them with its own title, description, accent colour and domain. The accent is one of seven named colours or a hex of your own, and every page of the site wears it; a reader picks a light or dark ground (or the system's) with the header's theme toggle. One corpus can therefore publish several public archives without any data being duplicated — and a channel can appear on more than one.

Within a site, channels can be arranged into named groups, which is what drives the navigation on the published pages.

When a video disappears

The archive periodically re-checks whether recordings still exist where they came from, and records the answer per video: available, unlisted, private, members-only, or deleted. That state is visible in the published site and can be filtered on.

One caveat that matters, and that the software is careful about: a recording counts as available until something re-checks it and finds otherwise. Nobody re-checks tens of thousands of videos continuously. "Available" means not known to be gone, never "confirmed still there". Only the recordings marked as gone are evidence of anything.

Audio for recordings that have disappeared upstream is protected from cleanup, on the reasoning that a local copy of something no longer available anywhere is the one thing you cannot re-fetch.

From a shell or an agent

Everything above can be run without a browser, three ways:

  • pnpm ops <action> sends the editor the same action a click does — add a channel, sync, download, import, transcribe, tag, fetch posts, prepare reports, publish — over HTTP. It needs the running editor's address (ARCHILYZER_EDITOR_URL) and its WORKER_TOKEN. A job-starting action returns once the job is queued; --wait follows it to the end.
  • pnpm archilyzer <command> is the core's own command line: the index, publish stages (publish now, publish build <id>, publish deploy <id>), reports, offline refreshes, and doctor, which says what the machine can do.
  • The MCP server lets an AI assistant read a published archive — search, transcripts, reports — and ask the editor for the media behind a cited moment. See Use with AI.
pnpm ops create-channel --json '{"fields":{"name":"Example","handling":"youtube","url":"https://www.youtube.com/@example"}}'
pnpm ops sync --json '{"slug":"example"}' --wait
pnpm ops publish --json '{"verb":"now"}' --wait

Every fetch still goes through the editor's paced, per-platform queues; nothing here downloads around them. Recipes for each task are in OPERATING.md, and every command and action in COMMANDS.md.

Next: Deploy to Cloudflare.