Archilyzer
← All docs

What Archilyzer is

The shape of the thing: three programs, one corpus, and a static site at the end of it.

Archilyzer takes a channel's back catalogue, downloads it, transcribes it on your own hardware, and builds a static website you host yourself. The result is a permanent, searchable record of what someone said on video — searchable to the second, and still there after the original comes down.

It is a program you run, not a service you sign up for. There is no account, no API key, no server of ours in the path, and nothing phones home.

The shape of it

Three programs share one library and one pile of data.

  • The editor is a local admin app. You add channels, watch the download and transcription queues, and press the button that builds and deploys a site. It runs on your machine and is never exposed to the public.
  • The corpus is a directory on disk — one folder per channel, holding the downloaded media, the transcripts, and a search index. It is deliberately kept outside the code: a fresh copy of the software has no data, and updating the software never touches your archive.
  • The export is the public artefact: a static site, pre-rendered to plain HTML and JSON. No database, no runtime, no server-side code. It can be hosted on anything that serves files.

A fourth piece, the project site you are reading, and a fifth, the MCP server described in AI and MCP, round out the workspace.

How a video becomes a searchable page

  1. List. Ask the platform what a channel has published. Archilyzer keeps this list separately from what it has already fetched, so it always knows the difference between "new" and "already have it".
  2. Fetch. Download what is missing. For channels where the platform already publishes captions, those are taken directly. For everything else, only the audio is fetched.
  3. Transcribe. Audio is transcribed locally, by whisper.cpp or one of the other supported backends. This is the slow step and the one that wants a GPU; nothing is sent to a third-party transcription service.
  4. Index. Transcripts are cut into pages and written as a paginated JSON index, so a browser can search a corpus of tens of thousands of videos without downloading it.
  5. Compose and publish. A static site is built from that index and deployed.

Steps 1–3 can run unattended on a schedule. See Running an archive.

What you get at the end

  • Search across every transcript, with results that jump to the exact second of the recording they came from.
  • A player that follows the transcript, and a transcript that follows the player.
  • Multiple sites from one corpus. Channels are grouped into sites, so one archive can publish several public faces without duplicating the data.
  • Bulk downloads of the transcripts, for anyone who wants the raw material.
  • A machine-readable index at /corpus.json, so an LLM or a script can navigate the whole archive without scraping it.

What it is not

  • Not a hosted service. You install it, you run it, you pay for your own hosting.
  • Not a forge. The source is a read-only git mirror on this site — clone and pull, but nothing takes a push or a pull request.
  • Not a downloader you point at one video. It is built around back catalogues: thousands of recordings, kept current.
  • Not automatic transcription in the cloud. The transcribing happens on your machine, at your machine's speed.

Next: Install.