# Archilyzer **Self-hosted, searchable video-transcript archives.** Archilyzer takes a channel's back catalogue, downloads it, transcribes it on your own hardware, and builds a static website you host yourself. The result is a permanent, searchable record of what someone said on video — searchable to the second, and still there after the original comes down. It is a program you run, not a service you sign up for. There is no account, no API key, no server of ours in the path, and nothing phones home. MIT licensed. Project site: [archilyzer.pages.dev](https://archilyzer.pages.dev) --- ## What do you want to do? | I want to… | Start here | Do I need to host anything? | |---|---|---| | **Search or study an archive that already exists** | [Use a public instance](#1-use-an-archive-that-already-exists) | No | | **Archive somebody else's public channel** | [Run your own archive](#2-run-your-own-archive) | Yes — a machine that stays on | | **Make my own channel searchable** (transcripts, metadata, a search page) | [Publish your own channel](#3-publish-your-own-channel) | Yes, plus somewhere to serve static files | | **Research a corpus with Claude Code** — cited answers and sweep reports | [Use it with Claude Code](#4-use-it-with-claude-code) | No — point it at a public instance | | **Pull clips / build a video about a subject** | [Clips and video](#clips-and-video-download-the-moment-not-the-movie) | No — search text, then fetch just the cited seconds | | **Work on Archilyzer itself** | [CONTRIBUTING.md](CONTRIBUTING.md) | — | These are not exclusive. Plenty of people research a public instance for a while before deciding to host their own. --- ## 1. Use an archive that already exists You do not need to install Archilyzer to use one. Every published archive is a static site with a real search UI, and every archive is **machine-navigable by design**: - **`/corpus.json`** — a machine-readable index describing how to navigate the shards (where each channel's manifest is, how a video id maps to a page, what a record looks like). It is served with CORS headers, so anything can fetch it. - **`/llms.txt`** — the same contract as prose, following the llmstxt.org convention. The most popular public instance is the **[Jeralyzer](https://jeralyzer.pages.dev)** — 30 channels and 30,886 videos as of its 2026-08-07 build. Point a browser at it and search; that is the whole of it. Instances cross-link to each other in their footers, so one is a way in to the rest. You can see the machine contract for yourself without installing anything: ```bash curl https://jeralyzer.pages.dev/corpus.json | head -c 400 curl https://jeralyzer.pages.dev/llms.txt ``` To let an LLM read an archive properly, run the bundled **MCP server** against it. It is a local tool you run yourself, it only reads already-published JSON, and it never modifies the archive: ```bash # point it at a published instance over HTTP (from the repo root) TRANSCRIPT_SITE_URL=https://jeralyzer.pages.dev pnpm --silent archilyzer mcp ``` There is nothing to compile — it runs from source through `tsx`. `archilyzer mcp` is the repo's CLI starting `mcp/`'s server. Keep `--silent`: stdout is the MCP protocol's channel, and without it pnpm 9 and 10 print their `> …` script banner there first (pnpm 11 prints its `$ …` to stderr). To register it with Claude Code, from the repo's root: ```bash claude mcp add archilyzer \ --env TRANSCRIPT_SITE_URL=https://jeralyzer.pages.dev \ --env ARCHILYZER_EDITOR_URL=http://localhost:3001 \ --env WORKER_TOKEN=… \ -- pnpm --silent -C "$PWD" archilyzer mcp ``` The two editor lines are optional: they let `fetch_clip` ask a local editor for clip media. Use `TRANSCRIPT_HUB_URL` instead to federate a whole hub of sites, or `TRANSCRIPT_LOCAL_DIR` to read a local build off disk. It gives a client full-text search with clickable second-level citations, complete match enumeration (so you can *count* rather than sample), batch transcript reads, and metadata including AI chapters where they exist. See **[mcp/README.md](mcp/README.md)** for the full tool list, the `mcp.json` form for other clients, and the `/sweep` and `/ask` commands. This path needs Node and pnpm. It does not need yt-dlp, ffmpeg, a GPU, or a corpus. --- ## 2. Run your own archive The full pipeline: list a channel, fetch what is missing, transcribe it locally, index it, and publish a static site. ``` list → fetch → transcribe → index → compose → publish ``` Steps 1–3 can run unattended on a schedule. Transcription is the slow step and the one that wants a GPU; nothing is sent to a third-party transcription service. ### What it costs you - **Disk.** The corpus grows to whatever your channels amount to. Point `TRANSCRIPTS_DIR` at a large disk *before* you start — moving it later means moving everything. - **Time.** A back catalogue of thousands of recordings is a multi-day first pass. After that it is incremental. - **Hosting.** The output is plain files. Anything that serves static files will do. ### Quick start ```bash corepack enable # once, so the pinned pnpm version is used pnpm install # installs every workspace package pnpm dev:editor # the editor, at http://localhost:3001 ``` The editor **starts fine with no data** — a fresh checkout has no corpus, and that is expected. Add your first channel from the Channels page. Then, to build and serve the public site: ```bash pnpm build # publish index + publish build $SITE_ID; export/out links to the bundle pnpm start:export # serve it at http://localhost:3000 ``` **[SETUP.md](SETUP.md)** has the full per-OS instructions. Below is the short version. ### archive.org items A channel can hold recordings from archive.org — a whole item (`https://archive.org/details/`, when it holds one media file) or single files of a multi-file item (`…/details//`, e.g. one video of a channel archive). They are imported, never listed, and transcribed like any transcribe channel: 1. **Create the channel**: Platform *archive.org*, handling *Transcribe*, and no URL (an archive.org channel is never auto-synced). 2. **Import one recording**: the channel's *Import video* with the item or file URL, or `pnpm ops import-video --json '{"slug":"","url":"https://archive.org/details/"}'`. An item with several media files is refused and its files are named. 3. **Import chosen files of one item**: `pnpm ops import-archive-org --json '{"slug":"","item":"","files":["",…]}'` — or `"match": ""` over the item's file names; add `"dryRun": true` to see the list first. One job, one file at a time. A whole item's video id is its identifier; a file's is `__-`. Each record keeps `archiveorg.json` beside its metadata: the item's title, date, creator and collections, the item's torrent, and — for a mirror of a YouTube upload — the original's id, URL, title and upload date, read from the `.info.json` uploaded with it. The video page shows both; a citation of the record links **archive.org** and the **torrent** (and, for a mirror, the original **YouTube** upload at the cited second), so a reader can fetch the file and check it. A file with no uploaded `.info.json` is dated by the `YYYYMMDD` its name starts with and, when the item gives it no title of its own, titled from its name — the leading date, the `[ views]` count, the YouTube id and the extension off — and `pnpm archilyzer archive-org refresh [--dry-run]` brings a channel's existing records to that title and date, offline, printing old → new. **Files come over BitTorrent when possible** — to be extra polite to archive.org. Every item has a torrent (`_archive.torrent`) that lists archive.org itself as a web seed, so a torrent client takes what other peers have from them and only the rest from archive.org. With [aria2](https://aria2.github.io/) installed (`aria2c` on PATH, or `ARIA2C_BIN`; the Docker images include it), an import fetches just the chosen file from the item's torrent, then **seeds it for 10 minutes or to a ratio of 1, whichever comes first** — the job holds archive.org's queue while it seeds. Without aria2c, when the torrent does not carry the file, or when the torrent makes no progress for 5 minutes, the file is downloaded directly from `archive.org/download/…` (resumed with a Range request, backing off on 429/503) and the log says why ("fell back to direct download: …"). Either way the file is checked against the item's sha1/md5; a mismatch is downloaded once more directly, and a second mismatch fails the record. No yt-dlp is involved: the record's metadata is written from the item's metadata API, and an audio item that yt-dlp could not read imports like any other. The defaults are settings.json's `archiveOrg` block (`torrent`, `seedMinutes`, `seedRatio`, `stallMinutes`, `maxPeers`, `maxDownloadKiBps`, `maxUploadKiBps` — [SETTINGS.md](SETTINGS.md#archiveorg)); `"torrent": false` always downloads directly. `archilyzer doctor` reports whether aria2c is there. **It is polite to archive.org**: its own job queue, one download at a time with a jittered pause of at least 8 s between files, the item's metadata and torrent asked once (cached), an identifying User-Agent, `Retry-After` and exponential backoff honoured, a stop after repeated failures, and a file on disk never fetched again (`common/lib/archiveOrgClient.ts`, `common/controller/archiveOrgImport.ts`, `common/controller/archiveOrgDownload.ts`). Wayback Machine captures (web.archive.org pages, WARC records) are a different kind of record and are not handled by this; they would be a source of their own beside `common/lib/archiveOrg.ts`. ### BitChute BitChute is a platform like YouTube, Rumble and Odysee (`platform: "bitchute"`): - **A channel**: make a channel whose URL is the BitChute channel (`https://www.bitchute.com/channel//`) and sync it — yt-dlp lists its videos. - **One video**: **Import video** on any channel with its page (`https://www.bitchute.com/video//`), or `pnpm ops import-video --json '{"slug":"","url":"https://www.bitchute.com/video//"}'`. A record plays its mp4 in the archive's own player, which seeks to a cited second (BitChute's embed does not); a citation links the BitChute page, which takes no start time. **It is polite to BitChute**, which rate-limits: its own job queue, one transfer at a time; `--sleep-requests 3` and an exponential `--retry-sleep` on every yt-dlp spawn; at least 60 s, jittered, between two videos (`sleepBetweenDownloadsSeconds` when longer); nothing asked while BitChute is in a rate-limit cooldown, which a 429 starts; and a video on disk never fetched again (`common/ytdlp/platformArgs.mjs`, `common/controller/bitchuteImport.ts`). ### Wayback Machine captures **Import video** takes a Wayback Machine capture (`https://web.archive.org/web//`, with or without a replay modifier such as `id_`) and yt-dlp downloads it: an archived YouTube page through its Wayback extractor, a raw media file as the file. The record is named by what the capture is OF: an archived YouTube page by its YouTube id, a JW Player file (`…/videos/-.mp4` on `cdn.jwplayer.com`, `content.jwplatform.com`, `videos-fms.jwpsrv.com`) by its media id. Every download of a capture writes `wayback.json` beside its metadata: the original URL (an archived YouTube page's watch URL), the capture's timestamp, the capture as a page that plays and as its raw bytes. The video page says "Archived copy (Wayback Machine, ) of ". A citation of one links the original, marked as the original and as possibly gone, and the Wayback copy; its moment link is the capture, which plays. `pnpm archilyzer wayback refresh [--titles ] [--dry-run]` brings records imported before this up to it, offline: `wayback.json`, the dir renamed to its id through the snapshot's own reconcile pass (the roster entry moves with it), and with `--titles` (a JSON file of `id → {title, upload_date}`) the title and date of a raw file that has none. A record a running job holds is skipped and named. It prints old → new; a second run changes nothing. ### Requirements Always needed, to install and run the apps: | Tool | Version | Notes | |---|---|---| | Node.js | 20.9+ | LTS 20 or 22. | | pnpm | 9+ | Easiest via `corepack enable`. | | git | any | For the source, and the corpus keeps its own repo. | Needed only for the pipeline — install what you will actually use: | Tool | Needed for | |---|---| | yt-dlp | Listing, downloading and syncing channels. | | ffmpeg + ffprobe | Audio transcode and duration checks. | | A transcription backend | Channels you transcribe yourself (whisper.cpp by default; `chough` and `parakeet.cpp` also supported). | | rsync | Backing up the saved-video store, if you enable it. | | Docker or podman | Running the whole stack in containers ([RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md)), and the parallel multi-site export build. | **The apps run with none of the second table installed.** You just cannot fetch anything yet, which makes it easy to try the UI first and commit to the toolchain later. ### Linux ```sh sudo apt install nodejs yt-dlp ffmpeg rsync build-essential # Debian / Ubuntu sudo dnf install nodejs yt-dlp ffmpeg rsync gcc-c++ make # Fedora sudo pacman -S nodejs yt-dlp ffmpeg rsync base-devel # Arch ``` If your distribution's Node is older than 20.9, use `fnm` or `nvm`. The build-tools package is the compiler pnpm falls back to when a native module has no prebuilt binary for your platform. ### macOS ```sh xcode-select --install brew install node pnpm yt-dlp ffmpeg rsync whisper-cpp corepack enable ``` Homebrew's `whisper-cpp` saves you building the default transcription backend. ### Windows **To run an archive, use Docker Desktop.** One command, no Linux userland to set up by hand, and nothing published beyond your own machine: ```powershell git clone archilyzer cd archilyzer copy .env.example .env docker compose up -d --build ``` Then open **http://localhost:8081**. The first build compiles whisper.cpp and both Next.js apps, so it takes a while; after that, starting is seconds. On first boot it seeds a settings file with a working transcription worker and downloads a whisper model, so the editor is usable the moment it comes up. **[RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md)** covers the rest: the exposure model (every port binds `127.0.0.1` by default, including the public ones), how to serve your archive to the world without exposing the admin app, GPU transcription, and backup. The same stack runs on Linux and macOS. **To develop the project — or to run Claude Code against it — use WSL2 instead.** The transcription backends and several helper scripts are Unix-oriented, so the reliable path for working *on* the code is a Linux userland: ```powershell wsl --install -d Ubuntu ``` Then open Ubuntu and follow the [Linux](#linux) steps inside it. Two things that will bite you if you skip them: - **Keep the files inside the WSL filesystem** (`~/archilyzer`), *not* under `/mnt/c/`. Cross-filesystem I/O and file watching are dramatically slower, and this pipeline is I/O-heavy. - **Point `TRANSCRIPTS_DIR` at somewhere with room.** WSL2's virtual disk grows on demand but does not shrink on its own. The editor is a normal web app on `localhost:3001`, so you drive it from your Windows browser while it runs inside WSL — WSL2 forwards localhost, so no X server and no remote desktop. The apps and the yt-dlp pipeline *do* also run natively on Windows (install Node, yt-dlp and ffmpeg with winget or Scoop), but building the transcription backends and running the shell scripts is not supported there. #### Claude Code on Windows **Install Claude Code inside WSL2 as well, not on Windows.** It should sit in the same Linux userland as the repo, so it runs the same `pnpm`, `node`, `yt-dlp` and `ffmpeg` the project expects and sees ordinary Linux paths. Everything below runs in the Ubuntu terminal, not PowerShell: ```sh # inside Ubuntu sudo apt install nodejs # or fnm/nvm if the distro Node is < 20.9 corepack enable npm install -g @anthropic-ai/claude-code git clone ~/archilyzer # keep it in ~, NOT /mnt/c cd ~/archilyzer pnpm install claude ``` Then follow [Use it with Claude Code](#4-use-it-with-claude-code). Two Windows-specific things to keep in mind: - **Paths in the MCP registration must be WSL paths** (`/home/you/archilyzer`), never `C:\...`. Using `"$PWD"` from inside the repo, as the quickstart does, gets this right by construction. - **Keep the repo out of `/mnt/c/`.** It is the single biggest performance mistake here: `pnpm install`, file watching and the corpus I/O all get dramatically slower across the Windows filesystem boundary. > **Which of the two do you want?** The container stack *runs* an archive: it is the > shortest path from a clone to a working editor, and it is what to reach for if you > want the thing rather than the toolchain. WSL2 gives you the repo itself — a dev > server with hot reload, the e2e suite, the MCP server, and Claude Code sitting in > the same filesystem as the corpus. Running both is fine; they share nothing but the > source. --- ## 3. Publish your own channel If the channel is *yours*, Archilyzer is a way to give your back catalogue the things video platforms do not: a real transcript per video, searchable to the second, plus metadata, chapters and tags — on a site you own and can point a domain at. The mechanics are the same as [running an archive](#2-run-your-own-archive); what differs is emphasis: - **You already have the media**, so the download step is a formality and the transcription quality is the thing worth tuning. Pick a backend and a model size deliberately. - **Captions you already publish are reused.** Channels whose platform publishes captions need no transcription backend at all; those are taken directly. There is also an opt-in lane that *replaces* platform auto-captions with your own transcripts where you would rather have the better text. Where YouTube serves both, the original-audio captions (`en-orig`) are read before the served `en` track, which can reword what was said; a track with no text falls through to the next, and a video's page can pin another (`common/lib/videoStatus.ts`, "the caption-track rule"). The other English tracks are kept where their words differ — uploaded captions are not always what was said: search reads them too and says which track a hit is in, and the transcript reader switches to them (`common/lib/captionTracks.ts`). - **The output is yours to brand.** Site title, header, description, tagline, social links and channel grouping are all per-site configuration. - **One corpus can publish several sites.** Channels are grouped into sites, so a network of related channels can have both individual faces and a combined one without storing anything twice. - **Bulk transcript downloads** are generated for anyone who wants the raw material, and the `/corpus.json` contract above means AI tools can cite your work accurately instead of hallucinating it. ## 4. Use it with Claude Code Research, cited reports, and video — against your corpus or somebody else's. **You do not need a corpus of your own to get value out of this repo.** Pull it down with no data at all, point it at a public archive, and what you have is a toolkit: the tools an agent needs to read someone else's corpus properly, plus the accumulated knowledge of how to do that without getting it wrong. This is a first-class path alongside running an archive, not a lesser one. Three things ship in the repo for exactly this: - **The MCP server** (`mcp/`) — search, complete match enumeration, batch reads and metadata over any published archive, with citations that resolve to an exact second. - **The discipline that goes with it** (`mcp/src/instructions.ts`) — citation format, extractor budget, base-link expansion and the coverage rule are returned *by the tools as a plan*, so an agent is told how to work the corpus rather than having to infer it. That file is the single source of truth for it. - **Two slash commands** (`.claude/commands/{ask,sweep}.md`) — tracked in git, so you get them by cloning: - **`/ask `** — answer a question from the corpus, with citations, in the conversation. - **`/sweep `** — sweep the whole corpus for something and build a cited report file. The upshot: citations land on an exact second in a real recording rather than a vague "he said this somewhere", and a sweep can tell you it covered *everything* rather than quietly sampling. ### Quickstart (no corpus required) ```bash git clone https://archilyzer.pages.dev/source/archilyzer.git archilyzer # or the tarball on archilyzer.pages.dev/downloads/ cd archilyzer && pnpm install claude mcp add archilyzer \ --env TRANSCRIPT_SITE_URL=https://jeralyzer.pages.dev \ --env ARCHILYZER_EDITOR_URL=http://localhost:3001 \ --env WORKER_TOKEN=… \ -- pnpm --silent -C "$PWD" archilyzer mcp claude # then try: /ask what has he said about magic tournaments? ``` The two editor lines are optional: they let `fetch_clip` ask a local editor for clip media (`WORKER_TOKEN` is the editor's own). Leave them out for research alone. `archilyzer mcp` is `pnpm --filter yt-dlp-transcript-mcp exec tsx src/index.ts` through the repo's CLI. `--silent` (spelled out: `claude mcp add` has its own `-s`) keeps pnpm's script banner off stdout, which is the protocol's channel — pnpm 9 and 10 print it there without it. > **Register the server as `archilyzer`.** The shipped commands call > `mcp__archilyzer__ask_plan` / `mcp__archilyzer__sweep_plan`, and that tool name > embeds the server name **as you registered it**. Under any other name the commands > break; either use `archilyzer` or edit the `mcp____` prefix in those two > files. (`foo:bar` namespacing is plugin-only, so what you type stays `/ask`.) That is the whole setup for research. You need Node and pnpm — **no corpus, no yt-dlp, no GPU, nothing hosted.** Point `TRANSCRIPT_SITE_URL` at any instance, or use `TRANSCRIPT_HUB_URL` to federate several, or `TRANSCRIPT_LOCAL_DIR` for your own build. **On Windows, run all of this inside WSL2** — including Claude Code itself. See [Claude Code on Windows](#claude-code-on-windows). ### Clips and video: download the moment, not the movie This is the part that makes an archive worth more than a search box. A transcript gives you the **exact second** something was said, and the editor can fetch just those seconds. So the loop is: 1. Point the MCP server at an instance and `/ask` or `/sweep` about a subject. 2. Get back citations that resolve to precise moments in real recordings. 3. Ask for **just those clips** with the `fetch_clip` MCP tool — the editor fetches the window through its paced, cookie-aware, provenanced job; the file lands in the corpus beside the video (`channels//data//clips/`). Seconds of media, not hours. `full: true` fetches the whole recording into the saved-video store instead (it needs a video the editor already knows). `maxHeight` caps the source height; at 720 or less a whole recording is saved as the 720p H.264 preset. 4. Optionally, render them into a finished video. You are never downloading a back catalogue to find a quote. You search text, then fetch the few seconds you actually want. Steps 1–2 need no corpus and no media at all; step 3 needs a local Archilyzer editor (the MCP registered with `ARCHILYZER_EDITOR_URL` and `WORKER_TOKEN`) that already archives the cited channel: the editor fetches only for a channel it has under its `transcripts/`, so an MCP pointed at a public site with a fresh editor gets a 404 (`Channel "" not found`) on every clip. Step 4 adds `ffmpeg`/`ffprobe` and **ImageMagick with Pango** for the chrome. With no editor (a public-only setup), or none that archives the channel, the fallback is running yt-dlp yourself — `yt-dlp --download-sections` fetches just the cited seconds. A window fetched through the editor is kept in the corpus and reused by every later ask and render; the editor prunes clips by age. A render keeps its own `segments/`, `cards/` and the finished `.mp4` under the report's `out/`. "No corpus" means you are not mirroring a channel's back catalogue, not that nothing is stored: your disk use scales with the clips you actually pull, which for a report is minutes of video rather than years of it. ### Rendering a report to video `umtool/report-to-video/` turns a cited sweep report into an mp4: clips in chronological order, thin chrome carrying the citation, and a timeline of where you are. The report is *not* the regeneration source — a per-report `video.manifest.json` is, because a report's citations carry a start second and no clip length, so the real windows are recovered by matching each quote back to its covering caption cues. ```sh node umtool/report-to-video/resolve-windows.mjs /video.manifest.json --write node umtool/report-to-video/build-video.mjs /video.manifest.json ``` **Clip boundaries, and why this runs without a corpus.** Cutting on the raw cue span cuts mid-thought, because a cue boundary is just where the caption line wrapped. `resolve-windows.mjs` widens each clip outward to a whole sentence, and that needs cue **end** times — which neither a report nor an MCP snippet carries. Those come from the archive itself. A published instance serves the same record a local corpus holds, documented under `shardScheme` in `/corpus.json`: ```bash curl -s https://jeralyzer.pages.dev/transcripts/chrissie-mayr/page-0000.json | head -c 220 # [{"slug":…,"cues":[{"start":8.12,"end":10.31,"text":"squirrels move fast but that's just the"}… ``` So the scripts read cues from a local corpus when there is one and from the published archive when there is not — **no corpus required, and no configuration either**, since a manifest already records the archive it was built against (`provenance.siteOrigin`). | Flag | For | |---|---| | `--site-origin ` | Read cues from a specific archive, overriding the manifest. | | `--cue-source auto\|local\|http` | Which source to trust. `auto` is local-first. | | `--resolve-site-ids` | Recover from an id mismatch by scanning a channel's shards. Slow; see below. | > **The two sources can disagree, and not by rounding.** An archive is a snapshot; a > corpus keeps moving. Measured on this corpus — local six days newer than the publish — > three of four videos were byte-identical and the fourth had **65 of its 84 cue texts > rewritten, with timings shifted by up to 2.24 s**. That is enough to cut in the wrong > place, which is why the source is a flag rather than an implementation detail. Use > `--cue-source http` when you want the clip to match what a reader following the > citation will actually see, and `local` to refuse to fall back at all. > **Rumble videos have two ids.** The archive keys a recording by its **embed** id while > a local cue directory is named for the **URL slug**, so a manifest authored against > local directories misses on every Rumble clip. That fails with the diagnosis rather > than silently; fix it by adding `siteVideo` to the clip, or pass `--resolve-site-ids` > to find the record by scanning the channel's shards (each up to 8 MB, which is why it > is opt-in). A clip's `citeUrl` is deliberately *not* used for this — it may point at a > different recording on purpose, and that recording's clock is not the same one. ## How it fits together Three programs share one library and one pile of data. - **The editor** is a local admin app. You add channels, watch the download and transcription queues, and press the button that builds and deploys a site. It runs on your machine and is **never exposed to the public**. - **The corpus** is a directory on disk — one folder per channel, holding media, transcripts and a search index. It is deliberately kept outside the code: a fresh copy of the software has no data, and updating the software never touches your archive. - **The export** is the public artefact: a static site, pre-rendered to plain HTML and JSON. No database, no runtime, no server-side code. Two more pieces round out the workspace: the **project site** (`homepage/`) and the **MCP server** (`mcp/`). Full layout in [CONTRIBUTING.md](CONTRIBUTING.md). ## Posts: X, Bluesky and forum threads Beside video channels, a channel can be a **posts source**: an X or Bluesky account, or a **forum thread** (XenForo — Kiwi Farms is the first host, any XenForo 2 forum works the same way). Paste the account's or the thread's URL into the new-channel form; the posts are fetched into the corpus and searched alongside the transcripts. A forum thread is read in a headless browser, newest page first, one page at a time with a 10–20 s pause, on a browser profile kept per forum host. A browser check (Kiwi Farms' KiwiFlare) normally clears by itself and stays cleared; when it does not, or the thread needs a login, the run stops and says so, and **Connect forum session** on the channel page opens the profile in a window for you to clear it. Pages you saved from your own browser can be imported instead: `pnpm archilyzer posts import-html …` (or **Import saved pages** on the channel page). `pnpm archilyzer posts fetch --slug --pages N` reads just the latest N pages. ## Where your data lives Transcripts and per-channel state live at `/transcripts/` — **its own git repo**, untouched by the workspace. Every path and binary resolves through `getPaths()` (`common/lib/paths.ts`), and each can be overridden by an environment variable — `TRANSCRIPTS_DIR` moves the whole corpus, `YTDLP_BIN` / `FFMPEG_BIN` / `WHISPER_BIN` name the tools. The full list, with every other variable the code reads, is **[ENVIRONMENT.md](ENVIRONMENT.md)** (generated from a list the tests hold to the code). `pnpm archilyzer doctor` prints which overrides are set and whether every tool this machine is configured to use is there. Everything else lives in the editor's **Settings** page and is optional — a missing or partial settings file falls back to defaults. Every key is in [SETTINGS.md](SETTINGS.md). ## Publishing The static export can be served by anything. The path with the most support is Cloudflare Pages, with download archives too large for Pages' 25 MB per-file limit overflowing to R2. **[PUBLISH.md](PUBLISH.md)** covers building and deploying a site, the hub and the homepage (from the editor, `pnpm ops` or `pnpm archilyzer`), previews, the R2 setup, the configuration that keeps public archive downloads from being abused to run up costs, and the opt-in pipeline that builds several sites in parallel in isolated containers. Channels can sync automatically on a per-channel cadence via a cron heartbeat — see **[SCHEDULED_SYNC.md](SCHEDULED_SYNC.md)**. ## Getting the source, and contributing Archilyzer is MIT licensed and deliberately **forge-neutral**: there is no `.github/` directory, no provider-specific CI configuration, and nothing in the build that assumes a particular host. However you obtained this tree — a clone, a release archive, or a fork on whichever platform — it is a complete, self-contained working copy. **The canonical public copy is a read-only git mirror on the project site:** ```sh git clone https://archilyzer.pages.dev/source/archilyzer.git ``` It is `main` only, with its whole history, served as static files over git's dumb HTTP protocol, and regenerated from the private repository by every homepage build (`archilyzer source publish`, which `archilyzer build homepage` runs). Its commit ids differ from the private repository's because machine paths are scrubbed on the way out; a gate refuses to publish anything that still carries a denied string. Every file is browsable raw at [archilyzer.pages.dev/source/tree/](https://archilyzer.pages.dev/source/tree/), and the same tree without history is a tarball at [archilyzer.pages.dev/downloads](https://archilyzer.pages.dev/downloads/), with a `snapshot.json` sidecar (commit, size, SHA-256). How it is built: [PUBLISH.md](PUBLISH.md#the-source-mirror-homepage). See [CONTRIBUTING.md](CONTRIBUTING.md) to work on the code. ## Documentation | Document | Covers | |---|---| | [SETUP.md](SETUP.md) | Full per-OS install, transcription backends. | | [ENVIRONMENT.md](ENVIRONMENT.md) | Every environment variable, by audience (generated). | | [CONTRIBUTING.md](CONTRIBUTING.md) | Workspace layout, tests, the `archilyzer` CLI, internals. | | [OPERATING.md](OPERATING.md) | Running an archive from a shell or an agent: recipes over `pnpm ops`, `archilyzer` and the MCP. | | [COMMANDS.md](COMMANDS.md) | Every `archilyzer` command and `pnpm ops` action (generated). | | [RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md) | `docker compose up` for the whole stack: exposure model, auth, GPU. | | [PUBLISH.md](PUBLISH.md) | Building and deploying sites: Pages + R2, cost-abuse protection, parallel builds in containers. | | [SCHEDULED_SYNC.md](SCHEDULED_SYNC.md) | Unattended per-channel syncing. | | [WORKTREES.md](WORKTREES.md) | Parallel checkouts and the port scheme. | | [SETTINGS.md](SETTINGS.md) | Every `settings.json` key (generated). | | [SITE.md](SITE.md) | Every `site.json` key (generated). | | [CHANNEL.md](CHANNEL.md) | Every key of a channel's `config.json` (generated). | | [REPORT.md](REPORT.md) | The cited report document, `report.json` (generated). | | [CITATIONS.md](CITATIONS.md) | The citation model every cited document shares (generated). | | [AGENTS.md](AGENTS.md) | Instructions for coding agents working in this repo. | | [mcp/README.md](mcp/README.md) | The MCP server: tools, links, client setup. | ## License [MIT](LICENSE).