# Questions ## What does it cost to run? The software is free and MIT licensed. Your costs are hardware, electricity, and hosting. Hosting a published archive is the cheap part: static files on Cloudflare Pages fit comfortably in the free tier, and bulk download archives too large for it can go behind a free worker. The reference deployment publishes tens of thousands of transcripts this way. The expensive part is transcription, and it is paid in machine time rather than money. See below. ## How long does transcribing take? It depends entirely on your hardware and which backend you use, and any number quoted without both is meaningless. What is safe to say: it is the bottleneck, it wants a GPU, and it scales with audio hours rather than video count. A back catalogue of several thousand hours is a project measured in days of machine time, not minutes. Channels that already publish captions skip this step entirely and cost almost nothing. One trap worth knowing about: running too many jobs at once can push a GPU backend into falling back to the CPU, at which point everything gets several times slower simultaneously. More parallelism is not reliably faster. ## How much disk does it take? Transcripts are small — text compresses well and a whole corpus of them is megabytes. Media is not. Whether you keep the audio after transcribing is a per-channel decision, and the software tracks which recordings are safe to clean up and which are not. Audio for recordings that have disappeared upstream is never cleaned up automatically, on the grounds that it is the one thing you cannot re-fetch. ## Do I need a GPU? No, but transcription without one is slow enough to change what is practical. A CPU-only machine is fine for a channel that publishes captions, or for a small back catalogue you are content to work through over weeks. ## Is any of this sent anywhere? No. Downloads go from the platform to your machine, transcription runs on your machine, and the build writes files on your machine. Nothing is sent to us — there is no "us" in the runtime path at all. The two exceptions are both yours to make: publishing sends the built site to whatever host you choose, and the optional in-browser chat on a published site sends queries to an AI provider using an API key the *visitor* supplies. ## Is this legal? Downloading and keeping a copy of publicly published video is a question with different answers in different jurisdictions, and it is not one a piece of software can answer for you. The tool does what you tell it to; deciding what to tell it is your responsibility, as is complying with whatever terms and laws apply where you are. What the software does do is make the record accurate: it says what it archived, when, and whether the original is still there — rather than presenting an archive as if it were the platform. ## What happens when a video is deleted? Your copy stays. The archive re-checks whether recordings still exist at their source and records the answer, so a deleted recording is marked as deleted, remains searchable, and keeps its transcript. The reverse claim is one the software is careful never to make. A recording counts as available until something re-checks it and finds otherwise, and nothing re-checks a whole corpus continuously. "Available" here means *not known to be gone* — never "confirmed still there". ## Where is the git repository? On this site, read-only: `git clone https://archilyzer.pages.dev/source/archilyzer.git`. It is the main branch with its whole history, regenerated from the private repository with every deploy — machine paths scrubbed on the way out, which is why its commit ids differ. There is no GitHub, and nothing takes a push or a pull request. [Source](/source/) also has every file to browse and the history with every commit's diff, and [Downloads](/downloads/) the same tree as a tarball. ## Can I use it for one video? You can, but nothing about it is designed for that. It is built around back catalogues — thousands of recordings, kept current on a schedule — and almost every design decision in it, from the paginated index to the incremental build, exists because of scale. ## What this is not - **Not a hosted service.** There is no account, no subscription, no upload. - **Not a forge.** A read-only mirror: clone and pull, but no pushes, issues or pull requests. - **Not a downloader.** It supplies no downloader — you install `yt-dlp`, `ffmpeg` and a transcription backend yourself, and it drives them. - **Not cloud transcription.** Your hardware, your electricity, your queue.