Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 9ca58f67f87e024a75d48c690373a71e419c12ac
parent 557da85c2551ee71dc143773a2d9bfa84a36486f
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sat, 18 Apr 2026 01:43:58 -0400

new README

Diffstat:
MREADME.md | 29+++++++++++++++++++++++++++++
1 file changed, 29 insertions(+), 0 deletions(-)

diff --git a/README.md b/README.md @@ -45,6 +45,17 @@ If you want to change the code, use the development server for hot-reloading and pnpm run dev ``` +## Preprocessed Transcripts in `public/` + +`scripts/build-index.ts` (run automatically via the `prebuild` hook, or manually with `pnpm build:index`) reads everything under `transcripts/data/`, caches parsed results in an LMDB store at `transcripts/index.mdb`, and emits static JSON into `public/`: + +- `public/summaries/manifest.json` + `public/summaries/page-NNNN.json` — paginated list of video summaries. +- `public/transcripts/<slug>.json` — per-video cues and metadata. + +The app fetches these files at runtime as plain static assets, which keeps `next build` fast even with 10k+ videos — Next.js never has to parse VTT files or walk the transcripts directory during `next build` itself. Incremental reruns of the index script short-circuit via mtime checks, so a rerun after adding a handful of new videos only processes the new ones. + +The trade-off: **Next.js does not watch `public/` for changes**, so adding or editing transcripts while `pnpm run dev` is running won't trigger a hot-reload. Re-run `pnpm build:index` and refresh the browser (or restart the dev server) to pick up new transcripts. + ## Downloading Transcripts Use this command to download transcripts that end up in the expected format: @@ -52,3 +63,21 @@ Use this command to download transcripts that end up in the expected format: ``` yt-dlp --ignore-config --skip-download --restrict-filenames --write-info-json -o "subtitle:%(upload_date)s_%(id)s-%(fulltitle,title)s/transcript" -o "infojson:%(upload_date)s_%(id)s-%(fulltitle,title)s/metadata" --write-auto-subs -- <VIDEO_OR_PLAYLIST> ``` + +## Downloading Transcripts for a Whole Channel + +yt-dlp is far from bulletproof, and trying to scrape a channel with thousands of videos can be flaky. This multi-step process makes it more achievable: + +1. **Download a playlist (including a channel's "Videos" or "Streams" page) as a list of URLs** + `yt-dlp --flat-playlist --skip-download --print url <CHANNEL_OR_PLAYLIST_URL> > ../playlist.txt` + +2. **Start downloading from the saved playlist, writing completed downloads to an archive** + `yt-dlp --ignore-config --skip-download --restrict-filenames -fb --write-info-json -o "%(upload_date)s_%(id)s/video" -o "subtitle:%(upload_date)s_%(id)s/transcript" -o "infojson:%(upload_date)s_%(id)s/metadata" --write-auto-subs --force-write-archive --download-archive ../archive.txt --abort-on-error -t sleep -a ../playlist.txt` + - `--force-write-archive --download-archive ../archive.txt` is the main point — yt-dlp will locally track completed videos in a text file and completely skip the download if it has already been done. This makes resuming the download for even a massive playlist much easier. + - `-a ../playlist.txt` is how you load a text file line-by-line; just passing the path to the text file will not work. + - `-t sleep` is a preset that defines multiple small wait periods to prevent rate-limiting. + - `--abort-on-error` is nice to not keep churning while rate-limited, but could be omitted for more hands-off progress. + - `--cookies-from-browser` can be used to download age-restricted videos, though you should only use it when needed to minimize risk to your YouTube account. + +3. **Verify subtitles are present for each video in the archive** + Rarely, a video's subtitle can be missed in error and the download still marked as complete.