Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 36bcb340e4dce8a2b0161a827e24f14a06ada6c7
parent 6a0debffcb701b9b9dc632f6c1ad9fc3aa5c6663
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Mon,  6 Jul 2026 22:16:29 -0400

Merge branch 'main' into worktree-feat+archive-downloads

Diffstat:
Mexport/CHANGELOG.md | 2+-
Mr2-proxy/wrangler.toml | 4++--
2 files changed, 3 insertions(+), 3 deletions(-)

diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md @@ -1,6 +1,6 @@ # Changelog -## [Unreleased] +## [0.7.0] - 2026-07-06 - **Bring-your-own-AI: the archive is now machine-navigable for AI tools.** Every site publishes a small fixed set of discovery files — `llms.txt` (an LLM-readable overview) and `corpus.json` (a documented index of the channels and *how to fetch any transcript* from the existing paginated JSON shards), plus `robots.txt` and a `sitemap.xml`. Nothing is generated per video (the shard scheme is documented instead), so the file count stays constant no matter how large the corpus grows. This lets Claude Code and other tools browse and answer questions about the archive by fetching a couple of URLs. The federated hub publishes an aggregate `corpus.json`/`llms.txt` spanning every member site. - **New MCP server (`mcp/`) for Claude Code, Cursor, and other MCP clients.** A local tool that exposes the archive as MCP tools — `list_channels`, `search_transcripts` (timestamped snippets), `get_transcript`, `get_video_metadata` — reading the same static shards over disk or HTTP. It can point at one site or federate a whole hub. It never changes or hosts the site; see `mcp/README.md` for setup. - **New "Use with AI" page + copy-for-AI buttons.** A `/use-with-ai` page (linked from the header and footer) explains the in-browser chat, the `llms.txt`/`corpus.json` discovery files, and the MCP server. The player toolbar gains a "Copy as Markdown" control that copies the current transcript or live chat as clean, timestamped markdown for pasting into any AI chat, and the search results header gains a "Copy for AI" button that copies the matched videos and their hit snippets as context. diff --git a/r2-proxy/wrangler.toml b/r2-proxy/wrangler.toml @@ -6,7 +6,7 @@ # One bucket + one Worker serves every export site (keys are namespaced by site # id). See ../DEPLOY_CLOUDFLARE.md for the full setup. -name = "ytdlp-archive-proxy" +name = "archilyzer-exports" main = "src/index.ts" compatibility_date = "2025-06-01" workers_dev = true @@ -15,7 +15,7 @@ workers_dev = true # you entered in the editor's "Archive overflow storage" setting. [[r2_buckets]] binding = "BUCKET" -bucket_name = "your-archives-bucket" +bucket_name = "archilyzer-exports" # Native per-edge rate limiting (free, no extra infrastructure). 60 requests per # 60s per client IP + file. Tune `limit` to taste; `period` must be 10 or 60.