Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 2714869140a12212ddd7a4ecb98f139c28086d47
parent b41f613dbad5a13aa4aa243fb98e821d72fbc7b8
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sun,  5 Jul 2026 15:10:44 -0400

R2 archive downloads: Cache-Control on uploads + Cloudflare deploy/security guide

Set Cache-Control: public, max-age=3600 on every archive uploaded to R2 so a
Cloudflare custom domain caches downloads at the edge (the main cost defense —
R2 egress is free, only origin reads bill, cached hits skip the origin).

Add DEPLOY_CLOUDFLARE.md documenting R2 bucket setup and the Cloudflare custom
domain / caching / rate-limiting / bot-mode / billing-alert configuration for
instance operators; link it from the README.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

Diffstat:
ADEPLOY_CLOUDFLARE.md | 192+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
MREADME.md | 4++++
Meditor/CHANGELOG.md | 2+-
Meditor/app/deploy/buildDeployCore.ts | 9+++++++++
4 files changed, 206 insertions(+), 1 deletion(-)

diff --git a/DEPLOY_CLOUDFLARE.md b/DEPLOY_CLOUDFLARE.md @@ -0,0 +1,192 @@ +# Deploying to Cloudflare (Pages + R2 archive overflow) + +The static site is deployed to **Cloudflare Pages**; large download archives that +exceed Pages' per-file limit overflow to **Cloudflare R2**. This guide covers the R2 +setup and — importantly — how to configure Cloudflare so that **public archive +downloads can't be abused to drive up your bill**. + +Everything here fits inside Cloudflare's **free tier**. + +- [How archives are served](#how-archives-are-served) +- [One-time R2 setup](#one-time-r2-setup) +- [Securing downloads against cost-abuse](#securing-downloads-against-cost-abuse) +- [Cost expectations](#cost-expectations) + +--- + +## How archives are served + +At build time, `compose-site.ts` generates one transcript zip and one live-chat zip +**per channel** into `export/public/archives/`, and records them in +`public/archives/manifest.json`. The `/downloads` page and the header **Downloads** +link read that manifest. + +Cloudflare Pages rejects any single asset larger than **25 MB**. Real channels blow +past that easily (a channel's live-chat zip can be hundreds of MB). So the pipeline +splits archives by size: + +| Archive size | Where it's served from | Manifest entry | +|---|---|---| +| ≤ 25 MB | Cloudflare **Pages** (shipped in `out/`, free) | `filename`, no `url` | +| > 25 MB, R2 configured | Cloudflare **R2**, uploaded on deploy | `url` → R2 | +| > 25 MB, R2 **not** configured | not served | `oversize: true`, shown as "Too large to host" | + +Oversize archives are staged during compose into `export/.r2-staging/<siteId>/archives/` +(gitignored, kept out of `public/`), then uploaded by the editor's **Deploy** / +**Build & deploy** actions to `<bucket>/<siteId>/archives/<file>.zip` **before** the +Pages deploy runs, so the manifest URLs resolve immediately. + +> **Note:** only the **editor's** deploy actions upload to R2 (that's where the host +> `wrangler` credentials live). A raw `pnpm run build` / `pnpm deploy` from the +> `export` package only *stages* the files — it does not upload them. + +--- + +## One-time R2 setup + +### 1. Create the bucket + +One bucket serves **all** your sites — objects are namespaced by `siteId`, so you do +**not** need one bucket per site. In the dashboard (**R2 → Create bucket**) or via the +same `wrangler` CLI the deploy already uses: + +```sh +wrangler r2 bucket create my-archives-bucket +``` + +### 2. Expose it publicly + +R2 buckets are private by default, and the Downloads page links to plain +`https://…/<file>.zip` URLs, so the bucket needs public read access. **Prefer a custom +domain** — it's the foundation for every cost-defense below: + +- **Custom domain (recommended):** bucket **Settings → Public access → Custom Domains + → Connect Domain**, e.g. `archives.example.com` (the domain must be on your + Cloudflare account). This routes downloads through Cloudflare's CDN, WAF, caching, + and rate limiting. +- **`r2.dev` subdomain (quick, but avoid for production):** bucket **Settings → Public + access → Allow Access**. You get a URL like `https://pub-abc123.r2.dev`. Cloudflare + throttles `r2.dev` and it gives you **none** of the CDN / WAF / rate-limiting + controls below. Fine for a quick test; not for a real instance. + +### 3. Authenticate `wrangler` + +Uploads run as `wrangler r2 object put … --remote` from the host, reusing whatever +auth your Pages deploy already uses. If `wrangler pages deploy` works today, this will +too. Otherwise run `wrangler login`, or set `CLOUDFLARE_API_TOKEN` with R2 write scope. + +### 4. Point the editor at the bucket + +In the editor, open **Settings** and set: + +| Field | Value | +|---|---| +| **Archive overflow storage (R2 bucket)** | the bucket name, e.g. `my-archives-bucket` | +| **Archive overflow public URL** | the public base from step 2, e.g. `https://archives.example.com` (no trailing slash needed) | + +Leave both blank to keep the old behavior (oversize archives dropped, shown as "Too +large to host"). + +On the next **Build & deploy**, oversize archives upload to R2 and the Downloads page +(and the header link, which reappears once anything is hostable) point at them. + +--- + +## Securing downloads against cost-abuse + +The threat: someone scripts repeated downloads of large archives to run up your bill. + +**The reassuring part — R2 egress is free.** Unlike S3, Cloudflare R2 charges **$0** +for bandwidth/egress. An attacker looping downloads of a 480 MB zip cannot run up a +bandwidth bill. The *only* metered cost from reads is **Class B operations** (10M free +per month, then $0.36/M) — and the defenses below make even that hard to reach. + +Set these up once, in order of impact: + +### 1. Serve through a custom domain, not `r2.dev` — the master switch + +Everything else in this section requires it. A custom domain puts downloads behind +Cloudflare's CDN and unlocks WAF, caching, rate limiting, and bot controls. See +[step 2 above](#2-expose-it-publicly). + +### 2. Edge caching (already configured in code) + +Every archive is uploaded with `Cache-Control: public, max-age=3600` (set in +`editor/app/deploy/buildDeployCore.ts`, constant `ARCHIVE_CACHE_CONTROL`). Served +through a custom domain, Cloudflare's CDN caches each archive at the edge, so repeated +downloads of the same file are served from cache and **never hit R2** — they don't +count as Class B ops and cost nothing. This is the single most effective cost defense: +a flood against one file mostly just warms one cache entry. + +To make caching aggressive, add a **Cache Rule** (dashboard: **Caching → Cache Rules → +Create**): + +- **When:** `URI Path` contains `/archives/` +- **Then:** *Eligible for cache*, **Edge TTL → Override → 1 day** (or longer). + +Archive filenames are stable and overwritten in place on re-deploy, so if you raise the +TTL a lot, **purge the cache on deploy** (dashboard **Caching → Purge**, or +`wrangler`/API) so updated archives aren't served stale. The 1-hour `Cache-Control` +default already bounds staleness without a purge; tune `ARCHIVE_CACHE_CONTROL` up if +your archives rarely change. + +### 3. Rate limiting — the hard backstop + +A Rate Limiting rule caps how fast any single client can pull archives, stopping a +flood that misses cache. Dashboard: **Security → WAF → Rate limiting rules → Create** +(the free plan includes one rule): + +- **When incoming requests match:** `URI Path` contains `/archives/` +- **Rate:** e.g. **20 requests per 1 minute** per client IP +- **Then:** *Block* for 10 minutes (or *Managed Challenge*). + +Tune the threshold to real usage — legitimate users download a handful of files, not +dozens per minute. + +### 4. Bot Fight Mode + WAF managed rules + +Dashboard: **Security → Bots → Bot Fight Mode** (free). Blocks the low-effort scripted +abuse that makes up most of this traffic. The free **WAF managed ruleset** adds a +baseline of protection at no cost. + +### 5. Hotlink protection (optional) + +Stops other sites embedding your archives and spending your ops budget serving their +audience. A WAF custom rule (**Security → WAF → Custom rules**): + +- **When:** `URI Path` contains `/archives/` **and** `Referer` does not contain your + domain **and** `Referer` is not empty +- **Then:** *Block*. + +(Allow an empty `Referer` so direct clicks and privacy-conscious browsers still work.) + +### 6. Billing / usage alerts + +R2 has no hard spend cap, but Cloudflare **Notifications** (dashboard: **Notifications +→ Add**) can email you when R2 storage or Class A/B operations cross a threshold — +cheap insurance so nothing surprises you. + +### What to skip + +**Signed URLs / token-gated downloads via a Worker** are the heavyweight option. Given +egress is free and caching neutralizes the ops cost, they're overkill for *cost* +defense — they add a Worker, key management, and friction for legitimate users. Only +reach for them if you want *access control* (private archives), not cost control. + +--- + +## Cost expectations + +Measured across all sites in this project, total compressed archives are **≈ 2–4.5 GB**. +Against R2's free tier: + +| Resource | Free tier / month | This project's usage | +|---|---|---| +| **Egress / bandwidth** | unlimited, **$0** | irrelevant — no egress charge exists | +| **Storage** | 10 GB-month | ~2–4.5 GB — comfortable | +| **Class A ops** (writes) | 1,000,000 | ~one PUT per archive per deploy — negligible | +| **Class B ops** (reads) | 10,000,000 | mostly absorbed by CDN cache | + +The realistic bill for hosting these archives is **$0**. The configuration above exists +to keep it that way under adversarial traffic, not because normal usage is close to any +limit. diff --git a/README.md b/README.md @@ -94,3 +94,7 @@ pnpm --filter yt-dlp-transcript-common exec tsx bin/verify-transcripts.ts --chan The export app reads from the LMDB-backed paginated JSON in `export/public/{summaries,transcripts}/` and renders a search/filter UI as a static `out/` directory. It's the read-only public face of the data. The export build uses `transpilePackages: ["yt-dlp-transcript-common"]` so editing common is hot-reloadable without a separate build step. `lmdb`, `msgpackr`, and `msgpackr-extract` are listed in `serverExternalPackages` so Next.js doesn't try to bundle them. + +## Deploying to Cloudflare + +The site deploys to Cloudflare Pages; download archives too large for Pages' 25 MB per-file limit overflow to Cloudflare R2. For R2 setup and the Cloudflare configuration that keeps public archive downloads from being abused to run up costs, see [DEPLOY_CLOUDFLARE.md](DEPLOY_CLOUDFLARE.md). diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -4,7 +4,7 @@ - **Search results scroll smoothly again on large result sets.** The results list windows one card per matching video (only the on-screen cards are mounted), but each visible card and every one of its hit rows was re-rendering on *every* scroll frame — and each hit row re-ran its `<mark>` highlighting, so a single video with hundreds of hits meant hundreds of redundant highlight passes per frame while scrolling. The result cards and individual hit rows are now memoized so an unchanged card/row is skipped during scroll, and opening the modal on a hit only re-renders the two rows whose highlight state actually changes. No visible/behavioral change — same DOM, same results, just far less work per frame. See `common/components/TranscriptSearch.tsx` (`ResultCard`/`HitRow` memoization, `openWithMode` stabilized via `useCallback`). - **The monitor widget gains a needs-work channel list, more interaction buttons, and an in-place settings gear.** Three additions, all driveable from the widget builder. **(1) A "Needs work" list** (URL flag `act=1`) — a compact, per-channel worklist of videos to download (`↓ N`) or transcribe (`✎ N`), reusing the same `loadActionableSummary` that powers the `/actionable` page via a new `/api/widget/actionable` route; it polls on a 15s floor (the backlog changes on job completions, not seconds) and caps at 6 channels with a `+N more` line. **(2) More interactions** behind the existing `controls=1` switch: each needs-work row gains the same per-channel **Download missing** / **Transcribe pending** buttons as the actionable page (reusing `InlineActionButton`), and the controls row adds **Retry all failed** alongside Pause/Resume + Drain. **(3) An in-place settings gear** (on by default; URL flag `gear=0` to hide, or a **Show settings gear** builder checkbox) — clicking it opens the builder's own form *inside the widget window*, so a pinned widget can be reconfigured live without opening the builder page; edits apply immediately and mirror into the address bar via `history.replaceState`, so a reload preserves them and the link stays copyable. The builder form is extracted into a shared `WidgetConfigForm` used by both the builder and the overlay, and the widget's poller now fetches immediately on (re)subscribe instead of after one interval, so newly-enabled sections render at once. Existing links render unchanged (the two new flags default to their old behavior; the gear is the one new default-visible affordance and is read-only — it mutates no server state). See `editor/app/widget/lib/config.ts`, the new `editor/app/widget/components/WidgetConfigForm.tsx` and `editor/app/api/widget/actionable/route.ts`, `editor/app/widget/components/{MonitorWidget,WidgetControls}.tsx`, `editor/app/widget/builder/components/WidgetBuilder.tsx`, and `editor/e2e/widget.spec.ts`. - **Every site build now bundles downloadable per-channel transcript & live-chat archive zips.** The archive builders (per-channel `<slug>.zip` / `<slug>.live_chat.zip`) previously only ran as standalone actions that wrote to a non-served directory; now `compose-site` generates them for the site's own channels straight into the served `public/archives/` and writes a `manifest.json` (sizes + counts) that the site's new **Downloads** page reads. **`zip` is now the default archive format** everywhere (was `tar.gz`), and the Build page's format help text tracks the selected format. Generation is **on by default with three opt-out levels**: a global **Generate archive zips on build** toggle in Settings, a per-site **Generate archive zips** toggle (plus an optional **Archive size cap (MB)**) on the site's page, and a per-build **Skip archive zips** checkbox on the Build and Build & Deploy controls (`BUILD_ARCHIVES=0`). See `common/bin/compose-site.ts` (`composeArchives`), `common/controller/archive{Transcripts,LiveChat}.ts` (new `outDir` option), `common/lib/archiveOptions.ts` (default + manifest types), `common/lib/{site,settings}.ts` (opt-out flags), and `editor/app/{deploy/buildDeployCore.ts,build/buildAction.ts,deploy/components/Build{Export,Deploy}Button.tsx,sites/components/SiteForm.tsx,settings/components/SettingsForm.tsx}`. -- **Oversize archives now overflow to Cloudflare R2 instead of being dropped.** A single file over 25 MB breaks a Cloudflare Pages deploy, so a channel zip over the cap (default 25 MB; `0` = no cap) used to be removed from what's served and flagged `oversize`. Now, when **Archive overflow storage** is configured in Settings (an R2 **bucket** + its **public URL**), `compose-site` stages each oversize archive to `export/.r2-staging/<siteId>/` and records its future public URL in the manifest; the deploy step then uploads it via `wrangler r2 object put <bucket>/<siteId>/archives/<file>` (reusing the host's wrangler auth) **before** the Pages deploy, so the Downloads page links straight to R2. With no bucket configured the old drop-and-flag behavior is unchanged (uploads run only in the editor's Deploy / Build & deploy actions, not a raw `pnpm deploy`). The **combined "whole site" archives were removed** — they duplicated the per-channel content and were always the first to blow the cap. See `common/lib/settings.ts` (`archiveStorage`), `common/bin/compose-site.ts` (staging + manifest `url`), `common/lib/archiveOptions.ts` (`ArchiveManifestEntry.url`), and `editor/app/deploy/buildDeployCore.ts` (`runArchiveUploadIntoLog`) / `deploy/deployAction.ts` / `build/buildAction.ts`. +- **Oversize archives now overflow to Cloudflare R2 instead of being dropped.** A single file over 25 MB breaks a Cloudflare Pages deploy, so a channel zip over the cap (default 25 MB; `0` = no cap) used to be removed from what's served and flagged `oversize`. Now, when **Archive overflow storage** is configured in Settings (an R2 **bucket** + its **public URL**), `compose-site` stages each oversize archive to `export/.r2-staging/<siteId>/` and records its future public URL in the manifest; the deploy step then uploads it via `wrangler r2 object put <bucket>/<siteId>/archives/<file>` (reusing the host's wrangler auth) **before** the Pages deploy, so the Downloads page links straight to R2. With no bucket configured the old drop-and-flag behavior is unchanged (uploads run only in the editor's Deploy / Build & deploy actions, not a raw `pnpm deploy`). The **combined "whole site" archives were removed** — they duplicated the per-channel content and were always the first to blow the cap. Each R2 upload also now sets `Cache-Control: public, max-age=3600` so a Cloudflare custom domain caches downloads at the edge — the main defense against download-abuse cost (R2 egress is free; only origin reads are billable, and cached hits skip the origin). New **[DEPLOY_CLOUDFLARE.md](../DEPLOY_CLOUDFLARE.md)** documents the full R2 setup plus the Cloudflare custom-domain / caching / rate-limiting / bot config for instance operators. See `common/lib/settings.ts` (`archiveStorage`), `common/bin/compose-site.ts` (staging + manifest `url`), `common/lib/archiveOptions.ts` (`ArchiveManifestEntry.url`), and `editor/app/deploy/buildDeployCore.ts` (`runArchiveUploadIntoLog`, `ARCHIVE_CACHE_CONTROL`) / `deploy/deployAction.ts` / `build/buildAction.ts`. - **The Duplicates page is now a per-site toggle and hides itself when empty.** Each site's editor page gains a **Show the Duplicates page** checkbox (on by default). `compose-site` writes the site-filtered `duplicates.json` only when the toggle is on *and* there's at least one in-scope cluster, and the export Header keys its Duplicates nav link off a new `hasDuplicates()` — so the link and page disappear both when a site opts out and when it simply has no detected duplicates. See `common/lib/site.ts` (`duplicates` flag), `common/bin/compose-site.ts` (gated write), `export/app/lib/duplicates.ts` (new), `export/app/components/Header.tsx`, and `editor/app/sites/{components/SiteForm.tsx,actions.ts}`. - **You can now change a channel's slug (its id) — deliberately, from the Danger zone.** A channel's slug *is* its on-disk directory name (`transcripts/channels/<slug>/`), so it used to be fixed at creation ("Slug is fixed once a channel is created"). A new **Rename** form in the channel's Danger zone lifts that: enter a new slug and **type the current slug to confirm** (same friction as delete), and the rename is blocked while the channel has running/queued jobs (the in-memory registry keys by slug). Because the slug is a directory name, the rename does a **full migration** of every slug-keyed store so nothing silently breaks: it moves the channel dir (config, data, playlist, snapshot, shards, failed lists) **and** the saved-video store dir — rewriting each `saved-video.json` pointer's absolute `dir` so persisted source videos still resolve — then retargets every site.json membership, the sync scheduler's per-channel backoff state, and any job bookmarks. The two filesystem moves run first and roll back on failure; the metadata updates that follow are atomic and best-effort (surfaced as warnings). Renaming **changes the channel's public URL** (the old one 404s), which the form warns about. The slug grammar is also now validated on create. See `common/controller/renameChannel.ts`, `common/controller/channels.ts` (`isValidChannelSlug`), `common/lib/savedVideo-server.ts` (`rewriteSavedVideoDir`), `common/jobs/bookmarks.ts` (`renameChannelInBookmarks`), `editor/app/channels/{actions.ts,components/RenameChannelForm.tsx,[slug]/page.tsx}`, and `editor/e2e/channel-rename.spec.ts`. - **New Queue diagnostics page (`/jobs/queue`): see & force-release stuck jobs.** The job system has two sources of truth that can drift — the registry owns each job's `status`, the scheduler owns the running SLOT per queue. A cancel that never finalizes (a child that ignored SIGTERM, a crashed finalizer) leaves a job "cancelled" in the registry while the scheduler still marks its slot running, silently blocking every job behind it on that queue — and the Active Jobs page hides it (it filters to running/queued). The new **Queue** page reconciles the two: it builds from the **scheduler** as the source of truth for slots, cross-checks each against its registry record, and flags a running head as **stuck** when the record is terminal-but-holding-slot, evicted, or (softer) a live job idle past 10 minutes. It **auto-heals** the hard cases on every view/poll (frees terminal/evicted slots), shows a health strip (active queues, running, queued, **stuck**, workers), per-queue cards with the held-for duration / PID (`kill -9` hint) / last log line, and a **Force-release** button per slot (SIGKILLs the child and frees the slot unconditionally) plus a **Reap all stuck** action. Force-release is also available on any running job in Active Jobs, and Active Jobs links to Queue with a stuck-count badge. See `common/jobs/registry.ts` (`forceRelease`), `editor/app/jobs/queue/*`, `editor/app/jobs/{actions.ts,components/ForceReleaseJobButton.tsx}`, and `editor/e2e/queue.spec.ts`. diff --git a/editor/app/deploy/buildDeployCore.ts b/editor/app/deploy/buildDeployCore.ts @@ -85,6 +85,14 @@ function archiveStagingDir(siteId: string, paths: Paths): string { ); } +// Cache-Control set on every uploaded archive. Served through a Cloudflare custom +// domain, this lets the CDN absorb repeated/abusive downloads at the edge instead +// of hitting R2 (each origin GET is a billable Class B op), which is the main cost +// defense for public archives — see DEPLOY_CLOUDFLARE.md. 1h balances flood +// absorption against re-deployed archives (stable filenames, overwritten in place) +// going stale; raise it if your archives rarely change. +const ARCHIVE_CACHE_CONTROL = "public, max-age=3600"; + // Upload this site's staged oversize archives to the configured R2 bucket, so the // remote URLs the served manifest points at actually resolve. Reuses the host's // wrangler auth (process.env), same as the Pages deploy. No-op (returns 0) when @@ -126,6 +134,7 @@ export async function runArchiveUploadIntoLog( `${bucket}/${key}`, `--file=${path.join(stagingDir, file)}`, "--content-type=application/zip", + `--cache-control=${ARCHIVE_CACHE_CONTROL}`, "--remote", ], cwd: paths.exportDir,