commit 2714869140a12212ddd7a4ecb98f139c28086d47
parent b41f613dbad5a13aa4aa243fb98e821d72fbc7b8
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Sun, 5 Jul 2026 15:10:44 -0400
R2 archive downloads: Cache-Control on uploads + Cloudflare deploy/security guide
Set Cache-Control: public, max-age=3600 on every archive uploaded to R2 so a
Cloudflare custom domain caches downloads at the edge (the main cost defense —
R2 egress is free, only origin reads bill, cached hits skip the origin).
Add DEPLOY_CLOUDFLARE.md documenting R2 bucket setup and the Cloudflare custom
domain / caching / rate-limiting / bot-mode / billing-alert configuration for
instance operators; link it from the README.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Diffstat:
4 files changed, 206 insertions(+), 1 deletion(-)
diff --git a/DEPLOY_CLOUDFLARE.md b/DEPLOY_CLOUDFLARE.md
@@ -0,0 +1,192 @@
+# Deploying to Cloudflare (Pages + R2 archive overflow)
+
+The static site is deployed to **Cloudflare Pages**; large download archives that
+exceed Pages' per-file limit overflow to **Cloudflare R2**. This guide covers the R2
+setup and — importantly — how to configure Cloudflare so that **public archive
+downloads can't be abused to drive up your bill**.
+
+Everything here fits inside Cloudflare's **free tier**.
+
+- [How archives are served](#how-archives-are-served)
+- [One-time R2 setup](#one-time-r2-setup)
+- [Securing downloads against cost-abuse](#securing-downloads-against-cost-abuse)
+- [Cost expectations](#cost-expectations)
+
+---
+
+## How archives are served
+
+At build time, `compose-site.ts` generates one transcript zip and one live-chat zip
+**per channel** into `export/public/archives/`, and records them in
+`public/archives/manifest.json`. The `/downloads` page and the header **Downloads**
+link read that manifest.
+
+Cloudflare Pages rejects any single asset larger than **25 MB**. Real channels blow
+past that easily (a channel's live-chat zip can be hundreds of MB). So the pipeline
+splits archives by size:
+
+| Archive size | Where it's served from | Manifest entry |
+|---|---|---|
+| ≤ 25 MB | Cloudflare **Pages** (shipped in `out/`, free) | `filename`, no `url` |
+| > 25 MB, R2 configured | Cloudflare **R2**, uploaded on deploy | `url` → R2 |
+| > 25 MB, R2 **not** configured | not served | `oversize: true`, shown as "Too large to host" |
+
+Oversize archives are staged during compose into `export/.r2-staging/<siteId>/archives/`
+(gitignored, kept out of `public/`), then uploaded by the editor's **Deploy** /
+**Build & deploy** actions to `<bucket>/<siteId>/archives/<file>.zip` **before** the
+Pages deploy runs, so the manifest URLs resolve immediately.
+
+> **Note:** only the **editor's** deploy actions upload to R2 (that's where the host
+> `wrangler` credentials live). A raw `pnpm run build` / `pnpm deploy` from the
+> `export` package only *stages* the files — it does not upload them.
+
+---
+
+## One-time R2 setup
+
+### 1. Create the bucket
+
+One bucket serves **all** your sites — objects are namespaced by `siteId`, so you do
+**not** need one bucket per site. In the dashboard (**R2 → Create bucket**) or via the
+same `wrangler` CLI the deploy already uses:
+
+```sh
+wrangler r2 bucket create my-archives-bucket
+```
+
+### 2. Expose it publicly
+
+R2 buckets are private by default, and the Downloads page links to plain
+`https://…/<file>.zip` URLs, so the bucket needs public read access. **Prefer a custom
+domain** — it's the foundation for every cost-defense below:
+
+- **Custom domain (recommended):** bucket **Settings → Public access → Custom Domains
+ → Connect Domain**, e.g. `archives.example.com` (the domain must be on your
+ Cloudflare account). This routes downloads through Cloudflare's CDN, WAF, caching,
+ and rate limiting.
+- **`r2.dev` subdomain (quick, but avoid for production):** bucket **Settings → Public
+ access → Allow Access**. You get a URL like `https://pub-abc123.r2.dev`. Cloudflare
+ throttles `r2.dev` and it gives you **none** of the CDN / WAF / rate-limiting
+ controls below. Fine for a quick test; not for a real instance.
+
+### 3. Authenticate `wrangler`
+
+Uploads run as `wrangler r2 object put … --remote` from the host, reusing whatever
+auth your Pages deploy already uses. If `wrangler pages deploy` works today, this will
+too. Otherwise run `wrangler login`, or set `CLOUDFLARE_API_TOKEN` with R2 write scope.
+
+### 4. Point the editor at the bucket
+
+In the editor, open **Settings** and set:
+
+| Field | Value |
+|---|---|
+| **Archive overflow storage (R2 bucket)** | the bucket name, e.g. `my-archives-bucket` |
+| **Archive overflow public URL** | the public base from step 2, e.g. `https://archives.example.com` (no trailing slash needed) |
+
+Leave both blank to keep the old behavior (oversize archives dropped, shown as "Too
+large to host").
+
+On the next **Build & deploy**, oversize archives upload to R2 and the Downloads page
+(and the header link, which reappears once anything is hostable) point at them.
+
+---
+
+## Securing downloads against cost-abuse
+
+The threat: someone scripts repeated downloads of large archives to run up your bill.
+
+**The reassuring part — R2 egress is free.** Unlike S3, Cloudflare R2 charges **$0**
+for bandwidth/egress. An attacker looping downloads of a 480 MB zip cannot run up a
+bandwidth bill. The *only* metered cost from reads is **Class B operations** (10M free
+per month, then $0.36/M) — and the defenses below make even that hard to reach.
+
+Set these up once, in order of impact:
+
+### 1. Serve through a custom domain, not `r2.dev` — the master switch
+
+Everything else in this section requires it. A custom domain puts downloads behind
+Cloudflare's CDN and unlocks WAF, caching, rate limiting, and bot controls. See
+[step 2 above](#2-expose-it-publicly).
+
+### 2. Edge caching (already configured in code)
+
+Every archive is uploaded with `Cache-Control: public, max-age=3600` (set in
+`editor/app/deploy/buildDeployCore.ts`, constant `ARCHIVE_CACHE_CONTROL`). Served
+through a custom domain, Cloudflare's CDN caches each archive at the edge, so repeated
+downloads of the same file are served from cache and **never hit R2** — they don't
+count as Class B ops and cost nothing. This is the single most effective cost defense:
+a flood against one file mostly just warms one cache entry.
+
+To make caching aggressive, add a **Cache Rule** (dashboard: **Caching → Cache Rules →
+Create**):
+
+- **When:** `URI Path` contains `/archives/`
+- **Then:** *Eligible for cache*, **Edge TTL → Override → 1 day** (or longer).
+
+Archive filenames are stable and overwritten in place on re-deploy, so if you raise the
+TTL a lot, **purge the cache on deploy** (dashboard **Caching → Purge**, or
+`wrangler`/API) so updated archives aren't served stale. The 1-hour `Cache-Control`
+default already bounds staleness without a purge; tune `ARCHIVE_CACHE_CONTROL` up if
+your archives rarely change.
+
+### 3. Rate limiting — the hard backstop
+
+A Rate Limiting rule caps how fast any single client can pull archives, stopping a
+flood that misses cache. Dashboard: **Security → WAF → Rate limiting rules → Create**
+(the free plan includes one rule):
+
+- **When incoming requests match:** `URI Path` contains `/archives/`
+- **Rate:** e.g. **20 requests per 1 minute** per client IP
+- **Then:** *Block* for 10 minutes (or *Managed Challenge*).
+
+Tune the threshold to real usage — legitimate users download a handful of files, not
+dozens per minute.
+
+### 4. Bot Fight Mode + WAF managed rules
+
+Dashboard: **Security → Bots → Bot Fight Mode** (free). Blocks the low-effort scripted
+abuse that makes up most of this traffic. The free **WAF managed ruleset** adds a
+baseline of protection at no cost.
+
+### 5. Hotlink protection (optional)
+
+Stops other sites embedding your archives and spending your ops budget serving their
+audience. A WAF custom rule (**Security → WAF → Custom rules**):
+
+- **When:** `URI Path` contains `/archives/` **and** `Referer` does not contain your
+ domain **and** `Referer` is not empty
+- **Then:** *Block*.
+
+(Allow an empty `Referer` so direct clicks and privacy-conscious browsers still work.)
+
+### 6. Billing / usage alerts
+
+R2 has no hard spend cap, but Cloudflare **Notifications** (dashboard: **Notifications
+→ Add**) can email you when R2 storage or Class A/B operations cross a threshold —
+cheap insurance so nothing surprises you.
+
+### What to skip
+
+**Signed URLs / token-gated downloads via a Worker** are the heavyweight option. Given
+egress is free and caching neutralizes the ops cost, they're overkill for *cost*
+defense — they add a Worker, key management, and friction for legitimate users. Only
+reach for them if you want *access control* (private archives), not cost control.
+
+---
+
+## Cost expectations
+
+Measured across all sites in this project, total compressed archives are **≈ 2–4.5 GB**.
+Against R2's free tier:
+
+| Resource | Free tier / month | This project's usage |
+|---|---|---|
+| **Egress / bandwidth** | unlimited, **$0** | irrelevant — no egress charge exists |
+| **Storage** | 10 GB-month | ~2–4.5 GB — comfortable |
+| **Class A ops** (writes) | 1,000,000 | ~one PUT per archive per deploy — negligible |
+| **Class B ops** (reads) | 10,000,000 | mostly absorbed by CDN cache |
+
+The realistic bill for hosting these archives is **$0**. The configuration above exists
+to keep it that way under adversarial traffic, not because normal usage is close to any
+limit.
diff --git a/README.md b/README.md
@@ -94,3 +94,7 @@ pnpm --filter yt-dlp-transcript-common exec tsx bin/verify-transcripts.ts --chan
The export app reads from the LMDB-backed paginated JSON in `export/public/{summaries,transcripts}/` and renders a search/filter UI as a static `out/` directory. It's the read-only public face of the data.
The export build uses `transpilePackages: ["yt-dlp-transcript-common"]` so editing common is hot-reloadable without a separate build step. `lmdb`, `msgpackr`, and `msgpackr-extract` are listed in `serverExternalPackages` so Next.js doesn't try to bundle them.
+
+## Deploying to Cloudflare
+
+The site deploys to Cloudflare Pages; download archives too large for Pages' 25 MB per-file limit overflow to Cloudflare R2. For R2 setup and the Cloudflare configuration that keeps public archive downloads from being abused to run up costs, see [DEPLOY_CLOUDFLARE.md](DEPLOY_CLOUDFLARE.md).
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -4,7 +4,7 @@
- **Search results scroll smoothly again on large result sets.** The results list windows one card per matching video (only the on-screen cards are mounted), but each visible card and every one of its hit rows was re-rendering on *every* scroll frame — and each hit row re-ran its `<mark>` highlighting, so a single video with hundreds of hits meant hundreds of redundant highlight passes per frame while scrolling. The result cards and individual hit rows are now memoized so an unchanged card/row is skipped during scroll, and opening the modal on a hit only re-renders the two rows whose highlight state actually changes. No visible/behavioral change — same DOM, same results, just far less work per frame. See `common/components/TranscriptSearch.tsx` (`ResultCard`/`HitRow` memoization, `openWithMode` stabilized via `useCallback`).
- **The monitor widget gains a needs-work channel list, more interaction buttons, and an in-place settings gear.** Three additions, all driveable from the widget builder. **(1) A "Needs work" list** (URL flag `act=1`) — a compact, per-channel worklist of videos to download (`↓ N`) or transcribe (`✎ N`), reusing the same `loadActionableSummary` that powers the `/actionable` page via a new `/api/widget/actionable` route; it polls on a 15s floor (the backlog changes on job completions, not seconds) and caps at 6 channels with a `+N more` line. **(2) More interactions** behind the existing `controls=1` switch: each needs-work row gains the same per-channel **Download missing** / **Transcribe pending** buttons as the actionable page (reusing `InlineActionButton`), and the controls row adds **Retry all failed** alongside Pause/Resume + Drain. **(3) An in-place settings gear** (on by default; URL flag `gear=0` to hide, or a **Show settings gear** builder checkbox) — clicking it opens the builder's own form *inside the widget window*, so a pinned widget can be reconfigured live without opening the builder page; edits apply immediately and mirror into the address bar via `history.replaceState`, so a reload preserves them and the link stays copyable. The builder form is extracted into a shared `WidgetConfigForm` used by both the builder and the overlay, and the widget's poller now fetches immediately on (re)subscribe instead of after one interval, so newly-enabled sections render at once. Existing links render unchanged (the two new flags default to their old behavior; the gear is the one new default-visible affordance and is read-only — it mutates no server state). See `editor/app/widget/lib/config.ts`, the new `editor/app/widget/components/WidgetConfigForm.tsx` and `editor/app/api/widget/actionable/route.ts`, `editor/app/widget/components/{MonitorWidget,WidgetControls}.tsx`, `editor/app/widget/builder/components/WidgetBuilder.tsx`, and `editor/e2e/widget.spec.ts`.
- **Every site build now bundles downloadable per-channel transcript & live-chat archive zips.** The archive builders (per-channel `<slug>.zip` / `<slug>.live_chat.zip`) previously only ran as standalone actions that wrote to a non-served directory; now `compose-site` generates them for the site's own channels straight into the served `public/archives/` and writes a `manifest.json` (sizes + counts) that the site's new **Downloads** page reads. **`zip` is now the default archive format** everywhere (was `tar.gz`), and the Build page's format help text tracks the selected format. Generation is **on by default with three opt-out levels**: a global **Generate archive zips on build** toggle in Settings, a per-site **Generate archive zips** toggle (plus an optional **Archive size cap (MB)**) on the site's page, and a per-build **Skip archive zips** checkbox on the Build and Build & Deploy controls (`BUILD_ARCHIVES=0`). See `common/bin/compose-site.ts` (`composeArchives`), `common/controller/archive{Transcripts,LiveChat}.ts` (new `outDir` option), `common/lib/archiveOptions.ts` (default + manifest types), `common/lib/{site,settings}.ts` (opt-out flags), and `editor/app/{deploy/buildDeployCore.ts,build/buildAction.ts,deploy/components/Build{Export,Deploy}Button.tsx,sites/components/SiteForm.tsx,settings/components/SettingsForm.tsx}`.
-- **Oversize archives now overflow to Cloudflare R2 instead of being dropped.** A single file over 25 MB breaks a Cloudflare Pages deploy, so a channel zip over the cap (default 25 MB; `0` = no cap) used to be removed from what's served and flagged `oversize`. Now, when **Archive overflow storage** is configured in Settings (an R2 **bucket** + its **public URL**), `compose-site` stages each oversize archive to `export/.r2-staging/<siteId>/` and records its future public URL in the manifest; the deploy step then uploads it via `wrangler r2 object put <bucket>/<siteId>/archives/<file>` (reusing the host's wrangler auth) **before** the Pages deploy, so the Downloads page links straight to R2. With no bucket configured the old drop-and-flag behavior is unchanged (uploads run only in the editor's Deploy / Build & deploy actions, not a raw `pnpm deploy`). The **combined "whole site" archives were removed** — they duplicated the per-channel content and were always the first to blow the cap. See `common/lib/settings.ts` (`archiveStorage`), `common/bin/compose-site.ts` (staging + manifest `url`), `common/lib/archiveOptions.ts` (`ArchiveManifestEntry.url`), and `editor/app/deploy/buildDeployCore.ts` (`runArchiveUploadIntoLog`) / `deploy/deployAction.ts` / `build/buildAction.ts`.
+- **Oversize archives now overflow to Cloudflare R2 instead of being dropped.** A single file over 25 MB breaks a Cloudflare Pages deploy, so a channel zip over the cap (default 25 MB; `0` = no cap) used to be removed from what's served and flagged `oversize`. Now, when **Archive overflow storage** is configured in Settings (an R2 **bucket** + its **public URL**), `compose-site` stages each oversize archive to `export/.r2-staging/<siteId>/` and records its future public URL in the manifest; the deploy step then uploads it via `wrangler r2 object put <bucket>/<siteId>/archives/<file>` (reusing the host's wrangler auth) **before** the Pages deploy, so the Downloads page links straight to R2. With no bucket configured the old drop-and-flag behavior is unchanged (uploads run only in the editor's Deploy / Build & deploy actions, not a raw `pnpm deploy`). The **combined "whole site" archives were removed** — they duplicated the per-channel content and were always the first to blow the cap. Each R2 upload also now sets `Cache-Control: public, max-age=3600` so a Cloudflare custom domain caches downloads at the edge — the main defense against download-abuse cost (R2 egress is free; only origin reads are billable, and cached hits skip the origin). New **[DEPLOY_CLOUDFLARE.md](../DEPLOY_CLOUDFLARE.md)** documents the full R2 setup plus the Cloudflare custom-domain / caching / rate-limiting / bot config for instance operators. See `common/lib/settings.ts` (`archiveStorage`), `common/bin/compose-site.ts` (staging + manifest `url`), `common/lib/archiveOptions.ts` (`ArchiveManifestEntry.url`), and `editor/app/deploy/buildDeployCore.ts` (`runArchiveUploadIntoLog`, `ARCHIVE_CACHE_CONTROL`) / `deploy/deployAction.ts` / `build/buildAction.ts`.
- **The Duplicates page is now a per-site toggle and hides itself when empty.** Each site's editor page gains a **Show the Duplicates page** checkbox (on by default). `compose-site` writes the site-filtered `duplicates.json` only when the toggle is on *and* there's at least one in-scope cluster, and the export Header keys its Duplicates nav link off a new `hasDuplicates()` — so the link and page disappear both when a site opts out and when it simply has no detected duplicates. See `common/lib/site.ts` (`duplicates` flag), `common/bin/compose-site.ts` (gated write), `export/app/lib/duplicates.ts` (new), `export/app/components/Header.tsx`, and `editor/app/sites/{components/SiteForm.tsx,actions.ts}`.
- **You can now change a channel's slug (its id) — deliberately, from the Danger zone.** A channel's slug *is* its on-disk directory name (`transcripts/channels/<slug>/`), so it used to be fixed at creation ("Slug is fixed once a channel is created"). A new **Rename** form in the channel's Danger zone lifts that: enter a new slug and **type the current slug to confirm** (same friction as delete), and the rename is blocked while the channel has running/queued jobs (the in-memory registry keys by slug). Because the slug is a directory name, the rename does a **full migration** of every slug-keyed store so nothing silently breaks: it moves the channel dir (config, data, playlist, snapshot, shards, failed lists) **and** the saved-video store dir — rewriting each `saved-video.json` pointer's absolute `dir` so persisted source videos still resolve — then retargets every site.json membership, the sync scheduler's per-channel backoff state, and any job bookmarks. The two filesystem moves run first and roll back on failure; the metadata updates that follow are atomic and best-effort (surfaced as warnings). Renaming **changes the channel's public URL** (the old one 404s), which the form warns about. The slug grammar is also now validated on create. See `common/controller/renameChannel.ts`, `common/controller/channels.ts` (`isValidChannelSlug`), `common/lib/savedVideo-server.ts` (`rewriteSavedVideoDir`), `common/jobs/bookmarks.ts` (`renameChannelInBookmarks`), `editor/app/channels/{actions.ts,components/RenameChannelForm.tsx,[slug]/page.tsx}`, and `editor/e2e/channel-rename.spec.ts`.
- **New Queue diagnostics page (`/jobs/queue`): see & force-release stuck jobs.** The job system has two sources of truth that can drift — the registry owns each job's `status`, the scheduler owns the running SLOT per queue. A cancel that never finalizes (a child that ignored SIGTERM, a crashed finalizer) leaves a job "cancelled" in the registry while the scheduler still marks its slot running, silently blocking every job behind it on that queue — and the Active Jobs page hides it (it filters to running/queued). The new **Queue** page reconciles the two: it builds from the **scheduler** as the source of truth for slots, cross-checks each against its registry record, and flags a running head as **stuck** when the record is terminal-but-holding-slot, evicted, or (softer) a live job idle past 10 minutes. It **auto-heals** the hard cases on every view/poll (frees terminal/evicted slots), shows a health strip (active queues, running, queued, **stuck**, workers), per-queue cards with the held-for duration / PID (`kill -9` hint) / last log line, and a **Force-release** button per slot (SIGKILLs the child and frees the slot unconditionally) plus a **Reap all stuck** action. Force-release is also available on any running job in Active Jobs, and Active Jobs links to Queue with a stuck-count badge. See `common/jobs/registry.ts` (`forceRelease`), `editor/app/jobs/queue/*`, `editor/app/jobs/{actions.ts,components/ForceReleaseJobButton.tsx}`, and `editor/e2e/queue.spec.ts`.
diff --git a/editor/app/deploy/buildDeployCore.ts b/editor/app/deploy/buildDeployCore.ts
@@ -85,6 +85,14 @@ function archiveStagingDir(siteId: string, paths: Paths): string {
);
}
+// Cache-Control set on every uploaded archive. Served through a Cloudflare custom
+// domain, this lets the CDN absorb repeated/abusive downloads at the edge instead
+// of hitting R2 (each origin GET is a billable Class B op), which is the main cost
+// defense for public archives — see DEPLOY_CLOUDFLARE.md. 1h balances flood
+// absorption against re-deployed archives (stable filenames, overwritten in place)
+// going stale; raise it if your archives rarely change.
+const ARCHIVE_CACHE_CONTROL = "public, max-age=3600";
+
// Upload this site's staged oversize archives to the configured R2 bucket, so the
// remote URLs the served manifest points at actually resolve. Reuses the host's
// wrangler auth (process.env), same as the Pages deploy. No-op (returns 0) when
@@ -126,6 +134,7 @@ export async function runArchiveUploadIntoLog(
`${bucket}/${key}`,
`--file=${path.join(stagingDir, file)}`,
"--content-type=application/zip",
+ `--cache-control=${ARCHIVE_CACHE_CONTROL}`,
"--remote",
],
cwd: paths.exportDir,