Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit c398971931039cf55c7b7395507118cda4d3bd86
parent b28e844cba464b6bd24b850da0a13a172077583f
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Tue,  6 Oct 2026 10:50:37 -0400

plans: S5 third round — the wrangler check, cloudflare-auth through the preflight, the deploy refusals in the container; the doctor's changelog bullet

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
Meditor/CHANGELOG.md | 2+-
Mplans/release-18.md | 45++++++++++++++++++++++++++++++++++++++++++---
2 files changed, 43 insertions(+), 4 deletions(-)

diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -10,7 +10,7 @@ - **`build site`, `build all` and `deploy site` are aliases of the publish commands** and print what they run: `build site <id>` is `publish index` (skipped with `--nodata`) then `publish build <id> --force`; `build all` is `publish index` then `publish build all --runner auto` (containers when an engine answers, else one site at a time on the host); `deploy site <id>` is `publish deploy <id>`, which now ships the site's own bundle and refuses a site never built that way. `publish build all --runner docker` builds every stale site in containers on a Linux host and refuses with "the docker runner needs an engine on this host" where there is none. - **Substitute your own yt-dlp in Docker.** Point `YTDLP_BIN` at a zipapp you built, or set `YTDLP_SOURCE_HOST_DIR` to a yt-dlp checkout and start with `docker-compose.ytdlp.yml`: the image runs it with its own python, and nothing is rebuilt. Every editor boot logs `yt-dlp: <path> <version> (image|override)` (`MISSING` when it does not run; the editor still starts), and `YTDLP_AUTO_UPDATE` updates the image's yt-dlp only, warning instead of touching yours. - **The Docker image can publish.** It carries python, `pipx` and a pinned `git-filter-repo`, so the homepage's `/source` mirror builds in the container; `docker-compose.source.yml` mounts your repository read-only for it, and the scrub rules and denylist live in the config volume (`/data/config/archilyzer`). Cloudflare and R2 credentials come from `.env`. Run publish commands with `docker compose exec editor pnpm archilyzer …`, not `run --rm`. The `homepage` service serves a local deploy from the builds volume once there is one. RUNNING_IN_DOCKER.md has a Windows checklist. -- **`archilyzer doctor` checks what a publish needs.** Which yt-dlp runs (the image's, the host's or an override, and whether it runs), whether the Cloudflare token and the R2 keys are set (never their values; R2 only when a bucket is configured), free space for the site bundles, the publish lock (free, held by a running stage, or left by one that is gone — with the command to clear it; never cleared for you), the index stamp's age and which sites were built from an older one, the repository the source mirror reads, the private config dir, and whether this Node is new enough for the pinned wrangler (deploys need 22). +- **`archilyzer doctor` checks what a publish needs.** Which yt-dlp runs (the image's, the host's or an override, and whether it runs), whether the Cloudflare token and the R2 keys are set (never their values; R2 only when a bucket is configured) — judged exactly as a deploy judges them —, the wrangler a deploy runs (the pinned one or your `WRANGLER_BIN`, and that it starts and is the expected major), free space for the site bundles, the publish lock (free, held by a running stage, or left by one that is gone — with the command to clear it; never cleared for you), the index stamp's age and which sites were built from an older one, the repository the source mirror reads, the private config dir, and whether this Node is new enough for the pinned wrangler (deploys need 22). - **A cited moment at the very end of a recording prepares.** Prepare evidence media cuts a clip whose padding runs past the recording's end at the end (the recording's duration from its metadata), where it found no media for the padded span; a span that starts past the end is still refused. report-to-video keeps its strict rule. - **Exporting a changed report records a new revision of it.** `reports export` (and **Export reports** on a site's Reports tab, and the end of a prepare) commits a revision to the report's own git history, `sites/<site>/reports/<id>/history-git/`, whenever its `report.json` changed since the last one: the `report.json`, its Markdown export and the checksums of every export file, with a message of `Revision N` and a summary of the change. A re-export of an unchanged report records nothing. The commits carry the site's name and a `noreply@<site>.invalid` address with dates in UTC, never your git name, email or time zone. The Reports tab shows each report's revision, its commit and the last change under **Exports**, and the site's next build publishes the history. Add `history-git/` to the corpus repository's `.gitignore`. - **archive.org files come over BitTorrent when possible, else straight from archive.org — never through yt-dlp.** The chosen file of an archive.org import is fetched from the item's own torrent (`<identifier>_archive.torrent`, which lists archive.org as a web seed, so other peers take load off archive.org) with aria2c, only that file of the item, and seeded afterwards for 10 minutes or to a ratio of 1, whichever comes first; the log shows "torrent: <file> (n of m pieces, peers p, web seed yes)" and "seeding 10 min…". With no aria2c, a torrent that does not carry the file, or no progress for 5 minutes, it is downloaded directly from `archive.org/download/…` instead (resumable, backing off on 429/503), and the log says "fell back to direct download: <reason>". Every file is checked against archive.org's sha1/md5: a mismatch is downloaded once more directly, a second one fails the record. The record is written from the item's metadata: `metadata.info.json` with the file's page, the canonical id, the duration ffprobe measures and archive.org's playable copies of the file, the `archiveorg.json` provenance (a mirror's original title, date and uploader), and `audio.<fmt>` — an audio file already in the channel's format is used as is, anything else goes through the app's audio extraction, a video kept in the saved-video store when the channel keeps sources. An .avi/.mpeg/.flac/.wav original is fetched as archive.org's mp4 or mp3 of it. aria2c runs in its own process group: cancelling the job stops it and everything it started, and it stops itself if the editor exits. New settings block `archiveOrg` (`torrent`, `seedMinutes`, `seedRatio`, `stallMinutes`, `maxPeers`, `maxDownloadKiBps`, `maxUploadKiBps`), `ARIA2C_BIN`, an aria2c row in `archilyzer doctor`, and `aria2` in the runtime Docker images. diff --git a/plans/release-18.md b/plans/release-18.md @@ -1065,9 +1065,48 @@ nothing left): - `publish status` and `publish now` (RUNNING_IN_DOCKER.md names both) are S3's rows, not on this branch yet: `archilyzer: unknown command "publish status"` here. -**Third round** (after S2): the `wrangler` check (`wranglerBin` / `WRANGLER_MAJOR`), `cloudflare-auth` -through `cloudflareCredentialProblem` + `wranglerOAuthConfigFiles`, the envVars `readBy` touch-ups and -dedupe, the no-token and bogus-token deploy smoke. +**Third round** (S2 merged: `r18/integration` `4f3daeea`, a fast-forward of this branch; `pnpm install +--frozen-lockfile` brought in the pinned wrangler 4.147.0) + +| Commit | What | +|---|---| +| `056dfac9` | `doctor:` **`publish/wrangler`** — `wranglerBin(paths, env)` (the pin, or `WRANGLER_BIN`, named when set): not there → a note (a warning when a site names a project) naming `pnpm install`, a missing override FAILS; not executable → a warning; run with `--version`, whose major must be `WRANGLER_MAJOR` (else a warning), and a binary that does not start (Node below its floor) warns with its first stderr line. **`wrangler --version` writes a debug log under `~/.config/.wrangler/logs` on every run**, so the doctor sends it to a temp dir (`WRANGLER_LOG_PATH`) it removes — the doctor's header says so. **`cloudflare-auth`** grades with `cloudflareCredentialProblem` over `wranglerOAuthConfigFiles` (located, never read) and quotes the deploy's own sentence, so the two cannot disagree; the doctor's own login-path helper is gone. envVars `readBy`: `WRANGLER_BIN` adds `doctor.ts`; `ARCHILYZER_COMMIT`/`BRANCH` name `stageBodies.ts`'s `imageBuildFacts` (the duplicate rows were already resolved at the S2 merge). 1 test (6 states; the log dir is under the OS temp dir and gone after; nothing under HOME) | +| this one | `plans:` this table and the smoke; the doctor's changelog bullet names the wrangler check | + +Gates at `056dfac9`: tsc (all workspaces) clean (167 s). **common: 3,281/3,282** — 1 failed, +`controller/fetchPosts.test.ts` "a drain mid-page waits for the page's cursor…", a timing case, run while +the image built beside it; the file alone afterwards: 8/8 (S5 touches nothing it imports). **Editor unit:** +142/142. `doctor.test.ts` + `buildImage.test.ts` 35/35 — **the wrangler-floor drift test now runs** +(wrangler installed): Node 22 ≥ `>=22.0.0`, green, nothing skipped. + +**Compose smoke, third round** (`--target runtime` rebuilt from `056dfac9`, 351 s; the image is now +**2.00 GB** — wrangler and its workerd in `node_modules`; the same fixture, `s5site` given a +`cloudflareProject` so it is deployable; editor + `site`; `down -v` after): +- In the container: `node --version` v22.23.3; `common/node_modules/.bin/wrangler --version` and `node + common/node_modules/wrangler/bin/wrangler.js --version` both **4.147.0**, exit 0. No OAuth login files + (`/root/.wrangler/…`, `/root/.config/.wrangler/config/default.toml` absent), no token. +- `publish index` / `build s5site` / `deploy s5site --to local` → exit 0; `deployed.json` written + (sha256 `086ca1ee23883d29…`, mtime noted). +- `publish deploy s5site --preview smoke`, no credential → **exit 1**, `[deploy] REFUSED — no Cloudflare + credentials: set CLOUDFLARE_API_TOKEN in .env (or run \`wrangler login\` on this machine). Nothing was + sent to Cloudflare.`; wrangler never started; `deployed.json` byte- and mtime-identical. +- The same with `CLOUDFLARE_API_TOKEN=bogus` → exit 1, `deployed.json` identical — but the log ends + `[deploy] FAILED — wrangler exited 1.`, **not** the REFUSED sentence: Cloudflare answers a token that is + not token-SHAPED with `Invalid request headers [code: 6003]` / `Invalid format for Authorization header + [code: 6111]`, which `pagesDeploy.ts`'s `AUTH_FAILURE_RES` does not list. One more request with a + well-formed wrong token (40 random characters): Cloudflare says `Invalid access token [code: 9109]` + and the log ends on the exact **`[deploy] REFUSED by Cloudflare — the API token was not accepted`**, + exit 1, no `deployed.json`. Two requests to Cloudflare in all. Finding for S2's file (not changed + here): add codes 6003 and 6111 to the classifier — a mistyped or truncated token in `.env` is the likely + real case. +- `doctor` with no token: `node` ok "deploys: >= 22.0.0, wrangler 4.147.0"; `cloudflare-auth` WARN, quoting + the preflight's sentence, naming `s5site`; `wrangler` ok "/repo/common/node_modules/.bin/wrangler 4.147.0 + (the pin in common/package.json)"; `export-builds`, `publish-lock`, `index-stamp` ok. With the token set: + `cloudflare-auth` ok "is set (never printed)". The count of wrangler log files under the container's + `/root/.config/.wrangler/logs` was 3 before the doctor and 3 after (the deploys wrote those). + +S5 is complete with this round. Left for the rollout, as recorded above: building `runtime-vulkan` and +`runtime-cuda` once. ## Rollout