commit 02778c25f480a3ae837cab28dd5ac32b1a38f299
parent 8b39d63fb75f67c5ee216df409ea3b08faec7e0a
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Thu, 20 Aug 2026 01:06:10 -0400
docs: make the README an entrypoint, and say what you can do with no corpus
The README opened with the monorepo layout and spent its middle on e2e shard
wall-times — a contributor document wearing an entrypoint's name. It is now
organised around what someone actually arrives wanting: use an archive that
already exists, run one, publish your own channel, or work the corpus with
Claude Code. The internals it used to lead with move to CONTRIBUTING.md, intact
(including the measured e2e table and the note about why test counts must be
quoted with times).
The path that was missing entirely: you can pull this repo down with NO corpus,
point the MCP server at a published instance, and have a research environment
over tens of thousands of transcripts. Three things ship for that — the server,
the sweep discipline the *_plan tools hand back (mcp/src/instructions.ts), and
the two tracked slash commands — and none of it needed saying anywhere before.
From there the loop that makes an archive worth more than a search box:
transcripts give you an exact second, so yt-dlp --download-sections fetches the
moment rather than the movie.
Windows gets a real section rather than a footnote, including installing Claude
Code inside WSL2 alongside the repo, and the two mistakes that actually hurt
(the repo under /mnt/c, and Windows paths in the MCP registration).
SETUP.md said "there is no public git repository", which the README now
contradicts; it is reworded to be forge-neutral, matching the fact that this
tree carries no .github/ and no provider-specific CI. Also dropped an em dash
from a heading: it slugs differently on GitHub than on renderers that collapse
whitespace, so any anchor to it breaks on some platforms.
AGENTS.md gains what a future session in a fresh clone has to be told: that a
corpus-less clone is a supported way to work, that the MCP server must be
registered as `archilyzer` or the shipped commands silently break, where site
config actually lives (transcripts/sites/<id>/site.json, key `siteUrl` — not
under channels/, and not `publicUrl`, which is only the form's label), and the
live instances.
Verified against the live corpus: jeralyzer.pages.dev serves 30 channels /
30,886 videos, /corpus.json and /llms.txt both 200, CORS *.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Diffstat:
| M | AGENTS.md | | | 114 | +++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ |
| A | CONTRIBUTING.md | | | 169 | +++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ |
| M | README.md | | | 510 | +++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++---------------- |
| M | SETUP.md | | | 14 | ++++++++------ |
4 files changed, 699 insertions(+), 108 deletions(-)
diff --git a/AGENTS.md b/AGENTS.md
@@ -19,6 +19,120 @@ names the offending pid, instead of silently driving another session's server. B
See [WORKTREES.md](WORKTREES.md) for the port scheme, the queue, and the shared-data caveat.
+# Working this repo with no local corpus
+
+A clone of this repo with **no `transcripts/` directory at all** is a complete research
+environment over any *published* archive. That is a supported, first-class way to use
+it — not a degraded one — and it is how most people who pull the repo will start.
+
+Register the MCP server against a public instance and the corpus is readable over HTTP:
+
+```sh
+claude mcp add archilyzer \
+ --env TRANSCRIPT_SITE_URL=https://jeralyzer.pages.dev \
+ -- pnpm -C "$PWD" --filter yt-dlp-transcript-mcp exec tsx src/index.ts
+```
+
+`TRANSCRIPT_HUB_URL` federates several sites; `TRANSCRIPT_LOCAL_DIR` reads a local
+build off disk. The server never writes to an archive.
+
+**Register it as `archilyzer`.** `.claude/commands/{ask,sweep}.md` are tracked in git
+and call `mcp__archilyzer__ask_plan` / `mcp__archilyzer__sweep_plan`. That tool name
+embeds the server name as registered on this machine, so any other name silently breaks
+both commands. Either use `archilyzer` or edit the `mcp__<name>__` prefix in those two
+files.
+
+**The corpus discipline lives in `mcp/src/instructions.ts`, not in prose.** Citation
+format, extractor budget, base-link expansion and the coverage rule are returned *by
+the `*_plan` tools as a plan an agent then follows*. Fix guidance there, not by adding
+another markdown file that will drift from it.
+
+Use `/ask` for a question answered in the conversation and `/sweep` for a cited report
+written to a file.
+
+## Clips and report-to-video
+
+The high-value loop for a repo with no corpus: point the MCP at a public instance, ask
+about a subject, then use `yt-dlp --download-sections` to pull **just the cited
+seconds** rather than whole videos. Searching text first is what makes fetching cheap.
+
+`scripts/report-to-video/` renders a cited sweep report to an mp4. What it needs:
+
+- **yt-dlp** — clip media is fetched over the network per clip. Local media does not
+ help: most archived video dirs hold captions and metadata, not video.
+- **ffmpeg / ffprobe**, and **ImageMagick with Pango** for the timeline and cards.
+- **Cue end times**, for widening a clip from a cue span to a whole sentence. A cue
+ boundary is where the caption line wrapped, so cutting there ends mid-thought.
+ Neither a report nor an MCP snippet carries an end.
+
+The scripts currently read `transcript.cues.json` from local disk, so point
+`CHANNELS_DIR` at a channels directory holding the cited videos; both otherwise default
+it to an absolute path on the original author's machine.
+
+**That is a tooling limit, not a data one.** Published archives already serve
+`cues: [{ start, end, text }]` on their transcript pages — documented under
+`shardScheme` in `/corpus.json`, and verified 2026-08-20 against
+`https://jeralyzer.pages.dev/transcripts/chrissie-mayr/page-0000.json`
+(`{"start":8.12,"end":10.31,…}`). Teaching `resolve-windows.mjs` and `build-video.mjs`
+to fetch windows over HTTP — reusing the shard walk in `mcp/` — is what would let the
+whole video path run with **no local corpus at all**.
+
+Fetched clips are cached on disk under the report's `out/` (`clips-raw/`, `segments/`,
+`cards/`, `<slug>.mp4`) and reused on rebuild. The per-report `video.manifest.json` is
+the regeneration source of truth, not the report.
+
+As of 2026-08-20 this tooling is on the `feat/umtool-projects` branch, not on `main`.
+
+# The corpus, and where the live sites are configured
+
+`transcripts/` is **its own git repo**, gitignored by the workspace (`/transcripts` in
+`.gitignore`). It holds real production data — do not treat it as scratch.
+
+| Path | What it is |
+|---|---|
+| `transcripts/channels/<slug>/` | One channel: `config.json`, `playlist`, `archive`, `snapshot.json`, and `data/<videoId>/` holding media, transcripts and sidecars. |
+| `transcripts/sites/<id>/site.json` | **Per-site config, including the public URL.** This is where deployed-site facts live — *not* under `channels/`. |
+| `transcripts/index.mdb` | The LMDB transcript index. Key-only range scans over its `byChannel` sub-DB are cheap; see `common/controller/recencyIndex.ts`. |
+| `transcripts/saved-videos/` | Persisted source-video store. |
+| `transcripts/search-aliases.json`, `duplicates*.json` | Corpus-wide curated data. |
+
+**The public-URL key in `site.json` is `siteUrl`.** The editor form labels the field
+"Public URL", so grepping for `publicUrl` finds the UI hint and misses the data.
+
+## The live instances
+
+Read these out of `transcripts/sites/*/site.json` rather than hardcoding them — this
+table is a convenience and will drift:
+
+| Site id | Title | Public URL |
+|---|---|---|
+| `jeralyzer` | Jeralyzer | https://jeralyzer.pages.dev |
+| `rekietalyzer` | Rekietalyzer | https://rekietalyzer.pages.dev |
+| `hasanalyzer` | Hasanalyzer | https://hasanalyzer.pages.dev |
+| `anilyzer` | Anilyzer | https://anilyzer.pages.dev |
+| `bonnellyzer` | Bonnellyzer | https://bonnellyzer.pages.dev |
+| `jasolyzer` | Jasolyzer | *(no `siteUrl` — not published)* |
+
+Jeralyzer is the largest and most popular instance (30 channels / 30,886 videos at its
+2026-08-07 build). The project's own site is `https://archilyzer.pages.dev`
+(`PROJECT_URL` in `common/lib/project.ts`).
+
+Every published archive serves a machine contract you can check without installing
+anything — useful for verifying a claim about the export format against reality:
+
+```sh
+curl https://jeralyzer.pages.dev/corpus.json # spec/site/totals/channels + manifest URLs
+curl https://jeralyzer.pages.dev/llms.txt # the same contract as prose
+```
+
+Both are served with `access-control-allow-origin: *`.
+
+## Do not boot a second editor against the real corpus
+
+The editor's `instrumentation.ts` arms the runners on boot, which resumes long-running
+sweeps against production data. To exercise `common/` controllers over the real corpus,
+run them offline with `tsx` instead of starting a server.
+
# Roadmap
Long-running work on the local-AI derived corpus is tracked in [PLAN.md](PLAN.md).
diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md
@@ -0,0 +1,169 @@
+# Working on Archilyzer
+
+This is the developer's side of the project: how the workspace is laid out, how to run
+the test suites, and the internals worth knowing before you change something. If you
+just want to *run* an archive, you want [README.md](README.md) and
+[SETUP.md](SETUP.md) instead.
+
+Archilyzer is MIT-licensed. It is deliberately **forge-neutral**: there is no
+`.github/` directory, no provider-specific CI config, and nothing in the build that
+assumes a particular host. Patches are welcome however you can get them here — a
+mirror, a fork on any platform, or a plain `git format-patch` series.
+
+## Workspace layout
+
+A pnpm-workspace monorepo. Every package consumes `common/` via `workspace:*`.
+
+| Package | What it is |
+|---|---|
+| `common/` | The shared library — data layer, controllers, components, types. Nearly all the logic lives here. |
+| `editor/` | The private admin app (Next.js, port 3001): channels, the download/transcribe pipeline, jobs, settings, build + deploy. |
+| `export/` | The public static site (`output: "export"`). Read-only, no runtime. |
+| `homepage/` | The project's own site: marketing, docs, downloads, cross-site stats at `/stats/`. |
+| `mcp/` | The MCP server that exposes a published archive to LLM clients. See [mcp/README.md](mcp/README.md). |
+| `umtool/` | A local bench for the report-to-video work (port 3050). Not part of an archive install. |
+
+Your corpus lives at `<repo>/transcripts/` and is **its own git repo**, untouched by
+the workspace. That separation is deliberate: updating the software never touches
+your archive, and a fresh checkout has no data.
+
+## Everyday commands
+
+```bash
+pnpm install
+pnpm dev:editor # editor at http://localhost:3001
+pnpm build # build the static site under export/out/
+pnpm start:export # serve export/out/ at http://localhost:3000
+pnpm build:index # rebuild the search index in-process
+```
+
+`pnpm lint` at the root only lints `export/`. **The editor has no eslint config**, so
+`pnpm lint` inside `editor/` always fails — typecheck it with
+`pnpm exec tsc --noEmit` instead.
+
+## Parallel checkouts
+
+`pnpm wt add <branch>` creates a git worktree with its own non-colliding port block;
+`pnpm wt list` shows them. `pnpm dev:editor` auto-assigns ports per worktree, so
+several dev servers run side by side. See [WORKTREES.md](WORKTREES.md) for the port
+scheme and the shared-data caveat.
+
+## Tests
+
+Unit tests are plain `node:test` files run through `tsx`:
+
+```bash
+node_modules/.bin/tsx --test common/jobs/autoQueuePolicy.test.ts
+node --test scripts/*.test.mjs
+```
+
+Playwright covers the editor UI from `editor/e2e/`. Two ways to run it:
+
+- **`pnpm e2e`** — sequential, on the host, against `pnpm dev:test` (Next.js dev mode)
+ on port 3001. No Docker. Slow, but it picks up source edits without a rebuild and
+ gives you one readable log. Use this while iterating.
+- **`pnpm e2e:sharded`** — containerized, splitting the suite across N shards via
+ Playwright's `--shard=i/N`. Each shard runs in its own copy of a test image
+ (`Dockerfile.test`) that bakes the monorepo plus a prebuilt editor (`next start`).
+ `playwright merge-reports` then recombines the per-shard blob reports into
+ `editor/playwright-report/index.html`. Use this to verify a branch.
+
+Run `npx playwright install chromium` once before your first e2e run.
+
+**e2e does not run in parallel with itself.** Every entry point takes a
+machine-global lock, so one suite runs at a time and the rest wait. If you see
+`waiting for the e2e queue — held by …` and it sits there, that is working as
+intended, not a hang. Bypasses: `E2E_QUEUE=0`, `E2E_PORT_CHECK=0`,
+`E2E_QUEUE_TIMEOUT=<seconds>`.
+
+### Sharded runs in detail
+
+The npm script rebuilds the image first (layer cache covers unchanged deps, so
+repeat builds are seconds), then runs `scripts/run-sharded-e2e.mjs`.
+
+Shard count defaults to `min(max(2, cpus/2), 8)`. Override with
+`SHARDS=N pnpm e2e:sharded` or `node scripts/run-sharded-e2e.mjs --shards N`.
+
+**Retries default to 0**, so a sharded run reports the same failures a sequential one
+does. (The containers set `CI=true` because `playwright.config`'s
+`reuseExistingServer: !process.env.CI` needs it — each shard must start its own
+servers — but that would otherwise also switch on `retries: 2` and quietly hide flaky
+tests. The runner passes an explicit `--retries=0` to override it.) Raise it
+deliberately with `--retries N` when you actually want to measure flakiness.
+
+Anything else on the command line is forwarded to `playwright test` in every shard:
+
+```bash
+pnpm e2e:sharded -- --grep "digest"
+node scripts/run-sharded-e2e.mjs --shards 4 e2e/deploy-page.spec.ts
+```
+
+After merging the blob reports the runner prints a combined `N passed / M failed`,
+not just per-shard exit codes.
+
+### Wall-time comparison
+
+**Always quote the test count with the time.** An earlier version of this table was
+measured at 90 tests and stayed in the README unchanged while the suite grew to 392 —
+which made its "104 s sequential" wrong by more than an order of magnitude, invisibly.
+
+Measured 2026-07-30 on an 8-CPU machine, **392 tests**, `--retries=0`:
+
+| Mode | Wall time | Notes |
+| --- | --- | --- |
+| `pnpm e2e` (sequential) | **23.8 min** (1428 s) | one worker, `next dev` |
+| `pnpm e2e:sharded --shards 4` (image cached) | **7.1 min** (424 s) | 4 shards × `next start`, **~3.4× speedup** |
+| `docker build` (cached layers) | ~40–110 s | one-off, paid on dependency or source changes |
+
+Caveat on precision: the box was **not** quiet during these runs (an unrelated project
+was running its own Playwright suite at load ~15–23). A second sequential run under
+heavier load took 29.7 min for the same result, so treat these as an envelope rather
+than a constant.
+
+4 shards were used to bound memory, since each shard runs its own editor + export
+servers. One spec legitimately disagrees between the two routes: `widget.spec`'s "no
+buttons" is red under `next dev` (which injects a Dev Tools button) and green under
+`next start`.
+
+**The editor e2e suite runs in dev mode and therefore never prerenders.** A change to
+a layout or a client component is not verified until `pnpm build` passes too.
+
+## CLI shims
+
+The same controllers the editor uses are exposed as terminal shims under
+`common/bin/`, which is how you drive the pipeline headlessly or from cron:
+
+```bash
+pnpm --filter yt-dlp-transcript-common exec tsx bin/build-index.ts
+pnpm --filter yt-dlp-transcript-common exec tsx bin/transform.ts --channel <slug>
+pnpm --filter yt-dlp-transcript-common exec tsx bin/retry-failures.ts --channel <slug>
+pnpm --filter yt-dlp-transcript-common exec tsx bin/verify-transcripts.ts --channel <slug>
+```
+
+## Internals worth knowing
+
+**This is not the Next.js you may know.** The workspace tracks a recent major and its
+APIs, conventions and file structure differ from older releases. Read the relevant
+guide under `node_modules/next/dist/docs/` before writing app-router code, and heed
+deprecation notices.
+
+- The export build uses `transpilePackages: ["yt-dlp-transcript-common"]`, so editing
+ `common/` is hot-reloadable without a separate build step. `lmdb`, `msgpackr` and
+ `msgpackr-extract` are in `serverExternalPackages` so Next.js does not try to bundle
+ them.
+- The export app reads the LMDB-backed paginated JSON in
+ `export/public/{summaries,transcripts}/` and renders a search/filter UI into a
+ static `out/`.
+- `getPaths()` (`common/lib/paths.ts`) is the single resolver for every path and
+ binary. Nothing should hardcode a location; add an env override there instead.
+- `common/lib/project.ts` holds **product** identity (the name, the project URL) and
+ is deliberately import-free. **Operator** identity — what a given deployment calls
+ itself — lives in settings and per-site config. A string that should change when
+ someone else deploys this belongs in the latter.
+
+## Long-running work
+
+Multi-phase efforts are tracked in [PLAN.md](PLAN.md), with current status and
+verified codebase facts under `plans/`. Read `plans/STATE.md` and `plans/FACTS.md`
+before starting phase work — they are maintained precisely so you do not have to
+re-derive them.
diff --git a/README.md b/README.md
@@ -1,155 +1,461 @@
# Archilyzer
-Self-hosted, searchable video-transcript archives. Archilyzer downloads a channel's back
-catalogue, transcribes it locally, and builds a static site you host yourself.
+**Self-hosted, searchable video-transcript archives.**
-The project site — documentation and the source snapshot — is
-[archilyzer.pages.dev](https://archilyzer.pages.dev).
+Archilyzer takes a channel's back catalogue, downloads it, transcribes it on your own
+hardware, and builds a static website you host yourself. The result is a permanent,
+searchable record of what someone said on video — searchable to the second, and still
+there after the original comes down.
-A pnpm-workspace monorepo, split into five packages:
+It is a program you run, not a service you sign up for. There is no account, no API
+key, no server of ours in the path, and nothing phones home. MIT licensed.
-- **`common/`** — shared library (data layer, controllers, components, types). Consumed by the others via `workspace:*`.
-- **`editor/`** — dynamic Next.js app on port 3001 with admin UIs for channels, the yt-dlp pipeline, the whisper queue, and the build trigger.
-- **`export/`** — static Next.js app (`output: "export"`) that produces the read-only public site.
-- **`homepage/`** — static Next.js app for the project's own site: marketing home, docs, downloads, and a cross-site stats dashboard at `/stats/`.
-- **`mcp/`** — an MCP server exposing a published archive to Claude Code, Claude Desktop, Cursor and other clients. See [mcp/README.md](mcp/README.md).
+Project site: [archilyzer.pages.dev](https://archilyzer.pages.dev)
-Transcripts and per-channel state live at `<repo>/transcripts/` (its own git repo, untouched by the workspace).
+---
-Licensed under the [MIT License](LICENSE).
+## What do you want to do?
-## Requirements
+| I want to… | Start here | Do I need to host anything? |
+|---|---|---|
+| **Search or study an archive that already exists** | [Use a public instance](#1-use-an-archive-that-already-exists) | No |
+| **Archive somebody else's public channel** | [Run your own archive](#2-run-your-own-archive) | Yes — a machine that stays on |
+| **Make my own channel searchable** (transcripts, metadata, a search page) | [Publish your own channel](#3-publish-your-own-channel) | Yes, plus somewhere to serve static files |
+| **Research a corpus with Claude Code** — cited answers and sweep reports | [Use it with Claude Code](#4-use-it-with-claude-code) | No — point it at a public instance |
+| **Pull clips / build a video about a subject** | [Clips and video](#clips-and-video-download-the-moment-not-the-movie) | No — search text, then fetch just the cited seconds |
+| **Work on Archilyzer itself** | [CONTRIBUTING.md](CONTRIBUTING.md) | — |
-- **Node.js ≥ 20.9** (LTS 20 or 22) and **pnpm 9+** (easiest via `corepack enable`) and **git** — needed to install and run the web apps.
-- The download/transcribe **pipeline** additionally needs **yt-dlp**, **ffmpeg**/**ffprobe**, and a transcription backend (whisper.cpp by default). The web apps run without them.
+These are not exclusive. Plenty of people research a public instance for a while
+before deciding to host their own.
-**See [SETUP.md](SETUP.md) for full per-OS install instructions (Linux, macOS, Windows).**
+---
-## Getting started
+## 1. Use an archive that already exists
+
+You do not need to install Archilyzer to use one. Every published archive is a static
+site with a real search UI, and every archive is **machine-navigable by design**:
+
+- **`/corpus.json`** — a machine-readable index describing how to navigate the shards
+ (where each channel's manifest is, how a video id maps to a page, what a record
+ looks like). It is served with CORS headers, so anything can fetch it.
+- **`/llms.txt`** — the same contract as prose, following the llmstxt.org convention.
+
+The most popular public instance is the **[Jeralyzer](https://jeralyzer.pages.dev)** —
+30 channels and 30,886 videos as of its 2026-08-07 build. Point a browser at it and
+search; that is the whole of it. Instances cross-link to each other in their footers,
+so one is a way in to the rest.
+
+You can see the machine contract for yourself without installing anything:
```bash
-pnpm install
-pnpm dev:editor # editor at http://localhost:3001
-pnpm build # build the static site under export/out/
-pnpm start:export # serve export/out/ at http://localhost:3000
-npx playwright install chromium # one-time, before the first e2e run
-pnpm e2e # run the editor's Playwright suite (sequential)
-pnpm e2e:sharded # same suite, sharded across N Docker containers in parallel
+curl https://jeralyzer.pages.dev/corpus.json | head -c 400
+curl https://jeralyzer.pages.dev/llms.txt
```
-## Editor UI
+To let an LLM read an archive properly, run the bundled **MCP server** against it.
+It is a local tool you run yourself, it only reads already-published JSON, and it
+never modifies the archive:
-`pnpm dev:editor` (or `pnpm --filter editor run dev`) starts the editor on port 3001:
+```bash
+# point it at a published instance over HTTP
+TRANSCRIPT_SITE_URL=https://jeralyzer.pages.dev \
+ pnpm --filter yt-dlp-transcript-mcp exec tsx src/index.ts
+```
-- **/channels** — list, create, edit, delete channels. The form covers handling, url, audioFormat, cookiesFromBrowser, and ytdlpExtraArgs.
-- **/channels/<slug>** — pipeline panel (Store playlist / Download from playlist / Sync) and, for `transcribe` channels, a whisper panel (Transcribe missing / Retry failures / Verify).
-- **/build** — Build index (in-process) or Build static export (spawns `pnpm run build` in `export/`). Both stream logs live.
-- **/jobs** — recent jobs, running and archived. Click an id to tail its log.
-- **/settings** — site-settings form persisted to `settings.json`, plus a read-only view of `getPaths()` for env-resolution debugging.
+There is nothing to compile — it runs from source through `tsx`. To register it with
+Claude Code:
-Channels can also sync automatically on a per-channel cadence via a cron heartbeat — see [SCHEDULED_SYNC.md](SCHEDULED_SYNC.md).
+```bash
+claude mcp add archilyzer \
+ --env TRANSCRIPT_SITE_URL=https://jeralyzer.pages.dev \
+ -- pnpm -C /ABS/PATH/TO/this/repo --filter yt-dlp-transcript-mcp exec tsx src/index.ts
+```
-## Configuration & environment variables
+Use `TRANSCRIPT_HUB_URL` instead to federate a whole hub of sites, or
+`TRANSCRIPT_LOCAL_DIR` to read a local build off disk.
-Most configuration lives in the editor's **/settings** page (persisted to `settings.json`); it's optional and falls back to defaults. All paths and binaries resolve through `getPaths()` (`common/lib/paths.ts`) and can be overridden by environment variables. Common ones:
+It gives a client full-text search with clickable second-level citations, complete
+match enumeration (so you can *count* rather than sample), batch transcript reads, and
+metadata including AI chapters where they exist. See **[mcp/README.md](mcp/README.md)**
+for the full tool list, the `mcp.json` form for other clients, and the `/sweep` and
+`/ask` commands.
-| Variable | Default | Purpose |
-| --- | --- | --- |
-| `TRANSCRIPTS_DIR` | `<repo>/transcripts` | Channels, archives, LMDB cache, job logs. |
-| `EXPORT_PUBLIC_DIR` | `<repo>/export/public` | Where the index writes paginated JSON. |
-| `SETTINGS_FILE` | `<repo>/settings.json` | Site-settings file. |
-| `YTDLP_BIN` | `yt-dlp` (PATH lookup) | Binary used by the pipeline. |
-| `WHISPER_BIN` | `whisper-cli` (PATH lookup) | Default transcription binary. |
-| `WHISPER_MODEL` | `~/whispercpp/whisper.cpp/models/ggml-base.en.bin` | Model file passed to whisper-cli. |
-| `PARALLEL_TRANSCRIBE_LIMIT` | `4` | Max parallel transcription jobs. |
+This path needs Node and pnpm. It does not need yt-dlp, ffmpeg, a GPU, or a corpus.
-See **[SETUP.md](SETUP.md#configuration--environment-variables)** for the full list (ffmpeg/ffprobe, rsync, the `chough`/`parakeet` backends, saved-video and sites dirs, and the remote-worker token).
+---
-## Pipeline modes
+## 2. Run your own archive
-The editor's pipeline panel and `common/ytdlp/runYtdlp.ts` support three modes per channel:
+The full pipeline: list a channel, fetch what is missing, transcribe it locally, index
+it, and publish a static site.
-- **Store playlist** — `yt-dlp --flat-playlist --skip-download --print url <url>`. Atomically writes `<channel>/playlist`.
-- **Download from playlist** — reads `<channel>/playlist`, parses `<channel>/archive`, writes `<channel>/playlist.tofetch` containing only URLs whose ID is *not* already archived (this app-side prefilter avoids forcing yt-dlp to re-fetch metadata for known entries on platforms like Odysee), then runs yt-dlp with `--download-archive ./archive --break-on-existing -a playlist.tofetch` plus per-handling args.
-- **Sync** — `yt-dlp --lazy-playlist --download-archive ./archive --break-on-existing <url>`. Quick incremental fetch.
+```
+list → fetch → transcribe → index → compose → publish
+```
-YouTube channels (`handling: "youtube"`) use `--write-auto-subs --skip-download`. Transcribe channels (`handling: "transcribe"`) download audio only (`-f bestaudio -x --audio-format <m4a|mp3|opus>`); a transcription backend transcribes them later via the whisper panel. whisper.cpp is the default; `chough` and `parakeet.cpp` are selectable in **Settings → Transcription** (see [SETUP.md](SETUP.md#transcription-backends-optional)).
+Steps 1–3 can run unattended on a schedule. Transcription is the slow step and the one
+that wants a GPU; nothing is sent to a third-party transcription service.
-## End-to-end tests
+### What it costs you
-Playwright covers the editor UI from `editor/e2e/`. Two ways to run it:
+- **Disk.** The corpus grows to whatever your channels amount to. Point
+ `TRANSCRIPTS_DIR` at a large disk *before* you start — moving it later means moving
+ everything.
+- **Time.** A back catalogue of thousands of recordings is a multi-day first pass.
+ After that it is incremental.
+- **Hosting.** The output is plain files. Anything that serves static files will do.
-- **`pnpm e2e`** — sequential run on the host. Uses `pnpm dev:test` (Next.js dev mode) on port 3001. Simple, no Docker, but slow.
-- **`pnpm e2e:sharded`** — containerized run that splits the suite across N shards in parallel using Playwright's `--shard=i/N` mechanism. Each shard runs in its own copy of a test-specific Docker image (`Dockerfile.test`) that bakes the monorepo plus a prebuilt editor (`next start`). After all shards finish, `playwright merge-reports` recombines the per-shard blob reports into a single `editor/playwright-report/index.html`.
+### Quick start
- The npm script first rebuilds the image (Docker layer cache covers unchanged deps, so subsequent rebuilds are ~seconds), then runs `scripts/run-sharded-e2e.mjs` to orchestrate the shards.
+```bash
+corepack enable # once, so the pinned pnpm version is used
+pnpm install # installs every workspace package
+pnpm dev:editor # the editor, at http://localhost:3001
+```
- Shard count defaults to `min(max(2, cpus/2), 8)`. Override with `SHARDS=N pnpm e2e:sharded` or `node scripts/run-sharded-e2e.mjs --shards N`.
+The editor **starts fine with no data** — a fresh checkout has no corpus, and that is
+expected. Add your first channel from the Channels page.
- **Retries default to 0**, so a sharded run reports the same failures a sequential one does.
- (The containers set `CI=true` because `playwright.config`'s
- `reuseExistingServer: !process.env.CI` needs it — each shard must start its own servers — but
- that would otherwise also switch on `retries: 2` and quietly hide flaky tests. The runner
- passes an explicit `--retries=0` to override it.) Raise it deliberately with
- `--retries N` when you actually want to measure flakiness.
+Then, to build and serve the public site:
- **Anything else on the command line is forwarded to `playwright test` in every shard**, so a
- subset can be sharded too:
+```bash
+pnpm build # static site under export/out/
+pnpm start:export # serve it at http://localhost:3000
+```
- ```bash
- pnpm e2e:sharded -- --grep "digest"
- node scripts/run-sharded-e2e.mjs --shards 4 e2e/deploy-page.spec.ts
- ```
+**[SETUP.md](SETUP.md)** has the full per-OS instructions. Below is the short version.
- After merging the blob reports the runner prints a combined `N passed / M failed`, not just
- per-shard exit codes.
+### Requirements
-### Which route to use when
+Always needed, to install and run the apps:
-Use **`pnpm e2e`** while iterating — it reuses a running dev server, picks up source edits
-without a rebuild, and gives you a single readable log. Use **`pnpm e2e:sharded`** to verify a
-branch: it is several times faster over the whole suite, and it runs against `next start`
-(a real build) rather than `next dev`.
+| Tool | Version | Notes |
+|---|---|---|
+| Node.js | 20.9+ | LTS 20 or 22. |
+| pnpm | 9+ | Easiest via `corepack enable`. |
+| git | any | For the source, and the corpus keeps its own repo. |
-### Wall-time comparison
+Needed only for the pipeline — install what you will actually use:
-**Always quote the test count with the time.** The previous version of this table was measured
-at 90 tests and stayed in the README unchanged while the suite grew to 392 — which made its
-"104 s sequential" wrong by more than an order of magnitude, invisibly.
+| Tool | Needed for |
+|---|---|
+| yt-dlp | Listing, downloading and syncing channels. |
+| ffmpeg + ffprobe | Audio transcode and duration checks. |
+| A transcription backend | Channels you transcribe yourself (whisper.cpp by default; `chough` and `parakeet.cpp` also supported). |
+| rsync | Backing up the saved-video store, if you enable it. |
+| Docker or podman | Only for the parallel multi-site export build. |
-Measured 2026-07-30 on this machine (8 CPUs, **392 tests**, `--retries=0`):
+**The apps run with none of the second table installed.** You just cannot fetch
+anything yet, which makes it easy to try the UI first and commit to the toolchain
+later.
-| Mode | Wall time | Notes |
-| --- | --- | --- |
-| `pnpm e2e` (sequential) | **23.8 min** (1428 s) | one worker, `next dev` |
-| `pnpm e2e:sharded --shards 4` (image cached) | **7.1 min** (424 s) | 4 shards × `next start`, **~3.4× speedup** |
-| `docker build` (cached layers) | ~40–110 s | one-off, paid on dependency or source changes |
+### Linux
-Caveat on precision: the box was **not** quiet during these runs (an unrelated project was
-running its own Playwright suite at load ~15–23). A second sequential run under heavier load
-took 29.7 min for the same result, so treat these as an envelope rather than a constant.
+```sh
+sudo apt install nodejs yt-dlp ffmpeg rsync build-essential # Debian / Ubuntu
+sudo dnf install nodejs yt-dlp ffmpeg rsync gcc-c++ make # Fedora
+sudo pacman -S nodejs yt-dlp ffmpeg rsync base-devel # Arch
+```
-4 shards were used above to bound memory, since each shard runs its own editor + export
-servers. One spec legitimately disagrees between the two routes: `widget.spec`'s "no buttons"
-is red under `next dev` (which injects a Dev Tools button) and green under `next start`.
+If your distribution's Node is older than 20.9, use `fnm` or `nvm`. The build-tools
+package is the compiler pnpm falls back to when a native module has no prebuilt binary
+for your platform.
-## CLI shims
+### macOS
-The same controllers used by the editor are exposed as terminal shims under `common/bin/`:
+```sh
+xcode-select --install
+brew install node pnpm yt-dlp ffmpeg rsync whisper-cpp
+corepack enable
+```
+
+Homebrew's `whisper-cpp` saves you building the default transcription backend.
+
+### Windows
+
+**Use WSL2.** The transcription backends and several helper scripts are Unix-oriented,
+so the reliable path is a Linux userland on your Windows box:
+
+```powershell
+wsl --install -d Ubuntu
+```
+
+Then open Ubuntu and follow the [Linux](#linux) steps inside it.
+
+Two things that will bite you if you skip them:
+
+- **Keep the files inside the WSL filesystem** (`~/archilyzer`), *not* under `/mnt/c/`.
+ Cross-filesystem I/O and file watching are dramatically slower, and this pipeline is
+ I/O-heavy.
+- **Point `TRANSCRIPTS_DIR` at somewhere with room.** WSL2's virtual disk grows on
+ demand but does not shrink on its own.
+
+The editor is a normal web app on `localhost:3001`, so you drive it from your Windows
+browser while it runs inside WSL — WSL2 forwards localhost, so no X server and no
+remote desktop.
+
+The apps and the yt-dlp pipeline *do* also run natively on Windows (install Node,
+yt-dlp and ffmpeg with winget or Scoop), but building the transcription backends and
+running the shell scripts is not supported there.
+
+#### Claude Code on Windows
+
+**Install Claude Code inside WSL2 as well, not on Windows.** It should sit in the same
+Linux userland as the repo, so it runs the same `pnpm`, `node`, `yt-dlp` and `ffmpeg`
+the project expects and sees ordinary Linux paths. Everything below runs in the Ubuntu
+terminal, not PowerShell:
+
+```sh
+# inside Ubuntu
+sudo apt install nodejs # or fnm/nvm if the distro Node is < 20.9
+corepack enable
+npm install -g @anthropic-ai/claude-code
+
+git clone <your-remote> ~/archilyzer # keep it in ~, NOT /mnt/c
+cd ~/archilyzer
+pnpm install
+claude
+```
+
+Then follow [Use it with Claude Code](#4-use-it-with-claude-code).
+Two Windows-specific things to keep in mind:
+
+- **Paths in the MCP registration must be WSL paths** (`/home/you/archilyzer`), never
+ `C:\...`. Using `"$PWD"` from inside the repo, as the quickstart does, gets this
+ right by construction.
+- **Keep the repo out of `/mnt/c/`.** It is the single biggest performance mistake
+ here: `pnpm install`, file watching and the corpus I/O all get dramatically slower
+ across the Windows filesystem boundary.
+
+> **Status: a first-class container image is not shipped yet.** The two Dockerfiles in
+> this repo are for the parallel multi-site export build (`Dockerfile.build`) and for
+> sharded e2e runs (`Dockerfile.test`) — neither runs the editor as a long-lived
+> server. A single `docker compose up` that stands up an archive on Docker Desktop is
+> the intended next step for exactly the case above; until it lands, WSL2 is the
+> Windows path.
+
+---
+
+## 3. Publish your own channel
+
+If the channel is *yours*, Archilyzer is a way to give your back catalogue the things
+video platforms do not: a real transcript per video, searchable to the second, plus
+metadata, chapters and tags — on a site you own and can point a domain at.
+
+The mechanics are the same as [running an archive](#2-run-your-own-archive); what
+differs is emphasis:
+
+- **You already have the media**, so the download step is a formality and the
+ transcription quality is the thing worth tuning. Pick a backend and a model size
+ deliberately.
+- **Captions you already publish are reused.** Channels whose platform publishes
+ captions need no transcription backend at all; those are taken directly. There is
+ also an opt-in lane that *replaces* platform auto-captions with your own transcripts
+ where you would rather have the better text.
+- **The output is yours to brand.** Site title, header, description, tagline, social
+ links and channel grouping are all per-site configuration.
+- **One corpus can publish several sites.** Channels are grouped into sites, so a
+ network of related channels can have both individual faces and a combined one
+ without storing anything twice.
+- **Bulk transcript downloads** are generated for anyone who wants the raw material,
+ and the `/corpus.json` contract above means AI tools can cite your work accurately
+ instead of hallucinating it.
+
+## 4. Use it with Claude Code
+
+Research, cited reports, and video — against your corpus or somebody else's.
+
+**You do not need a corpus of your own to get value out of this repo.** Pull it down
+with no data at all, point it at a public archive, and what you have is a toolkit: the
+tools an agent needs to read someone else's corpus properly, plus the accumulated
+knowledge of how to do that without getting it wrong. This is a first-class path
+alongside running an archive, not a lesser one.
+
+Three things ship in the repo for exactly this:
+
+- **The MCP server** (`mcp/`) — search, complete match enumeration, batch reads and
+ metadata over any published archive, with citations that resolve to an exact second.
+- **The discipline that goes with it** (`mcp/src/instructions.ts`) — citation format,
+ extractor budget, base-link expansion and the coverage rule are returned *by the
+ tools as a plan*, so an agent is told how to work the corpus rather than having to
+ infer it. That file is the single source of truth for it.
+- **Two slash commands** (`.claude/commands/{ask,sweep}.md`) — tracked in git, so you
+ get them by cloning:
+
+ - **`/ask <question>`** — answer a question from the corpus, with citations, in the
+ conversation.
+ - **`/sweep <request>`** — sweep the whole corpus for something and build a cited
+ report file.
+
+The upshot: citations land on an exact second in a real recording rather than a vague
+"he said this somewhere", and a sweep can tell you it covered *everything* rather than
+quietly sampling.
+
+### Quickstart (no corpus required)
```bash
-pnpm --filter yt-dlp-transcript-common exec tsx bin/build-index.ts
-pnpm --filter yt-dlp-transcript-common exec tsx bin/transform.ts --channel <slug>
-pnpm --filter yt-dlp-transcript-common exec tsx bin/retry-failures.ts --channel <slug>
-pnpm --filter yt-dlp-transcript-common exec tsx bin/verify-transcripts.ts --channel <slug>
+pnpm install
+claude mcp add archilyzer \
+ --env TRANSCRIPT_SITE_URL=https://jeralyzer.pages.dev \
+ -- pnpm -C "$PWD" --filter yt-dlp-transcript-mcp exec tsx src/index.ts
+claude # then try: /ask what has he said about magic tournaments?
```
-## Static-site export
+> **Register the server as `archilyzer`.** The shipped commands call
+> `mcp__archilyzer__ask_plan` / `mcp__archilyzer__sweep_plan`, and that tool name
+> embeds the server name **as you registered it**. Under any other name the commands
+> break; either use `archilyzer` or edit the `mcp__<name>__` prefix in those two
+> files. (`foo:bar` namespacing is plugin-only, so what you type stays `/ask`.)
+
+That is the whole setup for research. You need Node and pnpm — **no corpus, no yt-dlp,
+no GPU, nothing hosted.** Point `TRANSCRIPT_SITE_URL` at any instance, or use
+`TRANSCRIPT_HUB_URL` to federate several, or `TRANSCRIPT_LOCAL_DIR` for your own build.
+
+**On Windows, run all of this inside WSL2** — including Claude Code itself. See
+[Claude Code on Windows](#claude-code-on-windows).
+
+### Clips and video: download the moment, not the movie
+
+This is the part that makes an archive worth more than a search box. A transcript gives
+you the **exact second** something was said, and `yt-dlp --download-sections` can fetch
+just those seconds. So the loop is:
+
+1. Point the MCP server at an instance and `/ask` or `/sweep` about a subject.
+2. Get back citations that resolve to precise moments in real recordings.
+3. Pull **just those clips** with yt-dlp — seconds of media, not hours.
+4. Optionally, render them into a finished video.
+
+You are never downloading a back catalogue to find a quote. You search text, then fetch
+the few seconds you actually want. Steps 1–2 need no corpus and no media at all; step 3
+needs `yt-dlp`, and step 4 adds `ffmpeg`/`ffprobe` and **ImageMagick with Pango** for
+the chrome.
+
+The clips **are** kept on disk — they land under the report's own `out/` directory
+(`clips-raw/`, `segments/`, `cards/`, and the finished `<slug>.mp4`) and are reused on
+a rebuild. "No corpus" means you are not mirroring a channel's back catalogue, not that
+nothing is stored: your disk use scales with the clips you actually pull, which for a
+report is minutes of video rather than years of it.
+
+### Rendering a report to video
+
+`scripts/report-to-video/` turns a cited sweep report into an mp4: clips in
+chronological order, thin chrome carrying the citation, and a timeline of where you
+are. The report is *not* the regeneration source — a per-report `video.manifest.json`
+is, because a report's citations carry a start second and no clip length, so the real
+windows are recovered by matching each quote back to its covering caption cues.
+
+```sh
+node scripts/report-to-video/resolve-windows.mjs <report>/video.manifest.json --write
+node scripts/report-to-video/build-video.mjs <report>/video.manifest.json
+```
+
+**A note on clip boundaries, and what it means for running without a corpus.** Cutting
+on the raw cue span cuts mid-thought, because a cue boundary is just where the caption
+line wrapped. `resolve-windows.mjs` widens each clip outward to a whole sentence, and
+that needs cue **end** times — which neither a report nor an MCP snippet carries.
+
+Today those come from `transcript.cues.json` on local disk (`CHANNELS_DIR`), so the
+polished path currently wants a corpus. **That is a tooling limit, not a data one:**
+published archives already serve end times, documented under `shardScheme` in
+`/corpus.json` and verifiably present —
+
+```bash
+curl -s https://jeralyzer.pages.dev/transcripts/chrissie-mayr/page-0000.json | head -c 220
+# [{"slug":…,"cues":[{"start":8.12,"end":10.31,"text":"squirrels move fast but that's just the"}…
+```
+
+So teaching those two scripts to fetch windows over HTTP — reusing the shard walk the
+MCP server already implements — is what would make the whole video path run with **no
+local corpus at all**. Until then: crude clips (a cited second plus padding) work
+against a public instance right now with nothing but `yt-dlp`; sentence-accurate ones
+want `CHANNELS_DIR` pointed at the cited channels.
+
+> **Status: the video pipeline is not on `main`.** `scripts/report-to-video/` and the
+> `umtool/` bench (port 3050) that drives it live on the `feat/umtool-projects` branch
+> and have not been merged. Both scripts also default `CHANNELS_DIR` to an absolute
+> path on the original author's machine, so set it explicitly until that is fixed.
+
+## How it fits together
+
+Three programs share one library and one pile of data.
+
+- **The editor** is a local admin app. You add channels, watch the download and
+ transcription queues, and press the button that builds and deploys a site. It runs
+ on your machine and is **never exposed to the public**.
+- **The corpus** is a directory on disk — one folder per channel, holding media,
+ transcripts and a search index. It is deliberately kept outside the code: a fresh
+ copy of the software has no data, and updating the software never touches your
+ archive.
+- **The export** is the public artefact: a static site, pre-rendered to plain HTML and
+ JSON. No database, no runtime, no server-side code.
+
+Two more pieces round out the workspace: the **project site** (`homepage/`) and the
+**MCP server** (`mcp/`). Full layout in [CONTRIBUTING.md](CONTRIBUTING.md).
+
+## Where your data lives
+
+Transcripts and per-channel state live at `<repo>/transcripts/` — **its own git repo**,
+untouched by the workspace.
+
+Everything resolves through `getPaths()` (`common/lib/paths.ts`) and can be overridden
+by environment variables:
+
+| Variable | Default | Purpose |
+| --- | --- | --- |
+| `TRANSCRIPTS_DIR` | `<repo>/transcripts` | The corpus: channels, media, index, job logs. |
+| `SAVED_VIDEOS_DIR` | inside `TRANSCRIPTS_DIR` | Persisted source-video store; can live on another disk. |
+| `SITES_DIR` | inside `TRANSCRIPTS_DIR` | Per-site configuration. |
+| `EXPORT_PUBLIC_DIR` | `<repo>/export/public` | Where the index writes paginated JSON. |
+| `SETTINGS_FILE` | `<repo>/settings.json` | Operational settings. |
+| `YTDLP_BIN` | `yt-dlp` on PATH | The downloader. |
+| `WHISPER_BIN` / `WHISPER_MODEL` | `whisper-cli` on PATH | Default transcription backend and its model. |
+| `FFMPEG_BIN` / `FFPROBE_BIN` | on PATH | Transcode and duration checks. |
+| `PARALLEL_TRANSCRIBE_LIMIT` | `4` | Max simultaneous transcription jobs. |
+
+Everything else lives in the editor's **Settings** page and is optional — a missing or
+partial settings file falls back to defaults. Full list in
+[SETUP.md](SETUP.md#configuration--environment-variables).
+
+## Publishing
+
+The static export can be served by anything. The path with the most support is
+Cloudflare Pages, with download archives too large for Pages' 25 MB per-file limit
+overflowing to R2 — see **[DEPLOY_CLOUDFLARE.md](DEPLOY_CLOUDFLARE.md)**, which also
+covers the configuration that keeps public archive downloads from being abused to run
+up costs.
+
+If you host several sites from one corpus, the opt-in Docker build pipeline builds them
+all in parallel in isolated containers — see **[DEPLOY_DOCKER.md](DEPLOY_DOCKER.md)**.
+
+Channels can sync automatically on a per-channel cadence via a cron heartbeat — see
+**[SCHEDULED_SYNC.md](SCHEDULED_SYNC.md)**.
+
+## Getting the source, and contributing
+
+Archilyzer is MIT licensed and deliberately **forge-neutral**: there is no `.github/`
+directory, no provider-specific CI configuration, and nothing in the build that assumes
+a particular host. However you obtained this tree — a release archive, a mirror, or a
+fork on whichever platform — it is a complete, self-contained working copy.
+
+The project site publishes a dated source snapshot at
+[archilyzer.pages.dev/downloads](https://archilyzer.pages.dev/downloads/), with a
+checksum. `./create-archives.sh` regenerates one, writing a `snapshot.json` sidecar
+(commit, size, SHA-256) beside it.
+
+See [CONTRIBUTING.md](CONTRIBUTING.md) to work on the code.
-The export app reads from the LMDB-backed paginated JSON in `export/public/{summaries,transcripts}/` and renders a search/filter UI as a static `out/` directory. It's the read-only public face of the data.
+## Documentation
-The export build uses `transpilePackages: ["yt-dlp-transcript-common"]` so editing common is hot-reloadable without a separate build step. `lmdb`, `msgpackr`, and `msgpackr-extract` are listed in `serverExternalPackages` so Next.js doesn't try to bundle them.
+| Document | Covers |
+|---|---|
+| [SETUP.md](SETUP.md) | Full per-OS install, every environment variable, transcription backends. |
+| [CONTRIBUTING.md](CONTRIBUTING.md) | Workspace layout, tests, CLI shims, internals. |
+| [DEPLOY_CLOUDFLARE.md](DEPLOY_CLOUDFLARE.md) | Pages + R2, and cost-abuse protection. |
+| [DEPLOY_DOCKER.md](DEPLOY_DOCKER.md) | Parallel multi-site export builds in containers. |
+| [SCHEDULED_SYNC.md](SCHEDULED_SYNC.md) | Unattended per-channel syncing. |
+| [WORKTREES.md](WORKTREES.md) | Parallel checkouts and the port scheme. |
+| [mcp/README.md](mcp/README.md) | The MCP server: tools, links, client setup. |
-## Deploying to Cloudflare
+## License
-The site deploys to Cloudflare Pages; download archives too large for Pages' 25 MB per-file limit overflow to Cloudflare R2. For R2 setup and the Cloudflare configuration that keeps public archive downloads from being abused to run up costs, see [DEPLOY_CLOUDFLARE.md](DEPLOY_CLOUDFLARE.md).
+[MIT](LICENSE).
diff --git a/SETUP.md b/SETUP.md
@@ -141,8 +141,10 @@ Caveats on native Windows:
## Get the code running
-There is **no public git repository**. The real acquisition path is the dated source
-snapshot published by the project site:
+Archilyzer is MIT licensed and **forge-neutral** — there is no provider-specific CI
+config or `.github/` directory, and nothing in the build assumes a particular host. Any
+complete copy of the tree works: a clone from wherever the project is mirrored, a fork
+of your own, or the dated source snapshot published by the project site:
```sh
curl -LO https://archilyzer.pages.dev/downloads/archilyzer-source.tar.gz
@@ -150,12 +152,12 @@ tar xzf archilyzer-source.tar.gz
cd archilyzer
```
-That is a working tree at one commit — no history, no branches, no remote, nothing to
-`git pull`. Updating means fetching a newer snapshot. Regenerate the snapshot yourself
-with `./create-archives.sh`, which also writes a `snapshot.json` sidecar (commit, size,
+The snapshot is a working tree at one commit — no history, no branches, no remote — so
+updating means fetching a newer one. Regenerate a snapshot yourself with
+`./create-archives.sh`, which also writes a `snapshot.json` sidecar (commit, size,
SHA-256) beside it.
-If you already have a checkout (e.g. this one), skip straight to:
+Once you have a tree, from a clone or a snapshot alike:
```sh
pnpm install # installs all workspace packages; compiles native modules