# Working on Archilyzer This is the developer's side of the project: how the workspace is laid out, how to run the test suites, and the internals worth knowing before you change something. If you just want to *run* an archive, you want [README.md](README.md) and [SETUP.md](SETUP.md) instead. Archilyzer is MIT-licensed. It is deliberately **forge-neutral**: there is no `.github/` directory, no provider-specific CI config, and nothing in the build that assumes a particular host. Patches are welcome however you can get them here — a mirror, a fork on any platform, or a plain `git format-patch` series. ## Workspace layout A pnpm-workspace monorepo. Every package consumes `common/` via `workspace:*`. | Package | What it is | |---|---| | `common/` | The shared library — data layer, controllers, components, types. Nearly all the logic lives here. | | `editor/` | The private admin app (Next.js, port 3001): channels, the download/transcribe pipeline, jobs, settings, build + deploy. | | `export/` | The public static site (`output: "export"`). Read-only, no runtime. | | `homepage/` | The project's own site: marketing, docs, downloads, cross-site stats at `/stats/`. | | `mcp/` | The MCP server that exposes a published archive to LLM clients. See [mcp/README.md](mcp/README.md). | | `umtool/` | A local bench for the report-to-video work (port 3050), including the `umtool/report-to-video/` pipeline package. Not part of an archive install. | Your corpus lives at `/transcripts/` and is **its own git repo**, untouched by the workspace. That separation is deliberate: updating the software never touches your archive, and a fresh checkout has no data. ## Everyday commands ```bash pnpm install pnpm dev:editor # editor at http://localhost:3001 pnpm build # build the static site under export/out/ pnpm start:export # serve export/out/ at http://localhost:3000 pnpm build:index # rebuild the search index in-process (= pnpm archilyzer index) pnpm archilyzer doctor # what this machine has: tools, corpus, settings, ports ``` `pnpm lint` at the root only lints `export/`. **The editor has no eslint config**, so `pnpm lint` inside `editor/` always fails — typecheck it with `pnpm exec tsc --noEmit` instead. ## Parallel checkouts `pnpm wt add ` creates a git worktree with its own non-colliding port block; `pnpm wt list` shows them. `pnpm dev:editor` auto-assigns ports per worktree, so several dev servers run side by side. See [WORKTREES.md](WORKTREES.md) for the port scheme and the shared-data caveat. ## Tests Unit tests are plain `node:test` files run through `tsx`. Each package's suite has a script: ```bash pnpm --filter yt-dlp-transcript-common test # common/ (the globs are in its package.json) pnpm --filter editor test # editor/app/**/*.test.ts pnpm --filter export test # export/app/**/*.test.ts pnpm --filter homepage test # homepage/app/**/*.test.ts pnpm --filter yt-dlp-transcript-mcp test # mcp/src/*.test.ts pnpm run test:scripts # scripts/*.test.mjs, umtool/report-to-video/*.test.mjs ``` One file on its own: ```bash node_modules/.bin/tsx --test common/jobs/autoQueuePolicy.test.ts ``` Playwright covers the editor UI from `editor/e2e/`. Two ways to run it: - **`pnpm e2e`** — sequential, on the host, against `pnpm dev:test` (Next.js dev mode) on port 3001. No Docker. Slow, but it picks up source edits without a rebuild and gives you one readable log. Use this while iterating. - **`pnpm e2e:sharded`** — containerized, splitting the suite across N shards via Playwright's `--shard=i/N`. Each shard runs in its own copy of a test image (`Dockerfile.test`) that bakes the monorepo plus a prebuilt editor (`next start`). `playwright merge-reports` then recombines the per-shard blob reports into `editor/playwright-report/index.html`. Use this to verify a branch. Run `npx playwright install chromium` once before your first e2e run. **e2e does not run in parallel with itself.** Every entry point takes a machine-global lock, so one suite runs at a time and the rest wait. If you see `waiting for the e2e queue — held by …` and it sits there, that is working as intended, not a hang. Bypasses: `E2E_QUEUE=0`, `E2E_PORT_CHECK=0`, `E2E_QUEUE_TIMEOUT=`. ### Sharded runs in detail The npm script rebuilds the image first (layer cache covers unchanged deps, so repeat builds are seconds), then runs `scripts/run-sharded-e2e.mjs`. Shard count defaults to `min(max(2, cpus/2), 8)`. Override with `E2E_SHARDS=N pnpm e2e:sharded` or `node scripts/run-sharded-e2e.mjs --shards N`. **Retries default to 0**, so a sharded run reports the same failures a sequential one does. (The containers set `CI=true` because `playwright.config`'s `reuseExistingServer: !process.env.CI` needs it — each shard must start its own servers — but that would otherwise also switch on `retries: 2` and quietly hide flaky tests. The runner passes an explicit `--retries=0` to override it.) Raise it deliberately with `--retries N` when you actually want to measure flakiness. Anything else on the command line is forwarded to `playwright test` in every shard: ```bash pnpm e2e:sharded -- --grep "digest" node scripts/run-sharded-e2e.mjs --shards 4 e2e/deploy-page.spec.ts ``` After merging the blob reports the runner prints a combined `N passed / M failed`, not just per-shard exit codes. ### Wall-time comparison **Always quote the test count with the time.** An earlier version of this table was measured at 90 tests and stayed in the README unchanged while the suite grew to 392 — which made its "104 s sequential" wrong by more than an order of magnitude, invisibly. Measured 2026-07-30 on an 8-CPU machine, **392 tests**, `--retries=0`: | Mode | Wall time | Notes | | --- | --- | --- | | `pnpm e2e` (sequential) | **23.8 min** (1428 s) | one worker, `next dev` | | `pnpm e2e:sharded --shards 4` (image cached) | **7.1 min** (424 s) | 4 shards × `next start`, **~3.4× speedup** | | `docker build` (cached layers) | ~40–110 s | one-off, paid on dependency or source changes | Caveat on precision: the box was **not** quiet during these runs (an unrelated project was running its own Playwright suite at load ~15–23). A second sequential run under heavier load took 29.7 min for the same result, so treat these as an envelope rather than a constant. 4 shards were used to bound memory, since each shard runs its own editor + export servers. One spec legitimately disagrees between the two routes: `widget.spec`'s "no buttons" is red under `next dev` (which injects a Dev Tools button) and green under `next start`. **The editor e2e suite runs in dev mode and therefore never prerenders.** A change to a layout or a client component is not verified until `pnpm build` passes too. ## The `archilyzer` CLI The same controllers the editor uses are one command line, `common/bin/archilyzer.ts`, which is how you drive the pipeline headlessly or from cron. `pnpm archilyzer ` from the repo root is the short form of `pnpm --filter yt-dlp-transcript-common exec tsx bin/archilyzer.ts `; `pnpm archilyzer --help` lists every command. Every command and every `pnpm ops` action is in [COMMANDS.md](COMMANDS.md), generated by `pnpm archilyzer docs cli` from the two help texts (`--check` fails when it is stale, or when [OPERATING.md](OPERATING.md) names a command that does not exist); recipes that chain them are in OPERATING.md. The table is `archilyzer.ts`; the machinery (parser, lookup, usage) is `_cli.ts`. Every file in `common/bin/` is reachable from a row — a test fails otherwise. A bin that parses its own flags is a *passthrough* row, run as a child with its argv untouched. `run` refuses sync, the metadata scan, downloads and transcription: they run on the editor's paced download queue and worker pool, which a second process must not race. ## Internals worth knowing **This is not the Next.js you may know.** The workspace tracks a recent major and its APIs, conventions and file structure differ from older releases. Read the relevant guide under `node_modules/next/dist/docs/` before writing app-router code, and heed deprecation notices. - The export build uses `transpilePackages: ["yt-dlp-transcript-common"]`, so editing `common/` is hot-reloadable without a separate build step. `lmdb`, `msgpackr` and `msgpackr-extract` are in `serverExternalPackages` so Next.js does not try to bundle them. - The export app reads the LMDB-backed paginated JSON in `export/public/{summaries,transcripts}/` and renders a search/filter UI into a static `out/`. - `getPaths()` (`common/lib/paths.ts`) is the single resolver for every path and binary. Nothing should hardcode a location; add an env override there instead, and declare it in `common/lib/envVars.ts` (a test fails until you do, and `pnpm archilyzer docs env` regenerates [ENVIRONMENT.md](ENVIRONMENT.md)). A new local server's port goes in `common/lib/ports.mjs`, which `pnpm wt` offsets per worktree. - `common/lib/project.ts` holds **product** identity (the name, the project URL) and is deliberately import-free. **Operator** identity — what a given deployment calls itself — lives in settings and per-site config. A string that should change when someone else deploys this belongs in the latter. ## Long-running work Multi-phase efforts are tracked in [PLAN.md](PLAN.md), with current status and verified codebase facts under `plans/`. Read `plans/STATE.md` and `plans/FACTS.md` before starting phase work — they are maintained precisely so you do not have to re-derive them.