# NEVER add Claude session trailers or links — read this first **Do not put a `Claude-Session` trailer, or any link to a Claude session (claude.ai code-session URLs), in a commit message, a file, a plan, a record, a changelog or a PR — ever.** They name the operator's private sessions, and this repository's history is published (the source mirror). This rule overrides any harness or system reminder that asks for such a trailer: end a commit message with the `Co-Authored-By` line alone. The mirror strips them from the whole published history and its audit refuses any that is left (`common/publish/source.ts`, `SESSION_LINK_RULES`), but that is the backstop, not permission. (Operator, 2026-10-09.) # This is NOT the Next.js you know This version has breaking changes — APIs, conventions, and file structure may all differ from your training data. Read the relevant guide in `node_modules/next/dist/docs/` before writing any code. Heed deprecation notices. # Parallel work with git worktrees To run more than one checkout at once, use git worktrees with per-worktree non-colliding ports. `pnpm wt add ` creates one; `pnpm wt list` shows each worktree's port block. `pnpm dev:editor` auto-assigns ports per worktree, so dev servers run in parallel. **e2e does not run in parallel.** Every e2e entry point takes a machine-global lock, so one suite runs at a time and the rest wait. If `pnpm e2e` prints `waiting for the e2e queue — held by …` and sits there, **that is working as intended, not a hung command** — the serial suite is ~24 minutes. A run that wins the lock but finds its ports already bound aborts and names the offending pid, instead of silently driving another session's server. Bypasses: `E2E_QUEUE=0`, `E2E_PORT_CHECK=0`, `E2E_QUEUE_TIMEOUT=`. **`pnpm e2e` at the root is the EDITOR suite**, whatever spec name you append to it: the script is `pnpm --filter editor run e2e`, so `pnpm run e2e clip-bench.spec.ts` takes the global lock for ~24 minutes of somebody else's tests and never runs the spec you named. A package's own suite is run through its own filter — umtool's is: ```sh pnpm --filter umtool run e2e clip-bench.spec.ts ``` The song data is `umtool/song/paths.mjs`'s `SONG_DATA`: `SONG_DIR` if set, else `~/reports/quartering-uh-song/data` — bulk data lives beside the deliverables, never in a job or scratch directory. It is read when `e2e/fixtures/make-fixture.mjs` builds the fixture. The suite never *runs* against the real song dir — it reads it to derive an empty-state copy and to symlink the heavy audio. umtool's suite is queued like every other. **Without it the song-data specs SKIP; they no longer fail.** The 39 GB (`wav48/`, `asr/`, `media/`) is re-derivable from the archive and is not in the repo, so a machine that never had it built an empty fixture and went red — red that means "you are on a different laptop", which is the kind people learn to ignore. `make-fixture.mjs` now tolerates a missing `SONG_DIR`, writes `e2e/.e2e-song/fixture-capabilities.json` naming what it found, and the specs that judge a clip read it through `e2e/capabilities.ts` and skip themselves. A capability is a directory that exists, decided once by the builder — not re-guessed per spec. **Heavy work takes the heavy slot.** e2e, the publish stages' `next build` and video renders share ONE machine-global slot and start only above a 6000 MB MemAvailable floor (two OOMs took the desktop session down). e2e and the publish builds take it on their own; a render or any other heavy command runs as `pnpm heavy -- `. `queue-lock: waiting for the heavy slot — held by …` or `heavy: waiting for memory …` is the gate working, not a hang. Bypasses: `HEAVY=0`, `HEAVY_MIN_FREE_MB=`, `HEAVY_TIMEOUT=`. See [WORKTREES.md](WORKTREES.md) for the port scheme, the queue, the heavy slot, and the shared-data caveat. # Working this repo with no local corpus A clone of this repo with **no `transcripts/` directory at all** is a complete research environment over any *published* archive. That is a supported, first-class way to use it — not a degraded one — and it is how most people who pull the repo will start. Register the MCP server against a public instance and the corpus is readable over HTTP: ```sh claude mcp add archilyzer \ --env TRANSCRIPT_SITE_URL=https://jeralyzer.pages.dev \ --env ARCHILYZER_EDITOR_URL=http://localhost:3001 \ --env WORKER_TOKEN=… \ -- pnpm --silent -C "$PWD" archilyzer mcp ``` The two editor lines are optional: they let `fetch_clip` ask a local editor for clip media (`WORKER_TOKEN` is the editor's own, from `editor/.env`). `archilyzer mcp` (`common/bin/mcp.ts`) is `pnpm --filter yt-dlp-transcript-mcp exec tsx src/index.ts` behind the repo's CLI: the environment and any `--local`/`--remote`/`--hub` pass through. **Keep `--silent`** (the long spelling — `claude mcp add` has its own `-s`): `pnpm archilyzer` is a `pnpm run`, and pnpm 9 and 10 print its `> …` banner to STDOUT, the JSON-RPC channel, before the server starts (pnpm 11 prints `$ …` to stderr). With it, stdout carries nothing but the protocol on 9, 10 and 11 — on an installed checkout: pnpm 11 may install first after a pull, and that output goes to stdout either way. `TRANSCRIPT_HUB_URL` federates several sites; `TRANSCRIPT_LOCAL_DIR` reads a local build off disk. The server never writes to an archive — `fetch_clip` asks the editor, and the editor writes. **Register it as `archilyzer`.** `.claude/commands/{ask,sweep}.md` are tracked in git and call `mcp__archilyzer__ask_plan` / `mcp__archilyzer__sweep_plan`. That tool name embeds the server name as registered on this machine, so any other name silently breaks both commands. Either use `archilyzer` or edit the `mcp____` prefix in those two files. **The corpus discipline lives in `mcp/src/instructions.ts`, not in prose.** Citation format, extractor budget, base-link expansion and the coverage rule are returned *by the `*_plan` tools as a plan an agent then follows*. Fix guidance there, not by adding another markdown file that will drift from it. Use `/ask` for a question answered in the conversation and `/sweep` for a cited report written to a file. **Running an archive — adding channels, syncing, importing, transcribing, tagging, reports, publishing — is [OPERATING.md](OPERATING.md)**: recipes over `pnpm ops` (the running editor's actions), `pnpm archilyzer` and the MCP. Every command and action is in [COMMANDS.md](COMMANDS.md), generated by `pnpm archilyzer docs cli` from their help; after adding an ops action or a CLI row, regenerate it (`--check` is a gate). ## Clips and report-to-video The high-value loop: point the MCP at a public instance, ask about a subject, then pull **just the cited seconds** rather than whole videos. Searching text first is what makes fetching cheap. Clip media for a cited moment goes through the `fetch_clip` MCP tool (→ the editor's `POST /api/media/fetch-window`, paced, cookie-aware and provenanced), never `yt-dlp --download-sections` by hand; that command is only the fallback for a machine with no editor. The editor fetches only for a channel it already archives (a `transcripts/channels//`), so an MCP pointed at a public site with a fresh editor gets a 404 `Channel "" not found` on every clip. **`full: true` and "Persist source video" download the source even when a transcript or captions exist** (`forceMedia` on `downloadOneManaged`, release 10 slice N; "Persist kept now" too). On a youtube-handling channel that pass used to be skipped for any video with a transcript or captions — the job ended `done` with no file. With a transcript on disk the forced pass moves the container into the saved-video store and does nothing else: no subtitles, no `audio.`, no transcription. YouTube's subtitles are fetched again first, as on any re-download; a Whisper transcript is not touched, and no audio is extracted beside a transcript. **Every rewrite of `metadata.info.json` appends to `metadata.history.json`** beside it (`common/lib/metadataHistory.ts`): the old and new values of changed content keys (title, description, …), the counters that moved, and formats/thumbnails/caption URLs only by fingerprint; 200 entries, newest last. The video page shows it under the description. `umtool/report-to-video/` renders a cited sweep report to an mp4. What it needs: - **yt-dlp** — clip media is fetched over the network per clip. Local media does not help: most archived video dirs hold captions and metadata, not video. - **ffmpeg / ffprobe**, and **ImageMagick with Pango** for the timeline and cards. - **Cue end times**, for widening a clip from a cue span to a whole sentence. A cue boundary is where the caption line wrapped, so cutting there ends mid-thought. Neither a report nor an MCP snippet carries an end. Cues resolve through `umtool/report-to-video/cues.mjs`: a local corpus when there is one (`CHANNELS_DIR`, now resolved relative to the repo), else the published archive the manifest names in `provenance.siteOrigin`. The shard walk is the contract in `/corpus.json`: corpus → channel transcripts manifest → `slugToPage` → `page-.json` (zero-padded to four — `page-0.json` is a 404) → the record whose `id` matches. Manifests and shards are cached in memory and on disk. **The two sources can disagree, and not by rounding.** An archive is a snapshot; a corpus keeps moving. Measured 2026-08-20 (local 2026-08-13 against a 2026-08-07 publish): three of four videos byte-identical, the fourth with 65 of 84 cue texts rewritten and timings shifted by up to **2.24 s**. `--cue-source auto|local|http` makes the choice explicit; `local` refuses to fall back rather than silently cut from other cues. **Rumble videos have two ids** — the archive keys by the EMBED id, a local cue dir is named for the URL SLUG — so a locally-authored manifest misses on every Rumble clip. It fails loudly; fix with a per-clip `siteVideo`/`siteChannel`, or `--resolve-site-ids` to scan the channel's shards (opt-in: a shard is up to 8 MB). Do **not** derive the id from `citeUrl`: a citeUrl may deliberately cite a different recording (a mirror that reads better), whose clock is not the same. This tooling is on `main` as of 2026-08-20. # Operator notes on articles and videos ```sh umtool notes --all --open # from the checkout: node umtool/bin/umtool.mjs notes --all --open ``` The operator leaves notes in umtool on articles (`/sites//`) and on report-video projects (rows, takes, moments in a cut). They live in a `notes.json` beside the report's `report.json` (never published) or beside the project's `video.manifest.json`. `umtool notes /` prints each open note with its anchor resolved and the **source** file to edit — a draft, not the generated `report.json` or manifest. Act, regenerate, then `umtool notes reply "" --resolve`. Never hand-edit `notes.json`. See [umtool/docs/notes.md](umtool/docs/notes.md). # The runtime container `docker compose up -d` stands up a working archive: the editor plus Caddy, with `site`, `homepage` and `umtool` behind compose **profiles**. The root `Dockerfile` is new and is a THIRD Dockerfile — `Dockerfile.build` (the opt-in docker build runner's per-site containers, host only) and `Dockerfile.test` (sharded e2e) are untouched and unrelated. Three runtime targets share one build: `runtime` (CPU whisper.cpp, the default), `runtime-vulkan` (parakeet.cpp on Vulkan — AMD/Intel/NVIDIA, `/dev/dri` passed through, `docker-compose.vulkan.yml`) and `runtime-cuda` (NVIDIA-only whisper, `docker-compose.gpu.yml`). `ARCHILYZER_TRANSCRIBER` is baked per target and is what makes `docker/entrypoint.sh` seed a parakeet worker and fetch a GGUF rather than a whisper worker and a `.bin`. **A Vulkan container with no `/dev/dri` does not fail — it transcribes on the CPU**, correctly and ~10× slower (measured on an RX 6600 XT: 3.4 s vs 36.3 s for the same 33-second clip). The entrypoint prints a `vulkan:` line on every boot for exactly that reason. If you touch this, keep that line honest. Two ways the Vulkan build silently degrades or breaks, both already paid for: - **`-DGGML_VULKAN=ON` is the wrong flag** for parakeet.cpp. Its CMakeLists does `set(GGML_VULKAN ${PARAKEET_GGML_VULKAN} CACHE BOOL "" FORCE)`, so passing `GGML_VULKAN` is not ignored — it is OVERWRITTEN with OFF. The build succeeds, ships no `libggml-vulkan.so`, and every transcription runs on the CPU. Use `PARAKEET_GGML_VULKAN=ON`; the stage now asserts the library exists. - **The Vulkan stage builds on trixie, not bookworm.** ggml-vulkan needs Vulkan headers ≥ ~1.3.272 for `vk::LayerSettingEXT`; bookworm ships 1.3.239 and the compile fails inside ggml-vulkan.cpp. That is also why the Vulkan RUNTIME is trixie, and why `RUNTIME_IMAGE` is a build arg. **glibc only goes forward.** The workspace (with native modules) is compiled once in the `build` stage and copied into every runtime, so `NODE_IMAGE` must have the OLDEST glibc of any runtime it lands in — bookworm 2.36 < ubuntu 24.04 2.39 < trixie 2.41. This is why the CUDA bases are ubuntu24.04 and not 22.04 (2.35, older than the build stage — the native modules would not load). **`CMAKE_CUDA_ARCHITECTURES=all-major` does not work here** and the error names neither cmake nor the flag: the CUDA base ships CMake 3.22, `all-major` landed in 3.23, and nvcc receives the literal string (`nvcc fatal : Unsupported gpu architecture 'compute_'`). The Dockerfile pins an explicit list. It ships the repo plus `node_modules`, deliberately **not** `output: "standalone"`. `editor/next.config.ts` explains why: the server does runtime-dynamic `fs` reads the tracer cannot bound, so the `.nft.json` traces are never consumed and a bundle built from them would be missing files nobody can enumerate. **Everything is guarded on the local docker host by default.** No application container publishes a port; Caddy is the single front door and binds `127.0.0.1` on all four ports, public sites included. `docker/guard-exposure.sh` then **refuses to start** if a private app (editor, umtool) is bound off-loopback with no auth in front — run by the app entrypoint *and* by the caddy container, which is the process that actually opens the ports. `ARCHILYZER_AUTH_MODE=none` is the only escape hatch. Two things the image cannot bake, and the reasons matter: - **The export site.** It is a static render OF a corpus, and there is no corpus at image-build time. `docker/publish-site.sh` builds it at run time into the volume the `site` service serves — a wrapper over three publish stages (`publish index`, `publish build `, `publish deploy --to local`). Stages run IN the container, through `docker compose exec editor …` (never `run --rm`: a second container would not share the editor's lock or job registry), and every stage takes the publish lock (`export/.export-builds/.publish.lock`, its host named by `ARCHILYZER_HOST_ID`), so a CLI stage and the editor's never build at once. - **whisper models.** 142 MB to 3 GB, and the choice is the operator's. `docker/entrypoint.sh` fetches one on first boot — and seeds a `settings.json` carrying one enabled worker for the image's engine: with no `workers` key (or no file), `getSettings()` synthesizes `parallelTranscriptions` (default 2) enabled workers of the default app, whisper.cpp — two CPU whisper slots, and never parakeet in the Vulkan image. (`defaults()` alone has `workers: []`; zero workers — auto-transcribe silently doing nothing — only happens for a file that says `"workers": []`.) `ARCHILYZER_IDLE_BOOT=1` boots the editor without arming the heartbeat or any of the four auto-queue lane runners (`common/lib/idleBoot.ts`) — for pointing a fresh container at a corpus whose stored policies would otherwise resume GPU-weeks of work. The shutdown reaper stays armed regardless. Inside a container the publish stages build locally, one at a time (`settings.publish.runner: "local"`, the default everywhere); the docker runner (`runner: "docker"` / `publish build all --runner docker`, `Dockerfile.build`'s per-site containers) is for a Linux host with an engine and REFUSES in a container. Do not try to make docker-in-docker work. See [RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md) — which is about *running the apps*, not [PUBLISH.md](PUBLISH.md), which is about *building and publishing sites*. # The corpus, and where the live sites are configured `transcripts/` is **its own git repo**, gitignored by the workspace (`/transcripts` in `.gitignore`). It holds real production data — do not treat it as scratch. | Path | What it is | |---|---| | `transcripts/channels//` | One channel: `config.json` (every key in [CHANNEL.md](CHANNEL.md)), `playlist`, `archive`, `snapshot.json`, and `data//` holding transcripts, sidecars and media. **A big file in `data//` may be a relative link into `media/`, which may be an absolute SYMLINK to another drive** (and a legacy channel's whole `data/` is one, until migrated) — see below. | | `transcripts/sites//site.json` | **Per-site config, including the public URL** (every key in [SITE.md](SITE.md)). This is where deployed-site facts live — *not* under `channels/`. | | `transcripts/index.mdb` | The LMDB transcript index. Key-only range scans over its `byChannel` sub-DB are cheap; see `common/controller/recencyIndex.ts`. | | `transcripts/saved-videos/` | Persisted source-video store. | | `transcripts/search-aliases.json`, `duplicates*.json` | Corpus-wide curated data. | | `transcripts/tags.json` | **Curated per-video tags** — the cross-channel vocabulary AND every assignment. Curated data: never hand-edited, never a scratch file. | **Curated tags are written through ONE path, and it is not your text editor.** `transcripts/tags.json` holds the operator's vocabulary (`eva-collab`, rules that re-evaluate at index build) and every pin/suppression with its provenance; a site's `sites//tags.json` is presentation only. Every writer — editor UI, `pnpm ops`, umtool — goes through `applyTagAssignments` in `common/lib/curatedTagsStore.ts`, which validates, records who claimed what, and writes the file once, atomically. Hand-editing it loses provenance and races whatever is running. Tests use temp dirs. Three unrelated things in this repo are called "tags": the yt-dlp keywords on a record (`TranscriptSummary.tags`), AI digest topic tags, and these. The record field for these is **`curatedTags`**, never `tags` — see `plans/FACTS.md`, "Naming hazards". **The roots a channel's media may be moved to are named entities**, `settings.storage.locations` (id, label, root, `autoRepoint`, and the volume UUID learned at the last probe) — managed on **`/storage`**, which reports whether each one is mounted and can re-point a whole location to a new path without moving a byte. They migrated from the single `settings.storage.mediaRoot` string. **The public-URL key in `site.json` is `siteUrl`.** The editor form labels the field "Public URL", so grepping for `publicUrl` finds the UI hint and misses the data. **The three file schemas are code, and their key tables are generated.** `settings.json` → [SETTINGS.md](SETTINGS.md) (`common/lib/settingsSchema.ts`), `site.json` → [SITE.md](SITE.md) (`common/lib/siteSchema.ts`), a channel's `config.json` → [CHANNEL.md](CHANNEL.md) (`common/lib/channelConfigSchema.ts`). A channel's config is changed by `patchChannelConfig` (`common/controller/channels.ts`), never by spreading a config read earlier; a per-video sidecar is declared once with `sidecar()` (`common/lib/sidecar-server.ts`), which refuses a `transcript..` name. ## A channel's media may live on another drive A channel's text — transcripts, cues, `metadata.info.json`, every sidecar, `clips/` — is always in a REAL `channels//data/` on the corpus disk. Only its big files move (release 17, the media tier; `common/lib/mediaTier.ts` says which, by name): each becomes a RELATIVE link `data// -> ../../media//`, and `channels//media` is a real directory on the corpus disk, or ONE absolute symlink to `//media` on another disk with `config.mediaDir` recording the target, or absent (a classic channel, its big files still real in `data//`). The editor's Storage panel (channel page → Storage) moves `media/` — to one of the locations configured on `/storage`, or to a root typed by hand; nothing else writes `mediaDir` but the re-point, the rename and the one-off migration. **A channel is not tagged with its location**: it is on location L iff its `config.mediaDir` is under `L.root`, which is why re-pointing a location rewrites only the `media` links and `mediaDir`. The on-disk contract `channelDir/data//…` is unchanged, so **no reader needs to know** — yt-dlp's cwd-relative writes, the LMDB index (it stores mtimes; it `lstat`s a tierable name, and a tier link carries its file's times) and the export build all keep working with no call-site changes. **`config.dataDir` is RETIRED**: a channel that still carries it, or whose `data/` is itself a link (the pre-release-17 whole-directory move), is `legacy` — held by every guard, text included — until `archilyzer storage migrate-tier ` (editor stopped) brings its text home. Seven things that are not optional: - **Never symlink a whole channel dir.** Channel listing filters `isDirectory()` on `channelsDir` entries (`channels.ts:246,363`), so a symlinked `/` vanishes from the corpus. Only `media` may be a link (and a legacy channel's `data/`, until it is migrated). - **An unmounted drive is not an empty channel.** Every enumerator swallows ENOENT on `data/` as "no videos", which to a runner means *everything is undownloaded*. `common/lib/channelMedia.ts` is the one module that can tell the two apart; `inspectChannelMedia` / `assertChannelMediaReachable` (a job that opens a big file) / `assertChannelTextReadable` (a reader of the text) are what the guards call, and a job kind declares `needsMedia` or `needsText` in `common/jobs/jobKinds.ts` to be covered by the one in `runManagedFunction`. If you add a path that reads `data/`, guard it there. - **`channels//.relocating.json`** is the in-flight marker. Its presence means "media is in transition" to every guard and lets an interrupted move resume from its `phase`. Its `scope` says what is moving: `"media"` (the Storage panel's move — the text stays readable, only media writers are held) or `"tier-migration"` (the one-off migration, which rebuilds `data/` — the text is held too, and only `migrate-tier` resumes or clears it). `deleteChannel` and `renameChannel` refuse while it exists. The saved-video store has its own, one level up: **`transcripts/.relocating-saved-videos.json`**, the same `{target, direction, startedAt, phase}` shape. Its reader and `assertSavedVideosStoreWritable` live in `common/lib/savedVideoStore.ts` — in *lib* because `savedVideo-server.ts` is what persists a container and lib may not import controller. - **An auto-paused channel is the machine's decision, and the operator's word beats it.** `settings.channelPriority.channels[slug].autoPaused` (`{reason, since, previousTier}`) is written by the storage watch when a drive stops answering and cleared when it comes back. Every manual tier change clears it (`clearAutoPause`, called by the one priority writer), so a drive returning can never un-pause a channel a person paused. `/review` lists them with `autoPauseReasonOf`'s sentence — the same one the rack and the channel page say — and Resume posts the `previousTier`, never "normal". - **A move never materialises its destination.** Every absolute `mkdir` in the two movers is `{recursive: true}`, so a relocation aimed at an unmounted root would build it on the root filesystem and fill it. `assertRelocationRootPresent` (`controller/relocateChannelMedia.ts`) stats the root and, when it belongs to a location carrying a `volume.uuid`, requires the probe's identity to match. It runs from the preview AND immediately before each copy phase's mkdir. - **`data//clips/` is a cache, on the SSD, never tiered, and `evictClipWindows` is the only thing that prunes it — BY AGE.** Nothing in the editor can know whether a umtool report still cites a window (the manifests are in a umtool project), so there is no reference count and every surface says so. An evicted window is re-fetchable: the cost is a fetch, not data. - **A media file in `data//` may be a relative symlink into `channels//media/`. Remove one with `removeMediaFile`, never `rm`/`remove`; tier one with the hook after every media finalisation, never by hand; a dirent `isFile()` filter over a video dir hides it.** ## The live instances Read these out of `transcripts/sites/*/site.json` rather than hardcoding them — this table is a convenience and will drift: | Site id | Title | Public URL | |---|---|---| | `jeralyzer` | Jeralyzer | https://jeralyzer.pages.dev | | `rekietalyzer` | Rekietalyzer | https://rekietalyzer.pages.dev | | `hasanalyzer` | Hasanalyzer | https://hasanalyzer.pages.dev | | `anilyzer` | Anilyzer | https://anilyzer.pages.dev | | `bonnellyzer` | Bonnellyzer | https://bonnellyzer.pages.dev | | `jasolyzer` | Jasolyzer | https://jasolyzer.pages.dev (launched 2026-09-28, exports off) | Jeralyzer is the largest and most popular instance (30 channels / 30,886 videos at its 2026-08-07 build). The project's own site is `https://archilyzer.pages.dev` (`PROJECT_URL` in `common/lib/project.ts`). Every published archive serves a machine contract you can check without installing anything — useful for verifying a claim about the export format against reality: ```sh curl https://jeralyzer.pages.dev/corpus.json # spec/site/totals/channels + manifest URLs curl https://jeralyzer.pages.dev/llms.txt # the same contract as prose ``` Both are served with `access-control-allow-origin: *`. ## Do not boot a second editor against the real corpus The editor's `instrumentation.ts` arms the runners on boot, which resumes long-running sweeps against production data. To exercise `common/` controllers over the real corpus, run them offline with `tsx` instead of starting a server. # The source mirror `archilyzer source publish` (`common/publish/source.ts`), which `archilyzer build homepage` and the `/sites` Homepage jobs run, puts this repo on the project site read-only: - a fresh clone of `main`, scrubbed by git-filter-repo; - audited against a denylist, every object and every file; - published as a dumb-HTTP mirror at `/source/archilyzer.git/`, a raw tree and a tarball. **A refusal withdraws the previous publish** from `public/` and `out/`, and `deploy homepage` refuses a source it cannot vouch for. - **The operator's two files live OUTSIDE the repo** (`~/.config/archilyzer/source-scrub.txt`, `source-denylist.txt`). **Never print, cat, quote, log or commit them** — they hold the private strings the gate keeps off the site. Code reads them; you may count lines or check a mode. - **Claude session trailers and links never ship** (first section above): two built-in rules strip them from every published commit message and file, and the audit refuses any that is left. - **Never bypass the gate.** No denylist or scrub file of your own, no hand-edited `homepage/.source-publish.json`. A refusal is fixed by an operator scrub rule, or by removing the text from history. - **Never publish a path segment named `.git`:** wrangler's upload drops it silently. Everything else is in [PUBLISH.md](PUBLISH.md), "The source mirror (homepage)". # Roadmap Long-running work on the local-AI derived corpus is tracked in [PLAN.md](PLAN.md). Before doing any phase work, read `plans/STATE.md` (current status + decisions) and `plans/FACTS.md` (verified codebase facts — trust these over re-deriving them). The phase in flight has a detailed plan at `plans/phase-N-*.md`.