commit 831f55a285af485e6632284b51c959024d6e1fd9
parent a145483403af6cac7a8c07665244b35242abe73e
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Thu, 13 Aug 2026 00:51:00 -0400
Give Archilyzer a site, and stop shipping the source inside every archive
Three live public archives existed and nothing said what built them. The root
README was still titled "yt-dlp transcript browser" and claimed three packages
where there are five, and every deployed archive served its own copy of the
whole workspace as code.tar.gz — undated, unchecksummed, N times over.
homepage/ stops being "our hub" and becomes the product's site:
/ what the software is, and a download
/docs/ seven hand-written pages, ordered by a typed manifest
/downloads/ one source snapshot, with a commit and a SHA-256
/stats/ the old dashboard, moved verbatim (the name it had before
acc9d76 deleted it)
The docs are hand-written rather than rendered from the root ones because
SETUP.md says `git clone <this-repo-url>` and there is no repository — that
instruction is actively false on a public site. content/README.md records which
root doc each page derives from, so the drift is tracked rather than pretended
away.
The site now builds WITHOUT a corpus, which is the point: prebuild chained a
multi-GB LMDB index onto every build, so the project's own marketing site could
not be built by someone who had merely unpacked the source. `build` and
`build:nodata` run the identical command; only pnpm's name-matched prebuild hook
differs. Routes degrade honestly — loadSummary() returns null rather than a
zero-filled summary, because "0 transcripts · 0 sites" reads as a broken product
rather than an unconfigured build.
A visual identity of its own: the new `archilyzer` theme family is graphite and
bone with no decorative brand colour at all. Colour appears in exactly two
places and both carry information — a recording's state at its source, and the
archives themselves. The tool should stop dressing as its own output.
Nothing on the page is invented. Every figure is computed at build time (the
public-archive count went 3 → 4 during this work), the rail shows real
acquisitions with their real states, and absent data renders an empty state.
Fabricating transcript rows on the marketing site for an archiving tool would
undercut the entire product.
Also here:
- common/lib/project.ts — the canonical product identity, zero imports so it is
safe from server components, client components and plain tsx alike. PRODUCT
identity lives there; OPERATOR identity stays in homepage.ts + settings.json.
- Every export site's footer drops code.tar.gz and gains "Built with
Archilyzer" — ungated by instance mode, independent of any configured hub,
same tab. The Downloads eyebrow is now conditional on something being behind
it. NOTE: PROJECT_URL is baked in at build time, so publish the project site
before rebuilding any archive.
- corpus.json and llms.txt name the software that produced them. The spec stays
at 3 on purpose: prior bumps announced a fetchable layer, and a credit line
breaks no reader. A test asserts both the field and the non-bump.
- MIT LICENSE. Without one, a published tarball granted its recipients nothing.
- archilyzer.com — which is NOT ours, and 301s to an unrelated page — purged
from six files in favour of archilyzer.pages.dev.
- Two pre-existing e2e bugs fixed: the homepage Playwright config read the
EDITOR's PORT instead of the HOMEPAGE_E2E_PORT allocated for it, and its e2e
script never took the machine-global queue lock.
Verified: tsc clean across all five packages; 392/392 common unit tests; 21/21
homepage e2e; 6/6 export site-branding; 5/5 hub federation; 2/2 editor export
footer; both build modes, with the no-data build checked against a tree with the
corpus artifacts genuinely removed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Diffstat:
82 files changed, 3508 insertions(+), 395 deletions(-)
diff --git a/.dockerignore b/.dockerignore
@@ -15,10 +15,11 @@ settings.json
transcripts/
export/public/summaries/
export/public/transcripts/
-export/public/code.tar.gz
-export/public/code.tar.xz
export/public/transcripts.tar.gz
export/public/transcripts.tar.xz
+# The source snapshot served by the project site — regenerated by
+# create-archives.sh on the host, never needed inside a build container.
+homepage/public/downloads/
# Build staging / caches / per-site outputs — bind-mounted at run time, never baked.
export/.export-index/
diff --git a/.gitignore b/.gitignore
@@ -43,9 +43,7 @@ yarn-error.log*
# nested transcripts repo
/transcripts
-# downloadable repo archives (generate with create-archives.sh)
-/export/public/code.tar.gz
-/export/public/code.tar.xz
+# downloadable transcript archives (see the commented block in create-archives.sh)
/export/public/transcripts.tar.gz
/export/public/transcripts.tar.xz
@@ -87,6 +85,10 @@ yarn-error.log*
/homepage/public/channel-sites.json
/homepage/public/chart-templates.json
/homepage/public/homepage-summary.json
+# source snapshot tarball + its sidecar (regenerate with ./create-archives.sh).
+# Note this does NOT cover /homepage/public/_headers — the `_headers` rule above
+# is scoped to /export/public, so the homepage's stays tracked.
+/homepage/public/downloads
# editor e2e fixtures and ephemeral state
/editor/test-transcripts/
@@ -108,6 +110,10 @@ yarn-error.log*
/export/test-results-2origin/
/export/.2origin/
/export/blob-report/
+# homepage's own Playwright output (its suite runs from homepage/)
+/homepage/test-results/
+/homepage/playwright-report/
+/homepage/out/
# site config
/settings.json
diff --git a/DEPLOY_CLOUDFLARE.md b/DEPLOY_CLOUDFLARE.md
@@ -1,6 +1,6 @@
# Deploying to Cloudflare (Pages + R2 archive overflow)
-The static site is deployed to **Cloudflare Pages**; large download archives that
+An **Archilyzer** site is deployed to **Cloudflare Pages**; large download archives that
exceed Pages' per-file limit overflow to **Cloudflare R2**. This guide covers the R2
setup and — importantly — how to configure Cloudflare so that **public archive
downloads can't be abused to drive up your bill**.
diff --git a/DEPLOY_DOCKER.md b/DEPLOY_DOCKER.md
@@ -1,6 +1,6 @@
# Docker export build pipeline
-Docker build mode builds **every site in parallel** in isolated containers, then
+Archilyzer's Docker build mode builds **every site in parallel** in isolated containers, then
deploys them serially — a large speedup when you host several sites, and stronger
isolation than the basic single-process build. This is opt-in: set **Build
pipeline → Docker** in Settings (or the toggle on the Deploy page). Basic mode is
diff --git a/LICENSE b/LICENSE
@@ -0,0 +1,21 @@
+MIT License
+
+Copyright (c) 2026 I Mean I'm Just Saying
+
+Permission is hereby granted, free of charge, to any person obtaining a copy
+of this software and associated documentation files (the "Software"), to deal
+in the Software without restriction, including without limitation the rights
+to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
+copies of the Software, and to permit persons to whom the Software is
+furnished to do so, subject to the following conditions:
+
+The above copyright notice and this permission notice shall be included in all
+copies or substantial portions of the Software.
+
+THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
+IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
+FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
+AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
+LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
+OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
+SOFTWARE.
diff --git a/README.md b/README.md
@@ -1,13 +1,23 @@
-# yt-dlp transcript browser
+# Archilyzer
-A pnpm-workspace monorepo for archiving and browsing video transcripts. The project is split into three packages:
+Self-hosted, searchable video-transcript archives. Archilyzer downloads a channel's back
+catalogue, transcribes it locally, and builds a static site you host yourself.
-- **`common/`** — shared library (data layer, controllers, components, types). Consumed by the other two via `workspace:*`.
+The project site — documentation and the source snapshot — is
+[archilyzer.pages.dev](https://archilyzer.pages.dev).
+
+A pnpm-workspace monorepo, split into five packages:
+
+- **`common/`** — shared library (data layer, controllers, components, types). Consumed by the others via `workspace:*`.
- **`editor/`** — dynamic Next.js app on port 3001 with admin UIs for channels, the yt-dlp pipeline, the whisper queue, and the build trigger.
- **`export/`** — static Next.js app (`output: "export"`) that produces the read-only public site.
+- **`homepage/`** — static Next.js app for the project's own site: marketing home, docs, downloads, and a cross-site stats dashboard at `/stats/`.
+- **`mcp/`** — an MCP server exposing a published archive to Claude Code, Claude Desktop, Cursor and other clients. See [mcp/README.md](mcp/README.md).
Transcripts and per-channel state live at `<repo>/transcripts/` (its own git repo, untouched by the workspace).
+Licensed under the [MIT License](LICENSE).
+
## Requirements
- **Node.js ≥ 20.9** (LTS 20 or 22) and **pnpm 9+** (easiest via `corepack enable`) and **git** — needed to install and run the web apps.
diff --git a/SCHEDULED_SYNC.md b/SCHEDULED_SYNC.md
@@ -1,6 +1,6 @@
# Scheduled channel sync
-Channels can sync automatically on a per-channel cadence (hourly, daily, every
+Archilyzer channels can sync automatically on a per-channel cadence (hourly, daily, every
N minutes, …), inspired by the Laravel scheduler: a single dumb heartbeat runs
often, and the **server** decides which channels are actually due.
diff --git a/SETUP.md b/SETUP.md
@@ -1,8 +1,13 @@
# Setup & requirements
-This guide takes you from a fresh clone to a running editor and static export on
-**Linux, macOS, or Windows**. It covers what to install, in what order, and which
-pieces you only need for specific features.
+This guide takes you from a fresh copy of **Archilyzer** to a running editor and static
+export on **Linux, macOS, or Windows**. It covers what to install, in what order, and
+which pieces you only need for specific features.
+
+> An operator-facing version of this guide is published at
+> [archilyzer.pages.dev/docs/install/](https://archilyzer.pages.dev/docs/install/). This
+> one is the contributor's copy and additionally covers worktrees, the e2e queue and the
+> sharded test run.
- [Requirements at a glance](#requirements-at-a-glance)
- [Install the toolchain](#install-the-toolchain)
@@ -136,10 +141,23 @@ Caveats on native Windows:
## Get the code running
+There is **no public git repository**. The real acquisition path is the dated source
+snapshot published by the project site:
+
```sh
-git clone <this-repo-url> yt-dlp-transcript-browser
-cd yt-dlp-transcript-browser
+curl -LO https://archilyzer.pages.dev/downloads/archilyzer-source.tar.gz
+tar xzf archilyzer-source.tar.gz
+cd archilyzer
+```
+That is a working tree at one commit — no history, no branches, no remote, nothing to
+`git pull`. Updating means fetching a newer snapshot. Regenerate the snapshot yourself
+with `./create-archives.sh`, which also writes a `snapshot.json` sidecar (commit, size,
+SHA-256) beside it.
+
+If you already have a checkout (e.g. this one), skip straight to:
+
+```sh
pnpm install # installs all workspace packages; compiles native modules
# (lmdb, msgpackr-extract, sharp, …) per the build allowlist
# in pnpm-workspace.yaml
diff --git a/common/components/themeConfig.ts b/common/components/themeConfig.ts
@@ -9,7 +9,12 @@ export const ACCENT_KEY = "ytdlp-tb:accent";
// Theme FAMILY (palette personality) — orthogonal to light/dark MODE.
// "base" is the neutral default (no data-theme attribute). Each family is a pure
// CSS token swap (a [data-theme="…"] block in tokens.css); no markup differs.
-export type ThemeFamily = "base" | "archive" | "selenized" | "swiss";
+export type ThemeFamily =
+ | "base"
+ | "archive"
+ | "selenized"
+ | "swiss"
+ | "archilyzer";
export type ThemeMode = "light" | "dark" | "system";
@@ -19,6 +24,10 @@ export const THEME_FAMILIES: { id: ThemeFamily; label: string }[] = [
{ id: "archive", label: "Archive" },
{ id: "selenized", label: "Selenized" },
{ id: "swiss", label: "Swiss" },
+ // The instrument face — graphite chrome, no decorative brand hue. The project
+ // site's default; selectable everywhere so an operator can dress an archive in
+ // it, though it is deliberately the tool's voice, not an archive's.
+ { id: "archilyzer", label: "Archilyzer" },
];
export const THEME_MODES: { id: ThemeMode; label: string }[] = [
diff --git a/common/lib/corpus.test.ts b/common/lib/corpus.test.ts
@@ -9,6 +9,7 @@ import {
renderSitemapXml,
CORPUS_SPEC_VERSION,
} from "./corpus";
+import { PROJECT_GENERATOR } from "./project";
import type { PublicSiteDescriptor } from "./siteDescriptor";
// Run with:
@@ -139,6 +140,37 @@ test("buildHubCorpus: drops members with no siteUrl, links each corpus", () => {
assert.match(txt, /A: https:\/\/a\.example/);
});
+test("both builders stamp the project generator, and llms.txt trails it", () => {
+ const site = buildSiteCorpus(descriptor({ siteUrl: "https://demo.example" }), {
+ hasArchives: false,
+ });
+ const hub = buildHubCorpus(
+ [{ siteId: "a", siteTitle: "A", siteUrl: "https://a.example" }],
+ { hubTitle: "The Hub", hubUrl: "https://hub.example", generatedAt: "t" },
+ );
+ assert.equal(site.generator, PROJECT_GENERATOR);
+ assert.equal(hub.generator, PROJECT_GENERATOR);
+ assert.match(site.generator, /^Archilyzer \(https:\/\//);
+ // The credit is the LAST line of both llms.txt renderings, so a human or an
+ // LLM reading top-down finds the corpus content before the colophon.
+ for (const txt of [renderSiteLlmsTxt(site), renderHubLlmsTxt(hub)]) {
+ const lines = txt.trimEnd().split("\n");
+ assert.equal(lines[lines.length - 1], `Generated by ${PROJECT_GENERATOR}`);
+ }
+});
+
+test("adding `generator` did NOT bump the corpus spec", () => {
+ // Guard on the reasoning, not just the number: prior bumps announced a new
+ // fetchable layer. An informational credit string breaks no reader, so a
+ // client pinned to spec 3 must keep validating. If this assertion is ever
+ // updated, the version-history comment in corpus.ts must justify why.
+ assert.equal(CORPUS_SPEC_VERSION, 3);
+ assert.equal(
+ buildSiteCorpus(descriptor(), { hasArchives: false }).spec,
+ 3,
+ );
+});
+
test("renderRobotsTxt: sitemap line only with an absolute siteUrl", () => {
assert.match(
renderRobotsTxt({ siteUrl: "https://demo.example/" }),
diff --git a/common/lib/corpus.ts b/common/lib/corpus.ts
@@ -1,4 +1,5 @@
import type { PublicSiteDescriptor } from "./siteDescriptor";
+import { PROJECT_GENERATOR } from "./project";
// The machine-readable corpus index emitted at `/corpus.json` on every export
// bundle (and an aggregate variant on a hub). It does NOT contain transcripts —
@@ -15,6 +16,14 @@ import type { PublicSiteDescriptor } from "./siteDescriptor";
// (postScheme + per-channel posts manifest pointers).
// v3: …and the DERIVED layer — AI digests (chapters + topic tags) served under
// the same shard scheme (digestScheme + per-channel digests manifest pointers).
+//
+// NOT a v4: the `generator` field added below is deliberately unversioned. Every
+// prior bump announced a new FETCHABLE LAYER — a reader that ignored it would
+// miss data it could otherwise have retrieved. `generator` is an informational
+// string that points at the software, not at any content; no reader's behaviour
+// changes by not knowing about it, and the only in-repo consumer
+// (mcp/src/source.ts) casts the parsed corpus loosely. Bumping the spec would
+// force every client to re-evaluate compatibility for a credit line.
export const CORPUS_SPEC_VERSION = 3;
// How to resolve a single transcript from the paginated shards, described once
@@ -124,6 +133,10 @@ export type SiteCorpus = {
spec: number;
kind: "site";
generatedAt: string;
+ // The software that produced this bundle, e.g. "Archilyzer (https://…)".
+ // Informational: it tells a machine reader where to find the tool that built
+ // the archive it's navigating. Not versioned — see CORPUS_SPEC_VERSION.
+ generator: string;
site: {
id: string;
title: string;
@@ -157,6 +170,8 @@ export type HubCorpus = {
spec: number;
kind: "hub";
generatedAt: string;
+ // See SiteCorpus.generator.
+ generator: string;
hub: { title: string; url?: string };
sites: HubCorpusSite[];
federation: { description: string };
@@ -224,6 +239,7 @@ export function buildSiteCorpus(
spec: CORPUS_SPEC_VERSION,
kind: "site",
generatedAt: descriptor.generatedAt,
+ generator: PROJECT_GENERATOR,
site: {
id: descriptor.siteId,
title: descriptor.siteTitle,
@@ -278,6 +294,7 @@ export function buildHubCorpus(
spec: CORPUS_SPEC_VERSION,
kind: "hub",
generatedAt: opts.generatedAt,
+ generator: PROJECT_GENERATOR,
hub: { title: opts.hubTitle, ...(opts.hubUrl ? { url: opts.hubUrl } : {}) },
sites,
federation: {
@@ -349,6 +366,8 @@ export function renderSiteLlmsTxt(corpus: SiteCorpus): string {
`- …and ${corpus.channels.length - shown.length} more — see corpus.json for the full list.`,
);
}
+ out.push("");
+ out.push(`Generated by ${corpus.generator}`);
return out.join("\n") + "\n";
}
@@ -380,6 +399,8 @@ export function renderHubLlmsTxt(corpus: HubCorpus): string {
for (const s of corpus.sites.slice(0, LLMS_INLINE_LIMIT)) {
out.push(`- ${s.title}: ${s.url} (corpus: ${s.corpus})`);
}
+ out.push("");
+ out.push(`Generated by ${corpus.generator}`);
return out.join("\n") + "\n";
}
diff --git a/common/lib/homepage.ts b/common/lib/homepage.ts
@@ -1,5 +1,6 @@
import fs from "node:fs";
import { getPaths, type Paths } from "./paths";
+import { PROJECT_NAME, PROJECT_TAGLINE } from "./project";
import {
getSettings,
normalizeSocialSvg,
@@ -26,17 +27,21 @@ export type HomepageConfig = {
// SiteSettings.socialLinks; an array (even empty) overrides it. Same semantics
// as Site.socialLinks — resolve with resolveHomepageSocialLinks() at render.
socialLinks?: SocialLink[];
- // Absolute public URL of the deployed hub, e.g. "https://archilyzer.com".
+ // Absolute public URL of the deployed hub, e.g. "https://archilyzer.pages.dev".
siteUrl?: string;
// Cloudflare Pages project the hub deploys to.
cloudflareProject?: string;
};
+// Neutral defaults for an unconfigured install. These are the PRODUCT strings
+// (common/lib/project.ts) standing in until the operator names their own
+// deployment in the editor — which is why they're pulled from the constant
+// rather than spelled out here twice.
function defaults(): HomepageConfig {
return {
- siteTitle: "Archilyzer",
- siteDescription: "Searchable video-transcript archives.",
- headerTitle: "Archilyzer",
+ siteTitle: PROJECT_NAME,
+ siteDescription: PROJECT_TAGLINE,
+ headerTitle: PROJECT_NAME,
homeTagline: "",
};
}
diff --git a/common/lib/homepageSummary.ts b/common/lib/homepageSummary.ts
@@ -1,6 +1,7 @@
import type { Platform } from "./platform";
import type { VideoStat } from "./stats";
import type { Site } from "./site";
+import { VIDEO_STATES, type VideoState } from "./availability";
// Pre-computed, lightweight cross-site summary for the hub (homepage) landing.
// Built once at compose time (compose-homepage.ts) and embedded into the SSG
@@ -16,7 +17,12 @@ import type { Site } from "./site";
// Two metrics are pre-binned at two granularities; Cumulative and Share (100%)
// are derived client-side from these, so no extra precompute is needed.
-export const HOMEPAGE_SUMMARY_VERSION = 3;
+// v4 adds `availability` (instance-wide counts by source-platform state) and
+// `status` on each recent item. Both are ADDITIVE and both are declared optional
+// on the type, because nothing gates on this number — it is a provenance marker,
+// not a compatibility check — and a summary written by an older build must keep
+// deserializing. Readers must treat their absence as "unknown", never as zero.
+export const HOMEPAGE_SUMMARY_VERSION = 4;
// Day buckets are capped to this many trailing days so the embedded summary stays
// small regardless of archive age (daily detail is only useful recently).
@@ -58,6 +64,10 @@ export type HomepageSummarySite = {
siteUrl: string; // public sites only — always present
transcribed: SiteMetricStat;
downloaded: SiteMetricStat;
+ // The site's own brand accent ("#rrggbb"), when it defines one. Absent for a
+ // site that hasn't set one — callers fall back to the chart palette rather
+ // than inventing a brand colour. Added in v4.
+ accent?: string;
};
export type HomepageChannelMeta = { slug: string; name: string };
@@ -71,6 +81,28 @@ export type HomepageRecentItem = {
siteTitle: string;
siteUrl: string;
transcribedDate: string; // YYYYMMDD
+ // Where this recording stands on the platform it came from, as of the last
+ // time we looked. Optional: summaries written before v4 don't carry it, and a
+ // renderer must show nothing rather than guess "available".
+ status?: VideoState;
+};
+
+// Instance-wide census of archived recordings by their state on the source
+// platform. This is the archive's whole point stated as a number, so the
+// honesty rules matter:
+//
+// • `available` is a FLOOR, not a fact. A recording counts as available until
+// someone re-checks it and finds otherwise; nobody re-checks 60,000 videos
+// continuously. It means "not known to be gone", never "confirmed safe".
+// • `deleted` is the opposite: every one of those was individually re-checked
+// and found gone. It is the only count here that is evidence.
+// Copy rendered from this must not launder the first into the second.
+export type HomepageAvailability = {
+ // Counts keyed by VideoState. Every state is present, zero included, so a
+ // renderer can iterate a stable set instead of probing for keys.
+ byState: Record<VideoState, number>;
+ // Total records censused — the denominator for any percentage.
+ counted: number;
};
export type HomepageSummary = {
@@ -91,9 +123,11 @@ export type HomepageSummary = {
series: Record<MetricKey, MetricSeries>;
// Public sites, sorted by transcribed total desc.
sites: HomepageSummarySite[];
- // Newest transcriptions (public universe), newest first. Built but not rendered
- // on the current landing — kept so the feed can be re-enabled with no rework.
+ // Newest transcriptions (public universe), newest first. Rendered as "the
+ // rail" on the project site's home page.
recent: HomepageRecentItem[];
+ // Instance-wide state census. Optional — absent from pre-v4 summaries.
+ availability?: HomepageAvailability;
};
const RECENT_LIMIT = 24;
@@ -292,6 +326,7 @@ export function buildHomepageSummary(
siteUrl: s.siteUrl,
transcribed: siteStatFrom(series.transcribed.month, s.siteId, nowMonth),
downloaded: siteStatFrom(series.downloaded.month, s.siteId, nowMonth),
+ ...(s.accent ? { accent: s.accent } : {}),
}))
// Keep a public site only if it has any activity in either metric.
.filter((s) => s.transcribed.total > 0 || s.downloaded.total > 0)
@@ -328,9 +363,22 @@ export function buildHomepageSummary(
siteTitle: site.siteTitle,
siteUrl: site.siteUrl,
transcribedDate: s.transcribedDate,
+ status: s.status,
};
});
+ // State census over every record, matching `totals`' instance-wide scope
+ // (pool-only channels included) rather than the charts' public-site universe.
+ // Seeded with every state at zero so the shape is stable across corpora.
+ const byState = Object.fromEntries(
+ VIDEO_STATES.map((s) => [s, 0]),
+ ) as Record<VideoState, number>;
+ for (const s of stats) byState[s.status] = (byState[s.status] ?? 0) + 1;
+ const availability: HomepageAvailability = {
+ byState,
+ counted: stats.length,
+ };
+
return {
version: HOMEPAGE_SUMMARY_VERSION,
generatedAt: now.toISOString(),
@@ -347,5 +395,6 @@ export function buildHomepageSummary(
series,
sites: summarySites,
recent,
+ availability,
};
}
diff --git a/common/lib/project.ts b/common/lib/project.ts
@@ -0,0 +1,38 @@
+// The project's own identity — the one place the product is named and located.
+//
+// This is deliberately PURE: zero imports, zero I/O, no framework types. That
+// makes it safe to pull into server components, `"use client"` trees, plain
+// `tsx` scripts and the MCP server alike, without dragging `node:fs` (or a
+// settings read) along with it.
+//
+// IDENTITY SPLIT — read this before adding a consumer. There are two identities
+// in this codebase and they are not the same thing:
+// • PRODUCT identity (here): what the software is called, where it lives, how
+// it describes itself. Same on every install. Used by the project site's
+// chrome, the "Built with Archilyzer" back-link every export site renders,
+// and the `generator` string in corpus.json.
+// • OPERATOR identity (common/lib/homepage.ts + settings.json): what THIS
+// operator calls their own deployment, their social links, their site
+// descriptions. Editable in the editor; differs per install.
+// A string that should change when someone else deploys this belongs there,
+// not here.
+export const PROJECT_NAME = "Archilyzer";
+
+// The canonical public home of the project site. Baked into every export
+// bundle's footer at build time, so changing it does not retroactively fix
+// already-deployed archives. If a bespoke domain is ever acquired, DO NOT swap
+// this and rebuild the world: point the new domain at Cloudflare Pages and add
+// a 301 from `archilyzer.pages.dev`, which keeps every deployed back-link
+// valid. Only then is changing this string a cosmetic follow-up.
+export const PROJECT_URL = "https://archilyzer.pages.dev";
+
+export const PROJECT_TAGLINE = "Self-hosted, searchable video-transcript archives.";
+
+// Where a visitor gets the source. A dated snapshot tarball — there is no
+// public git repository (see homepage/app/downloads).
+export const PROJECT_DOWNLOADS_URL = `${PROJECT_URL}/downloads/`;
+
+// The `generator` string stamped into corpus.json and llms.txt, so anything
+// that reads an archive machine-side can find the software that built it.
+// Shape mirrors the HTML <meta name="generator"> convention.
+export const PROJECT_GENERATOR = `${PROJECT_NAME} (${PROJECT_URL})`;
diff --git a/common/lib/settings.ts b/common/lib/settings.ts
@@ -196,7 +196,7 @@ export type SiteSettings = {
// common/lib/site.ts. The one presentation field that lives globally so a
// shared footer doesn't have to be repeated per site.
socialLinks: SocialLink[];
- // Absolute public URL of the family hub/homepage (e.g. "https://archilyzer.com").
+ // Absolute public URL of the family hub/homepage (e.g. "https://archilyzer.pages.dev").
// Every export site links back to it ("the family" backlink) when set. Empty =
// no hub link rendered. Normalized to a trailing-slash-free http(s) URL.
homepageUrl: string;
diff --git a/common/styles/fonts.ts b/common/styles/fonts.ts
@@ -15,8 +15,10 @@
// `.variable` classes must be present on <html> for the @font-face to be active,
// which is why every face is joined into `fontVars`), but only referenced when a
// family opts in:
-// • swiss → Archivo (a neo-grotesque; International Typographic Style)
-// • selenized → JetBrains Mono display+mono over IBM Plex Sans body (terminal)
+// • swiss → Archivo (a neo-grotesque; International Typographic Style)
+// • selenized → JetBrains Mono display+mono over IBM Plex Sans body (terminal)
+// • archilyzer → Archivo at 125% width (expanded, poster-scale display) over
+// IBM Plex Sans body + IBM Plex Mono data (the instrument face)
//
// `next/font` is transformed by Next's SWC loader across the whole compiled
// module graph, and `transpilePackages: ["yt-dlp-transcript-common"]` puts this
@@ -30,6 +32,7 @@ import {
Archivo,
JetBrains_Mono,
IBM_Plex_Sans,
+ IBM_Plex_Mono,
} from "next/font/google";
// --- Base / archive voice (the default super-family) ------------------------
@@ -52,9 +55,15 @@ export const fontMono = Source_Code_Pro({
// --- Per-family voices (opted into by tokens.css [data-theme] blocks) --------
// Swiss: a neo-grotesque for both display and body — tight, objective, Helvetica-
// adjacent. Used for --font-display + --font-sans under [data-theme="swiss"].
+// Archilyzer additionally drives Archivo's WIDTH axis, so `font-stretch: 125%`
+// yields a genuinely expanded cut for poster-scale display type — visibly a
+// different animal from the default-width Archivo that swiss uses. Declaring
+// the axis here (rather than loading Archivo a second time) costs swiss
+// nothing: wdth defaults to 100%, so its rendering is unchanged.
export const fontGrotesk = Archivo({
variable: "--font-grotesk",
subsets: ["latin"],
+ axes: ["wdth"],
});
// Selenized: a code face for display + mono (the "terminal" voice).
@@ -63,13 +72,22 @@ export const fontCode = JetBrains_Mono({
subsets: ["latin"],
});
-// Selenized: a humanist sans body that pairs with the code face.
+// Selenized + Archilyzer: a humanist sans body. Institutional, built for long
+// documentation — which is half of what the project site is.
export const fontPlex = IBM_Plex_Sans({
variable: "--font-plex",
subsets: ["latin"],
weight: ["400", "500", "600", "700"],
});
+// Archilyzer: tabular figures for counts, timecodes and recording states. The
+// data layer is a first-class voice in the instrument face, not an accent.
+export const fontPlexMono = IBM_Plex_Mono({
+ variable: "--font-plex-mono",
+ subsets: ["latin"],
+ weight: ["400", "500", "600"],
+});
+
// Space-joined `.variable` classes for <html className={...}>. EVERY face must be
// here so its @font-face is registered even if only some families reference it.
-export const fontVars = `${fontDisplay.variable} ${fontSans.variable} ${fontMono.variable} ${fontGrotesk.variable} ${fontCode.variable} ${fontPlex.variable}`;
+export const fontVars = `${fontDisplay.variable} ${fontSans.variable} ${fontMono.variable} ${fontGrotesk.variable} ${fontCode.variable} ${fontPlex.variable} ${fontPlexMono.variable}`;
diff --git a/common/styles/tokens.css b/common/styles/tokens.css
@@ -124,6 +124,19 @@ html[data-theme="swiss"] {
--font-display: var(--font-grotesk); /* Archivo neo-grotesque */
--font-sans: var(--font-grotesk);
}
+html[data-theme="archilyzer"] {
+ --radius: 0.125rem; /* 2px — a faceplate's edge break, not a rounded card */
+ --font-display: var(--font-grotesk); /* Archivo, widened below */
+ --font-sans: var(--font-plex); /* IBM Plex Sans — documentation body */
+ --font-mono: var(--font-plex-mono); /* IBM Plex Mono — counts, states, time */
+}
+/* Archivo carries a `wdth` axis (declared in common/styles/fonts.ts). Pushing
+ display type to 125% is what makes this family's headline voice EXPANDED —
+ engineered and poster-scale — rather than the same neo-grotesque swiss uses
+ at default width. Scoped to the display utility so body copy is unaffected. */
+html[data-theme="archilyzer"] .font-display {
+ font-stretch: 125%;
+}
/* ---------------------------------------------------------------------------
BASE family (default — no data-theme). Neutral zinc; blue brand accent.
@@ -182,6 +195,14 @@ html[data-theme="swiss"] {
--chart-grid: rgba(0, 0, 0, 0.08);
--chart-axis: #71717a;
--chart-tooltip-bg: #ffffff;
+
+ /* Recording state: "this no longer exists on its platform". Declared on the
+ base family (and on .dark below) rather than in each family block, so every
+ family inherits a usable value — a family that redeclares it (archilyzer)
+ simply wins on the same element. Without this, switching the project site
+ to another family via the ThemeMenu would leave state chips uncoloured. */
+ --state-gone: #dc2626;
+ --state-gone-soft: rgba(220, 38, 38, 0.12);
}
.dark {
@@ -233,6 +254,10 @@ html[data-theme="swiss"] {
--chart-grid: rgba(255, 255, 255, 0.08);
--chart-axis: #a1a1aa;
--chart-tooltip-bg: #18181b;
+
+ /* See :root — the dark default every family inherits unless it overrides. */
+ --state-gone: #f87171;
+ --state-gone-soft: rgba(248, 113, 113, 0.16);
}
/* ---------------------------------------------------------------------------
@@ -571,3 +596,144 @@ html[data-theme="swiss"] {
--chart-axis: #a3a3a3;
--chart-tooltip-bg: #161616;
}
+
+/* ---------------------------------------------------------------------------
+ ARCHILYZER family — the instrument face. This is the TOOL's voice, and it is
+ deliberately not the archives' voice: the warm brass "reading room" belongs
+ to the things Archilyzer builds, and dressing the instrument as its own
+ output flattens a hierarchy that should read instantly.
+
+ THE ARGUMENT, IN TOKENS: chrome is achromatic — graphite and bone, no
+ decorative brand hue anywhere. Colour appears in exactly two places and both
+ carry information: a recording's STATE (--state-gone, the one hot hue, worn
+ only by recordings that are gone), and the archives themselves (each site
+ card wears its own accent from site.json). The tool has no colour; the
+ recordings do.
+
+ --brand is therefore a cool signal, near-unused in chrome. It exists because
+ every family must set it — ThemeScript can override it per-user, shadcn
+ components reach for it — not because this family wants an accent. CTAs are
+ bone-on-graphite solid fills, not coloured buttons.
+
+ The ground is a LIFTED cool slate, not near-black: painted equipment under
+ light, rather than the void that every "dark mode with an acid accent"
+ landing page reaches for.
+ --------------------------------------------------------------------------- */
+[data-theme="archilyzer"] {
+ /* Light — bone ground, graphite ink. Anodized rather than papery. */
+ --background: #f3f6f7;
+ --surface: #e8edef;
+ --foreground: #161c21;
+ --card: #ffffff;
+ --card-foreground: #161c21;
+ --popover: #ffffff;
+ --popover-foreground: #161c21;
+ --primary: #202a31;
+ --primary-foreground: #f3f6f7;
+ --secondary: #e2e8eb;
+ --secondary-foreground: #202a31;
+ --muted: #e8edef;
+ --muted-foreground: #55646e;
+ --accent: #dde5e8;
+ --accent-foreground: #161c21;
+ --destructive: #a8412d;
+ --destructive-foreground: #ffffff;
+ --destructive-soft: rgba(168, 65, 45, 0.12);
+ --border: #d2dade;
+ --border-strong: #aeb9c0;
+ --input: #c4ced4;
+ --ring: #2f7f77;
+ --faint: #78868f;
+ --panel: rgba(255, 255, 255, 0.78);
+ --panel-2: rgba(232, 237, 239, 0.78);
+
+ --success: #1f7a4d;
+ --success-foreground: #ffffff;
+ --success-soft: rgba(31, 122, 77, 0.12);
+ --warning: #96650b;
+ --warning-foreground: #ffffff;
+ --warning-soft: rgba(150, 101, 11, 0.14);
+ --info: #1c5f96;
+ --info-foreground: #ffffff;
+ --info-soft: rgba(28, 95, 150, 0.12);
+
+ --brand: #2f7f77;
+ --brand-strong: #226058;
+ --brand-soft: rgba(47, 127, 119, 0.12);
+ --brand-ink: #ffffff;
+
+ /* Recording state. The ONLY hues the chrome permits, and they are readable as
+ a claim: "gone" is worn exclusively by recordings re-checked and found
+ absent from their platform. "kept" is deliberately achromatic — a recording
+ still being there is the unremarkable case, and colouring it would imply a
+ guarantee nobody re-verified. */
+ --state-gone: #a8412d;
+ --state-gone-soft: rgba(168, 65, 45, 0.12);
+
+ --chart-1: #2f7f77;
+ --chart-2: #1c5f96;
+ --chart-3: #a8412d;
+ --chart-4: #6a4f9c;
+ --chart-5: #8a6a1f;
+ --chart-surface: #ffffff;
+ --chart-grid: rgba(22, 28, 33, 0.09);
+ --chart-axis: #55646e;
+ --chart-tooltip-bg: #ffffff;
+}
+
+[data-theme="archilyzer"].dark {
+ /* Dark — the primary look. Lifted cool slate, bone text, hairline geometry. */
+ --background: #151b20;
+ --surface: #1c242b;
+ --foreground: #e7edf1;
+ --card: #1c242b;
+ --card-foreground: #e7edf1;
+ --popover: #202932;
+ --popover-foreground: #e7edf1;
+ --primary: #e7edf1;
+ --primary-foreground: #151b20;
+ --secondary: #242e37;
+ --secondary-foreground: #e7edf1;
+ --muted: #1c242b;
+ --muted-foreground: #a4b3bd;
+ --accent: #242e37;
+ --accent-foreground: #e7edf1;
+ --destructive: #c4553f;
+ --destructive-foreground: #f7eae7;
+ --destructive-soft: rgba(196, 85, 63, 0.18);
+ --border: #2a343c;
+ --border-strong: #3d4a54;
+ --input: rgba(231, 237, 241, 0.16);
+ --ring: #5fa8a0;
+ --faint: #8496a2;
+ --panel: rgba(28, 36, 43, 0.72);
+ --panel-2: rgba(36, 46, 55, 0.6);
+
+ --success: #5ab97f;
+ --success-foreground: #0c1f14;
+ --success-soft: rgba(90, 185, 127, 0.16);
+ --warning: #d3a03f;
+ --warning-foreground: #1d1605;
+ --warning-soft: rgba(211, 160, 63, 0.16);
+ --info: #6aa5d8;
+ --info-foreground: #071723;
+ --info-soft: rgba(106, 165, 216, 0.16);
+
+ --brand: #5fa8a0;
+ --brand-strong: #7cc2ba;
+ --brand-soft: rgba(95, 168, 160, 0.16);
+ --brand-ink: #0d1a19;
+
+ --state-gone: #c4553f;
+ --state-gone-soft: rgba(196, 85, 63, 0.16);
+
+ --chart-1: #5fa8a0;
+ --chart-2: #6aa5d8;
+ --chart-3: #c4553f;
+ --chart-4: #a98ede;
+ --chart-5: #d3a03f;
+ --chart-surface: #1a2229;
+ --chart-grid: rgba(231, 237, 241, 0.08);
+ --chart-axis: #8496a2;
+ --chart-tooltip-bg: #202932;
+}
diff --git a/create-archives.sh b/create-archives.sh
@@ -1,9 +1,69 @@
#!/bin/bash
-git archive -9 --format tar.gz -o "./export/public/code.tar.gz" main
+# Build the downloadable source snapshot served at /downloads/ on the project
+# site, plus a sidecar of facts about it.
+#
+# WHY A SNAPSHOT AND NOT A REPOSITORY. There is no public git remote, by choice.
+# `git archive` of `main` produces the working tree at one commit: no history, no
+# branches, no remote, nothing to `git pull`. The /downloads/ page says exactly
+# that rather than implying a clone.
+#
+# WHY THE SIDECAR. The filename is deliberately STABLE — a dated filename would
+# break every link the moment a new snapshot shipped. So the date, commit, size
+# and checksum live in snapshot.json beside it, and the page states the truth it
+# reads there instead of hardcoding facts that go stale. No sidecar → the page
+# renders an explanation and NO download link, rather than a link that lies.
+#
+# Safe to publish: `git archive` only includes TRACKED files, so the corpus
+# (transcripts/, gitignored), settings.json (untracked; only .example is
+# tracked) and node_modules are all structurally excluded.
+set -euo pipefail
+
+cd "$(dirname "$0")"
+
+REF="${1:-main}"
+OUT_DIR="homepage/public/downloads"
+TARBALL="$OUT_DIR/archilyzer-source.tar.gz"
+SIDECAR="$OUT_DIR/snapshot.json"
+
+mkdir -p "$OUT_DIR"
+
+# --prefix so the tarball unpacks into `archilyzer/` rather than spilling into
+# the current directory. The docs tell people to `cd archilyzer` afterwards.
+git archive -9 --format tar.gz --prefix=archilyzer/ -o "$TARBALL" "$REF"
+
+BYTES=$(stat -c %s "$TARBALL" 2>/dev/null || stat -f %z "$TARBALL")
+SHA=$(sha256sum "$TARBALL" | cut -d' ' -f1)
+COMMIT=$(git rev-parse "$REF")
+SUBJECT=$(git log -1 --format=%s "$REF")
+GENERATED_AT=$(date -u +%Y-%m-%dT%H:%M:%SZ)
+
+# node rather than a heredoc so the commit subject is JSON-escaped properly —
+# it is arbitrary human text and routinely contains quotes.
+GENERATED_AT="$GENERATED_AT" COMMIT="$COMMIT" SUBJECT="$SUBJECT" \
+BYTES="$BYTES" SHA="$SHA" node -e '
+ const fs = require("fs");
+ fs.writeFileSync(process.argv[1], JSON.stringify({
+ generatedAt: process.env.GENERATED_AT,
+ commit: process.env.COMMIT,
+ subject: process.env.SUBJECT,
+ bytes: Number(process.env.BYTES),
+ sha256: process.env.SHA,
+ }, null, 2) + "\n");
+' "$SIDECAR"
+
+echo "create-archives: $TARBALL ($(( BYTES / 1024 )) KiB) at ${COMMIT:0:12}"
+
+# Cloudflare Pages rejects any single asset over 25 MB. The snapshot is an order
+# of magnitude under that, but say so loudly if that ever stops being true —
+# silently shipping an asset Pages will refuse is a deploy that "succeeds" and
+# then 404s.
+if [ "$BYTES" -gt 26214400 ]; then
+ echo "create-archives: WARNING — $TARBALL exceeds Cloudflare Pages' 25 MB per-file limit." >&2
+fi
# Transcript archives are too big for CF pages (25mb)
# The comments are kept since they'll likely be useful later
-
+#
# cd transcripts
# # use xz over gz to get the file below 25MB
# git archive --format tar main | xz -9 > ../export/public/transcripts.tar.xz
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -1,6 +1,8 @@
# Changelog
## [Unreleased]
+- **The Homepage page is about the project's site now, not "the hub".** The `homepage` package stopped being *our shelf of archives* and became **Archilyzer's own site** — marketing home, documentation, source download — with the cross-site dashboard moved to `/stats/`. The editor page says so, and names the split it now sits on: the wordmark and page title are the **product's** identity and are no longer editable here, while the social links and the deploy target remain **yours**. The practical consequence to know about: the deployed site's meta description no longer comes from `homepage.json`, so the corpus-flavoured blurb that used to describe it (*"Search the transcripts of your favorite and least favorite creators!"*) has been replaced by a description of the software.
+- **Placeholder URLs stopped advertising a domain that isn't ours.** Four form hints and two library comments offered `https://archilyzer.com` as the example hub URL. That domain **belongs to someone else** — it redirects to an unrelated page — so anyone who typed the example verbatim would have pointed their sites at a stranger. Every occurrence now reads `https://archilyzer.pages.dev`, which is the real one.
- **The backfill card now breaks its figure down by kind, and every lane card gained a second line.** Backfill read "77,952 reachable · 77,134 need media", which looks inverted and is not: three kinds are running, and summing them produces a number in no unit at all. 99.5% of that "reachable" is attribution-from-text, which costs about one model call per transcript *chunk*; the whole of "need media" is diarization, which is a few hundred audio passes. Worse, the two overlap — most of the videos needing media for diarization are also inside the text-attribution total — so the corpus read as simultaneously all-actionable and all-blocked, and the biggest single fact was invisible: **77,463 videos are blocked on diarization's output**, which is precisely why attribution can only run text-only. The card now prints a line per kind (`Speaker diarization 329 · 77,134 need media`), each clause dropped at zero, and never adds them together. The other three lanes stopped being a status word over a sentence: transcription and downloads name their backlog and how many channels it spans, transcription adds worker occupancy, downloads adds free disk against the floor, and digest names how many channels the layer has reached plus the videos waiting on a transcript or held for want of a normalized one.
- **A full disk now stops the downloads lane visibly.** `disk.low` blocks downloads exactly as the manual pause does, but the lane derived its state from the manual toggle alone — so a real disk stop left the card reading "Idle" while nothing could move. The lane is now held by either switch and the card says which one is shut. The disk and manual-pause readouts moved off the pipeline band and onto the downloads card, so each fact is stated once, next to the button that changes it; the band keeps only what is pipeline-wide (running/queued, the sync heartbeat, the scheduler).
- **Cleanup now says what is holding the audio it can't reclaim, and what to run to get it back.** The page led with one number — how much space you can free right now — and said nothing about the rest of the disk. The sweep skips videos for four different reasons and reported them only as a line in a job log after the fact, with no bytes attached and nothing ranked. A **sieve** now runs down the page: all the audio on disk enters at the top, each gate siphons off its share (no transcript yet, the keep-latest window, do-not-clean pins, awaiting diarization), and the remainder steps down to the reclaimable figure the page already led with. A video leaves at the *first* gate it hits, exactly as the sweep's own cascade does, so the five figures are an attribution and never overlap — a pinned, undiarized video is counted once, under the pin. Below it, the **release ledger** ranks the channels holding the most, split by what it costs to get the space back: a run of a lane that is already weeks deep, or a setting that frees it the moment it changes. "Not counted" sits in the second group and is not called a hold — excluding a channel hides its bytes from the total, it never protected them, which makes it the fastest win on the page.
diff --git a/editor/app/homepage/HomepageConfigForm.tsx b/editor/app/homepage/HomepageConfigForm.tsx
@@ -49,7 +49,7 @@ export function HomepageConfigForm({ config }: { config: HomepageConfig }) {
<input
className={input}
name="siteUrl"
- placeholder="https://archilyzer.com"
+ placeholder="https://archilyzer.pages.dev"
defaultValue={config.siteUrl ?? ""}
/>
</label>
diff --git a/editor/app/homepage/page.tsx b/editor/app/homepage/page.tsx
@@ -13,11 +13,14 @@ export default function HomepageAdminPage() {
return (
<div className="flex flex-col gap-8">
<div>
- <h1 className="text-2xl font-semibold">Homepage (hub)</h1>
+ <h1 className="text-2xl font-semibold">Homepage (project site)</h1>
<p className="mt-1 text-sm text-muted-foreground">
- The Archilyzer hub, built by the <code>homepage</code> package. Its
- home page is a cross-site stats landing built from the whole channel
- pool; only its branding is edited here.
+ The <code>homepage</code> package builds Archilyzer’s own site —
+ the marketing home, the documentation and the source download. Its{" "}
+ <code>/stats/</code> route is a cross-site dashboard built from the
+ whole channel pool. Only the operator-facing bits are edited here: the
+ wordmark and page title are the product’s own, but these social
+ links and the deploy target are yours.
</p>
</div>
diff --git a/editor/app/settings/components/SettingsForm.tsx b/editor/app/settings/components/SettingsForm.tsx
@@ -66,7 +66,7 @@ export function SettingsForm({ initial, apps, digestApps }: Props) {
name="homepageUrl"
defaultValue={initial.homepageUrl}
type="url"
- hint="Absolute URL of the family hub/homepage (e.g. https://archilyzer.com). Every export site links back to it. Leave blank for no hub link."
+ hint="Absolute URL of the family hub/homepage (e.g. https://archilyzer.pages.dev). Every export site links back to it. Leave blank for no hub link."
/>
<Field
label="Max transcript page bytes"
diff --git a/editor/app/sites/components/SiteForm.tsx b/editor/app/sites/components/SiteForm.tsx
@@ -248,7 +248,7 @@ export function SiteForm({ initial, channels, allSites, isNew }: Props) {
label="Hub URL"
name="hubUrl"
defaultValue={initial.hubUrl ?? ""}
- hint="The hub this site belongs under (e.g. https://archilyzer.com). Shows a 'Hub' backlink and lets the hub recognize this site as a member. Leave blank to inherit the family default from Settings."
+ hint="The hub this site belongs under (e.g. https://archilyzer.pages.dev). Shows a 'Hub' backlink and lets the hub recognize this site as a member. Leave blank to inherit the family default from Settings."
/>
<label className="flex items-center gap-2 text-sm">
<input
diff --git a/editor/e2e/export-search.spec.ts b/editor/e2e/export-search.spec.ts
@@ -1,4 +1,5 @@
import { test, expect, type Page } from "@playwright/test";
+import { PROJECT_URL } from "../../common/lib/project";
import { readJson, writeSettings } from "./helpers";
const EXPORT_BASE = `http://localhost:${process.env.EXPORT_PORT ?? 3010}`;
@@ -571,14 +572,25 @@ test.describe("export footer", () => {
await waitForHydration(page);
});
- test("footer holds downloads + social links sourced from the site", async ({
+ test("footer credits the project + carries social links from the site", async ({
page,
}) => {
const footer = page.locator("footer");
await expect(footer).toBeVisible();
- await expect(
- footer.getByRole("link", { name: "code.tar.gz" }),
- ).toBeVisible();
+
+ // Every deployed archive links back to the project that built it. This
+ // replaced a per-site `code.tar.gz` link — each archive used to ship its
+ // own copy of the source, which is now published once, centrally.
+ //
+ // PROJECT_URL is imported rather than typed out so the assertion cannot
+ // drift from the constant that is baked into every instance.
+ const builtWith = footer.getByRole("link", { name: "Archilyzer" });
+ await expect(builtWith).toBeVisible();
+ await expect(builtWith).toHaveAttribute("href", PROJECT_URL);
+ // Navigation, not a social link: same tab, matching the header's hub link.
+ await expect(builtWith).not.toHaveAttribute("target", "_blank");
+ // The accessible name is exactly the product, with "Built with" outside it.
+ await expect(footer).toContainText("Built with Archilyzer");
// The fixture site (editor/e2e/fixtures/sites/testsite) ships one GitHub
// social link.
diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md
@@ -1,5 +1,10 @@
# Changelog
+## [0.8.7] - 2026-08-12
+- **Every archive now says what built it, and stopped shipping its own copy of the source.** The footer carried a `code.tar.gz` link on every site — `create-archives.sh` wrote a tarball into `export/public/` and each deployed archive served its own duplicate of the whole workspace (verified live: `200 application/gzip` on jeralyzer). That link is **gone**, replaced by **"Built with [Archilyzer](https://archilyzer.pages.dev)"** beside the social links. The source is now published once, centrally, with a checksum and a commit — rather than N times, undated, with no way to tell two copies apart. The back-link is deliberately **ungated by instance mode** (a hub is Archilyzer too, and a credit that appeared on sites but not hubs would make them look like different software), **independent of any configured hub URL** (that is the operator's family link, a different relationship), and opens **in the same tab**, matching the header's hub link — only social icons open new ones. *"Built with"* sits outside the anchor so the accessible name is exactly `Archilyzer`. One knock-on: the **Downloads** eyebrow used to be unconditional because that link always followed it; it now appears only when the transcript-archive or offline link is actually there, instead of captioning empty space.
+- **`corpus.json` and `llms.txt` name the software that produced them.** Both the per-site and the hub builders stamp a `generator` string, and both `llms.txt` renderings end with a `Generated by …` line — so anything reading an archive machine-side can find the tool that built it. **The corpus spec version is deliberately unchanged at 3**, and a test asserts it: every prior bump announced a new *fetchable layer*, where a client that ignored it would miss retrievable data. An informational credit breaks no reader, so forcing every consumer to re-evaluate compatibility for a byline would be the wrong trade. The reasoning is recorded in the version-history comment rather than left for the next person to reconstruct.
+- **The project is MIT licensed.** No licence existed before, which meant a published tarball granted its recipients nothing at all. `LICENSE` at the root, `license` in the root manifest.
+
## [0.8.6] - 2026-08-11
- **MCP: the corpus's signature question is now askable, and it answers in under a second.** `runSearchSpec` — the engine that supports availability-state, upload-date, media-type and age filters, plus the chat/description/tags scopes — had **exactly one call site**: the `open_link` handler. Every filter was therefore unreachable unless a human pasted a viewer share link, which meant *"what did the videos that have since been DELETED say about X"* — the question that motivated the last sweep — had no path at all. `search_transcripts` **and** `enumerate_matches` now take the filters as flat arguments from **one shared schema constant**: `states` (available / maybe_missing / deleted / private / members_only / unlisted), `date_from`/`date_to`, `media_type`, `age`, `scopes` (transcripts / chat / description / tags / metadata / posts) and **`exclude`** for video-level NOT — `"cup"` but not `"world cup"`, the one boolean case that actually bites. Both tools gain them **together**, deliberately from the same constant, because a filter reachable from one and not the other would make the two disagree about coverage, which is precisely the failure the stateless rebuild set out to make impossible. Arbitrary boolean trees stay `open_link`'s job — that is what a share link is *for*, and asking a model to author a `qt=` tree in a tool call would trade a real capability for a new class of malformed input. Unrecognised tokens are **named in the footer** rather than dropped: a typo'd state would otherwise widen the search back to the whole corpus and return a perfectly legitimate-looking answer. Non-timed layers (description / tags / channel name) emit `[description]`-tagged snippets with **no timestamp link**, since citing a description line as `@ 0:00` would assert that someone said it.
- **MCP: a filtered query now reads a fraction of the corpus instead of all of it.** `summaries/` is a **global index of every video** — slug, channel, title, upload date, livestream, age-restricted, presence state — that parses in well under a second, against ~1.3 GB and tens of seconds for the transcripts. The MCP already read it in `availabilityMap()` and **threw away everything except the presence state**. It now keeps the whole record, and because each channel manifest carries `slugToPage`, **a filtered query's exact page set is computable before a single transcript byte is read**. Measured on the live 170-page corpus, on an idle box, in one run: the removed-videos question drops from a **170-page / 1,288 MB** full scan to **8 pages / 62.6 MB — 0.82 s against the 16.4 s** the same term costs unfiltered; a one-year date range to 32 pages / 202 MB (2.6 s). The invariant that makes it safe is that **the index only ever prunes pages; the record predicate still decides every hit** — both call the same `passesFilters`, so they cannot drift, and a video the index has never heard of gets its page read unconditionally. A stale, partial or missing summaries set therefore costs time, never correctness. It pays in proportion to how selective the filter is (a broad attribute spread across the corpus gets almost nothing) and **says which path ran** — *"filter-pruned: planned 8 of 170 page(s)"* — so a slow query is explicable. An unfiltered query never reads the index at all: a filter that excludes nothing is treated as no filter, so planning can never make a whole-corpus scan slower.
diff --git a/export/app/components/Footer.tsx b/export/app/components/Footer.tsx
@@ -7,6 +7,10 @@ import {
resolveRelatedSites,
resolveSocialLinks,
} from "yt-dlp-transcript-common/lib/site";
+import {
+ PROJECT_NAME,
+ PROJECT_URL,
+} from "yt-dlp-transcript-common/lib/project";
import { currentSite } from "../lib/site";
import { instanceMode } from "../lib/mode";
import { hasArchives } from "../lib/archives";
@@ -24,15 +28,18 @@ export default function Footer() {
<footer className="mt-auto border-t border-border bg-muted/30">
<div className="max-w-6xl mx-auto px-4 sm:px-6 py-4 flex items-center justify-between gap-4 text-sm text-muted-foreground">
<div className="flex items-center gap-3">
- <span className="font-mono text-xs uppercase tracking-[0.14em]">
- Downloads
- </span>
- <a
- href="/code.tar.gz"
- className="underline underline-offset-2 hover:text-foreground transition-colors"
- >
- code.tar.gz
- </a>
+ {/* The "Downloads" eyebrow labels the two links that may follow it, so
+ it is conditional on at least one existing. It used to be
+ unconditional because a `code.tar.gz` link always sat behind it —
+ every archive shipped its own copy of the source. That is now the
+ project site's job (one snapshot, one place, one checksum), and
+ leaving the label behind would caption an empty space on a site
+ with neither archives nor a PWA. */}
+ {(hasArchives() || site.pwa) && (
+ <span className="font-mono text-xs uppercase tracking-[0.14em]">
+ Downloads
+ </span>
+ )}
{/* Transcript & live-chat archive zips, generated per build. Shown only
when this build actually produced servable archives. */}
{hasArchives() && (
@@ -64,7 +71,24 @@ export default function Footer() {
Use with AI
</a>
</div>
- {socialLinks.length > 0 && (
+ <div className="flex items-center gap-4">
+ {/* "Built with Archilyzer" — deliberately NOT gated on instanceMode()
+ (being identical on every deployment is the point), NOT dependent
+ on a configured hubUrl (that is the operator's family link, a
+ different thing), and NOT target="_blank": the header's hub link
+ navigates in the same tab, and only the social icons open new ones.
+ "Built with" sits OUTSIDE the anchor so the accessible name is
+ exactly "Archilyzer". */}
+ <span className="text-xs whitespace-nowrap">
+ Built with{" "}
+ <a
+ href={PROJECT_URL}
+ className="underline underline-offset-2 hover:text-foreground transition-colors"
+ >
+ {PROJECT_NAME}
+ </a>
+ </span>
+ {socialLinks.length > 0 && (
<ul className="flex items-center gap-3 list-none">
{socialLinks.map((link, i) => (
<li key={`${link.url}-${i}`}>
@@ -80,7 +104,8 @@ export default function Footer() {
</li>
))}
</ul>
- )}
+ )}
+ </div>
</div>
{related.length > 0 && (
<nav
diff --git a/export/e2e-hub/federation.spec.ts b/export/e2e-hub/federation.spec.ts
@@ -1,4 +1,5 @@
import { expect, test, type Page, type Route } from "@playwright/test";
+import { PROJECT_URL } from "../../common/lib/project";
// Federation core: the hub reads a DIFFERENT origin's static JSON contract over
// (simulated) CORS and merges it into one search. Origin B is never really
@@ -125,6 +126,19 @@ test.describe("hub federation — cross-origin browse + search", () => {
).toBeVisible();
});
+ test("hub mode carries the same project back-link as a site", async ({
+ page,
+ }) => {
+ // The back-link is deliberately NOT gated on instanceMode(): a hub is still
+ // something Archilyzer built, and a credit that appears on sites but not on
+ // hubs would make the two look like different software.
+ const builtWith = page.locator("footer").getByRole("link", {
+ name: "Archilyzer",
+ });
+ await expect(builtWith).toBeVisible();
+ await expect(builtWith).toHaveAttribute("href", PROJECT_URL);
+ });
+
test("adds an archive by URL and shows it on the shelf", async ({ page }) => {
await mockOriginB(page);
await addArchive(page, ORIGIN_B);
diff --git a/export/e2e/site-branding.spec.ts b/export/e2e/site-branding.spec.ts
@@ -1,4 +1,5 @@
import { test, expect } from "@playwright/test";
+import { PROJECT_URL } from "../../common/lib/project";
import { installRoutes } from "./helpers";
// The export app is built per-site; its SSR branding (document title, header,
@@ -29,6 +30,48 @@ test("footer renders the site's social links", async ({ page }) => {
).toHaveAttribute("href", "https://github.com/example");
});
+test("footer credits the project, ungated and in the same tab", async ({
+ page,
+}) => {
+ await installRoutes(page);
+ await page.goto("/");
+ const footer = page.locator("footer");
+ const builtWith = footer.getByRole("link", { name: "Archilyzer" });
+ // Every archive carries this, on every instance mode, whether or not the
+ // operator configured a hub. That sameness is the point of the back-link.
+ await expect(builtWith).toBeVisible();
+ await expect(builtWith).toHaveAttribute("href", PROJECT_URL);
+ await expect(builtWith).not.toHaveAttribute("target", "_blank");
+ await expect(footer).toContainText("Built with Archilyzer");
+});
+
+test("the Downloads eyebrow appears only when something is behind it", async ({
+ page,
+}) => {
+ await installRoutes(page);
+ await page.goto("/");
+ const footer = page.locator("footer");
+
+ // The label used to be unconditional, because a `code.tar.gz` link always
+ // followed it. With that link gone, the invariant is that the eyebrow and
+ // the links it captions appear together or not at all.
+ //
+ // Asserted as an invariant rather than a fixed expectation on purpose:
+ // hasArchives() reads the real export/public on disk, so whether this build
+ // has archives depends on what was last composed. A test pinned to one
+ // answer would pass or fail on the machine's state, not on the code.
+ const archives = footer.getByRole("link", { name: "Transcripts & chat" });
+ const offline = footer.getByRole("link", { name: "Offline" });
+ const behind = (await archives.count()) + (await offline.count());
+ const eyebrow = footer.getByText("Downloads", { exact: true });
+
+ if (behind > 0) {
+ await expect(eyebrow).toBeVisible();
+ } else {
+ await expect(eyebrow).toHaveCount(0);
+ }
+});
+
test("per-site accent is baked onto <html> as a brand override", async ({
page,
}) => {
diff --git a/homepage/CHANGELOG.md b/homepage/CHANGELOG.md
@@ -1,5 +1,62 @@
# Homepage Changelog
+## 2026-08-12
+
+- **This package is the project's site now.** It was a shelf of one operator's archives
+ dressed in the same warm-brass costume as the archives themselves. It is now the place
+ that explains, documents and distributes **Archilyzer** — because three live public
+ archives existed and *nothing said what built them*, while the root README was still
+ titled "yt-dlp transcript browser" and claimed three packages when there are five.
+ - **`/` is the marketing home.** What the software is, what it produces, and a download.
+ - **`/stats/` is the dashboard**, moved verbatim — restoring the name it had before
+ `33ed2fe` deleted it. Nothing about the charts changed.
+ - **`/docs/`** renders hand-written Markdown from `content/docs/`, ordered and grouped by
+ a typed manifest (`content/docs.ts`) rather than frontmatter, so a typo is a `tsc`
+ error and the nav order is one reviewable list.
+ - **`/downloads/`** publishes the source snapshot; **`/changelog/`** renders the viewer's.
+
+- **The site builds without a corpus — which is the point.** `prebuild` chained a multi-GB
+ LMDB index and a whole-pool compose onto every `next build`, so the project's own
+ marketing site could not be built by someone who had merely unpacked the source. `build`
+ and **`build:nodata`** now run the *identical* command; the only difference is that
+ pnpm's lifecycle hook matches on the script **name**, so `prebuild` fires for one and not
+ the other. Routes that want numbers **degrade honestly**: `loadSummary()` returns `null`
+ rather than a zero-filled summary, `/` drops its numeric bands entirely, and `/stats/`
+ still exists (so the nav never 404s) and explains why it is empty. *"0 transcripts · 0
+ sites"* would have read as a broken product rather than an unconfigured build.
+
+- **A visual identity of its own — the instrument face.** New `archilyzer` theme family:
+ graphite and bone, `2px` radius, hairline geometry, Archivo at 125% width for display,
+ IBM Plex Sans for documentation body and IBM Plex Mono for every number. The warm radial
+ glow and film grain are **cut**. The argument it makes is that the tool should stop
+ dressing as its own output: **chrome is achromatic**, and colour appears in exactly two
+ places, both carrying information — a recording's state at its source (`--state-gone`,
+ the one hot hue, worn only by recordings found gone) and the archives themselves, each
+ card wearing its own accent. The tool has no colour; the recordings do. The family
+ declares the **complete** token set in **both** modes, since the theme toggle ships and a
+ half-declared palette silently inherits the base family's values in whichever mode nobody
+ checked — a spec now asserts all fifty tokens resolve in light and dark.
+
+- **The rail: real acquisitions, never invented ones.** The home page's signature object is
+ a log of genuinely recent recordings — date, channel, title, and state at source. It
+ revives `RecentAdditions`, which had been built and left unrendered. The hard rule is
+ that **nothing on this page may be fabricated**: absent data renders an empty state, and
+ a recording whose state is unknown gets no chip at all rather than a cheerful default.
+ Inventing transcript rows on the marketing site for an archiving tool would undercut the
+ entire product. For the same reason every figure is computed at build time and none is
+ hardcoded — the public-archive count went from three to four *during* this work.
+
+- **The proof point is stated carefully.** 471 recordings in the reference corpus no longer
+ exist where they came from, and that is the argument for the software. But `available` is
+ a **floor, not a fact**: a recording counts as available until something re-checks it and
+ finds otherwise, and nothing re-checks 60,000 videos continuously. The copy says
+ *"simply not known to be gone"* and never *"everything else is safe"*.
+
+- **Two pre-existing e2e bugs fixed.** The Playwright config read `process.env.PORT` — the
+ **editor's** port variable — instead of `HOMEPAGE_E2E_PORT`, which `scripts/worktree.mjs`
+ had been allocating for it all along and nothing read; and the `e2e` script never took the
+ machine-global queue lock, so it could race a running export suite for ports.
+
## 2026-06-29
- **Charts rebuilt around presets + a power-user explorer.** The single five-control
diff --git a/homepage/app/changelog/page.tsx b/homepage/app/changelog/page.tsx
@@ -0,0 +1,53 @@
+import { readFileSync } from "node:fs";
+import path from "node:path";
+import type { Metadata } from "next";
+import { Changelog } from "yt-dlp-transcript-common/components/Changelog";
+import { PageShell, PageHeading } from "../components/PageShell";
+
+export const metadata: Metadata = {
+ title: "Changelog",
+ description:
+ "What has changed in the published archive sites Archilyzer builds.",
+};
+
+// The viewer's changelog — the one that describes what a VISITOR to a published
+// archive sees change. Read from the export package at build time, mirroring
+// export/app/changelog/page.tsx's readFileSync-from-cwd approach; `output:
+// "export"` means this runs once at build, never per request.
+//
+// It reaches outside this package (to ../export), which is fine: the source
+// tarball is a `git archive` of the whole monorepo, so the file is always
+// alongside. It is read defensively anyway — a missing changelog should degrade
+// to a sentence, not fail the build.
+function loadChangelog(): string | null {
+ try {
+ return readFileSync(
+ path.join(process.cwd(), "..", "export", "CHANGELOG.md"),
+ "utf8",
+ );
+ } catch {
+ return null;
+ }
+}
+
+export default function ChangelogPage() {
+ const source = loadChangelog();
+ return (
+ <PageShell className="flex flex-col gap-8">
+ <PageHeading
+ eyebrow="Release notes"
+ title="Changelog"
+ standfirst="Changes to the archive sites Archilyzer publishes — the search, the viewer, the machine-readable corpus. Newest first."
+ />
+ {source ? (
+ <div className="doc-measure">
+ <Changelog source={source} />
+ </div>
+ ) : (
+ <p className="doc-measure text-[var(--muted-foreground)] leading-relaxed">
+ The changelog file wasn’t found in this build.
+ </p>
+ )}
+ </PageShell>
+ );
+}
diff --git a/homepage/app/components/ArchiveCards.tsx b/homepage/app/components/ArchiveCards.tsx
@@ -0,0 +1,70 @@
+import type { HomepageSummarySite } from "yt-dlp-transcript-common/lib/homepageSummary";
+import { seriesColor } from "yt-dlp-transcript-common/lib/homepageChart";
+
+// THE ONLY POLYCHROME MOMENT ON THE PAGE.
+//
+// The chrome of this family spends no colour at all. Here it does, and the
+// colour carries the argument: the tool is achromatic, the archives are not. A
+// visitor should be able to see at a glance that these are distinct
+// publications, not one product's sections.
+//
+// WHY NOT SiteGrid. SiteGrid is a `"use client"` component whose props ARE chart
+// state — metric, breakdown, a hidden-set and a toggle callback — because on
+// /stats/ its cards double as the chart legend. There is no chart here, so
+// reusing it would mean inventing chart state and shipping a client bundle to
+// render a static list. This is the same data, rendered flat and server-side.
+//
+// COLOUR SOURCE, in order: the site's own `accent` from its config, when it has
+// one. Otherwise the colour that site already wears in the dashboard's charts —
+// which is a legend assignment, deterministic per site, and not a brand claim
+// being invented on the site's behalf.
+export function ArchiveCards({ sites }: { sites: HomepageSummarySite[] }) {
+ if (sites.length === 0) return null;
+ return (
+ // Hairline geometry without phantom cells: each card draws its own right
+ // and bottom rule and the list draws the top and left, so a row that
+ // doesn't divide evenly leaves empty SPACE rather than an empty filled
+ // cell. (A `gap-px` grid over a coloured background paints the gaps of the
+ // missing cells too, which reads as a broken card.)
+ <ul className="grid grid-cols-1 border-t border-l border-[var(--border)] sm:grid-cols-2 lg:grid-cols-3 list-none">
+ {sites.map((site, i) => {
+ const color = site.accent ?? seriesColor(i);
+ return (
+ <li
+ key={site.siteId}
+ className="border-r border-b border-[var(--border)] bg-[var(--surface)]"
+ >
+ <a
+ href={site.siteUrl}
+ target="_blank"
+ rel="noopener noreferrer"
+ className="group flex h-full flex-col gap-3 p-5 transition-colors hover:bg-[var(--panel-2)]"
+ >
+ <div className="flex items-center gap-3">
+ <span
+ aria-hidden="true"
+ className="h-5 w-[3px] shrink-0"
+ style={{ backgroundColor: color }}
+ />
+ <h3 className="font-display text-base font-bold uppercase tracking-[0.05em] text-[var(--foreground)]">
+ {site.siteTitle}
+ </h3>
+ </div>
+ {site.siteDescription && (
+ <p className="line-clamp-2 text-sm leading-relaxed text-[var(--muted-foreground)]">
+ {site.siteDescription}
+ </p>
+ )}
+ <div className="mt-auto flex items-baseline gap-2 pt-2">
+ <span className="tabular text-lg text-[var(--foreground)]">
+ {site.transcribed.total.toLocaleString()}
+ </span>
+ <span className="label-machine">transcripts</span>
+ </div>
+ </a>
+ </li>
+ );
+ })}
+ </ul>
+ );
+}
diff --git a/homepage/app/components/CustomizePanel.tsx b/homepage/app/components/CustomizePanel.tsx
@@ -116,7 +116,7 @@ export function CustomizePanel({
aria-pressed={state.excludeTop}
onClick={() => onPatch({ excludeTop: !state.excludeTop })}
className={
- "rounded-full border px-3 py-1.5 text-[0.8125rem] font-medium transition-all duration-150 " +
+ "rounded-[var(--radius)] border px-3 py-1.5 text-[0.8125rem] font-medium transition-all duration-150 " +
(state.excludeTop
? "border-[var(--brand)] bg-[var(--brand-soft)] text-[var(--brand-strong)]"
: "border-[var(--border)] text-[var(--muted-foreground)] hover:border-[var(--border-strong)] hover:text-[var(--foreground)]")
diff --git a/homepage/app/components/DocBody.tsx b/homepage/app/components/DocBody.tsx
@@ -0,0 +1,194 @@
+import type { ReactNode, AnchorHTMLAttributes } from "react";
+import { Children } from "react";
+import Link from "next/link";
+import Markdown from "markdown-to-jsx";
+import { CopyLinkButton } from "yt-dlp-transcript-common/components/CopyLinkButton";
+
+// The documentation renderer.
+//
+// WHY NOT common/components/Markdown.tsx. Three reasons, each disqualifying on
+// its own: it hard-codes `target="_blank"` on EVERY anchor, so an internal
+// cross-link between doc pages would open a new tab; it is tuned for chat
+// density (small headings, tight rhythm), which is wrong for a document someone
+// reads top to bottom; and it has no heading anchors, so a section can't be
+// linked to. This borrows Changelog.tsx's anchored-heading pattern instead.
+//
+// A TRAP WORTH NAMING: markdown-to-jsx does NOT escape raw HTML by default
+// (`disableParsingRawHTML` is off), despite the header comment on Markdown.tsx
+// claiming it does. That is safe here only because these files are hand-written
+// and in-repo — never render untrusted Markdown through this component.
+
+function flattenText(children: ReactNode): string {
+ let out = "";
+ Children.forEach(children, (child) => {
+ if (child == null || typeof child === "boolean") return;
+ if (typeof child === "string" || typeof child === "number") {
+ out += String(child);
+ return;
+ }
+ if (typeof child === "object" && "props" in child) {
+ const props = (child as { props?: { children?: ReactNode } }).props;
+ if (props?.children !== undefined) out += flattenText(props.children);
+ }
+ });
+ return out;
+}
+
+function slugify(input: string): string {
+ return input
+ .toLowerCase()
+ .replace(/[^a-z0-9]+/g, "-")
+ .replace(/-+/g, "-")
+ .replace(/^-|-$/g, "");
+}
+
+function anchored(
+ Tag: "h2" | "h3",
+ className: string,
+): (p: { children?: ReactNode }) => ReactNode {
+ return function Heading({ children }) {
+ const anchor = slugify(flattenText(children).trim()) || null;
+ return (
+ <Tag id={anchor ?? undefined} className={className}>
+ <span>{children}</span>
+ {anchor && <CopyLinkButton anchor={anchor} />}
+ </Tag>
+ );
+ };
+}
+
+// A leading `# Title` is stripped by the loader, but a stray one shouldn't
+// out-shout the page's own title if a file grows a second top-level heading.
+function H1({ children }: { children?: ReactNode }) {
+ return (
+ <h2 className="scroll-mt-24 mt-12 mb-3 font-display text-xl font-bold uppercase tracking-[0.04em] text-[var(--foreground)]">
+ {children}
+ </h2>
+ );
+}
+
+// Internal links navigate in place; external ones open in a new tab. Getting
+// this backwards is exactly the bug that ruled out the shared Markdown
+// component, so the test suite asserts an internal cross-link has no `target`.
+function Anchor({
+ href,
+ children,
+ ...rest
+}: AnchorHTMLAttributes<HTMLAnchorElement>) {
+ const className =
+ "text-[var(--foreground)] underline decoration-[var(--border-strong)] underline-offset-[3px] hover:decoration-[var(--brand)] transition-colors";
+ if (!href) return <span className={className}>{children}</span>;
+ if (/^https?:\/\//i.test(href)) {
+ return (
+ <a
+ {...rest}
+ href={href}
+ target="_blank"
+ rel="noopener noreferrer"
+ className={className}
+ >
+ {children}
+ </a>
+ );
+ }
+ // In-page fragments stay plain anchors; next/link would push a history entry
+ // for what is just a scroll.
+ if (href.startsWith("#")) {
+ return (
+ <a {...rest} href={href} className={className}>
+ {children}
+ </a>
+ );
+ }
+ return (
+ <Link href={href} className={className}>
+ {children}
+ </Link>
+ );
+}
+
+const OVERRIDES = {
+ h1: { component: H1 },
+ h2: {
+ component: anchored(
+ "h2",
+ "scroll-mt-24 mt-12 mb-3 font-display text-xl font-bold uppercase tracking-[0.04em] text-[var(--foreground)] flex items-baseline",
+ ),
+ },
+ h3: {
+ component: anchored(
+ "h3",
+ "scroll-mt-24 mt-8 mb-2 font-sans text-base font-semibold text-[var(--foreground)] flex items-baseline",
+ ),
+ },
+ h4: {
+ props: {
+ className:
+ "mt-6 mb-2 label-machine",
+ },
+ },
+ p: { props: { className: "my-4 leading-[1.7] text-[var(--muted-foreground)]" } },
+ ul: {
+ props: {
+ className:
+ "list-disc pl-5 my-4 space-y-2 leading-[1.7] text-[var(--muted-foreground)] marker:text-[var(--faint)]",
+ },
+ },
+ ol: {
+ props: {
+ className:
+ "list-decimal pl-5 my-4 space-y-2 leading-[1.7] text-[var(--muted-foreground)] marker:text-[var(--faint)]",
+ },
+ },
+ li: { props: { className: "leading-[1.7]" } },
+ strong: { props: { className: "font-semibold text-[var(--foreground)]" } },
+ em: { props: { className: "italic" } },
+ code: {
+ props: {
+ className:
+ "px-1 py-0.5 rounded-[var(--radius)] bg-[var(--surface)] border border-[var(--border)] font-mono text-[0.85em] text-[var(--foreground)]",
+ },
+ },
+ pre: {
+ props: {
+ className:
+ "my-5 p-4 rounded-[var(--radius)] bg-[var(--surface)] border border-[var(--border)] overflow-x-auto text-[0.8125rem] leading-relaxed font-mono [&_code]:p-0 [&_code]:border-0 [&_code]:bg-transparent",
+ },
+ },
+ a: { component: Anchor },
+ hr: { props: { className: "my-10 border-t border-[var(--border)]" } },
+ blockquote: {
+ props: {
+ className:
+ "my-5 pl-4 border-l-2 border-[var(--border-strong)] text-[var(--muted-foreground)]",
+ },
+ },
+ table: {
+ props: {
+ className:
+ "my-5 block w-full overflow-x-auto text-left border-collapse text-sm",
+ },
+ },
+ th: {
+ props: {
+ className:
+ "border-b border-[var(--border-strong)] px-3 py-2 font-mono text-[0.6875rem] uppercase tracking-[0.12em] text-[var(--faint)] whitespace-nowrap",
+ },
+ },
+ td: {
+ props: {
+ className:
+ "border-b border-[var(--border)] px-3 py-2 align-top text-[var(--muted-foreground)]",
+ },
+ },
+};
+
+export function DocBody({ source }: { source: string }) {
+ return (
+ <div className="doc-measure">
+ <Markdown options={{ overrides: OVERRIDES, forceBlock: true }}>
+ {source}
+ </Markdown>
+ </div>
+ );
+}
diff --git a/homepage/app/components/Footer.tsx b/homepage/app/components/Footer.tsx
@@ -1,38 +1,73 @@
+import Link from "next/link";
import {
getSettings,
sizeSocialSvg,
} from "yt-dlp-transcript-common/lib/settings";
import { resolveHomepageSocialLinks } from "yt-dlp-transcript-common/lib/homepage";
+import {
+ PROJECT_NAME,
+ PROJECT_TAGLINE,
+} from "yt-dlp-transcript-common/lib/project";
import { currentHomepage } from "../lib/homepage";
+import { NAV } from "../lib/nav";
-// Hub footer: just social links (its own override or the global default). The
-// per-content-site "related sites" footer lives in the export app; the hub IS
-// the cross-site index, so it doesn't repeat that list here.
+// The project site's footer. Note the identity split it embodies: the wordmark
+// and tagline are PRODUCT strings (the same on every install), while the social
+// links remain OPERATOR config — they are this deployment's accounts, not the
+// software's.
export default function Footer() {
const config = currentHomepage();
const socialLinks = resolveHomepageSocialLinks(config, getSettings());
return (
- <footer className="mt-16 border-t border-[var(--border)]">
- <div className="max-w-6xl mx-auto px-5 sm:px-6 py-6 flex items-center justify-between gap-4">
- <span className="font-mono text-[0.6875rem] uppercase tracking-[0.22em] text-[var(--faint)]">
- {config.siteTitle}
- </span>
- {socialLinks.length > 0 && (
- <ul className="flex items-center gap-4 list-none">
- {socialLinks.map((link, i) => (
- <li key={`${link.url}-${i}`}>
- <a
- href={link.url}
- title={link.label}
- aria-label={link.label}
- target="_blank"
- rel="noopener noreferrer"
- className="inline-block w-5 h-5 text-[var(--faint)] hover:text-[var(--brand)] transition-colors [&_svg]:w-full [&_svg]:h-full"
- dangerouslySetInnerHTML={{ __html: sizeSocialSvg(link.svg) }}
- />
+ <footer className="mt-20 border-t border-[var(--border)]">
+ <div className="max-w-6xl mx-auto px-5 sm:px-6 py-10 flex flex-col gap-8 sm:flex-row sm:justify-between">
+ <div className="flex flex-col gap-2 max-w-xs">
+ <span className="font-display text-sm font-bold uppercase tracking-[0.16em] text-[var(--foreground)]">
+ {PROJECT_NAME}
+ </span>
+ <p className="text-sm leading-relaxed text-[var(--muted-foreground)]">
+ {PROJECT_TAGLINE}
+ </p>
+ <p className="mt-1 text-xs text-[var(--faint)]">
+ MIT licensed. You run it; you host it; the archive is yours.
+ </p>
+ </div>
+
+ <nav aria-label="Footer" className="flex flex-col gap-3">
+ <span className="label-machine">Sections</span>
+ <ul className="flex flex-col gap-2 list-none">
+ {NAV.map((item) => (
+ <li key={item.href}>
+ <Link
+ href={item.href}
+ className="text-sm text-[var(--muted-foreground)] hover:text-[var(--foreground)] transition-colors"
+ >
+ {item.label}
+ </Link>
</li>
))}
</ul>
+ </nav>
+
+ {socialLinks.length > 0 && (
+ <div className="flex flex-col gap-3">
+ <span className="label-machine">Elsewhere</span>
+ <ul className="flex items-center gap-4 list-none">
+ {socialLinks.map((link, i) => (
+ <li key={`${link.url}-${i}`}>
+ <a
+ href={link.url}
+ title={link.label}
+ aria-label={link.label}
+ target="_blank"
+ rel="noopener noreferrer"
+ className="inline-block w-5 h-5 text-[var(--faint)] hover:text-[var(--foreground)] transition-colors [&_svg]:w-full [&_svg]:h-full"
+ dangerouslySetInnerHTML={{ __html: sizeSocialSvg(link.svg) }}
+ />
+ </li>
+ ))}
+ </ul>
+ </div>
)}
</div>
</footer>
diff --git a/homepage/app/components/Header.tsx b/homepage/app/components/Header.tsx
@@ -1,26 +1,74 @@
import Link from "next/link";
+import { PROJECT_NAME } from "yt-dlp-transcript-common/lib/project";
import { ThemeToggle } from "yt-dlp-transcript-common/components/ThemeToggle";
import { ThemeMenu } from "yt-dlp-transcript-common/components/ThemeMenu";
-import { currentHomepage } from "../lib/homepage";
+import { NAV } from "../lib/nav";
-// The hub's header: a compact sticky bar with the brass mark + wordmark, the
-// same size on every page.
+function NavList({ className }: { className?: string }) {
+ return (
+ <ul className={`flex items-center list-none ${className ?? ""}`}>
+ {NAV.map((item) => (
+ <li key={item.href}>
+ <Link
+ href={item.href}
+ className="font-mono text-[0.7rem] uppercase tracking-[0.16em] whitespace-nowrap text-[var(--muted-foreground)] hover:text-[var(--foreground)] transition-colors"
+ >
+ {item.label}
+ </Link>
+ </li>
+ ))}
+ </ul>
+ );
+}
+
+// The project site's header: a hairline bar with the wordmark and the four
+// destinations. It carried NO links at all before this — the page it sat above
+// was the whole site.
+//
+// The wordmark is PRODUCT identity (common/lib/project.ts), not the operator's
+// editable homepage config: this bar says what the software is called, and that
+// is the same string on every install. The operator's own naming still governs
+// /stats/ and the social links in the footer.
+//
+// The mark is a solid triangle — a play head, the thing every recording here
+// began as. Bone, not brass: this family spends no colour on chrome.
export default function Header() {
- const { headerTitle } = currentHomepage();
return (
- <header className="sticky top-0 z-20 border-b border-[var(--border)] bg-[rgba(12,10,8,0.72)] backdrop-blur-md">
- <div className="max-w-6xl mx-auto px-5 sm:px-6 flex h-14 items-center">
- <Link href="/" className="group flex items-center gap-2.5">
- <span className="h-2.5 w-2.5 shrink-0 rotate-45 rounded-[3px] bg-[var(--brand)] shadow-[0_0_14px_var(--brand)] transition-colors group-hover:bg-[var(--brand-strong)]" />
- <span className="font-display text-[1.35rem] font-semibold leading-none tracking-tight text-[var(--foreground)]">
- {headerTitle}
+ <header className="sticky top-0 z-20 border-b border-[var(--border)] bg-[var(--background)]/85 backdrop-blur-md">
+ <div className="max-w-6xl mx-auto px-5 sm:px-6 flex h-14 items-center gap-6">
+ <Link
+ href="/"
+ className="group flex items-center gap-2.5 shrink-0"
+ aria-label={`${PROJECT_NAME} home`}
+ >
+ <span
+ aria-hidden="true"
+ className="h-0 w-0 border-y-[6px] border-y-transparent border-l-[9px] border-l-[var(--foreground)] transition-colors group-hover:border-l-[var(--brand)]"
+ />
+ <span className="font-display text-[1.05rem] font-bold uppercase leading-none tracking-[0.14em] text-[var(--foreground)]">
+ {PROJECT_NAME}
</span>
</Link>
- <div className="ml-auto flex items-center gap-2">
+ <nav aria-label="Main" className="ml-auto hidden sm:block">
+ <NavList className="gap-6" />
+ </nav>
+ <div className="ml-auto sm:ml-0 flex items-center gap-2">
<ThemeMenu />
<ThemeToggle />
</div>
</div>
+ {/* Below `sm` the bar has no room for four labels beside the wordmark and
+ the theme controls, so the nav drops to its own scrollable rule rather
+ than collapsing behind a menu button — four links do not earn a
+ disclosure widget. Only one of the two is ever in the accessibility
+ tree (the other is display:none), but they carry distinct labels so a
+ test or a screen reader can never conflate them. */}
+ <nav
+ aria-label="Main, compact"
+ className="sm:hidden border-t border-[var(--border)] overflow-x-auto"
+ >
+ <NavList className="gap-5 px-5 h-10" />
+ </nav>
</header>
);
}
diff --git a/homepage/app/components/HomeLanding.tsx b/homepage/app/components/HomeLanding.tsx
@@ -57,7 +57,7 @@ export function HomeLanding({ summary }: { summary: HomepageSummary }) {
style={{ "--d": "200ms" } as React.CSSProperties}
>
<PresetBar active={state.preset} onSelect={selectPreset} />
- <div className="panel-instrument rounded-2xl p-3 pt-4 sm:p-5">
+ <div className="panel-instrument rounded-[var(--radius)] p-3 pt-4 sm:p-5">
<ChartPanel
summary={summary}
state={state}
diff --git a/homepage/app/components/KpiHeader.tsx b/homepage/app/components/KpiHeader.tsx
@@ -1,6 +1,16 @@
import type { HomepageSummary } from "yt-dlp-transcript-common/lib/homepageSummary";
-function Stat({
+// The headline numbers, as a tabular strip rather than a row of cards.
+//
+// It used to be five rounded, blurred tiles with a brass gradient down the left
+// edge — the archive-room look. This family reads numbers as instrument
+// readout: hairline-divided cells, monospaced tabular figures, a machine label
+// underneath. Nothing here is decorated; the number IS the graphic.
+//
+// Server-safe (no "use client"), so it renders into the static HTML on both the
+// home page and /stats/.
+
+function Cell({
value,
label,
delay,
@@ -11,22 +21,17 @@ function Stat({
}) {
return (
<div
- className="reveal group relative flex flex-col gap-1 overflow-hidden rounded-xl border border-[var(--border)] bg-[var(--panel)] px-4 py-4 backdrop-blur-sm transition-colors hover:border-[var(--border-strong)]"
+ className="reveal flex flex-col gap-2 px-4 py-5 sm:px-5 border-t border-[var(--border)] sm:border-t-0 sm:border-l first:border-l-0 first:border-t-0"
style={{ "--d": `${delay}ms` } as React.CSSProperties}
>
- <span className="absolute left-0 top-0 h-full w-px bg-gradient-to-b from-[var(--brand)] to-transparent opacity-60" />
- <span className="font-display text-[2rem] sm:text-[2.5rem] font-semibold leading-none tracking-[-0.01em] text-[var(--foreground)] tabular-nums">
+ <span className="tabular text-[1.75rem] sm:text-[2rem] font-medium leading-none tracking-[-0.02em] text-[var(--foreground)]">
{value.toLocaleString()}
</span>
- <span className="font-mono text-[0.625rem] uppercase tracking-[0.18em] text-[var(--muted-foreground)]">
- {label}
- </span>
+ <span className="label-machine">{label}</span>
</div>
);
}
-// Headline KPI tiles for the hub landing. Pure presentational; numbers come
-// pre-computed in the summary, so this renders in the static HTML.
export function KpiHeader({ summary }: { summary: HomepageSummary }) {
const t = summary.totals;
const updated = summary.generatedAt
@@ -37,23 +42,21 @@ export function KpiHeader({ summary }: { summary: HomepageSummary }) {
})
: null;
const stats = [
- { value: t.transcripts, label: "Transcripts" },
- { value: t.downloads, label: "Downloads" },
- { value: t.sites, label: "Sites" },
+ { value: t.downloads, label: "Recordings" },
+ { value: t.transcripts, label: "Transcribed" },
+ { value: t.hoursArchived, label: "Hours" },
{ value: t.channels, label: "Channels" },
- { value: t.hoursArchived, label: "Hours archived" },
+ { value: t.sites, label: "Archives" },
];
return (
<section className="flex flex-col gap-3">
- <div className="grid grid-cols-2 gap-3 sm:grid-cols-3 lg:grid-cols-5">
+ <div className="grid grid-cols-2 sm:grid-cols-3 lg:grid-cols-5 border border-[var(--border)] bg-[var(--surface)]">
{stats.map((s, i) => (
- <Stat key={s.label} value={s.value} label={s.label} delay={i * 70} />
+ <Cell key={s.label} value={s.value} label={s.label} delay={i * 60} />
))}
</div>
{updated && (
- <p className="font-mono text-[0.625rem] uppercase tracking-[0.18em] text-[var(--faint)]">
- Index updated · {updated}
- </p>
+ <p className="label-machine">Index built · {updated}</p>
)}
</section>
);
diff --git a/homepage/app/components/PageShell.tsx b/homepage/app/components/PageShell.tsx
@@ -0,0 +1,48 @@
+// The standard content column for every route except `/`.
+//
+// The root layout's <main> is deliberately full-bleed: the home page is built
+// from edge-to-edge rules and banded sections that must reach the viewport
+// edges. Every other page opts back into a measured column through this, so the
+// container lives in one place rather than being re-typed per route.
+export function PageShell({
+ children,
+ className,
+}: {
+ children: React.ReactNode;
+ className?: string;
+}) {
+ return (
+ <div
+ className={`w-full max-w-6xl mx-auto px-5 sm:px-6 py-10 sm:py-14 ${className ?? ""}`}
+ >
+ {children}
+ </div>
+ );
+}
+
+// A page's title block: machine eyebrow, display title, optional standfirst.
+// Consistent across docs, downloads, stats and the 404 so the site reads as one
+// instrument rather than four pages that happen to share a header.
+export function PageHeading({
+ eyebrow,
+ title,
+ standfirst,
+}: {
+ eyebrow?: string;
+ title: string;
+ standfirst?: string;
+}) {
+ return (
+ <div className="flex flex-col gap-3 doc-measure">
+ {eyebrow && <span className="label-machine">{eyebrow}</span>}
+ <h1 className="font-display text-3xl sm:text-4xl font-bold uppercase leading-[1.05] tracking-[0.01em] text-[var(--foreground)]">
+ {title}
+ </h1>
+ {standfirst && (
+ <p className="text-[1.0625rem] leading-relaxed text-[var(--muted-foreground)]">
+ {standfirst}
+ </p>
+ )}
+ </div>
+ );
+}
diff --git a/homepage/app/components/PresetBar.tsx b/homepage/app/components/PresetBar.tsx
@@ -28,7 +28,7 @@ export function PresetBar({
title={PRESETS[id].hint}
onClick={() => onSelect(id)}
className={
- "rounded-full border px-4 py-2 text-sm font-medium transition-all duration-150 " +
+ "rounded-[var(--radius)] border px-4 py-2 text-sm font-medium transition-all duration-150 " +
(on
? "border-[var(--brand)] bg-[var(--brand-soft)] text-[var(--brand-strong)]"
: "border-[var(--border)] text-[var(--muted-foreground)] hover:border-[var(--border-strong)] hover:text-[var(--foreground)]")
@@ -39,7 +39,7 @@ export function PresetBar({
);
})}
{active === "custom" && (
- <span className="rounded-full border border-dashed border-[var(--border-strong)] px-4 py-2 text-sm text-[var(--faint)]">
+ <span className="rounded-[var(--radius)] border border-dashed border-[var(--border-strong)] px-4 py-2 text-sm text-[var(--faint)]">
Custom
</span>
)}
diff --git a/homepage/app/components/RecentAdditions.tsx b/homepage/app/components/RecentAdditions.tsx
@@ -1,5 +1,19 @@
import type { HomepageRecentItem } from "yt-dlp-transcript-common/lib/homepageSummary";
-import { dayLabel } from "yt-dlp-transcript-common/lib/homepageChart";
+import type { VideoState } from "yt-dlp-transcript-common/lib/availability";
+
+// THE RAIL — the product in one object.
+//
+// Real recordings, most recently acquired first: the date, the channel, the
+// title, and where the recording stands on the platform it came from. Most read
+// KEPT. The ones re-checked and found gone read GONE, and that chip is the only
+// hot colour above the fold.
+//
+// HARD RULE: nothing here may be invented. No placeholder titles, no sample
+// quotes, no assumed states. Inventing content on the marketing site for an
+// archiving tool would undercut the entire product, so absent data renders an
+// empty state and absent state renders no chip at all — never a cheerful
+// default. That is why `status` is optional all the way down: a summary written
+// before the field existed must show nothing rather than guess.
// Open the transcript on its destination site: the export apps open the modal
// from the `v` slug param (see common/components/PlayerProvider). siteUrl carries
@@ -9,52 +23,128 @@ function transcriptHref(item: HomepageRecentItem): string | null {
return `${item.siteUrl}/?v=${encodeURIComponent(item.slug)}`;
}
-// Newest transcriptions across every site, each badged with where it landed and
-// linked straight to the transcript.
-export function RecentAdditions({ items }: { items: HomepageRecentItem[] }) {
- if (items.length === 0) return null;
+// "20260701" -> "2026-07-01". The rail is a log, so it wants the sortable form,
+// not a friendly one.
+function isoDay(yyyymmdd: string): string {
+ if (!/^\d{8}$/.test(yyyymmdd)) return yyyymmdd;
+ return `${yyyymmdd.slice(0, 4)}-${yyyymmdd.slice(4, 6)}-${yyyymmdd.slice(6, 8)}`;
+}
+
+// The chip is a claim about a recording, so each state gets its own words.
+// "available" deliberately reads KEPT rather than anything stronger: it means we
+// have a copy and nothing has told us the original is gone — not that the
+// original is confirmed to still be there.
+const STATE_CHIP: Record<VideoState, { label: string; gone: boolean }> = {
+ available: { label: "Kept", gone: false },
+ maybe_missing: { label: "Missing?", gone: false },
+ unlisted: { label: "Unlisted", gone: false },
+ private: { label: "Private", gone: true },
+ members_only: { label: "Members", gone: true },
+ deleted: { label: "Gone", gone: true },
+};
+
+function StateChip({ status }: { status?: VideoState }) {
+ if (!status) return null;
+ const chip = STATE_CHIP[status];
+ if (!chip) return null;
return (
- <section className="flex flex-col gap-3">
- <h2 className="text-lg font-semibold text-foreground">
- Recent additions
- </h2>
- <ul className="flex flex-col divide-y divide-border rounded-xl border border-border bg-card">
- {items.map((item) => {
+ <span
+ className={
+ "shrink-0 px-2 py-0.5 font-mono text-[0.625rem] uppercase tracking-[0.14em] rounded-[var(--radius)] border " +
+ (chip.gone
+ ? "text-[var(--state-gone)] border-[var(--state-gone)] bg-[var(--state-gone-soft)]"
+ : "text-[var(--faint)] border-[var(--border)]")
+ }
+ >
+ {chip.label}
+ </span>
+ );
+}
+
+export function RecentAdditions({
+ items,
+ limit = 8,
+}: {
+ items: HomepageRecentItem[];
+ limit?: number;
+}) {
+ const rows = items.slice(0, limit);
+ if (rows.length === 0) {
+ return (
+ <div className="faceplate p-5">
+ <p className="text-sm text-[var(--muted-foreground)]">
+ No acquisitions to show — this build shipped without corpus data.
+ </p>
+ </div>
+ );
+ }
+ return (
+ <div className="faceplate overflow-hidden">
+ <div className="flex items-baseline justify-between gap-4 px-4 py-3 border-b border-[var(--border)]">
+ <span className="label-machine">Recent acquisitions</span>
+ <span className="label-machine">State at source</span>
+ </div>
+ <ul className="flex flex-col list-none">
+ {rows.map((item, i) => {
const href = transcriptHref(item);
- const title = (
- <span className="font-medium text-foreground group-hover:text-info">
- {item.title}
- </span>
+ // Two shapes, one row. Wide: a single log line — date, channel,
+ // title, state. Narrow: the metadata and the state chip share the
+ // top line and the title gets the full width beneath, because
+ // squeezing four columns into 375px truncates the title to nothing
+ // and orphans the chip on its own line.
+ const inner = (
+ <>
+ <div className="flex w-full items-center gap-3 sm:contents">
+ <time
+ dateTime={isoDay(item.transcribedDate)}
+ className="tabular text-[0.6875rem] text-[var(--faint)] shrink-0 sm:w-24"
+ >
+ {isoDay(item.transcribedDate)}
+ </time>
+ <span className="min-w-0 flex-1 font-mono text-[0.6875rem] uppercase tracking-[0.1em] text-[var(--faint)] truncate sm:flex-none sm:w-52">
+ {item.channel}
+ </span>
+ <span className="ml-auto sm:hidden">
+ <StateChip status={item.status} />
+ </span>
+ </div>
+ <span className="min-w-0 w-full text-sm text-[var(--muted-foreground)] group-hover:text-[var(--foreground)] transition-colors line-clamp-2 sm:w-auto sm:flex-1 sm:truncate sm:line-clamp-none">
+ {item.title}
+ </span>
+ <span className="hidden sm:block">
+ <StateChip status={item.status} />
+ </span>
+ </>
);
+ const rowClass =
+ "rail-row group flex flex-wrap items-center gap-x-4 gap-y-1.5 px-4 py-3 sm:flex-nowrap sm:py-2.5";
return (
- <li key={`${item.siteId}:${item.slug}`}>
- {/* eslint-disable-next-line @next/next/no-html-link-for-pages */}
- <a
- href={href ?? undefined}
- {...(href
- ? { target: "_blank", rel: "noopener noreferrer" }
- : { "aria-disabled": true })}
- className="group flex flex-col gap-1 px-4 py-3 sm:flex-row sm:items-center sm:justify-between"
- >
- <div className="min-w-0">
- {title}
- <p className="truncate text-xs text-muted-foreground">
- {item.channel}
- </p>
- </div>
- <div className="flex shrink-0 items-center gap-2 text-xs text-muted-foreground">
- <span className="rounded-full bg-muted px-2 py-0.5 font-medium text-muted-foreground">
- {item.siteTitle}
- </span>
- <time dateTime={item.transcribedDate}>
- {dayLabel(item.transcribedDate)}
- </time>
+ <li
+ key={`${item.siteId}:${item.slug}`}
+ className="border-b border-[var(--border)] last:border-b-0"
+ >
+ {href ? (
+ <a
+ href={href}
+ target="_blank"
+ rel="noopener noreferrer"
+ className={rowClass}
+ style={{ "--d": `${360 + i * 55}ms` } as React.CSSProperties}
+ >
+ {inner}
+ </a>
+ ) : (
+ <div
+ className={rowClass}
+ style={{ "--d": `${360 + i * 55}ms` } as React.CSSProperties}
+ >
+ {inner}
</div>
- </a>
+ )}
</li>
);
})}
</ul>
- </section>
+ </div>
);
}
diff --git a/homepage/app/components/Segmented.tsx b/homepage/app/components/Segmented.tsx
@@ -18,7 +18,7 @@ export function Segmented<T extends string>({
<div
role="group"
aria-label={label}
- className="inline-flex items-center rounded-full border border-[var(--border)] bg-[var(--panel)] p-0.5 backdrop-blur-sm"
+ className="inline-flex items-center rounded-[var(--radius)] border border-[var(--border)] bg-[var(--panel)] p-0.5 "
>
{options.map((o) => {
const active = o.value === value;
@@ -33,7 +33,7 @@ export function Segmented<T extends string>({
if (!disabled) onChange(o.value);
}}
className={
- "rounded-full px-3 py-1.5 text-[0.8125rem] font-medium transition-all duration-150 " +
+ "rounded-[var(--radius)] px-3 py-1.5 text-[0.8125rem] font-medium transition-all duration-150 " +
(active
? "bg-[var(--brand-soft)] text-[var(--brand-strong)] shadow-[inset_0_0_0_1px_var(--brand-soft)]"
: disabled
diff --git a/homepage/app/components/SiteGrid.tsx b/homepage/app/components/SiteGrid.tsx
@@ -88,7 +88,7 @@ function SiteCard({
</h3>
</div>
{isLeader && (
- <span className="shrink-0 rounded-full border border-[var(--brand-soft)] bg-[var(--brand-soft)] px-2 py-0.5 font-mono text-[0.625rem] uppercase tracking-[0.1em] text-[var(--brand-strong)]">
+ <span className="shrink-0 rounded-[var(--radius)] border border-[var(--brand-soft)] bg-[var(--brand-soft)] px-2 py-0.5 font-mono text-[0.625rem] uppercase tracking-[0.1em] text-[var(--brand-strong)]">
#1 this month
</span>
)}
@@ -118,7 +118,7 @@ function SiteCard({
);
const cardClass =
- "group relative flex flex-col gap-3 rounded-xl border border-[var(--border)] bg-[var(--panel)] p-5 backdrop-blur-sm transition duration-200 hover:-translate-y-0.5 hover:border-[var(--border-strong)] hover:bg-[var(--panel-2)]" +
+ "group relative flex flex-col gap-3 rounded-[var(--radius)] border border-[var(--border)] bg-[var(--panel)] p-5 transition duration-200 hover:-translate-y-0.5 hover:border-[var(--border-strong)] hover:bg-[var(--panel-2)]" +
(hidden ? " opacity-40" : "");
return (
@@ -132,7 +132,7 @@ function SiteCard({
onClick={onToggle}
aria-pressed={!hidden}
title={hidden ? "Show on chart" : "Hide from chart"}
- className="absolute right-3 top-3 z-10 rounded-full border border-[var(--border)] bg-[rgba(20,16,9,0.85)] px-2 py-0.5 font-mono text-[0.625rem] uppercase tracking-[0.1em] text-[var(--muted-foreground)] transition-colors hover:border-[var(--border-strong)] hover:text-[var(--foreground)]"
+ className="absolute right-3 top-3 z-10 rounded-[var(--radius)] border border-[var(--border)] bg-[var(--surface)] px-2 py-0.5 font-mono text-[0.625rem] uppercase tracking-[0.1em] text-[var(--muted-foreground)] transition-colors hover:border-[var(--border-strong)] hover:text-[var(--foreground)]"
>
{hidden ? "Show" : "Hide"}
</button>
diff --git a/homepage/app/docs/[slug]/page.tsx b/homepage/app/docs/[slug]/page.tsx
@@ -0,0 +1,82 @@
+import Link from "next/link";
+import { notFound } from "next/navigation";
+import type { Metadata } from "next";
+import { PageShell, PageHeading } from "../../components/PageShell";
+import { DocBody } from "../../components/DocBody";
+import { findDoc, listDocs, readDoc } from "../../lib/docs";
+
+// `dynamicParams = false` + params from the same existence-filtered list means a
+// manifest entry with a missing file is a 404, never a build crash.
+export const dynamicParams = false;
+
+export function generateStaticParams() {
+ return listDocs().map((doc) => ({ slug: doc.slug }));
+}
+
+export async function generateMetadata({
+ params,
+}: {
+ params: Promise<{ slug: string }>;
+}): Promise<Metadata> {
+ const { slug } = await params;
+ const doc = findDoc(slug);
+ if (!doc) return {};
+ return { title: doc.title, description: doc.blurb };
+}
+
+export default async function DocPage({
+ params,
+}: {
+ params: Promise<{ slug: string }>;
+}) {
+ const { slug } = await params;
+ const doc = findDoc(slug);
+ if (!doc) notFound();
+
+ const docs = listDocs();
+ const i = docs.findIndex((d) => d.slug === doc.slug);
+ const prev = i > 0 ? docs[i - 1] : null;
+ const next = i >= 0 && i < docs.length - 1 ? docs[i + 1] : null;
+
+ return (
+ <PageShell className="flex flex-col gap-8">
+ <Link
+ href="/docs/"
+ className="label-machine hover:text-[var(--foreground)] transition-colors w-fit"
+ >
+ ← All docs
+ </Link>
+ <PageHeading title={doc.title} standfirst={doc.blurb} />
+ <DocBody source={readDoc(doc)} />
+
+ {(prev || next) && (
+ <nav
+ aria-label="Document"
+ className="doc-measure mt-6 flex flex-col gap-4 border-t border-[var(--border)] pt-6 sm:flex-row sm:justify-between"
+ >
+ {prev ? (
+ <Link href={`/docs/${prev.slug}/`} className="flex flex-col gap-1">
+ <span className="label-machine">Previous</span>
+ <span className="text-sm text-[var(--muted-foreground)] hover:text-[var(--foreground)] transition-colors">
+ {prev.title}
+ </span>
+ </Link>
+ ) : (
+ <span />
+ )}
+ {next && (
+ <Link
+ href={`/docs/${next.slug}/`}
+ className="flex flex-col gap-1 sm:items-end"
+ >
+ <span className="label-machine">Next</span>
+ <span className="text-sm text-[var(--muted-foreground)] hover:text-[var(--foreground)] transition-colors">
+ {next.title}
+ </span>
+ </Link>
+ )}
+ </nav>
+ )}
+ </PageShell>
+ );
+}
diff --git a/homepage/app/docs/page.tsx b/homepage/app/docs/page.tsx
@@ -0,0 +1,54 @@
+import Link from "next/link";
+import type { Metadata } from "next";
+import { PageShell, PageHeading } from "../components/PageShell";
+import { groupedDocs } from "../lib/docs";
+
+export const metadata: Metadata = {
+ title: "Docs",
+ description:
+ "How to install, run, and publish an Archilyzer archive: setup, operation, deployment, and the machine-readable corpus.",
+};
+
+// The index renders from the same listDocs() that feeds generateStaticParams,
+// so it structurally cannot link to a page that doesn't exist.
+export default function DocsIndexPage() {
+ const groups = groupedDocs();
+ return (
+ <PageShell className="flex flex-col gap-12">
+ <PageHeading
+ eyebrow="Documentation"
+ title="Docs"
+ standfirst="Everything needed to take a channel URL and end up with a searchable archive you host. Written for the operator, not the contributor."
+ />
+
+ {groups.map((group) => (
+ <section key={group.label} className="flex flex-col gap-4">
+ <h2 className="label-machine">{group.label}</h2>
+ <ul className="flex flex-col list-none border-t border-[var(--border)]">
+ {group.docs.map((doc) => (
+ <li key={doc.slug} className="border-b border-[var(--border)]">
+ <Link
+ href={`/docs/${doc.slug}/`}
+ className="group flex flex-col gap-1 py-4 sm:flex-row sm:items-baseline sm:gap-6"
+ >
+ <span className="font-display text-base font-bold uppercase tracking-[0.04em] text-[var(--foreground)] sm:w-64 sm:shrink-0">
+ {doc.title}
+ </span>
+ <span className="text-sm leading-relaxed text-[var(--muted-foreground)] group-hover:text-[var(--foreground)] transition-colors">
+ {doc.blurb}
+ </span>
+ </Link>
+ </li>
+ ))}
+ </ul>
+ </section>
+ ))}
+
+ {groups.length === 0 && (
+ <p className="doc-measure text-[var(--muted-foreground)]">
+ No documentation shipped with this build.
+ </p>
+ )}
+ </PageShell>
+ );
+}
diff --git a/homepage/app/downloads/page.tsx b/homepage/app/downloads/page.tsx
@@ -0,0 +1,170 @@
+import Link from "next/link";
+import type { Metadata } from "next";
+import { PageShell, PageHeading } from "../components/PageShell";
+import { loadSnapshot, formatBytes, SNAPSHOT_HREF } from "../lib/snapshot";
+
+export const metadata: Metadata = {
+ title: "Downloads",
+ description:
+ "Get the Archilyzer source as a dated snapshot tarball — MIT licensed, with a checksum.",
+};
+
+function Fact({ label, children }: { label: string; children: React.ReactNode }) {
+ return (
+ <div className="flex flex-col gap-1 border-t border-[var(--border)] py-3 sm:flex-row sm:gap-6">
+ <span className="label-machine sm:w-32 sm:shrink-0 sm:pt-0.5">
+ {label}
+ </span>
+ <span className="min-w-0 break-all text-sm text-[var(--muted-foreground)]">
+ {children}
+ </span>
+ </div>
+ );
+}
+
+export default function DownloadsPage() {
+ const snapshot = loadSnapshot();
+ const date = snapshot
+ ? new Date(snapshot.generatedAt).toLocaleDateString(undefined, {
+ year: "numeric",
+ month: "long",
+ day: "numeric",
+ })
+ : null;
+
+ return (
+ <PageShell className="flex flex-col gap-10">
+ <PageHeading
+ eyebrow="Source"
+ title="Downloads"
+ standfirst="Archilyzer is MIT licensed and published as a dated snapshot of the source tree."
+ />
+
+ <div className="doc-measure flex flex-col gap-4">
+ <p className="leading-[1.7] text-[var(--muted-foreground)]">
+ <strong className="font-semibold text-[var(--foreground)]">
+ This is a snapshot, not a repository.
+ </strong>{" "}
+ It is a <code className="font-mono text-[var(--foreground)]">git archive</code>{" "}
+ of the main branch: the working tree at one commit, and nothing else.
+ There is <em>no history, no branches, no remote, and nothing to{" "}
+ <code className="font-mono">git pull</code></em>. Updating means
+ downloading a newer snapshot.
+ </p>
+ </div>
+
+ {snapshot ? (
+ <div className="flex flex-col gap-6">
+ <div>
+ <a
+ href={SNAPSHOT_HREF}
+ className="inline-flex items-center gap-3 bg-[var(--foreground)] px-6 py-3.5 font-mono text-xs uppercase tracking-[0.16em] text-[var(--background)] rounded-[var(--radius)] hover:bg-[var(--brand)] hover:text-[var(--brand-ink)] transition-colors"
+ >
+ archilyzer-source.tar.gz
+ <span className="tabular opacity-70">
+ {formatBytes(snapshot.bytes)}
+ </span>
+ </a>
+ </div>
+
+ <div className="doc-measure flex flex-col">
+ <Fact label="Snapshot of">{date}</Fact>
+ <Fact label="Commit">
+ <span className="tabular">{snapshot.commit}</span>
+ </Fact>
+ <Fact label="Subject">{snapshot.subject}</Fact>
+ <Fact label="Size">
+ <span className="tabular">
+ {snapshot.bytes.toLocaleString()} bytes
+ </span>
+ </Fact>
+ <Fact label="SHA-256">
+ <span className="tabular">{snapshot.sha256}</span>
+ </Fact>
+ </div>
+
+ <div className="doc-measure">
+ <p className="mb-2 label-machine">Verify it</p>
+ <pre className="p-4 rounded-[var(--radius)] bg-[var(--surface)] border border-[var(--border)] overflow-x-auto text-[0.8125rem] font-mono text-[var(--muted-foreground)]">
+ {`sha256sum archilyzer-source.tar.gz\ntar xzf archilyzer-source.tar.gz\ncd archilyzer && pnpm install`}
+ </pre>
+ </div>
+ </div>
+ ) : (
+ // No sidecar, or no tarball beside it: say so and render NO link. A
+ // download button pointing at a file that isn't there is worse than an
+ // honest absence.
+ <div className="doc-measure faceplate p-5 flex flex-col gap-3">
+ <p className="text-[var(--muted-foreground)] leading-relaxed">
+ This build has no source snapshot attached.
+ </p>
+ <p className="text-sm text-[var(--faint)] leading-relaxed">
+ The tarball and its sidecar are build artefacts, generated by{" "}
+ <code className="font-mono">./create-archives.sh</code> and not
+ committed. A site built without running it ships no download, and
+ says so here rather than linking to a file that isn’t there.
+ </p>
+ </div>
+ )}
+
+ <div className="doc-measure flex flex-col gap-8 border-t border-[var(--border)] pt-8">
+ <div className="flex flex-col gap-3">
+ <h2 className="font-display text-lg font-bold uppercase tracking-[0.05em] text-[var(--foreground)]">
+ What’s inside
+ </h2>
+ <p className="text-sm leading-[1.7] text-[var(--muted-foreground)]">
+ Every tracked file in the workspace: the editor, the static-site
+ generator, the shared library, the MCP server, this site, the
+ documentation, and the test suites. Enough to install, run, build and
+ publish.
+ </p>
+ </div>
+
+ <div className="flex flex-col gap-3">
+ <h2 className="font-display text-lg font-bold uppercase tracking-[0.05em] text-[var(--foreground)]">
+ What isn’t
+ </h2>
+ <ul className="list-disc pl-5 space-y-2 text-sm leading-[1.7] text-[var(--muted-foreground)] marker:text-[var(--faint)]">
+ <li>
+ <strong className="text-[var(--foreground)]">
+ The archive itself.
+ </strong>{" "}
+ No transcripts, no media, no index. The corpus lives outside the
+ source tree by design, so it is structurally impossible for it to
+ end up in here.
+ </li>
+ <li>
+ <strong className="text-[var(--foreground)]">
+ Any operator’s configuration.
+ </strong>{" "}
+ The settings file is untracked; only a minimal example ships.
+ </li>
+ <li>
+ <strong className="text-[var(--foreground)]">Dependencies.</strong>{" "}
+ No <code className="font-mono">node_modules</code>. Run{" "}
+ <code className="font-mono">pnpm install</code> after unpacking.
+ </li>
+ <li>
+ <strong className="text-[var(--foreground)]">
+ Models or binaries.
+ </strong>{" "}
+ yt-dlp, ffmpeg and a transcription backend are yours to install —
+ see <Link href="/docs/install/" className="underline decoration-[var(--border-strong)] underline-offset-2 hover:text-[var(--foreground)]">Install</Link>.
+ </li>
+ </ul>
+ </div>
+
+ <div className="flex flex-col gap-3">
+ <h2 className="font-display text-lg font-bold uppercase tracking-[0.05em] text-[var(--foreground)]">
+ License
+ </h2>
+ <p className="text-sm leading-[1.7] text-[var(--muted-foreground)]">
+ MIT. Use it, change it, run it, publish with it. The full text is in{" "}
+ <code className="font-mono text-[var(--foreground)]">LICENSE</code>{" "}
+ at the root of the archive.
+ </p>
+ </div>
+ </div>
+ </PageShell>
+ );
+}
diff --git a/homepage/app/favicon.ico b/homepage/app/favicon.ico
Binary files differ.
diff --git a/homepage/app/globals.css b/homepage/app/globals.css
@@ -2,10 +2,15 @@
@import "../../common/styles/tokens.css";
@source "../../common/components";
-/* The hub is committed to the warm-ink "archive" family (it defaults to
- data-theme="archive" + dark). Design tokens, the `dark` variant, and the
- theme palettes now live in common/styles/tokens.css; only the homepage's
- atmosphere + bespoke layer styles remain here. */
+/* This is the PROJECT's site — the instrument, not one of the archives it
+ builds. It commits to the "archilyzer" family (graphite chrome, achromatic,
+ colour reserved for recording state and for the archives' own accents).
+ Palettes and the `dark` variant live in common/styles/tokens.css; only the
+ faceplate geometry and the page-load reveal remain here.
+
+ The warm radial glow + film grain that used to live here is deliberately
+ GONE. It was archive-room dressing, correct back when this page was a shelf
+ of archives; a faceplate wants flat ground and crisp rules. */
html {
background-color: var(--background);
@@ -17,59 +22,38 @@ body {
-webkit-font-smoothing: antialiased;
}
-/* Atmosphere: a warm overhead glow + a faint engraved grid, fixed behind all
- content. Pointer-transparent and z-negative so nothing is obstructed. */
-body::before {
- content: "";
- position: fixed;
- inset: 0;
- z-index: -2;
- pointer-events: none;
- background:
- radial-gradient(
- 120% 80% at 50% -10%,
- rgba(227, 177, 92, 0.16),
- rgba(227, 177, 92, 0) 60%
- ),
- radial-gradient(
- 90% 60% at 85% 0%,
- rgba(110, 168, 255, 0.08),
- rgba(110, 168, 255, 0) 55%
- ),
- linear-gradient(180deg, #110d08 0%, var(--background) 38%, #08070500 100%);
+/* An instrument module: a flat faceplate with hairline geometry. No blur, no
+ gradient, no glow — the panel is a solid surface with a drawn edge, and its
+ top hairline is the only thing that "signs" it. */
+.panel-instrument {
+ background: var(--chart-surface);
+ border: 1px solid var(--border);
+ border-top-color: var(--border-strong);
}
-body::after {
- content: "";
- position: fixed;
- inset: 0;
- z-index: -1;
- pointer-events: none;
- opacity: 0.05;
- mix-blend-mode: soft-light;
- background-image: url("data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' width='180' height='180'%3E%3Cfilter id='n'%3E%3CfeTurbulence type='fractalNoise' baseFrequency='0.82' numOctaves='2' stitchTiles='stitch'/%3E%3C/filter%3E%3Crect width='100%25' height='100%25' filter='url(%23n)'/%3E%3C/svg%3E");
+/* A hairline module used for content panels (as opposed to the chart's
+ instrument panel). Square-ish by family radius; the top rule is the tell. */
+.faceplate {
+ background: var(--surface);
+ border: 1px solid var(--border);
+ border-radius: var(--radius);
}
-/* The chart "instrument panel": opaque so the atmospheric glow/grain stops at
- the plot edge (the data zone reads as a calm instrument, not the warm room),
- with a hairline brass-tinted top edge so brass still signs the panel. No
- backdrop-blur — the surface is solid by design. */
-.panel-instrument {
- background: var(--chart-surface);
- border: 1px solid var(--border-strong);
- border-top-color: var(--brand-soft);
- box-shadow: inset 0 1px 0 rgba(233, 220, 197, 0.04);
+/* Small-caps machine label: section eyebrows, column heads, unit suffixes. */
+.label-machine {
+ font-family: var(--font-mono), ui-monospace, monospace;
+ font-size: 0.6875rem;
+ line-height: 1;
+ text-transform: uppercase;
+ letter-spacing: 0.18em;
+ color: var(--faint);
}
-/* A faint hairline grid layer the page can opt into for the "ledger" feel. */
-.grid-ledger {
- background-image:
- linear-gradient(var(--border) 1px, transparent 1px),
- linear-gradient(90deg, var(--border) 1px, transparent 1px);
- background-size: 56px 56px;
- background-position: center;
- -webkit-mask-image: radial-gradient(120% 120% at 50% 0%, #000 35%, transparent 78%);
- mask-image: radial-gradient(120% 120% at 50% 0%, #000 35%, transparent 78%);
+/* Tabular figures everywhere a number can change length between renders, so
+ columns of counts and timecodes don't shift. */
+.tabular {
+ font-family: var(--font-mono), ui-monospace, monospace;
+ font-variant-numeric: tabular-nums;
}
/* One orchestrated page-load reveal. Children stagger via inline --d. */
@@ -90,18 +74,24 @@ body::after {
animation-delay: var(--d, 0ms);
}
+/* The rail's rows arrive in sequence, as acquisitions do. Same easing as the
+ page reveal so it reads as one orchestrated load, not a second animation. */
+.rail-row {
+ opacity: 0;
+ animation: rise 0.5s cubic-bezier(0.2, 0.7, 0.2, 1) both;
+ animation-delay: var(--d, 0ms);
+}
+
@media (prefers-reduced-motion: reduce) {
- .reveal {
+ .reveal,
+ .rail-row {
animation: none;
opacity: 1;
}
}
-/* Brass underline used under section eyebrows. */
-.rule-accent {
- background: linear-gradient(
- 90deg,
- var(--brand) 0%,
- rgba(227, 177, 92, 0) 70%
- );
+/* Documentation prose measure. Docs are a single generous column, never
+ broadsheet columns — the family's geometry is panels, not newspaper. */
+.doc-measure {
+ max-width: 68ch;
}
diff --git a/homepage/app/layout.tsx b/homepage/app/layout.tsx
@@ -1,21 +1,61 @@
-import type { Metadata } from "next";
+import type { Metadata, Viewport } from "next";
import { fontVars } from "yt-dlp-transcript-common/styles/fonts";
import { ThemeScript } from "yt-dlp-transcript-common/components/ThemeScript";
import { ThemeProvider } from "yt-dlp-transcript-common/components/ThemeProvider";
-import { currentHomepage } from "./lib/homepage";
+import {
+ PROJECT_NAME,
+ PROJECT_TAGLINE,
+ PROJECT_URL,
+} from "yt-dlp-transcript-common/lib/project";
import Header from "./components/Header";
import Footer from "./components/Footer";
import "./globals.css";
-export function generateMetadata(): Metadata {
- const c = currentHomepage();
- return {
- title: {
- default: c.siteTitle,
- template: `%s — ${c.siteTitle}`,
- },
- description: c.siteDescription,
- };
+// Browser-chrome color: the archilyzer family's dark ground (tokens.css). Not
+// an accent — this family has no decorative brand hue, and the phone's chrome
+// should continue the faceplate rather than announce a colour.
+const THEME_COLOR = "#151b20";
+
+// PRODUCT identity, not the operator's. This used to read `currentHomepage()`,
+// the editable config in sites/_homepage/homepage.json — correct when this was
+// "our hub", wrong now that it is the software's own site. One visible
+// consequence: the meta description no longer inherits the operator's
+// corpus-flavoured blurb. The operator's naming still governs /stats/ and the
+// footer's social links.
+export const metadata: Metadata = {
+ metadataBase: new URL(PROJECT_URL),
+ title: {
+ default: `${PROJECT_NAME} — ${PROJECT_TAGLINE}`,
+ template: `%s — ${PROJECT_NAME}`,
+ },
+ description:
+ "Archilyzer downloads a channel's back catalogue, transcribes it on your " +
+ "own machine, and builds a static, searchable site you host yourself.",
+ icons: {
+ icon: [
+ { url: "/icons/icon.svg", type: "image/svg+xml" },
+ { url: "/icons/icon-192.png", sizes: "192x192", type: "image/png" },
+ { url: "/icons/icon-512.png", sizes: "512x512", type: "image/png" },
+ ],
+ apple: [{ url: "/icons/apple-touch-icon.png", sizes: "180x180" }],
+ },
+ openGraph: {
+ type: "website",
+ siteName: PROJECT_NAME,
+ title: `${PROJECT_NAME} — ${PROJECT_TAGLINE}`,
+ description:
+ "Self-hosted video archiving: download a back catalogue, transcribe it " +
+ "locally, publish a static site that is searchable to the second.",
+ url: PROJECT_URL,
+ // No purpose-built OG image yet; the app icon is honest and legible at
+ // card size. A real one is a follow-up, not a placeholder to invent.
+ images: [{ url: "/icons/icon-512.png", width: 512, height: 512 }],
+ },
+ twitter: { card: "summary" },
+};
+
+export function generateViewport(): Viewport {
+ return { themeColor: THEME_COLOR };
}
export default function RootLayout({
@@ -29,13 +69,11 @@ export default function RootLayout({
suppressHydrationWarning
className={`${fontVars} h-full antialiased`}
>
- <body className="min-h-full flex flex-col bg-[var(--background)] text-[var(--foreground)] font-sans selection:bg-[var(--brand-soft)] selection:text-[var(--brand-strong)]">
- <ThemeScript defaultTheme="archive" defaultMode="dark" />
- <ThemeProvider defaultTheme="archive" defaultMode="dark">
+ <body className="min-h-full flex flex-col bg-[var(--background)] text-[var(--foreground)] font-sans selection:bg-[var(--brand-soft)] selection:text-[var(--foreground)]">
+ <ThemeScript defaultTheme="archilyzer" defaultMode="dark" />
+ <ThemeProvider defaultTheme="archilyzer" defaultMode="dark">
<Header />
- <main className="flex-1 w-full max-w-6xl mx-auto px-5 sm:px-6 py-8 sm:py-12">
- {children}
- </main>
+ <main className="flex-1 w-full">{children}</main>
<Footer />
</ThemeProvider>
</body>
diff --git a/homepage/app/lib/docs.ts b/homepage/app/lib/docs.ts
@@ -0,0 +1,57 @@
+import fs from "node:fs";
+import path from "node:path";
+import {
+ DOCS,
+ DOC_GROUP_ORDER,
+ type DocEntry,
+ type DocGroup,
+} from "../../content/docs";
+
+export type { DocEntry, DocGroup };
+export { DOC_GROUP_ORDER };
+
+// Build-time only. `output: "export"` means every one of these reads happens
+// during `next build`; nothing here runs per request. Mirrors the
+// readFileSync-from-cwd approach in export/app/changelog/page.tsx.
+function docPath(entry: DocEntry): string {
+ return path.join(process.cwd(), "content", "docs", entry.file);
+}
+
+// The manifest entries whose file actually exists.
+//
+// This filter is the whole safety property of the docs subsystem: it feeds BOTH
+// `generateStaticParams()` (with `dynamicParams = false`) AND the /docs/ index,
+// so a manifest entry with a missing file 404s quietly instead of crashing the
+// build, and the index can never link to a page that isn't there. Those two
+// facts have to come from the same call, or they will disagree.
+export function listDocs(): DocEntry[] {
+ return DOCS.filter((entry) => {
+ try {
+ return fs.statSync(docPath(entry)).isFile();
+ } catch {
+ return false;
+ }
+ });
+}
+
+export function findDoc(slug: string): DocEntry | null {
+ return listDocs().find((d) => d.slug === slug) ?? null;
+}
+
+// Read a doc's Markdown, minus its leading `# Title` line. The file keeps that
+// heading so it stands alone as a document; the page renders its own <h1> from
+// the manifest, and two titles in a row reads as a mistake.
+export function readDoc(entry: DocEntry): string {
+ const raw = fs.readFileSync(docPath(entry), "utf8");
+ return raw.replace(/^?\s*#[^\n#][^\n]*\n+/, "");
+}
+
+// The index, grouped and ordered by the manifest. Groups with no surviving
+// entries are dropped rather than rendered empty.
+export function groupedDocs(): { label: string; docs: DocEntry[] }[] {
+ const docs = listDocs();
+ return DOC_GROUP_ORDER.map((g) => ({
+ label: g.label,
+ docs: docs.filter((d) => d.group === g.id),
+ })).filter((g) => g.docs.length > 0);
+}
diff --git a/homepage/app/lib/nav.ts b/homepage/app/lib/nav.ts
@@ -0,0 +1,14 @@
+// The site's navigation, declared once and rendered by both the header and the
+// footer. Order is editorial: what the software is, how to get it, what this
+// deployment has done, what changed.
+//
+// Every entry must resolve on a `build:nodata` tree — /stats/ renders an honest
+// no-data panel rather than 404ing — so the nav never has to be conditional.
+export type NavItem = { href: string; label: string };
+
+export const NAV: NavItem[] = [
+ { href: "/docs/", label: "Docs" },
+ { href: "/downloads/", label: "Downloads" },
+ { href: "/stats/", label: "Stats" },
+ { href: "/changelog/", label: "Changelog" },
+];
diff --git a/homepage/app/lib/snapshot.ts b/homepage/app/lib/snapshot.ts
@@ -0,0 +1,57 @@
+import fs from "node:fs";
+import path from "node:path";
+
+// Facts about the published source snapshot, written by create-archives.sh
+// beside the tarball it describes.
+export type Snapshot = {
+ generatedAt: string;
+ commit: string;
+ subject: string;
+ bytes: number;
+ sha256: string;
+};
+
+// The tarball's public path. Stable by design: a dated filename would invalidate
+// every link that has ever been shared the moment a new snapshot ships. The date
+// lives in the sidecar instead.
+export const SNAPSHOT_HREF = "/downloads/archilyzer-source.tar.gz";
+
+// Read the sidecar, or null when this build has no snapshot.
+//
+// The null case is not hypothetical: the tarball and its sidecar are gitignored
+// build artefacts, so a fresh unpack of the source has neither until
+// ./create-archives.sh runs. The download page MUST render no link at all in
+// that case — a link to a file that isn't there is worse than an explanation of
+// why there isn't one.
+export function loadSnapshot(): Snapshot | null {
+ try {
+ const file = path.join(
+ process.cwd(),
+ "public",
+ "downloads",
+ "snapshot.json",
+ );
+ const parsed = JSON.parse(fs.readFileSync(file, "utf8")) as Snapshot;
+ if (!parsed || typeof parsed.bytes !== "number" || !parsed.sha256) {
+ return null;
+ }
+ // Believe the sidecar only if the file it describes is actually present —
+ // they are written together but deployed as separate assets.
+ const tarball = path.join(
+ process.cwd(),
+ "public",
+ "downloads",
+ path.basename(SNAPSHOT_HREF),
+ );
+ if (!fs.statSync(tarball).isFile()) return null;
+ return parsed;
+ } catch {
+ return null;
+ }
+}
+
+export function formatBytes(bytes: number): string {
+ const mb = bytes / (1024 * 1024);
+ if (mb >= 1) return `${mb.toFixed(2)} MB`;
+ return `${Math.round(bytes / 1024)} KB`;
+}
diff --git a/homepage/app/lib/summary.ts b/homepage/app/lib/summary.ts
@@ -0,0 +1,32 @@
+import fs from "node:fs";
+import path from "node:path";
+import type { HomepageSummary } from "yt-dlp-transcript-common/lib/homepageSummary";
+
+// The build-time cross-site summary, or null when this build shipped without
+// corpus data (`build:nodata`, or a source-only checkout that has never run
+// compose — see next.config.ts).
+//
+// NULL, NOT AN EMPTY SUMMARY. The previous loader synthesized a zero-filled
+// summary so the page always had something to render, which turned "we have no
+// data" into "0 transcripts · 0 sites" — a claim that the product does nothing,
+// printed in the same typeface as the real numbers. Callers must branch: render
+// the live numbers only when this returns non-null, and say plainly why they're
+// absent otherwise.
+//
+// Read with `readFileSync` at module scope of a server component, mirroring
+// export/app/changelog/page.tsx. `output: "export"` means this runs at build,
+// never at request time.
+export function loadSummary(): HomepageSummary | null {
+ try {
+ const file = path.join(process.cwd(), "public", "homepage-summary.json");
+ const parsed = JSON.parse(
+ fs.readFileSync(file, "utf8"),
+ ) as HomepageSummary;
+ // A file that exists but carries no sites/totals is as uninformative as no
+ // file at all; treat it the same rather than rendering an empty dashboard.
+ if (!parsed || typeof parsed.totals?.transcripts !== "number") return null;
+ return parsed;
+ } catch {
+ return null;
+ }
+}
diff --git a/homepage/app/not-found.tsx b/homepage/app/not-found.tsx
@@ -0,0 +1,32 @@
+import Link from "next/link";
+import { PageShell, PageHeading } from "./components/PageShell";
+import { NAV } from "./lib/nav";
+
+// Branded 404. Static export renders this to out/404.html, which Cloudflare
+// Pages serves for unmatched paths.
+export default function NotFound() {
+ return (
+ <PageShell className="flex flex-col gap-8">
+ <PageHeading
+ eyebrow="404"
+ title="No such page"
+ standfirst="That address isn't part of this site. It may have been renamed, or it may never have existed here."
+ />
+ <nav aria-label="Site sections" className="flex flex-col gap-3">
+ <span className="label-machine">Try one of these</span>
+ <ul className="flex flex-wrap gap-x-6 gap-y-2 list-none">
+ {[{ href: "/", label: "Home" }, ...NAV].map((item) => (
+ <li key={item.href}>
+ <Link
+ href={item.href}
+ className="font-mono text-sm uppercase tracking-[0.14em] text-[var(--muted-foreground)] underline decoration-[var(--border-strong)] underline-offset-4 hover:text-[var(--foreground)] transition-colors"
+ >
+ {item.label}
+ </Link>
+ </li>
+ ))}
+ </ul>
+ </nav>
+ </PageShell>
+ );
+}
diff --git a/homepage/app/page.tsx b/homepage/app/page.tsx
@@ -1,53 +1,263 @@
-import fs from "node:fs";
-import path from "node:path";
-import {
- HOMEPAGE_SUMMARY_VERSION,
- type HomepageSummary,
-} from "yt-dlp-transcript-common/lib/homepageSummary";
-import { HomeLanding } from "./components/HomeLanding";
-
-// The hub home route IS the cross-site landing: the wordmark, headline KPIs, one
-// stacked activity chart with a minimal control strip, and the site grid — no
-// section labels, kept deliberately sparse. All of it renders from a small
-// summary pre-computed at build (compose-homepage writes public/homepage-
-// summary.json), so the page is static HTML with no multi-MB client fetch.
-
-function emptySummary(): HomepageSummary {
- const emptyBucket = { buckets: [], total: [], bySite: {}, byChannel: {} };
- const emptyMetric = { day: emptyBucket, week: emptyBucket, month: emptyBucket };
- return {
- version: HOMEPAGE_SUMMARY_VERSION,
- generatedAt: "",
- totals: {
- transcripts: 0,
- downloads: 0,
- sites: 0,
- channels: 0,
- hoursArchived: 0,
- transcribedThisMonth: 0,
- downloadedThisMonth: 0,
- },
- channels: [],
- series: { transcribed: emptyMetric, downloaded: emptyMetric },
- sites: [],
- recent: [],
- };
+import Link from "next/link";
+import { RecentAdditions } from "./components/RecentAdditions";
+import { KpiHeader } from "./components/KpiHeader";
+import { ArchiveCards } from "./components/ArchiveCards";
+import { loadSummary } from "./lib/summary";
+
+// The project's front page. Its ONE job: make a visitor understand what
+// Archilyzer is and download it.
+//
+// Every number on this page is computed at build time from the operator's own
+// corpus and gated on that corpus existing — see loadSummary(). There are no
+// hardcoded figures, because they move: the public-archive count went from three
+// to four during this page's own construction. A build without data drops the
+// numeric bands entirely rather than printing zeroes.
+
+const CONTAINER = "w-full max-w-6xl mx-auto px-5 sm:px-6";
+
+function Band({
+ children,
+ className,
+}: {
+ children: React.ReactNode;
+ className?: string;
+}) {
+ return (
+ <section className={`border-t border-[var(--border)] ${className ?? ""}`}>
+ <div className={`${CONTAINER} py-14 sm:py-20`}>{children}</div>
+ </section>
+ );
}
-// Read the build-time summary from public/. Absent in plain `next dev` (compose
-// hasn't run) — fall back to an empty summary so the page still renders.
-function loadSummary(): HomepageSummary {
- try {
- const file = path.join(process.cwd(), "public", "homepage-summary.json");
- return JSON.parse(fs.readFileSync(file, "utf8")) as HomepageSummary;
- } catch {
- return emptySummary();
- }
+function Eyebrow({ children }: { children: React.ReactNode }) {
+ return <h2 className="label-machine mb-8">{children}</h2>;
}
export default function Home() {
const summary = loadSummary();
- // The compact wordmark lives in the sticky header (same on every page), so the
- // home route starts straight into the landing.
- return <HomeLanding summary={summary} />;
+ const hours = summary?.totals.hoursArchived ?? null;
+ const gone = summary?.availability?.byState.deleted ?? 0;
+ const counted = summary?.availability?.counted ?? 0;
+
+ return (
+ <>
+ {/* ── Hero ─────────────────────────────────────────────────────────── */}
+ <section className={`${CONTAINER} pt-14 pb-16 sm:pt-20 sm:pb-24`}>
+ <div className="flex flex-col gap-8">
+ <h1
+ className="reveal font-display text-[2.5rem] leading-[0.98] font-bold uppercase tracking-[0.005em] text-[var(--foreground)] sm:text-[4rem] lg:text-[4.75rem] max-w-5xl"
+ style={{ "--d": "60ms" } as React.CSSProperties}
+ >
+ {hours ? (
+ <>
+ <span className="tabular font-semibold">
+ {hours.toLocaleString()}
+ </span>{" "}
+ hours of speech,
+ <br />
+ searchable to the second.
+ </>
+ ) : (
+ <>
+ Every word a channel
+ <br />
+ ever said, searchable
+ <br />
+ to the second.
+ </>
+ )}
+ </h1>
+
+ <p
+ className="reveal max-w-2xl text-lg leading-relaxed text-[var(--muted-foreground)]"
+ style={{ "--d": "160ms" } as React.CSSProperties}
+ >
+ Archilyzer downloads a channel’s back catalogue, transcribes it
+ on your machine, and builds a static site you host yourself. No
+ account, no service, no server of ours in the path.
+ </p>
+
+ <div
+ className="reveal flex flex-wrap items-center gap-x-8 gap-y-4"
+ style={{ "--d": "240ms" } as React.CSSProperties}
+ >
+ <Link
+ href="/downloads/"
+ className="inline-flex items-center gap-2 bg-[var(--foreground)] px-6 py-3 font-mono text-xs uppercase tracking-[0.16em] text-[var(--background)] rounded-[var(--radius)] hover:bg-[var(--brand)] hover:text-[var(--brand-ink)] transition-colors"
+ >
+ Download the source
+ </Link>
+ <Link
+ href="/docs/install/"
+ className="font-mono text-xs uppercase tracking-[0.16em] text-[var(--muted-foreground)] underline decoration-[var(--border-strong)] underline-offset-[6px] hover:text-[var(--foreground)] transition-colors"
+ >
+ Read the setup guide →
+ </Link>
+ </div>
+
+ {/* The rail: real acquisitions, real states. The signature object. */}
+ <div
+ className="reveal mt-6"
+ style={{ "--d": "320ms" } as React.CSSProperties}
+ >
+ <RecentAdditions items={summary?.recent ?? []} />
+ </div>
+ </div>
+ </section>
+
+ {/* ── The numbers, only when there are numbers ─────────────────────── */}
+ {summary && (
+ <Band>
+ <Eyebrow>This deployment, at last build</Eyebrow>
+ <KpiHeader summary={summary} />
+ </Band>
+ )}
+
+ {/* ── What it does ─────────────────────────────────────────────────── */}
+ <Band>
+ <Eyebrow>What it does</Eyebrow>
+ <div className="grid grid-cols-1 gap-px bg-[var(--border)] border border-[var(--border)] md:grid-cols-3">
+ {[
+ {
+ title: "Archive",
+ body: "Point it at a channel and it takes the whole back catalogue — then keeps up with new uploads on a schedule you set. Where the platform publishes captions it takes those; where it doesn't, it downloads the audio and transcribes it locally.",
+ },
+ {
+ title: "Publish",
+ body: "A build produces plain HTML and JSON: no database, no runtime, nothing to keep running. One corpus can publish several sites, each with its own channels, name and domain, without duplicating a byte of data.",
+ },
+ {
+ title: "Search & ask",
+ body: "Search every transcript at once and land on the exact second a phrase was said. Each archive also publishes a machine-readable index, so an LLM can navigate it without scraping.",
+ },
+ ].map((panel) => (
+ <div
+ key={panel.title}
+ className="bg-[var(--surface)] p-6 flex flex-col gap-3"
+ >
+ <h3 className="font-display text-lg font-bold uppercase tracking-[0.05em] text-[var(--foreground)]">
+ {panel.title}
+ </h3>
+ <p className="text-sm leading-[1.7] text-[var(--muted-foreground)]">
+ {panel.body}
+ </p>
+ </div>
+ ))}
+ </div>
+ </Band>
+
+ {/* ── The proof point ──────────────────────────────────────────────── */}
+ {gone > 0 && (
+ <Band>
+ <div className="flex flex-col gap-5 max-w-3xl">
+ <p className="font-display text-2xl sm:text-3xl font-bold uppercase leading-[1.15] tracking-[0.01em] text-[var(--foreground)]">
+ <span className="tabular font-semibold text-[var(--state-gone)]">
+ {gone.toLocaleString()}
+ </span>{" "}
+ recordings in this archive no longer exist where they came from.
+ </p>
+ <p className="text-[var(--muted-foreground)] leading-relaxed">
+ Those are the ones re-checked and found gone — deleted, or pulled
+ behind a wall. Their transcripts are still here, still searchable,
+ still citable to the second.
+ </p>
+ <p className="text-sm text-[var(--faint)] leading-relaxed">
+ The other {(counted - gone).toLocaleString()} are simply not known
+ to be gone. Nothing re-checks tens of thousands of recordings
+ continuously, so that figure is a floor, not a guarantee — which is
+ rather the argument for keeping a copy.
+ </p>
+ </div>
+ </Band>
+ )}
+
+ {/* ── How it works ─────────────────────────────────────────────────── */}
+ <Band>
+ <Eyebrow>How it works</Eyebrow>
+ <ol className="grid grid-cols-1 gap-8 sm:grid-cols-2 lg:grid-cols-4 list-none">
+ {[
+ {
+ n: "01",
+ t: "Pick channels",
+ d: "Add a URL. Decide whether its captions are taken as published or its audio is transcribed here.",
+ },
+ {
+ n: "02",
+ t: "Transcribe",
+ d: "Audio is transcribed on your own hardware, by whisper.cpp or another local backend. Nothing is uploaded.",
+ },
+ {
+ n: "03",
+ t: "Compose",
+ d: "Transcripts become a paginated index and a static site — rebuilt incrementally, so one new video costs one new video.",
+ },
+ {
+ n: "04",
+ t: "Ship",
+ d: "Deploy the output anywhere that serves files. The reference deployment runs on a free hosting tier.",
+ },
+ ].map((step) => (
+ <li key={step.n} className="flex flex-col gap-2">
+ <span className="tabular text-xs text-[var(--faint)]">
+ {step.n}
+ </span>
+ <h3 className="font-display text-base font-bold uppercase tracking-[0.05em] text-[var(--foreground)]">
+ {step.t}
+ </h3>
+ <p className="text-sm leading-[1.7] text-[var(--muted-foreground)]">
+ {step.d}
+ </p>
+ </li>
+ ))}
+ </ol>
+ </Band>
+
+ {/* ── What this isn't. The highest-trust block on the page, and the
+ cheapest: it is all simply true. ──────────────────────────────── */}
+ <Band>
+ <Eyebrow>What this isn’t</Eyebrow>
+ <ul className="grid grid-cols-1 gap-x-10 gap-y-6 sm:grid-cols-2 list-none max-w-4xl">
+ {[
+ [
+ "Not a hosted service",
+ "There is nothing to sign up for. It runs on your machine and publishes to your hosting.",
+ ],
+ [
+ "No public repository",
+ "The source ships as a dated snapshot tarball. No history, no branches, nothing to pull.",
+ ],
+ [
+ "You supply the tools",
+ "yt-dlp, ffmpeg and a transcription backend are yours to install. Archilyzer drives them; it doesn't bundle them.",
+ ],
+ [
+ "Transcription is slow",
+ "It is bound by your GPU or CPU. A large back catalogue is days of machine time, and no wishing changes that.",
+ ],
+ ].map(([title, body]) => (
+ <li key={title} className="flex flex-col gap-1.5">
+ <h3 className="font-mono text-xs uppercase tracking-[0.14em] text-[var(--foreground)]">
+ {title}
+ </h3>
+ <p className="text-sm leading-[1.7] text-[var(--muted-foreground)]">
+ {body}
+ </p>
+ </li>
+ ))}
+ </ul>
+ </Band>
+
+ {/* ── The reference deployment ─────────────────────────────────────── */}
+ {summary && summary.sites.length > 0 && (
+ <Band>
+ <Eyebrow>The reference deployment</Eyebrow>
+ <p className="mb-8 max-w-2xl text-[var(--muted-foreground)] leading-relaxed">
+ Public archives built and hosted with Archilyzer, from one corpus on
+ one machine. This is what the output looks like in use.
+ </p>
+ <ArchiveCards sites={summary.sites} />
+ </Band>
+ )}
+ </>
+ );
}
diff --git a/homepage/app/stats/page.tsx b/homepage/app/stats/page.tsx
@@ -0,0 +1,60 @@
+import type { Metadata } from "next";
+import { HomeLanding } from "../components/HomeLanding";
+import { PageShell, PageHeading } from "../components/PageShell";
+import { loadSummary } from "../lib/summary";
+
+// The operator's dashboard: headline KPIs, the cross-site activity chart with
+// its preset/customize strip, and the site grid. This used to BE the home route
+// — it moved here when `/` became the project's marketing page, restoring the
+// name the dashboard had before 33ed2fe deleted it.
+//
+// Everything on this page is OPERATOR identity, computed from whoever's corpus
+// built the bundle. It renders from the small summary compose-homepage writes to
+// public/homepage-summary.json, so it is static HTML with no multi-MB client
+// fetch.
+
+export const metadata: Metadata = {
+ title: "Stats",
+ description:
+ "Cross-site activity for this deployment: transcription and download " +
+ "counts over time, by site and by channel.",
+};
+
+// A build with no corpus data still serves this route — the nav must never
+// 404 — but it says why the numbers are missing instead of drawing zeroes.
+function NoData() {
+ return (
+ <div className="doc-measure flex flex-col gap-4">
+ <p className="text-[var(--muted-foreground)] leading-relaxed">
+ This build shipped without dashboard data. The numbers here are
+ generated from the operator’s own corpus at build time, and a
+ source-only build has none.
+ </p>
+ <p className="text-sm text-[var(--faint)] leading-relaxed">
+ Running <code className="font-mono text-[var(--foreground)]">pnpm --filter homepage run build</code>{" "}
+ against a configured install indexes the corpus and composes the summary
+ first. <code className="font-mono text-[var(--foreground)]">build:nodata</code>{" "}
+ skips that step on purpose, so the site can be built from the source
+ tarball alone.
+ </p>
+ </div>
+ );
+}
+
+export default function StatsPage() {
+ const summary = loadSummary();
+ return (
+ <PageShell className="flex flex-col gap-10">
+ <PageHeading
+ eyebrow="This deployment"
+ title="Stats"
+ standfirst={
+ summary
+ ? "What this install has archived, and when. Every number is computed at build time from the operator's own corpus."
+ : undefined
+ }
+ />
+ {summary ? <HomeLanding summary={summary} /> : <NoData />}
+ </PageShell>
+ );
+}
diff --git a/homepage/content/README.md b/homepage/content/README.md
@@ -0,0 +1,55 @@
+# homepage/content — provenance and drift rules
+
+This directory holds the hand-written documentation rendered at `/docs/`. It is
+**not** a route: nothing under `content/` is served, and this file in particular
+exists for whoever edits the docs next.
+
+## Why these are hand-written rather than rendered from the root docs
+
+The obvious move is to render `SETUP.md`, `DEPLOY_CLOUDFLARE.md` and friends
+directly, and keep one copy. That was rejected for a decisive reason:
+
+> `SETUP.md` says `git clone <this-repo-url>`. **There is no public repository.**
+> Rendering that on a public site publishes an instruction that cannot be
+> followed — the acquisition path is a dated snapshot tarball from `/downloads/`.
+
+The root docs also lean on contributor furniture that actively confuses an
+operator: `pnpm wt` worktrees, the machine-global e2e queue, `plans/`, shard
+counts, wall-time tables for the test suite. Someone who wants to archive a
+channel does not need to know that the e2e suite serializes on a lock file.
+
+## The accepted cost
+
+These pages will drift from the root docs. That is the trade, taken knowingly.
+Symlinking or transcluding would only move the problem: the two audiences
+genuinely want different documents, and a shared file would end up serving
+neither. The mitigation is this table — when you change a root doc, check whether
+its public counterpart needs the same change.
+
+| Public page | Derived from | Watch for drift in |
+|---|---|---|
+| `docs/what-is-archilyzer.md` | `README.md` | package list, pipeline modes |
+| `docs/install.md` | `SETUP.md` | tool versions, env-var table, backend list |
+| `docs/operate.md` | `README.md`, `SCHEDULED_SYNC.md` | editor routes, scheduler settings |
+| `docs/deploy-cloudflare.md` | `DEPLOY_CLOUDFLARE.md` | the 25 MB Pages limit, R2 options |
+| `docs/deploy-docker.md` | `DEPLOY_DOCKER.md` | phase structure, settings names |
+| `docs/ai-and-mcp.md` | `mcp/README.md` | tool names, `corpus.json` shape |
+| `docs/faq.md` | — (written for this site) | claims about cost and hardware |
+
+## House rules for these files
+
+- **No frontmatter.** Title, blurb, group and order live in `content/docs.ts`.
+- **No raw HTML and no `<!-- comments -->`.** The renderer is `markdown-to-jsx`
+ with `disableParsingRawHTML` **off**, so raw markup would be parsed and
+ rendered, not escaped. (The header comment on
+ `common/components/Markdown.tsx` claims otherwise; it is wrong.)
+- **Relative links** to other doc pages (`/docs/install/`) render as client-side
+ navigation; absolute `http(s)://` links open in a new tab. That behaviour is
+ in `app/components/DocBody.tsx`, not in the Markdown.
+- **A leading `# Title` is stripped** by the loader, because the page renders its
+ own `<h1>` from the manifest. Keep it in the file anyway so each document is
+ valid standalone.
+- **Adding a page** means adding the file *and* the `DOCS` entry. A manifest
+ entry whose file is missing is filtered out of both the index and
+ `generateStaticParams`, so it 404s rather than breaking the build — which is
+ the safe failure, but it is still a mistake.
diff --git a/homepage/content/docs.ts b/homepage/content/docs.ts
@@ -0,0 +1,89 @@
+// The documentation manifest — the one reviewable list that decides what /docs
+// contains, in what order, under which headings.
+//
+// WHY A MANIFEST AND NOT FRONTMATTER. Nothing in this tree parses frontmatter,
+// and adding `gray-matter` would fight the workspace's `minimumReleaseAge` pin
+// for a dependency whose whole job is to hold four strings. Beyond that: order
+// and grouping are EDITORIAL decisions about the set, not facts about any one
+// file — a page can't sensibly declare that it comes third. Here a typo is a
+// `tsc` error rather than a page that quietly vanishes from the index.
+
+export type DocGroup = "start" | "operate" | "publish" | "extend";
+
+export type DocEntry = {
+ /** URL segment: /docs/<slug>/ */
+ slug: string;
+ /** File under content/docs/. */
+ file: string;
+ /** <h1> and nav label. */
+ title: string;
+ /** One line, shown on the index and as the page's meta description. */
+ blurb: string;
+ group: DocGroup;
+};
+
+export const DOC_GROUP_ORDER: { id: DocGroup; label: string }[] = [
+ { id: "start", label: "Start here" },
+ { id: "operate", label: "Running it" },
+ { id: "publish", label: "Publishing" },
+ { id: "extend", label: "Going further" },
+];
+
+export const DOCS: DocEntry[] = [
+ {
+ slug: "what-is-archilyzer",
+ file: "what-is-archilyzer.md",
+ title: "What Archilyzer is",
+ blurb:
+ "The shape of the thing: three programs, one corpus, and a static site at the end of it.",
+ group: "start",
+ },
+ {
+ slug: "install",
+ file: "install.md",
+ title: "Install",
+ blurb:
+ "What to install, in what order, and which parts you can skip until you need them.",
+ group: "start",
+ },
+ {
+ slug: "operate",
+ file: "operate.md",
+ title: "Running an archive",
+ blurb:
+ "Adding channels, downloading, transcribing, and keeping an archive current without babysitting it.",
+ group: "operate",
+ },
+ {
+ slug: "deploy-cloudflare",
+ file: "deploy-cloudflare.md",
+ title: "Deploy to Cloudflare",
+ blurb:
+ "Publishing the static site to Pages, and what to do about archives too large for it.",
+ group: "publish",
+ },
+ {
+ slug: "deploy-docker",
+ file: "deploy-docker.md",
+ title: "Building several sites at once",
+ blurb:
+ "The opt-in Docker pipeline: parallel per-site builds in isolated containers.",
+ group: "publish",
+ },
+ {
+ slug: "ai-and-mcp",
+ file: "ai-and-mcp.md",
+ title: "AI and MCP",
+ blurb:
+ "Every archive is machine-navigable. Point Claude Code, Cursor, or a browser chat at it.",
+ group: "extend",
+ },
+ {
+ slug: "faq",
+ file: "faq.md",
+ title: "Questions",
+ blurb:
+ "Cost, legality, hardware, what happens when a video disappears, and what this is not.",
+ group: "extend",
+ },
+];
diff --git a/homepage/content/docs/ai-and-mcp.md b/homepage/content/docs/ai-and-mcp.md
@@ -0,0 +1,77 @@
+# AI and MCP
+
+Every archive Archilyzer builds is machine-navigable by design. That is not a
+bolt-on: the same paginated JSON the site's own search reads is public, described
+by an index that tells a client how to navigate it, and served with the CORS
+headers that let anything fetch it.
+
+## The published contract
+
+Every archive serves two files that exist for machines:
+
+- **`/corpus.json`** — a machine-readable index. It does not contain transcripts;
+ a corpus of tens of thousands of videos cannot be enumerated per-video without
+ running into a host's file-count limits. Instead it *documents how to navigate
+ the shards*: where each channel's manifest is, how to turn a video's id into a
+ page number, and what a record looks like when you get there. It also names the
+ software that built the archive.
+- **`/llms.txt`** — the same thing as prose, following the llmstxt.org
+ convention, for a client that reads a page rather than an API.
+
+Between them, an LLM with a fetch tool can navigate a whole archive without
+scraping a single HTML page, and without any code of ours running on its side.
+
+## The MCP server
+
+Included in the source is an MCP server that exposes an archive — a single site,
+or several at once — to Claude Code, Claude Desktop, Cursor, or any other MCP
+client.
+
+It is **a local tool you run yourself**. It changes nothing: it only reads
+already-published static JSON, either from a directory on disk or over HTTP.
+
+What it can do:
+
+- **Search transcripts** for a term, a phrase or a regular expression, with
+ timestamped snippets. Every timestamp is a link to that exact second of the
+ recording, in the archive's own viewer.
+- **Filter before scanning** — by availability state, upload date range, media
+ type, or scope. This is not just convenience: a filtered query computes exactly
+ which pages it needs before reading a single transcript byte, which is the
+ difference between reading a whole corpus and reading a fraction of it.
+- **Enumerate a query's complete match set** as a worklist in one pass, so an
+ agent can cover everything rather than reporting a number it took from the
+ first page.
+- **Batch-read** many videos at once, as bounded excerpt windows around a query
+ rather than whole transcripts.
+- **Report everything known about one video** — metadata, engagement counts,
+ whether other copies of the same recording are archived, and how much of its
+ runtime the transcript actually covers.
+
+## The honesty features, which are the point
+
+Anything that lets a model summarise an archive can also let it summarise the
+archive *wrongly*, confidently. Several behaviours exist specifically to make
+that harder:
+
+- **A partial result says so first, not in a footnote.** If a scan hits a cap,
+ the first line of the response says the result is a sample and names which
+ channels went unread.
+- **Counts are of recordings, not uploads.** When the same recording exists on
+ two platforms, it collapses to one row — and says how many it collapsed, and
+ which. A count that double-counts mirrors looks exactly like a correct one.
+- **A truncated transcript is flagged as a warning**, not reported as a
+ percentage, because a transcript covering 40% of a video will otherwise support
+ a confident conclusion that something was never said.
+- **A timestamp is never translated across two copies** of a recording unless the
+ two were actually measured as aligned. A mirror with a different intro carries
+ the same words at different times, and a translated citation would look
+ perfectly plausible while pointing at the wrong moment.
+
+## In the browser, without any of this
+
+The published site also carries a chat interface that searches the archive and
+answers with citations, using an API key the visitor supplies themselves. The key
+stays in their browser; nothing is proxied through the archive.
+
+Next: [Questions](/docs/faq/).
diff --git a/homepage/content/docs/deploy-cloudflare.md b/homepage/content/docs/deploy-cloudflare.md
@@ -0,0 +1,84 @@
+# Deploy to Cloudflare
+
+The published site is static files, so it can be hosted anywhere that serves
+them. Cloudflare Pages is what the reference deployment uses, and everything
+described here fits inside Cloudflare's free tier.
+
+## The straightforward part
+
+A build produces an `out/` directory. Deploying it is one command:
+
+```sh
+wrangler pages deploy out --project-name my-archive --branch main
+```
+
+The editor can run this for you as part of a build-and-deploy action. Pass
+`--branch main` explicitly: wrangler otherwise infers the branch from git, and
+from anywhere that isn't the main branch it will quietly publish a *preview* URL
+instead of production — which looks like a successful deploy of a site nobody can
+see.
+
+## The part that needs thought: bulk archives
+
+Each build can generate one transcript zip and one live-chat zip per channel, so
+that people can take the raw material rather than scraping the site. These are
+listed in a manifest that the Downloads page reads.
+
+Cloudflare Pages rejects any single file larger than **25 MB**, and real channels
+blow through that easily — a busy channel's live-chat archive can run to hundreds
+of megabytes. So archives are split by size:
+
+| Archive size | Served from |
+|---|---|
+| Up to 25 MB | Pages, shipped inside the site |
+| Over 25 MB, object storage configured | Cloudflare R2 |
+| Over 25 MB, nothing configured | Not served; shown as "too large to host" |
+
+Oversize archives are staged during the build and uploaded before the site
+deploy, so the links resolve the moment the site goes live.
+
+Note that only the editor's deploy actions upload them — that is where the
+storage credentials live. Building and deploying by hand from the command line
+stages the files without uploading them.
+
+## Setting up object storage
+
+One bucket serves every site; objects are namespaced per site, so there is no
+need for one bucket each.
+
+Buckets are private by default, and the Downloads page links to plain URLs, so
+the bucket needs public read access. There are three ways to arrange that:
+
+- **A small worker on a free `workers.dev` subdomain** — recommended. It streams
+ objects from the bucket and gives you edge caching *and* rate limiting you
+ control in code. No domain purchase, no WHOIS record, and one worker serves
+ every site. A ready-made one is included in the source.
+- **A custom domain** on your Cloudflare account, which routes downloads through
+ the CDN and the dashboard's own rate-limiting rules.
+- **The built-in public subdomain**, which is fine for a smoke test but is
+ throttled by Cloudflare and gives you no cache or rate-limit control of your
+ own.
+
+## Why the rate limiting matters
+
+Public download links to multi-hundred-megabyte files are, left unguarded, a way
+for someone to run up your bill on purpose. The worker option exists precisely
+because it puts a limiter in front of them that you control, at no cost, without
+owning a domain.
+
+Uploads go through the S3-compatible API rather than the wrangler CLI, because
+wrangler caps a single upload at 300 MiB and real archives exceed that. That is
+why archive uploads need a second set of credentials beyond the ones the site
+deploy already uses.
+
+## If you host somewhere else
+
+Nothing in the published output is Cloudflare-specific. It is HTML, JSON, and
+static assets: any object store or web server will do. Two things to configure
+wherever you land:
+
+- **Serve the 404 page** for unmatched paths.
+- **Set a short cache lifetime on the downloads directory** if you publish
+ archives, so a rebuilt archive isn't shadowed by a stale cached copy.
+
+Next: [Building several sites at once](/docs/deploy-docker/).
diff --git a/homepage/content/docs/deploy-docker.md b/homepage/content/docs/deploy-docker.md
@@ -0,0 +1,52 @@
+# Building several sites at once
+
+If one corpus publishes several sites, the default build handles them one after
+another. The container pipeline builds them **in parallel**, in isolated
+containers, and then deploys them serially. It is opt-in, and the ordinary build
+is unchanged and remains the default.
+
+Turn it on in the editor's settings, or from the toggle on the deploy page.
+
+## What you need
+
+- A container engine: Docker, or podman. Rootless podman is a good fit, since it
+ maps container files to your host user automatically.
+- The host still needs Node and pnpm — two of the three phases run there — plus
+ your hosting credentials in the environment.
+
+The build image is created and cached automatically the first time you run it.
+
+## How a parallel build runs
+
+One job runs three ordered phases:
+
+1. **Shared work, on the host, once.** Build the search index and warm the
+ archive cache for the union of every site's channels. Only the host writes
+ this shared state, so the containers can never race each other for it. This
+ phase is the long pole on a cold build and nearly instant on a warm one,
+ because unchanged channels are skipped.
+2. **Per-site builds, in parallel containers.** Each site composes and builds in
+ its own container, up to a configured concurrency limit, writing to its own
+ output directory. Containers mount the corpus and index **read-only**.
+3. **Deploy, on the host, serially.** Once every build has finished, each site is
+ published in turn.
+
+## Why credentials never enter a container
+
+Deployment runs on the host, in phase three, precisely so that hosting
+credentials are never passed into a build container. A build container gets
+read-only mounts, a non-root user, and its own output directory — and nothing
+that could publish anything.
+
+Network access is left enabled inside the containers, because the site build
+fetches its web fonts. The isolation comes from the read-only mounts and the
+separate output directories, not from cutting the network.
+
+## When it is worth it
+
+The parallel pipeline pays off in proportion to how many sites you publish and
+how little they share. With one site it is strictly slower — you pay for image
+management and gain nothing. With several, phase two collapses from the sum of
+the build times to roughly the longest single one.
+
+Next: [AI and MCP](/docs/ai-and-mcp/).
diff --git a/homepage/content/docs/faq.md b/homepage/content/docs/faq.md
@@ -0,0 +1,98 @@
+# Questions
+
+## What does it cost to run?
+
+The software is free and MIT licensed. Your costs are hardware, electricity, and
+hosting.
+
+Hosting a published archive is the cheap part: static files on Cloudflare Pages
+fit comfortably in the free tier, and bulk download archives too large for it can
+go behind a free worker. The reference deployment publishes tens of thousands of
+transcripts this way.
+
+The expensive part is transcription, and it is paid in machine time rather than
+money. See below.
+
+## How long does transcribing take?
+
+It depends entirely on your hardware and which backend you use, and any number
+quoted without both is meaningless.
+
+What is safe to say: it is the bottleneck, it wants a GPU, and it scales with
+audio hours rather than video count. A back catalogue of several thousand hours
+is a project measured in days of machine time, not minutes. Channels that already
+publish captions skip this step entirely and cost almost nothing.
+
+One trap worth knowing about: running too many jobs at once can push a GPU
+backend into falling back to the CPU, at which point everything gets several
+times slower simultaneously. More parallelism is not reliably faster.
+
+## How much disk does it take?
+
+Transcripts are small — text compresses well and a whole corpus of them is
+megabytes. Media is not. Whether you keep the audio after transcribing is a
+per-channel decision, and the software tracks which recordings are safe to clean
+up and which are not.
+
+Audio for recordings that have disappeared upstream is never cleaned up
+automatically, on the grounds that it is the one thing you cannot re-fetch.
+
+## Do I need a GPU?
+
+No, but transcription without one is slow enough to change what is practical. A
+CPU-only machine is fine for a channel that publishes captions, or for a small
+back catalogue you are content to work through over weeks.
+
+## Is any of this sent anywhere?
+
+No. Downloads go from the platform to your machine, transcription runs on your
+machine, and the build writes files on your machine. Nothing is sent to us —
+there is no "us" in the runtime path at all.
+
+The two exceptions are both yours to make: publishing sends the built site to
+whatever host you choose, and the optional in-browser chat on a published site
+sends queries to an AI provider using an API key the *visitor* supplies.
+
+## Is this legal?
+
+Downloading and keeping a copy of publicly published video is a question with
+different answers in different jurisdictions, and it is not one a piece of
+software can answer for you. The tool does what you tell it to; deciding what to
+tell it is your responsibility, as is complying with whatever terms and laws
+apply where you are.
+
+What the software does do is make the record accurate: it says what it archived,
+when, and whether the original is still there — rather than presenting an
+archive as if it were the platform.
+
+## What happens when a video is deleted?
+
+Your copy stays. The archive re-checks whether recordings still exist at their
+source and records the answer, so a deleted recording is marked as deleted,
+remains searchable, and keeps its transcript.
+
+The reverse claim is one the software is careful never to make. A recording
+counts as available until something re-checks it and finds otherwise, and nothing
+re-checks a whole corpus continuously. "Available" here means *not known to be
+gone* — never "confirmed still there".
+
+## Where is the git repository?
+
+There isn't one. The source is published as a
+[dated snapshot tarball](/downloads/): a working tree with no history, no
+branches and no remote. Updating means downloading a newer snapshot.
+
+## Can I use it for one video?
+
+You can, but nothing about it is designed for that. It is built around back
+catalogues — thousands of recordings, kept current on a schedule — and almost
+every design decision in it, from the paginated index to the incremental build,
+exists because of scale.
+
+## What this is not
+
+- **Not a hosted service.** There is no account, no subscription, no upload.
+- **Not a public repository.** Dated snapshots, no `git pull`.
+- **Not a downloader.** It supplies no downloader — you install `yt-dlp`,
+ `ffmpeg` and a transcription backend yourself, and it drives them.
+- **Not cloud transcription.** Your hardware, your electricity, your queue.
diff --git a/homepage/content/docs/install.md b/homepage/content/docs/install.md
@@ -0,0 +1,154 @@
+# Install
+
+From an unpacked snapshot to a running editor. The web apps need very little; the
+download-and-transcribe pipeline needs the media tools, and you can add those
+later.
+
+## What you need
+
+Always required, to install dependencies and run the apps:
+
+| Tool | Version | Notes |
+|---|---|---|
+| Node.js | 20.9 or newer | LTS 20 or 22. Required by Next.js 16. |
+| pnpm | 9 or newer | Easiest via Corepack, which ships with Node. |
+| A C/C++ toolchain | platform default | Only if no prebuilt binary exists for a native module on your platform. Usually not needed. |
+
+Needed only for the pipeline — install what you will actually use:
+
+| Tool | Needed for |
+|---|---|
+| yt-dlp | Downloading and syncing channels. The whole pipeline. |
+| ffmpeg and ffprobe | Audio transcode and duration checks. |
+| A transcription backend | Channels you transcribe yourself. whisper.cpp by default. |
+| rsync | Backing up the saved-video store, if you enable it. |
+| Docker or podman | Only for the parallel multi-site build. |
+
+The apps start and run with **none** of the second table installed. You just
+can't fetch anything yet.
+
+## Get the code
+
+There is no public git repository. The source is published here as a dated
+snapshot tarball:
+
+```sh
+curl -LO https://archilyzer.pages.dev/downloads/archilyzer-source.tar.gz
+tar xzf archilyzer-source.tar.gz
+cd archilyzer
+```
+
+That is a working tree, not a clone — no history, no branches, no remote,
+nothing to `git pull`. Updating means downloading a newer snapshot. See
+[Downloads](/downloads/) for what is and isn't inside.
+
+## Install and start
+
+```sh
+corepack enable # once, so the pinned pnpm version is used
+pnpm install # installs every workspace package
+pnpm dev:editor # the editor, at http://localhost:3001
+```
+
+The editor **starts fine with no data**. A fresh unpack has no corpus directory,
+and that is expected — it is created as you use it. Add your first channel from
+the editor's Channels page.
+
+To build and serve the public site:
+
+```sh
+pnpm build # static site under export/out/
+pnpm start:export # serve it at http://localhost:3000
+```
+
+## Platform notes
+
+### Linux
+
+Install Node from your distribution or a version manager, then the media tools:
+
+```sh
+# Debian / Ubuntu
+sudo apt install nodejs yt-dlp ffmpeg rsync build-essential
+
+# Fedora
+sudo dnf install nodejs yt-dlp ffmpeg rsync gcc-c++ make
+
+# Arch
+sudo pacman -S nodejs yt-dlp ffmpeg rsync base-devel
+```
+
+If your distribution's Node is older than 20.9, install it with `fnm` or `nvm`
+instead. The build-tools package is the compiler pnpm falls back to when a
+native module has no prebuilt binary for your platform.
+
+### macOS
+
+```sh
+xcode-select --install # C/C++ toolchain
+brew install node pnpm yt-dlp ffmpeg rsync
+corepack enable
+```
+
+whisper.cpp can be installed straight from Homebrew (`brew install
+whisper-cpp`), which saves you building it.
+
+### Windows
+
+**Use WSL2.** The transcription toolchain and the helper shell scripts assume a
+Unix shell, so the smooth path is:
+
+```powershell
+wsl --install -d Ubuntu
+```
+
+then follow the Linux steps inside it. Keep the files **inside** the WSL
+filesystem rather than under `/mnt/c/`, or file watching and I/O will be slow.
+
+The apps and the yt-dlp pipeline do work natively on Windows — install Node,
+yt-dlp and ffmpeg with winget or Scoop — but building the transcription backends
+and running the shell scripts is Unix-oriented.
+
+## Transcription backends
+
+You only need one of these for channels whose audio you transcribe yourself.
+Channels where the platform publishes captions need no backend at all.
+
+- **whisper.cpp** — the default. Build it, download a `ggml` model, and either
+ put `whisper-cli` on your `PATH` or point `WHISPER_BIN` at it. Set
+ `WHISPER_MODEL` to the model file.
+- **chough** — set `CHOUGH_BIN`; it downloads a model when none is configured,
+ and can talk to a remote server via `CHOUGH_URL`.
+- **parakeet.cpp** — driven through a bundled overlapping-segment wrapper. Needs
+ `parakeet-cli` and a `.gguf` model. It is interruptible: stopping it stitches
+ together the partial transcript rather than throwing the work away.
+
+The active backend is chosen once, globally, in the editor's settings. All three
+call ffmpeg, so that must be installed either way.
+
+## Configuration
+
+Most configuration lives in the editor's settings page and is written to a
+settings file at the root. It is entirely optional — a missing or partial file
+falls back to built-in defaults, so the app runs out of the box.
+
+Paths and binaries can also be overridden by environment variables set before
+launch. The ones worth knowing:
+
+| Variable | Default | Purpose |
+|---|---|---|
+| `TRANSCRIPTS_DIR` | `<root>/transcripts` | The corpus: channels, media, index, job logs. |
+| `SAVED_VIDEOS_DIR` | inside `TRANSCRIPTS_DIR` | Persisted source-video store; can live on another disk. |
+| `SITES_DIR` | inside `TRANSCRIPTS_DIR` | Per-site configuration. |
+| `EXPORT_PUBLIC_DIR` | `<root>/export/public` | Where the index writes its paginated JSON. |
+| `YTDLP_BIN` | `yt-dlp` on `PATH` | The downloader. |
+| `WHISPER_BIN` | `whisper-cli` on `PATH` | whisper.cpp binary. |
+| `WHISPER_MODEL` | a path under your home directory | whisper.cpp model file. |
+| `FFMPEG_BIN` / `FFPROBE_BIN` | on `PATH` | Transcode and duration checks. |
+| `PARALLEL_TRANSCRIBE_LIMIT` | `4` | Maximum simultaneous transcription jobs. |
+
+Pointing `TRANSCRIPTS_DIR` at a large disk before you start is the one decision
+worth making early — the corpus grows to whatever your channels amount to, and
+moving it later means moving everything.
+
+Next: [Running an archive](/docs/operate/).
diff --git a/homepage/content/docs/operate.md b/homepage/content/docs/operate.md
@@ -0,0 +1,108 @@
+# Running an archive
+
+Day-to-day operation: adding channels, keeping them current, and getting from a
+pile of audio to a published site.
+
+## Adding a channel
+
+Everything starts on the editor's **Channels** page. A channel needs a URL and a
+decision about how it is handled:
+
+- **Captions** — for platforms that already publish machine captions, those are
+ fetched directly and no media is downloaded. This is fast and costs almost
+ nothing.
+- **Transcribe** — the audio is downloaded and transcribed on your machine. This
+ is the slow, expensive path, and the one that produces a transcript for
+ material that never had captions at all.
+
+You can also set per-channel download arguments, an audio format, and whether to
+pass cookies from a browser profile for material that needs an account.
+
+## The three fetch modes
+
+The pipeline offers three operations per channel, and the distinction matters
+once a catalogue gets large:
+
+- **Store playlist** — ask the platform for the channel's full list of videos and
+ write it down. Nothing is downloaded.
+- **Download from playlist** — compare that list against what has already been
+ fetched, and download only the difference. Filtering on our side rather than
+ the downloader's avoids re-fetching metadata for thousands of entries you
+ already have, which on some platforms is the difference between minutes and
+ hours.
+- **Sync** — a quick incremental pass for new uploads. This is what runs on a
+ schedule.
+
+## Keeping it current without watching it
+
+Each channel can carry its own sync cadence — hourly, daily, every N minutes.
+A heartbeat runs frequently and the server decides which channels are actually
+due, so nothing is scheduled twice and a missed tick simply runs at the next one.
+There are no catch-up storms after the machine has been asleep.
+
+There are two ways to drive the heartbeat, and they are interchangeable:
+
+- **Internal timer** — the editor arms it at startup. Set a cadence in settings
+ and there is nothing else to install. This is the recommended setup.
+- **External cron** — leave the internal timer off and have the system's cron
+ hit the tick endpoint instead. Useful when schedules are managed centrally.
+
+Scheduled syncs run through the same queue as a manual click, so they cannot
+collide with one, and they appear live in the editor exactly like manual work.
+
+## Transcribing
+
+Transcription is the bottleneck, and it is bounded by hardware rather than
+patience. A few things worth knowing:
+
+- Jobs run in parallel up to a configured limit. More is not always faster —
+ when the machine runs out of memory, a GPU backend silently falls back to the
+ CPU and everything gets slower at once.
+- Failures are retryable per channel, and there is a verification pass that
+ re-checks transcripts against their audio and flags the ones that came out
+ truncated.
+- Transcript coverage — how much of a recording's runtime the transcript
+ actually spans — is tracked per video, because a transcript that covers 40% of
+ a video will otherwise look exactly like a complete one.
+
+## Building and publishing
+
+A build has two halves. First the index: transcripts are read out of the corpus
+and written as paginated JSON. Then the compose-and-build step turns that into a
+static site for each configured site.
+
+Builds are incremental. Channels whose contents haven't changed are skipped, so a
+rebuild after one new video does not re-process the whole archive.
+
+Publishing is a separate action from building, and both can be triggered from the
+editor. See [Deploy to Cloudflare](/docs/deploy-cloudflare/) for hosting, and
+[Building several sites at once](/docs/deploy-docker/) if you run more than one.
+
+## Sites, groups, and one corpus
+
+Channels live in a single shared pool. A **site** is a selection of them with its
+own title, description, accent colour and domain. One corpus can therefore
+publish several public archives without any data being duplicated — and a channel
+can appear on more than one.
+
+Within a site, channels can be arranged into named groups, which is what drives
+the navigation on the published pages.
+
+## When a video disappears
+
+The archive periodically re-checks whether recordings still exist where they came
+from, and records the answer per video: available, unlisted, private,
+members-only, or deleted. That state is visible in the published site and can be
+filtered on.
+
+One caveat that matters, and that the software is careful about: a recording
+counts as *available* until something re-checks it and finds otherwise. Nobody
+re-checks tens of thousands of videos continuously. "Available" means **not known
+to be gone**, never "confirmed still there". Only the recordings marked as gone
+are evidence of anything.
+
+Audio for recordings that have disappeared upstream is protected from cleanup, on
+the reasoning that a local copy of something no longer available anywhere is the
+one thing you cannot re-fetch.
+
+Next: [Deploy to Cloudflare](/docs/deploy-cloudflare/).
diff --git a/homepage/content/docs/what-is-archilyzer.md b/homepage/content/docs/what-is-archilyzer.md
@@ -0,0 +1,70 @@
+# What Archilyzer is
+
+Archilyzer takes a channel's back catalogue, downloads it, transcribes it on your
+own hardware, and builds a static website you host yourself. The result is a
+permanent, searchable record of what someone said on video — searchable to the
+second, and still there after the original comes down.
+
+It is a program you run, not a service you sign up for. There is no account, no
+API key, no server of ours in the path, and nothing phones home.
+
+## The shape of it
+
+Three programs share one library and one pile of data.
+
+- **The editor** is a local admin app. You add channels, watch the download and
+ transcription queues, and press the button that builds and deploys a site. It
+ runs on your machine and is never exposed to the public.
+- **The corpus** is a directory on disk — one folder per channel, holding the
+ downloaded media, the transcripts, and a search index. It is deliberately kept
+ outside the code: a fresh copy of the software has no data, and updating the
+ software never touches your archive.
+- **The export** is the public artefact: a static site, pre-rendered to plain
+ HTML and JSON. No database, no runtime, no server-side code. It can be hosted
+ on anything that serves files.
+
+A fourth piece, the **project site** you are reading, and a fifth, the **MCP
+server** described in [AI and MCP](/docs/ai-and-mcp/), round out the workspace.
+
+## How a video becomes a searchable page
+
+1. **List.** Ask the platform what a channel has published. Archilyzer keeps
+ this list separately from what it has already fetched, so it always knows the
+ difference between "new" and "already have it".
+2. **Fetch.** Download what is missing. For channels where the platform already
+ publishes captions, those are taken directly. For everything else, only the
+ audio is fetched.
+3. **Transcribe.** Audio is transcribed locally, by whisper.cpp or one of the
+ other supported backends. This is the slow step and the one that wants a GPU;
+ nothing is sent to a third-party transcription service.
+4. **Index.** Transcripts are cut into pages and written as a paginated JSON
+ index, so a browser can search a corpus of tens of thousands of videos
+ without downloading it.
+5. **Compose and publish.** A static site is built from that index and deployed.
+
+Steps 1–3 can run unattended on a schedule. See
+[Running an archive](/docs/operate/).
+
+## What you get at the end
+
+- **Search across every transcript**, with results that jump to the exact second
+ of the recording they came from.
+- **A player** that follows the transcript, and a transcript that follows the
+ player.
+- **Multiple sites from one corpus.** Channels are grouped into sites, so one
+ archive can publish several public faces without duplicating the data.
+- **Bulk downloads** of the transcripts, for anyone who wants the raw material.
+- **A machine-readable index** at `/corpus.json`, so an LLM or a script can
+ navigate the whole archive without scraping it.
+
+## What it is not
+
+- Not a hosted service. You install it, you run it, you pay for your own hosting.
+- Not a public repository. The source ships as a
+ [dated snapshot](/downloads/) — there is no `git clone` and nothing to pull.
+- Not a downloader you point at one video. It is built around back catalogues:
+ thousands of recordings, kept current.
+- Not automatic transcription in the cloud. The transcribing happens on your
+ machine, at your machine's speed.
+
+Next: [Install](/docs/install/).
diff --git a/homepage/e2e/docs.spec.ts b/homepage/e2e/docs.spec.ts
@@ -0,0 +1,67 @@
+import { test, expect } from "@playwright/test";
+
+// The docs subsystem: the index, one rendered page, and the two renderer traps
+// that ruled out reusing common/components/Markdown.tsx.
+
+test("the index links every shipped page, and each one resolves", async ({
+ page,
+}) => {
+ await page.goto("/docs/");
+ const links = page.locator('a[href^="/docs/"]');
+ const count = await links.count();
+ expect(count).toBeGreaterThan(0);
+
+ // The index renders from the same existence-filtered list that generates the
+ // routes, so a link here that 404s means those two have come apart.
+ const hrefs = new Set<string>();
+ for (let i = 0; i < count; i++) {
+ const href = await links.nth(i).getAttribute("href");
+ if (href && href !== "/docs/") hrefs.add(href);
+ }
+ expect(hrefs.size).toBeGreaterThan(0);
+ for (const href of hrefs) {
+ const res = await page.request.get(href);
+ expect(res.status(), href).toBe(200);
+ }
+});
+
+test("a doc page renders prose, headings and anchors", async ({ page }) => {
+ await page.goto("/docs/install/");
+ await expect(page.getByRole("heading", { level: 1 })).toContainText("Install");
+
+ // A paragraph rendering as a paragraph guards the raw-HTML behaviour:
+ // markdown-to-jsx does NOT escape raw HTML here, so stray markup in a content
+ // file would come through as elements instead of text.
+ const paragraphs = page.locator(".doc-measure p");
+ await expect(paragraphs.first()).toBeVisible();
+ expect((await paragraphs.first().innerText()).length).toBeGreaterThan(20);
+
+ // Section headings get an id, so they can be linked to.
+ const h2 = page.locator(".doc-measure h2[id]").first();
+ await expect(h2).toBeVisible();
+ expect(await h2.getAttribute("id")).toBeTruthy();
+});
+
+test("an internal cross-link stays in the tab; an external one doesn't", async ({
+ page,
+}) => {
+ await page.goto("/docs/install/");
+
+ // THE TRAP: the shared Markdown component hard-codes target="_blank" on every
+ // anchor. Rendering docs through it would fling a reader into a new tab for
+ // following a cross-reference between two pages of the same document.
+ const internal = page.locator('.doc-measure a[href^="/docs/"]').first();
+ await expect(internal).toBeVisible();
+ await expect(internal).not.toHaveAttribute("target", "_blank");
+
+ const external = page.locator('.doc-measure a[href^="http"]').first();
+ if (await external.count()) {
+ await expect(external).toHaveAttribute("target", "_blank");
+ await expect(external).toHaveAttribute("rel", /noopener/);
+ }
+});
+
+test("an unknown doc slug is a 404, not a crash", async ({ page }) => {
+ const res = await page.request.get("/docs/not-a-real-page/");
+ expect(res.status()).toBe(404);
+});
diff --git a/homepage/e2e/downloads.spec.ts b/homepage/e2e/downloads.spec.ts
@@ -0,0 +1,57 @@
+import { test, expect } from "@playwright/test";
+
+// The downloads page has to be correct in BOTH states — with a snapshot
+// attached and without — because the tarball is a gitignored build artefact
+// that a fresh checkout does not have. The rule it must never break: no link
+// to a file that isn't there.
+
+test("says plainly that this is a snapshot, not a repository", async ({
+ page,
+}) => {
+ await page.goto("/downloads/");
+ await expect(page.getByRole("heading", { level: 1 })).toContainText(
+ "Downloads",
+ );
+ await expect(
+ page.getByText(/this is a snapshot, not a repository/i),
+ ).toBeVisible();
+ await expect(page.getByText(/nothing to/i)).toBeVisible();
+});
+
+test("the download link and its facts agree, or there is no link at all", async ({
+ page,
+}) => {
+ await page.goto("/downloads/");
+ const link = page.getByRole("link", { name: /archilyzer-source\.tar\.gz/i });
+
+ if ((await link.count()) === 0) {
+ // No snapshot in this build: explain, don't link.
+ await expect(
+ page.getByText(/no source snapshot attached/i),
+ ).toBeVisible();
+ return;
+ }
+
+ const href = await link.getAttribute("href");
+ expect(href).toBe("/downloads/archilyzer-source.tar.gz");
+
+ // The advertised file must actually be served. A stale sidecar beside a
+ // missing tarball is the exact failure the loader is written to prevent.
+ const res = await page.request.get(href!);
+ expect(res.status()).toBe(200);
+
+ // The sidecar's facts are rendered, not hardcoded: a checksum and a commit.
+ await expect(page.getByText("SHA-256")).toBeVisible();
+ const sha = await page
+ .locator("text=/^[0-9a-f]{64}$/")
+ .first()
+ .textContent();
+ expect(sha?.trim()).toMatch(/^[0-9a-f]{64}$/);
+});
+
+test("names what the tarball does not contain", async ({ page }) => {
+ await page.goto("/downloads/");
+ await expect(page.getByText(/The archive itself/)).toBeVisible();
+ await expect(page.getByText(/Any operator’s configuration/)).toBeVisible();
+ await expect(page.getByRole("heading", { name: "License" })).toBeVisible();
+});
diff --git a/homepage/e2e/landing.spec.ts b/homepage/e2e/landing.spec.ts
@@ -1,105 +0,0 @@
-import { test, expect, type Page } from "@playwright/test";
-
-// The landing's cross-site chart: the three site-based presets, plus the
-// power-user Customize toolbox (chart types, scale, exclude-outlier, indexed)
-// and its guard rails. Asserts against the live recharts SVG.
-
-const pressedPreset = (page: Page) =>
- page.locator('[aria-label="Chart preset"] button[aria-pressed="true"]');
-
-const yTicks = (page: Page) =>
- page
- .locator(".recharts-yAxis .recharts-cartesian-axis-tick-value")
- .allTextContents();
-
-const fullLines = (page: Page) =>
- // Lines that actually span the data (guards against the scale-undefined glitch
- // that collapsed lines to 1-2 points).
- page.evaluate(
- () =>
- [...document.querySelectorAll(".recharts-line-curve")].filter(
- (p) => (p.getAttribute("d") || "").length > 30,
- ).length,
- );
-
-test.beforeEach(async ({ page }) => {
- await page.goto("/");
- await page.waitForSelector(".recharts-surface");
-});
-
-test("Recent is the default: three full site lines on a linear axis", async ({
- page,
-}) => {
- await expect(pressedPreset(page)).toHaveText("Recent");
- await expect(page.locator(".recharts-line-curve")).toHaveCount(3);
- expect(await fullLines(page)).toBe(3);
- // Linear (not symlog): no "1K"/"10K" decade ticks.
- const ticks = (await yTicks(page)).map((t) => t.trim());
- expect(ticks).not.toContain("10K");
-});
-
-test("all three presets draw site lines; All-time is symlog", async ({ page }) => {
- for (const name of ["Recent", "Cumulative", "All-time"]) {
- await page.getByRole("button", { name, exact: true }).click();
- await expect(pressedPreset(page)).toHaveText(name);
- await expect(page.locator(".recharts-line-curve")).toHaveCount(3);
- expect(await fullLines(page)).toBe(3);
- }
- // All-time is the tamed symlog view.
- const ticks = (await yTicks(page)).map((t) => t.trim());
- for (const t of ["0", "10", "100", "1K", "10K"]) expect(ticks).toContain(t);
-});
-
-test("Customize still reaches the full toolbox (ranked / small / channel)", async ({
- page,
-}) => {
- await page.getByRole("button", { name: "Customize" }).click();
- await page.getByRole("button", { name: "By channel", exact: true }).click();
- await page.getByRole("button", { name: "Ranked", exact: true }).click();
- await expect(page.locator(".recharts-bar-rectangle").first()).toBeVisible();
- await page.getByRole("button", { name: "Small", exact: true }).click();
- await expect(page.locator(".recharts-responsive-container").first()).toBeVisible();
-});
-
-test("Customize greys invalid option combinations", async ({ page }) => {
- await page.getByRole("button", { name: "Customize" }).click();
- // Default Recent is a line chart: Share invalid, Symlog valid.
- await expect(page.getByRole("button", { name: "Share", exact: true })).toBeDisabled();
- await expect(page.getByRole("button", { name: "Symlog", exact: true })).toBeEnabled();
- // Switch to a stacked Area: Symlog invalid, Share valid.
- await page.getByRole("button", { name: "Area", exact: true }).click();
- await expect(page.getByRole("button", { name: "Symlog", exact: true })).toBeDisabled();
- await expect(page.getByRole("button", { name: "Share", exact: true })).toBeEnabled();
-});
-
-test("exclude-outlier drops a site line", async ({ page }) => {
- await page.getByRole("button", { name: "Customize" }).click();
- await expect(page.locator(".recharts-line-curve")).toHaveCount(3);
- await page.getByRole("button", { name: /^Exclude / }).click();
- await expect(page.locator(".recharts-line-curve")).toHaveCount(2);
-});
-
-test("Indexed rebases every series to a shared baseline", async ({ page }) => {
- await page.getByRole("button", { name: "Customize" }).click();
- await page.getByRole("button", { name: "Indexed", exact: true }).click();
- await expect(pressedPreset(page)).toHaveCount(0); // custom
- const ys = await page.evaluate(() =>
- [...document.querySelectorAll(".recharts-line-curve")]
- .map((p) => {
- const m = (p.getAttribute("d") || "").match(
- /^M\s*[-\d.]+[, ]\s*([-\d.]+)/,
- );
- return m ? parseFloat(m[1]) : null;
- })
- .filter((v): v is number => v != null),
- );
- expect(ys.length).toBeGreaterThanOrEqual(2);
- expect(Math.max(...ys) - Math.min(...ys)).toBeLessThan(2);
-});
-
-test("mobile: small multiples stack to one column", async ({ page }) => {
- await page.setViewportSize({ width: 375, height: 800 });
- await page.getByRole("button", { name: "Customize" }).click();
- await page.getByRole("button", { name: "Small", exact: true }).click();
- await expect(page.locator(".recharts-responsive-container")).toHaveCount(3);
-});
diff --git a/homepage/e2e/marketing.spec.ts b/homepage/e2e/marketing.spec.ts
@@ -0,0 +1,76 @@
+import { test, expect } from "@playwright/test";
+
+// The home page's job is to say what Archilyzer is and offer the download.
+// Everything asserted here is data-INDEPENDENT by construction: no count, no
+// site name, no headline number. The numeric bands are gated on corpus data, so
+// pinning one would make this suite fail on a source-only build — which is
+// exactly the build this page has to work on.
+
+test.beforeEach(async ({ page }) => {
+ await page.goto("/");
+});
+
+test("says what it is and offers the source", async ({ page }) => {
+ await expect(page.getByRole("heading", { level: 1 })).toContainText(
+ /searchable to the second/i,
+ );
+ await expect(
+ page.getByRole("link", { name: /download the source/i }),
+ ).toHaveAttribute("href", "/downloads/");
+ await expect(
+ page.getByRole("link", { name: /read the setup guide/i }),
+ ).toHaveAttribute("href", "/docs/install/");
+});
+
+test("the trust block states what it isn't", async ({ page }) => {
+ // The cheapest, highest-value block on the page — every claim in it is
+ // simply true, and each one pre-empts a wrong assumption.
+ for (const claim of [
+ "Not a hosted service",
+ "No public repository",
+ "You supply the tools",
+ "Transcription is slow",
+ ]) {
+ await expect(page.getByText(claim, { exact: true })).toBeVisible();
+ }
+});
+
+test("every nav destination resolves", async ({ page }) => {
+ // The nav must never point at a 404, including on a build with no corpus
+ // data — which is why /stats/ always exists and degrades in place.
+ for (const [label, path] of [
+ ["Docs", "/docs/"],
+ ["Downloads", "/downloads/"],
+ ["Stats", "/stats/"],
+ ["Changelog", "/changelog/"],
+ ]) {
+ const res = await page.request.get(path);
+ expect(res.status(), `${label} → ${path}`).toBe(200);
+ }
+});
+
+test("the rail never invents a recording", async ({ page }) => {
+ // Absent data must produce an empty state, not filler. An archiving tool
+ // that fabricates transcript rows on its own marketing page has undercut
+ // itself before a visitor has read a sentence.
+ const rail = page.getByText("Recent acquisitions");
+ const empty = page.getByText(/shipped without corpus data/i);
+ const hasRail = await rail.isVisible().catch(() => false);
+ if (!hasRail) {
+ await expect(empty).toBeVisible();
+ return;
+ }
+ // With data: every row carries a real date in the log's own format.
+ const times = page.locator("time[datetime]");
+ await expect(times.first()).toBeVisible();
+ for (const text of await times.allTextContents()) {
+ expect(text.trim()).toMatch(/^\d{4}-\d{2}-\d{2}$/);
+ }
+});
+
+test("branded 404 for an unknown path", async ({ page }) => {
+ await page.goto("/no-such-page/");
+ await expect(page.getByRole("heading", { name: /no such page/i })).toBeVisible();
+ // Branded, not bare: the site chrome is still there to get you somewhere.
+ await expect(page.getByRole("link", { name: /Archilyzer home/i })).toBeVisible();
+});
diff --git a/homepage/e2e/stats.spec.ts b/homepage/e2e/stats.spec.ts
@@ -0,0 +1,129 @@
+import { test, expect, type Page } from "@playwright/test";
+
+// The /stats/ dashboard's cross-site chart: the three site-based presets, plus
+// the power-user Customize toolbox (chart types, scale, exclude-outlier,
+// indexed) and its guard rails. Asserts against the live recharts SVG.
+//
+// This file HARD-REQUIRES corpus data — it asserts an exact site-line count, so
+// there is nothing meaningful to check on a `build:nodata` tree. The whole file
+// skips when /stats/ served its no-data panel instead of the dashboard, which
+// keeps a source-only checkout green without weakening the assertions for a
+// configured one. (The site-line count follows the corpus: it is read from the
+// rendered legend rather than pinned at 3, because the operator's public-site
+// list grows.)
+
+const pressedPreset = (page: Page) =>
+ page.locator('[aria-label="Chart preset"] button[aria-pressed="true"]');
+
+const yTicks = (page: Page) =>
+ page
+ .locator(".recharts-yAxis .recharts-cartesian-axis-tick-value")
+ .allTextContents();
+
+const fullLines = (page: Page) =>
+ // Lines that actually span the data (guards against the scale-undefined glitch
+ // that collapsed lines to 1-2 points).
+ page.evaluate(
+ () =>
+ [...document.querySelectorAll(".recharts-line-curve")].filter(
+ (p) => (p.getAttribute("d") || "").length > 30,
+ ).length,
+ );
+
+// How many public sites this build's summary actually has. The suite pins its
+// counts to THIS rather than a literal, because the operator's public-site list
+// grows (it went 3 → 4 when Bonnellyzer shipped) and a literal would turn a new
+// archive into a red suite.
+let siteCount = 0;
+
+test.beforeEach(async ({ page }) => {
+ await page.goto("/stats/");
+ // A build with no corpus data serves the honest no-data panel here instead of
+ // the dashboard. Nothing below is meaningful without numbers.
+ const noData = await page
+ .getByText("This build shipped without dashboard data")
+ .isVisible()
+ .catch(() => false);
+ test.skip(noData, "no corpus data in this build — /stats/ has no dashboard");
+ await page.waitForSelector(".recharts-surface");
+ siteCount = await page.locator(".recharts-line-curve").count();
+ expect(siteCount, "summary should carry at least two public sites").toBeGreaterThanOrEqual(2);
+});
+
+test("Recent is the default: every site line is full, on a linear axis", async ({
+ page,
+}) => {
+ await expect(pressedPreset(page)).toHaveText("Recent");
+ expect(await fullLines(page)).toBe(siteCount);
+ // Linear (not symlog): no "1K"/"10K" decade ticks.
+ const ticks = (await yTicks(page)).map((t) => t.trim());
+ expect(ticks).not.toContain("10K");
+});
+
+test("all three presets draw site lines; All-time is symlog", async ({ page }) => {
+ for (const name of ["Recent", "Cumulative", "All-time"]) {
+ await page.getByRole("button", { name, exact: true }).click();
+ await expect(pressedPreset(page)).toHaveText(name);
+ await expect(page.locator(".recharts-line-curve")).toHaveCount(siteCount);
+ expect(await fullLines(page)).toBe(siteCount);
+ }
+ // All-time is the tamed symlog view.
+ const ticks = (await yTicks(page)).map((t) => t.trim());
+ for (const t of ["0", "10", "100", "1K", "10K"]) expect(ticks).toContain(t);
+});
+
+test("Customize still reaches the full toolbox (ranked / small / channel)", async ({
+ page,
+}) => {
+ await page.getByRole("button", { name: "Customize" }).click();
+ await page.getByRole("button", { name: "By channel", exact: true }).click();
+ await page.getByRole("button", { name: "Ranked", exact: true }).click();
+ await expect(page.locator(".recharts-bar-rectangle").first()).toBeVisible();
+ await page.getByRole("button", { name: "Small", exact: true }).click();
+ await expect(page.locator(".recharts-responsive-container").first()).toBeVisible();
+});
+
+test("Customize greys invalid option combinations", async ({ page }) => {
+ await page.getByRole("button", { name: "Customize" }).click();
+ // Default Recent is a line chart: Share invalid, Symlog valid.
+ await expect(page.getByRole("button", { name: "Share", exact: true })).toBeDisabled();
+ await expect(page.getByRole("button", { name: "Symlog", exact: true })).toBeEnabled();
+ // Switch to a stacked Area: Symlog invalid, Share valid.
+ await page.getByRole("button", { name: "Area", exact: true }).click();
+ await expect(page.getByRole("button", { name: "Symlog", exact: true })).toBeDisabled();
+ await expect(page.getByRole("button", { name: "Share", exact: true })).toBeEnabled();
+});
+
+test("exclude-outlier drops a site line", async ({ page }) => {
+ await page.getByRole("button", { name: "Customize" }).click();
+ await expect(page.locator(".recharts-line-curve")).toHaveCount(siteCount);
+ await page.getByRole("button", { name: /^Exclude / }).click();
+ await expect(page.locator(".recharts-line-curve")).toHaveCount(siteCount - 1);
+});
+
+test("Indexed rebases every series to a shared baseline", async ({ page }) => {
+ await page.getByRole("button", { name: "Customize" }).click();
+ await page.getByRole("button", { name: "Indexed", exact: true }).click();
+ await expect(pressedPreset(page)).toHaveCount(0); // custom
+ const ys = await page.evaluate(() =>
+ [...document.querySelectorAll(".recharts-line-curve")]
+ .map((p) => {
+ const m = (p.getAttribute("d") || "").match(
+ /^M\s*[-\d.]+[, ]\s*([-\d.]+)/,
+ );
+ return m ? parseFloat(m[1]) : null;
+ })
+ .filter((v): v is number => v != null),
+ );
+ expect(ys.length).toBeGreaterThanOrEqual(2);
+ expect(Math.max(...ys) - Math.min(...ys)).toBeLessThan(2);
+});
+
+test("mobile: small multiples stack to one column", async ({ page }) => {
+ await page.setViewportSize({ width: 375, height: 800 });
+ await page.getByRole("button", { name: "Customize" }).click();
+ await page.getByRole("button", { name: "Small", exact: true }).click();
+ await expect(page.locator(".recharts-responsive-container")).toHaveCount(
+ siteCount,
+ );
+});
diff --git a/homepage/e2e/theme.spec.ts b/homepage/e2e/theme.spec.ts
@@ -1,9 +1,13 @@
import { test, expect, type Page } from "@playwright/test";
-// Phase 0 foundation: the shared theme system (common/styles/tokens.css +
-// ThemeScript + ThemeProvider + ThemeToggle). The homepage commits to the
-// "archive" family and defaults to ink (dark); the mode toggle must persist
-// across reloads and be applied before hydration (no flash of the wrong theme).
+// The shared theme system (common/styles/tokens.css + ThemeScript +
+// ThemeProvider + ThemeToggle) as this site uses it. The project site commits to
+// the "archilyzer" family — the instrument face — and defaults to dark. The mode
+// toggle must persist across reloads and be applied before hydration (no flash
+// of the wrong theme).
+//
+// This used to assert the "archive" family: the site wore the same warm-brass
+// costume as the archives it builds, back when it was a shelf of them.
async function htmlState(page: Page) {
return page.evaluate(() => ({
@@ -13,14 +17,14 @@ async function htmlState(page: Page) {
}));
}
-test("archive ink default; mode toggle persists with no FOUC", async ({
+test("archilyzer dark default; mode toggle persists with no FOUC", async ({
page,
}) => {
await page.goto("/");
- // Default: archive family, ink (dark) mode.
+ // Default: the instrument family, dark.
let s = await htmlState(page);
- expect(s.theme).toBe("archive");
+ expect(s.theme).toBe("archilyzer");
expect(s.dark).toBe(true);
const toggle = page.getByRole("button", { name: /switch to/i });
@@ -36,7 +40,7 @@ test("archive ink default; mode toggle persists with no FOUC", async ({
s = await htmlState(page);
expect(s.mode).toBe("light");
expect(s.dark).toBe(false);
- expect(s.theme).toBe("archive");
+ expect(s.theme).toBe("archilyzer");
// Reload: the pre-paint inline script must re-apply the persisted mode on the
// very first commit, before React hydrates.
@@ -48,3 +52,94 @@ test("archive ink default; mode toggle persists with no FOUC", async ({
expect(onCommit.mode).toBe("light");
expect(onCommit.dark).toBe(false);
});
+
+test("the family declares a complete palette in BOTH modes", async ({
+ page,
+}) => {
+ // The failure this guards is silent: a family that omits a token inherits the
+ // BASE family's value, so a half-declared palette looks merely "a bit off"
+ // rather than broken — and only in the mode nobody checked. Light mode ships
+ // here because the theme toggle does.
+ const TOKENS = [
+ "--background",
+ "--foreground",
+ "--card",
+ "--card-foreground",
+ "--popover",
+ "--popover-foreground",
+ "--primary",
+ "--primary-foreground",
+ "--secondary",
+ "--secondary-foreground",
+ "--muted",
+ "--muted-foreground",
+ "--accent",
+ "--accent-foreground",
+ "--destructive",
+ "--destructive-foreground",
+ "--destructive-soft",
+ "--border",
+ "--border-strong",
+ "--input",
+ "--ring",
+ "--surface",
+ "--faint",
+ "--panel",
+ "--panel-2",
+ "--success",
+ "--success-foreground",
+ "--success-soft",
+ "--warning",
+ "--warning-foreground",
+ "--warning-soft",
+ "--info",
+ "--info-foreground",
+ "--info-soft",
+ "--brand",
+ "--brand-strong",
+ "--brand-soft",
+ "--brand-ink",
+ "--state-gone",
+ "--state-gone-soft",
+ "--chart-1",
+ "--chart-2",
+ "--chart-3",
+ "--chart-4",
+ "--chart-5",
+ "--chart-surface",
+ "--chart-grid",
+ "--chart-axis",
+ "--chart-tooltip-bg",
+ "--radius",
+ ];
+
+ await page.goto("/");
+
+ const readAll = (names: string[]) =>
+ page.evaluate((tokens) => {
+ const cs = getComputedStyle(document.documentElement);
+ return Object.fromEntries(
+ tokens.map((t) => [t, cs.getPropertyValue(t).trim()]),
+ );
+ }, names);
+
+ const setMode = (dark: boolean) =>
+ page.evaluate((d) => {
+ localStorage.setItem("ytdlp-tb:mode", d ? "dark" : "light");
+ document.documentElement.classList.toggle("dark", d);
+ }, dark);
+
+ await setMode(true);
+ const dark = await readAll(TOKENS);
+ await setMode(false);
+ const light = await readAll(TOKENS);
+
+ for (const t of TOKENS) {
+ expect(dark[t], `${t} must be set in dark mode`).not.toBe("");
+ expect(light[t], `${t} must be set in light mode`).not.toBe("");
+ }
+ // …and the two modes must actually differ, or "light mode" is a label on the
+ // dark palette.
+ expect(light["--background"]).not.toBe(dark["--background"]);
+ expect(light["--foreground"]).not.toBe(dark["--foreground"]);
+});
diff --git a/homepage/next.config.ts b/homepage/next.config.ts
@@ -1,6 +1,27 @@
+import path from "node:path";
import type { NextConfig } from "next";
+// THE TWIN BUILD SCRIPTS (package.json). `build` and `build:nodata` run the
+// exact same command — `next build`. The difference is invisible in the command
+// and lives entirely in pnpm's lifecycle hook, which matches on the script
+// NAME: `prebuild` fires before `build` and pulls in `build:data` (the
+// multi-GB LMDB index + a whole-pool compose), and there is no `prebuild:nodata`
+// hook, so `build:nodata` renders the site from source alone.
+//
+// That is the point of the split: this is the project's own marketing + docs
+// site, and someone who has only unpacked the source tarball must be able to
+// build it. Routes that need corpus data degrade honestly (see
+// app/lib/summary.ts) rather than rendering zeroes. JSON can't hold comments,
+// hence this note here and in homepage/CHANGELOG.md.
const nextConfig: NextConfig = {
+ // Pin the workspace root. Turbopack infers it by walking up for a lockfile
+ // and taking the outermost one, so an unrelated pnpm-lock.yaml anywhere above
+ // the checkout (a stray one in $HOME is enough) silently relocates the root —
+ // after which `@import "../../common/styles/tokens.css"` in app/globals.css
+ // resolves outside the project and `next dev` fails to boot at all. The
+ // export app documents this as load-bearing for exactly this construct, which
+ // this app also uses.
+ turbopack: { root: path.join(__dirname, "..") },
output: "export",
trailingSlash: true,
images: { unoptimized: true },
diff --git a/homepage/package.json b/homepage/package.json
@@ -7,13 +7,15 @@
"dev": "next dev --port ${HOMEPAGE_DEV_PORT:-3030}",
"build:index": "NODE_OPTIONS=--max-old-space-size=8192 tsx ../common/bin/build-index.ts",
"compose": "tsx ../common/bin/compose-homepage.ts",
- "prebuild": "pnpm run build:index",
- "build": "pnpm run compose && next build",
+ "build:data": "pnpm run build:index && pnpm run compose",
+ "prebuild": "pnpm run build:data",
+ "build": "next build",
+ "build:nodata": "next build",
"start": "serve out -l ${HOMEPAGE_PORT:-3031}",
"lint": "eslint",
- "e2e": "playwright test",
+ "e2e": "node ../scripts/queue-lock.mjs --ports HOMEPAGE_E2E_PORT:3040 -- playwright test",
"e2e:ui": "playwright test --ui",
- "deploy": "pnpm dlx wrangler pages deploy out"
+ "deploy": "pnpm dlx wrangler pages deploy out --project-name archilyzer --branch main"
},
"dependencies": {
"@tanstack/react-query": "^5.99.1",
diff --git a/homepage/playwright.config.ts b/homepage/playwright.config.ts
@@ -1,9 +1,14 @@
import { defineConfig, devices } from "@playwright/test";
// Homepage e2e. Runs against `next dev` (default mode) so it reflects uncommitted
-// source — the landing reads its build-time summary from public/ on disk, which
-// is present in dev. Kill any stale dev server on the port between runs.
-const PORT = Number(process.env.PORT ?? 3032);
+// source — /stats/ reads its build-time summary from public/ on disk, which is
+// present in dev. Kill any stale dev server on the port between runs.
+//
+// The port comes from HOMEPAGE_E2E_PORT, the block scripts/worktree.mjs already
+// allocates per worktree (3040 on main). It used to read `PORT`, which is the
+// EDITOR's port variable — so a worktree run pointed this suite at whatever the
+// editor had been given, and the per-worktree offset never applied here.
+const PORT = Number(process.env.HOMEPAGE_E2E_PORT ?? 3040);
const baseURL = `http://localhost:${PORT}`;
export default defineConfig({
diff --git a/homepage/public/_headers b/homepage/public/_headers
@@ -0,0 +1,8 @@
+# Cloudflare Pages header rules for the project site.
+#
+# The source snapshot lives at a STABLE filename (archilyzer-source.tar.gz), so
+# a new snapshot replaces the bytes behind a URL that has already been cached.
+# A short, revalidating lifetime keeps the edge from serving yesterday's build
+# for hours while still absorbing a burst of downloads.
+/downloads/*
+ Cache-Control: public, max-age=300, must-revalidate
diff --git a/homepage/public/icons/apple-touch-icon.png b/homepage/public/icons/apple-touch-icon.png
Binary files differ.
diff --git a/homepage/public/icons/icon-192.png b/homepage/public/icons/icon-192.png
Binary files differ.
diff --git a/homepage/public/icons/icon-512.png b/homepage/public/icons/icon-512.png
Binary files differ.
diff --git a/homepage/public/icons/icon.svg b/homepage/public/icons/icon.svg
@@ -0,0 +1,10 @@
+<svg xmlns="http://www.w3.org/2000/svg" width="512" height="512" viewBox="0 0 512 512" role="img" aria-label="Archilyzer">
+ <!-- The tool's mark, not an archive's. The export sites wear a brass diamond
+ on warm black; this is a bone play head on graphite — the instrument that
+ builds them, in the instrument family's palette (see tokens.css,
+ [data-theme="archilyzer"]). The bar beneath is the transcript: the
+ recording resolved into a line of text. -->
+ <rect width="512" height="512" rx="96" fill="#151b20"/>
+ <polygon points="188,132 372,236 188,340" fill="#e7edf1"/>
+ <rect x="140" y="384" width="232" height="20" rx="4" fill="#8496a2"/>
+</svg>
diff --git a/homepage/public/icons/maskable-512.png b/homepage/public/icons/maskable-512.png
Binary files differ.
diff --git a/homepage/public/icons/maskable.svg b/homepage/public/icons/maskable.svg
@@ -0,0 +1,7 @@
+<svg xmlns="http://www.w3.org/2000/svg" width="512" height="512" viewBox="0 0 512 512" role="img" aria-label="Archilyzer">
+ <!-- Maskable variant: same mark, pulled into the safe zone (the inner 80%
+ circle) because a launcher may crop this to any shape. -->
+ <rect width="512" height="512" fill="#151b20"/>
+ <polygon points="206,164 350,246 206,328" fill="#e7edf1"/>
+ <rect x="176" y="356" width="160" height="16" rx="4" fill="#8496a2"/>
+</svg>
diff --git a/mcp/README.md b/mcp/README.md
@@ -454,7 +454,7 @@ means "this video's URL" in the output.
```
list_channels # the default corpus
list_channels source="remote:https://rekietalyzer.pages.dev"
-search_transcripts query="k cups" source="hub:https://archilyzer.com#jeralyzer,rekietalyzer"
+search_transcripts query="k cups" source="hub:https://archilyzer.pages.dev#jeralyzer,rekietalyzer"
resolve_source source="jeralyzer" # → remote:https://jeralyzer.pages.dev
```
@@ -505,7 +505,7 @@ the TypeScript path alias resolves):
pnpm --filter yt-dlp-transcript-mcp exec tsx src/index.ts --remote https://rekietalyzer.pages.dev
# a hub, federating every member site
-pnpm --filter yt-dlp-transcript-mcp exec tsx src/index.ts --hub https://archilyzer.com
+pnpm --filter yt-dlp-transcript-mcp exec tsx src/index.ts --hub https://archilyzer.pages.dev
# local shards on disk
pnpm --filter yt-dlp-transcript-mcp exec tsx src/index.ts --local ../export/public
diff --git a/package.json b/package.json
@@ -2,12 +2,14 @@
"name": "yt-dlp-transcript-browser-monorepo",
"version": "0.1.0",
"private": true,
+ "license": "MIT",
"type": "module",
"scripts": {
"build:index": "pnpm --filter yt-dlp-transcript-common exec tsx bin/build-index.ts",
"sync:tick": "pnpm --filter yt-dlp-transcript-common exec tsx bin/sync-tick.ts",
"build:export": "pnpm --filter export run build",
"build:homepage": "pnpm --filter homepage run build",
+ "build:homepage:nodata": "pnpm --filter homepage run build:nodata",
"build": "pnpm --filter export run build",
"start:export": "node scripts/worktree.mjs run -- pnpm --filter export run start",
"start:homepage": "node scripts/worktree.mjs run -- pnpm --filter homepage run start",