Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 4c8e658fbbea820d68e39447ecc71c01800a7e38
parent d76d2719c92e74f3e5661a9a0a446bc372369f2c
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Thu, 24 Sep 2026 16:24:03 -0400

plans: phase3-files-numbers tool, the slice 4b record, changelog

plans/tools/phase3-files-numbers.ts: one process, never writes the corpus.
Prints each site.json parse + what writeSite writes back, each config.json
parse + md5 of the strict write-back, and for a sorted sample of up to 300
dirs per sidecar filename the md5 of load() and of write(load()), plus an
unknown-key report; FREEZE_TO copies exactly those inputs so main and the
branch run over the same bytes. main vs branch: diff empty, 3,832 lines;
unknown keys: none.

The "Slice 4b, as shipped" record in plans/one-core-phase-3.md: commits,
the build:index plan correction (it writes $TRANSCRIPTS_DIR/index.mdb, so it
was timed over a scratch corpus of symlinks), deviations 1-10, behaviour
changes (including the one-time key-reorder diff in transcripts/), the
review fixes, the 14 JSON writers still on the per-pid temp name, the
numbers, build:index timing (1542.74 s loaded / 1699.98 s branch / 1691.98 s
quiet; +0.5 %, not a regression) and every gate. Changelog entry.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Diffstat:
Meditor/CHANGELOG.md | 1+
Mplans/one-core-phase-3.md | 157+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aplans/tools/phase3-files-numbers.ts | 317+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
3 files changed, 475 insertions(+), 0 deletions(-)

diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,6 +1,7 @@ # Changelog ## [Unreleased] +- **`site.json`, each channel's `config.json` and the per-video sidecars now have one schema each, and the two config files have generated key tables.** **`SITE.md`** and **`CHANNEL.md`** (new, repo root) list every key with its default and meaning, generated by `common/bin/file-schemas-docs.ts` and checked by a test. Nothing an operator has configured reads or saves differently: every live `site.json` and `config.json`, and a 1,763-file sample of sidecars, read and write back byte-for-byte as before. **Fixed:** a social-channel fetch no longer undoes Configure-form edits made while it was running (it used to write back the whole config it read when it started). Every change to a channel's config now re-reads the file at the moment it saves and changes only its own fields, so a sync stamping its time and a form save made at the same moment both land. Two writes to the same file from the editor no longer share one temporary file. - **`settings.json` has one schema and one writer, and its key table is generated.** Every key, its default, its clamp and its documentation is now one zod schema (`common/lib/settingsSchema.ts`); `getSettings`/`writeSettings` both parse through it, and every settings form saves through one helper (`editor/app/settings/saveSettings.ts`) that merges only what the form changed. **`SETTINGS.md`** (new, repo root) lists every key with its default and what it does, and `settings.json.example` is now the full default object — both generated by `common/bin/settings-example.ts` and checked by a test, so neither can drift. Nothing an operator has configured reads differently. **Fixed:** adding or editing a storage location on `/storage` no longer erases the record of which location the saved-video store is on (`storage.savedVideosLocationId`). - **Every live panel now polls one endpoint, `/api/view/<name>`, and the eight old addresses still answer.** The change token, the job head, workers, the operations board, the sync console and the widget's three strips were eight separate API routes that each did the same thing; they are one route serving eight named views (`pulse`, `activeJobs`, `workers`, `autoQueueStatus`, `schedulerStatus`, `widgetSync`, `widgetActionable`, `cleanable`), and the editor's own pages poll it there. `/api/pulse`, `/api/jobs/active`, `/api/workers`, `/api/auto-queue/status`, `/api/scheduler/status` and `/api/widget/{sync,actionable,cleanable}` are kept as **rewrites**, not redirects — same method, status, body and query string (`?rev=` included) — so a monitor widget pinned in a browser, or any script polling the old path, keeps working untouched. `/api/widget/presets` is unchanged. An unknown view name is a 404. **One behaviour change you might notice:** the operations board (every 3 s) and the sync console (every 5 s) now send their next poll only after the previous one answers, and abandon a poll that takes longer than 10–15 s — so a slow editor no longer piles requests up behind itself, and a hung request no longer stops the page updating. - **A lane's pause is one key on the lane, and the four old pause fields are gone from `settings.json`.** Holding a lane has been `autoQueue.<lane>.held` since the runner work landed; until now the file also still carried the four flags that used to mean it — `transcriptionsPaused`, `downloadsPaused`, `digest.digestsPaused` and the backwards `backfill.enabled` (where *enabled* meant *not held*) — which were read only when a lane had no `held` yet, to carry an older file's pause across. Every lane now carries its own key, so those four are **deleted**: nothing reads them, no form writes them, and the next settings save drops them from the file. A settings.json that still spells one of them holds nothing with it, so a hand-edited file (or a very old backup restored over a newer one) can no longer resurrect a pause you had lifted, or lift one you had set. "Run the backfill lane" on the diarization page and the Hold/Pause buttons write the one key, as they already did. **UPGRADING: boot once on the release that writes `held` before taking this one.** That release is the one that moved the gate onto the lane and carried the old fields across on read; a single boot of it (any settings save, or just starting the editor and pausing/resuming anything) puts `autoQueue.<lane>.held` in your settings.json, after which **nothing you can see changes here** — the same buttons, the same labels, the same pauses. An install that jumps straight from an older release to this one has no `held` keys at all and **loses its pauses**: transcription, downloads and digests come up running, and the backfill lane comes up held. Re-set them from the dashboard, or add the keys by hand before starting. diff --git a/plans/one-core-phase-3.md b/plans/one-core-phase-3.md @@ -782,3 +782,160 @@ docs) at `5b43dc8d`, **1650** after `81deae69` (+2 docs wiring tests); `test:scr diarization, attribution, scheduler, cadence-ui, storage-locations, channel-storage, workers, worker-remote, parakeet, parakeet-partial, chough, transcription-app-migration, disk-space, channel-priority, settings — **157/157 passed, exit 0, 9.9 min**, first run, nothing re-run. + +## Slice 4b, as shipped — the rest of the file schemas (2026-09-24) + +Branch `one-core/phase-3-s4b` off `main` `1e27f7c3`, eight commits, then `main` (`3241fed2`, +slice 3a) merged in — merge sha and the post-merge gates in the follow-up below. + +| sha | what | +|---|---| +| `03164485` | `lib/jsonFile-server.ts` (+test): `readJsonFile` (+sync twin) → `{ok, value} \| {ok:false, reason: absent\|unreadable\|unparseable}`; `writeJsonAtomic` — unique temp name, writes chained per absolute path; `writeJsonAtomicSync` for the three synchronous stores. The 7 private copies (digest-server, chartsStore, aliasesStore, curatedTagsStore, buildIndex, buildStats, compose-homepage) and the inline tmp JSON writes in site.ts, settings.ts, controller/channels.ts, runYtdlp.ts, channelSnapshot.ts, the 7 single-file sidecar servers and the write halves of savedVideo-/posts-/clipWindow-server folded. `getSettings` reads through `readJsonFileSync`. Every file keeps its bytes | +| `bb7fe82c` | `lib/sidecar-server.ts` (+test): `sidecar(filename, schema, {indent})` → `{filename, path, load, write, remove}`; throws at declaration on a `SUB_FILE_RE` match; `SIDECAR_FILENAMES`; `sidecarField(coerce)`. 8 pairs / 9 files declared, each shape check an exported `coerceX`, the old `loadX`/`writeX`/`xPath` names kept. `SUB_FILE_RE` exported (`videoStatus.ts:88`); `diarization.test.ts` uses it | +| `2bf65626` | `lib/siteSchema.ts` (+test): per-key `settingsField` object over the existing parsers + one object step (defaultGroupId, membership groupIds, re-emit every key); `parseSite`, `siteToDisk`; the Site types, `SITE_ID_RE` and the three parsers moved in; `site.ts` = I/O + resolvers, `export *`. `SITE_FIELD_DOCS`, `SITE_CHANNEL_MEMBERSHIP_FIELD_DOCS`, `RELATED_SITE_GROUP_FIELD_DOCS`; `CHANNEL_GROUP_FIELD_DOCS` in `channelGroups.ts` | +| `e75047f8` | `CHANNEL_CONFIG_COERCIONS` in `channelConfig.ts` (pure) — the one coercion set, extracted verbatim; `parseChannelConfig` composes it zod-free. `lib/channelConfigSchema.ts` (server-only zod, +test) composes the same functions + `stripUndefined`. `CHANNEL_CONFIG_FIELD_DOCS` (the type's comments moved in), `AUDIO_CHECK_FIELD_DOCS`, `DOWNLOAD_FILTER_FIELD_DOCS`, `CHANNEL_CONFIG_KEYS` (from the docs record), `CHANNEL_SYNC_STATE_KEYS` | +| `f6a08bd1` | `controller/channels.ts`: `readChannelConfigFile`, strict throwing `writeChannelConfig`, `patchChannelConfig(paths, slug, patch, {unset})` under `withJsonFileLock`. Repointed: runYtdlp's three stamps (`updateConfigField` deleted), fetchPosts, relocateChannelMedia ×2, storageLocations ×2, renameChannel, socialActions, the channel form (`{unset: CHANNEL_FORM_FIELDS}`), both exclude toggles, the scheduler's bulk cadence save. buildIndex/buildStats readers folded (+tests) | +| `7c03d7c9` | `bin/file-schemas-docs.ts` (+`--check`), `lib/fileSchemaDocs.ts` (+test) → root `SITE.md`, `CHANNEL.md`; settingsDocs' table helpers exported; SETUP.md, AGENTS.md links; FACTS `SUB_FILE_RE` anchor corrected | +| `973e59ee` | review fixes — below | +| (this) | `plans/tools/phase3-files-numbers.ts`, this record, changelog | + +**Plan correction — `build:index` is not read-only.** The plan said to time `pnpm --filter +export build:index` "on the real corpus … read-only over `transcripts/`". It is not: the build +writes its LMDB to `paths.lmdbPath = $TRANSCRIPTS_DIR/index.mdb` (14 GB on the live corpus), +which is not env-overridable. Pointed at the primary's `transcripts/` it would have rewritten +production's index under the live editor. Both timed runs instead used a scratch corpus +(`$CLAUDE_JOB_DIR/tmp/s4b-buildindex.sh`): a directory of SYMLINKS to the live `channels/`, +`sites/`, `saved-videos/` and the corpus-wide JSON files, with no `index.mdb`, and +`EXPORT_PUBLIC_DIR` in scratch — so every read is the real corpus, every write is scratch, and +each run is a cold full build (the incremental path would skip the very per-video coercions +being timed). + +**Deviations.** +1. *No mtime freshness in `sidecar()`* — no reader consumes one. +2. *The channel coercions live in `channelConfig.ts`, and the zod schema wraps them.* + `channelConfig.ts` is value-imported by six `"use client"` modules, so `parseChannelConfig` + cannot delegate to zod. `CHANNEL_CONFIG_COERCIONS` there is the one set; both + `parseChannelConfig` and `channelConfigSchema` compose it; a test pins that the two agree + over generated inputs. +3. *`writeJsonAtomic` takes `newline` as well as `indent`.* The plan's rule ("`\n` only at + indent 2") would have changed bytes: attribution/diarization are compact WITH a newline, the + chart/alias/tag stores indented WITHOUT one, the export pages compact without one. +4. *A synchronous `writeJsonAtomicSync`* for chartsStore/aliasesStore/curatedTagsStore (a + synchronous API cannot await the chain; nothing else runs while it does). +5. *The site types and parsers moved into `siteSchema.ts`* (4a's settings pattern, to avoid a + `site.ts` ↔ `siteSchema.ts` cycle). And zod 4 OMITS an input-absent key whose transform + returns `undefined`, so the object step re-emits every key in `SITE_KEYS` order — the + always-emit shape `parseSite` has always had. +6. *`migrateToSites.ts:101` is not on `readChannelConfigFile`*: it reads the legacy raw `group` + key the schema drops. +7. *`patchChannelConfig` returns what it wrote (or null) and holds a per-path lock.* The two + callers that used to fall back to their own copy of a missing config still do, through + `writeChannelConfig`: relocateChannelMedia (move out) and renameChannel. +8. *doNotClean / excludeTruncatedCheck keep their rule*: ANY parseable JSON — even `null` or + `3` — reads as a marker (`{setAt:""}`), so the coercion is not "non-null object". +9. *storageLocations' rollback restores only `dataDir`* (the one field the job changed), not + the whole config it found at its start. +10. *relocateChannelMedia's move-out does not throw on a null patch* (the review asked for a + throw at both sites; storageLocations throws). The move-out has no ledger to roll back and + the swap has already happened when the config is written, so a throw would leave a swapped + link and an unrecorded `dataDir`; it writes the job's own copy of the config instead — the + same thing its existing no-config branch did. + +**Behaviour changes (all intended).** +- A social fetch no longer reverts Configure-form edits made while it ran (`fetchPosts` wrote + back the config it read at its start; it now patches `lastSyncedAt`). +- A sync over a `config.json` that is not valid JSON used to throw at the stamp; the stamp is + now skipped (`readChannelConfig` already answered null, so the scheduler never picks such a + channel). +- `resolveEffectiveAvailability`'s download-outcome side reads through `loadDownloadOutcome`, + so it now needs the full shape check. A read-only scan of all 57,897 live + `download-outcome.json` files: 1,120 carry an `availabilityClass`, **0** answer differently. +- **Config writes now come out in SCHEMA key order.** Before, a form save or toggle wrote its + caller's spread order (form keys appended last); syncs already normalised it through + `updateConfigField`. Semantically nothing moves, but `transcripts/` is its own git repo: + the first form save or toggle per channel may show a one-time key-reorder diff there. The + numbers tool (parse → write) does not exercise this. + +**Review fixes (`973e59ee`).** +- The jsonFile state (`tmpSeq`, write chains, file locks) lives on `globalThis.__yttJsonFile__` + — the house pattern (`jobs/registry.ts`, `controller/autoRunner.ts`) — because Next can load + the module once per bundle layer (instrumentation-armed runners vs server actions), and two + copies would each start the counter at 0 and hold separate locks. The temp name gains 4 + random bytes (`${file}.tmp-${pid}-${seq}-${hex}`). A test imports a second module instance and + checks both share one chain. +- storageLocations' re-point THROWS on a null patch, so the ledger rolls the link back instead + of recording `configWritten` with no `dataDir` on disk. relocateChannelMedia's move-out, on a + null patch, writes the job's own copy instead (deviation 10). +- `writeChannelConfig` returns what it wrote, so the patch parses once; the site `z.object` is + built once at module load (`parseSite` fills `siteId`); a dead import removed. +- SITE.md / CHANNEL.md prose corrected (the always-written site keys named; "absent means + inherit" only for the overrides; the two whole-config fallbacks named; the doubled + `downloadFilter`/`audioCheck` headings removed), with two new pinning tests. + +**Not every JSON writer is on the shared writer.** The 14 below are JSON writers still on the +per-pid temp name `${file}.tmp-${process.pid}` — not "non-JSON", as a draft of this record +said: +`controller/failedTranscriptions.ts:33,54`, `controller/maybeMissingStore.ts:54`, +`controller/rosterStore.ts:235`, `controller/duplicateShorts.ts:640,734`, +`controller/scanCorruptMedia.ts:393,454`, `controller/shard.ts:56`, +`controller/backupSavedVideos.ts:118`, `controller/relocateDir.ts:139`, +`jobs/syncSchedulerState.ts:122`, `jobs/workerDefaults.ts:61`, `lib/widgetPresets.ts:83`, +`lib/homepage.ts:129`, `bin/migrate-channel-priority.ts:206`. +**`maybeMissingStore` and `rosterStore` are per-channel files written from several lanes — the +same race class this slice fixes for `config.json`. Owed, later.** The non-JSON writers +(normalizeTranscript/LiveChat cues, `runYtdlp` playlist, xSessionBroker, videoActions, +cutReleaseAction, transcode, transcribeOne, savedVideo's copy, the umtool ones) are out of +scope. The chain is per process: a CLI beside the live editor is not covered (the rename is +still atomic). + +Also noted by review, accepted: `channelConfig.ts` and `channelGroups.ts` are value-imported by +client modules, so the four channel/group `*_FIELD_DOCS` string records ship to the browser +(a few KB, no zod — the grep below). + +**Numbers** (`plans/tools/phase3-files-numbers.ts`, one process, never writes the corpus). The +inputs were FROZEN once (`FREEZE_TO`) into a scratch tree — 6 `site.json`, 71 `config.json`, +and the first ≤300 dirs per sidecar filename in sorted slug × sorted id order: 1,763 sidecar +files (attribution 259, diarization 300, availability 300, download-outcome 300, +transcribe-outcome 300, do-not-clean 1, exclude-truncated-check 3, ai-digest 300, +ai-digest.overrides 0 — none exist in the corpus) — because the live corpus moves under the +running editor. The script prints each site parse + the file `writeSite(getSite())` writes, +each config parse + md5 of `writeChannelConfig(readChannelConfig())`, and per sidecar md5 of +the canonical load + md5 of `write(load())` (load only for the two markers, whose only writer +stamps the clock). `main` (a detached worktree at `1e27f7c3`) vs the branch at `e75047f8`, +`f6a08bd1` and — after the review fixes — `973e59ee` (`s4b-numbers-fix.txt`): **diff empty, +3,832 lines**, each time. **Unknown-key report: empty** for all +77 files. 0 null loads. + +Parity beyond the corpus (ad hoc, main's function vs the branch's): `parseSite` over 20,000 +random inputs — deepStrictEqual and key order equal; `writeSite` over 3,000 random sites — +byte-identical files or the identical throw; `parseChannelConfig` over 20,000 random inputs — +deepStrictEqual and key order equal. + +**`build:index` timing** (cold full build, scratch corpus as above; the build's own +`Done in` figure — 77,624 transcripts, 6 sites built): + +| run | when | machine | time | +|---|---|---|---| +| before-main (`1e27f7c3`) #1 | 14:03–14:29 | I/O-loaded: the live editor's sync-all + metadata scan and an ffmpeg remux on the platter drive (`/proc/pressure/io` "some" 60–78 %) | 1542.74 s | +| branch (`973e59ee`) | 14:29–14:57 | same load, partly | 1699.98 s | +| before-main #2 | 15:54–16:22 | quiet | 1691.98 s | + +Branch vs before-main #2: **+0.5 %** — within the 5 % gate, **not a regression**; run #1 is the +outlier (the load shifted which run it slowed; the numbers are not a controlled benchmark). +No profiling was needed. Run #2 was taken by the coordinator with the same harness. + +**Gates on the branch tip before the merge (`973e59ee` + this commit):** + +| gate | result | +|---|---| +| `pnpm -r --workspace-concurrency=1 exec tsc --noEmit` | clean after every commit | +| common tests | **1709/1709** (1663 before; +46: jsonFile 9, sidecar 8, siteSchema 11, channelConfigSchema 7, channels +5, fileSchemaDocs 6) | +| editor unit | 67/67 | +| `pnpm run test:scripts` | 156 pass / 1 skip | +| mcp | 219/219 | +| `next build` editor / export (at `7c03d7c9`) | green; `├ ƒ /api/view/[name]`; `grep -rl 'ZodError\|_zod'` over both `.next/static` prints nothing | +| `file-schemas-docs --check` | clean | +| numbers tool vs `main` | diff empty (above) | +| e2e editor, the plan's 34 specs, at `7c03d7c9` | **190 passed**, 0 failed, 11.1 min, exit 0 — every spec the plan named exists; none dropped | +| e2e editor after the review fixes (`973e59ee`): storage-locations, channel-storage, channel-rename, sites-crud, channels, digest, availability | **63 passed**, 0 failed, 3.7 min | +| e2e export: site-branding, related-sites, pwa-search, ask-chat | **33 passed**, 0 failed, 1.0 min | diff --git a/plans/tools/phase3-files-numbers.ts b/plans/tools/phase3-files-numbers.ts @@ -0,0 +1,317 @@ +#!/usr/bin/env tsx +// The one-core Phase 3 slice 4b measurement: what the site.json, channel +// config.json and per-video sidecar readers ANSWER over the real corpus, and +// what a write of that answer puts on disk — printed deterministically so a run +// on `main` and a run on the branch can be diffed. +// +// Model: phase3-settings-numbers.ts next door, and the same rules. +// +// NEVER WRITES THE CORPUS. Every file is COPIED into a scratch tree under +// os.tmpdir() first; every read-for-measurement and every write-back happens on +// the copy, and the scratch tree is deleted at the end. The live corpus is only +// ever read (readdir, readFile, copyFile source). +// +// NEVER BOOTS A SERVER. The readers are called in-process; instrumentation.ts +// is not loaded, so nothing is armed. +// +// ONE PROCESS. `TRANSCRIPTS_DIR` is pointed at the scratch tree BEFORE any +// module that memoises `getPaths()` is imported, and every reader used here +// takes its location explicitly (a Paths, or a video dir). +// +// USES ONLY NAMES PRESENT ON BOTH SIDES of slice 4b (getSite, writeSite, +// readChannelConfig, writeChannelConfig, the loadX/writeX sidecar functions), +// so the SAME file runs on `main` and on the branch. The declared-key lists come +// from the branch's docs records when they exist and from a literal otherwise; +// the literal is checked against the records when both are present. +// +// Usage, from the repo root: +// LIVE_TRANSCRIPTS_DIR=/abs/transcripts node_modules/.bin/tsx plans/tools/phase3-files-numbers.ts > out.txt +// Default LIVE_TRANSCRIPTS_DIR: the primary checkout's, a sibling of this repo. +// SAMPLE_N (default 300): sidecar dirs sampled per filename. +// +// FREEZING THE INPUTS. The live corpus moves while a slice is in flight (the +// running editor stamps `lastSyncedAt`, writes outcomes), so a `main` run in the +// morning and a branch run in the evening would differ for reasons that are not +// this code. `FREEZE_TO=/abs/dir` copies exactly what a measurement reads — +// every site.json, every config.json, and the sampled sidecar files, in the same +// layout — into that directory and exits; both runs then take +// `LIVE_TRANSCRIPTS_DIR=/abs/dir`, and a frozen tree samples itself. + +import { createHash } from "node:crypto"; +import fs from "node:fs"; +import os from "node:os"; +import path from "node:path"; +import { fileURLToPath } from "node:url"; + +const HERE = path.dirname(fileURLToPath(import.meta.url)); +const REPO = path.resolve(HERE, "..", ".."); +const LIVE = + process.env.LIVE_TRANSCRIPTS_DIR ?? + path.join(path.dirname(REPO), "yt-dlp-transcript-browser", "transcripts"); +const SAMPLE_N = Number(process.env.SAMPLE_N ?? 300); + +const SCRATCH = fs.mkdtempSync(path.join(os.tmpdir(), "phase3-files-")); +process.env.TRANSCRIPTS_DIR = SCRATCH; +process.env.TZ = "UTC"; + +// The keys each file may carry, as of 2026-09-24. On the branch these must +// equal the docs records (asserted below). +const SITE_KEYS_LITERAL = [ + "siteId", "siteTitle", "siteDescription", "headerTitle", "homeTagline", + "accent", "socialLinks", "groups", "defaultGroupId", "channels", + "cloudflareProject", "siteUrl", "relatedSites", "pwa", "archives", + "archiveMaxBytes", "duplicates", "hubUrl", +]; +const CHANNEL_KEYS_LITERAL = [ + "handling", "sourceKind", "postFetcher", "socialHandle", "platform", "name", + "url", "audioFormat", "downloadFormat", "keepSourceVideo", "keepLatest", + "extractionMode", "savedVideosDir", "dataDir", "ytdlpExtraArgs", "subLangs", + "lastSyncedAt", "lastFullDownloadAt", "lastFullSweepAt", "excludeFromBuild", + "excludeFromCleanup", "syncIntervalMinutes", "fullSweepIntervalMinutes", + "skipLiveDownloads", "downloadFilter", "cookiesFromBrowser", "cookieMode", + "sleepBetweenDownloadsSeconds", "audioCheck", +]; + +function sortedKeys(_key: string, value: unknown): unknown { + if (!value || typeof value !== "object" || Array.isArray(value)) return value; + const src = value as Record<string, unknown>; + const out: Record<string, unknown> = {}; + for (const k of Object.keys(src).sort()) out[k] = src[k]; + return out; +} + +function canonical(v: unknown): string { + // `undefined` members are dropped by JSON; that is also what a write drops. + return JSON.stringify(v, sortedKeys, 2) ?? "undefined"; +} + +function md5(text: string | Buffer): string { + return createHash("md5").update(text).digest("hex"); +} + +function sortedDir(dir: string): string[] { + try { + return fs.readdirSync(dir).sort(); + } catch (e) { + return []; + } +} + +function rawJson(file: string): unknown { + try { + return JSON.parse(fs.readFileSync(file, "utf8")); + } catch { + return undefined; + } +} + +function rawKeys(raw: unknown): string[] { + return raw && typeof raw === "object" && !Array.isArray(raw) + ? Object.keys(raw as object) + : []; +} + +async function declaredKeys(): Promise<{ site: string[]; channel: string[] }> { + const siteMod = (await import("../../common/lib/site")) as Record<string, unknown>; + const chMod = (await import("../../common/lib/channelConfig")) as Record<string, unknown>; + const siteDocs = siteMod.SITE_FIELD_DOCS as Record<string, string> | undefined; + const chKeys = chMod.CHANNEL_CONFIG_KEYS as readonly string[] | undefined; + const same = (a: readonly string[], b: readonly string[]) => + [...a].sort().join() === [...b].sort().join(); + if (siteDocs && !same(Object.keys(siteDocs), SITE_KEYS_LITERAL)) { + throw new Error("SITE_FIELD_DOCS disagrees with this script's literal list"); + } + if (chKeys && !same(chKeys, CHANNEL_KEYS_LITERAL)) { + throw new Error("CHANNEL_CONFIG_KEYS disagrees with this script's literal list"); + } + return { site: SITE_KEYS_LITERAL, channel: CHANNEL_KEYS_LITERAL }; +} + +async function sites(declared: string[]): Promise<void> { + const { getPaths } = await import("../../common/lib/paths"); + const { getSite, writeSite, siteConfigFile } = await import("../../common/lib/site"); + const paths = getPaths(); + console.log("# site.json"); + for (const id of sortedDir(path.join(LIVE, "sites"))) { + const live = path.join(LIVE, "sites", id, "site.json"); + if (!fs.existsSync(live)) continue; + const file = siteConfigFile(paths, id); + fs.mkdirSync(path.dirname(file), { recursive: true }); + fs.copyFileSync(live, file); + console.log(`## ${id}`); + const unknown = rawKeys(rawJson(file)).filter((k) => !declared.includes(k)); + console.log(`unknown keys: ${JSON.stringify(unknown.sort())}`); + const site = getSite(id, paths); + console.log(canonical(site)); + try { + await writeSite(site, paths); + console.log("### written by writeSite(getSite())"); + console.log(fs.readFileSync(file, "utf8").trimEnd()); + } catch (e) { + console.log(`WRITE THREW: ${(e as Error).message}`); + } + } +} + +async function channels(declared: string[]): Promise<string[]> { + const { getPaths } = await import("../../common/lib/paths"); + const { readChannelConfig, writeChannelConfig } = await import( + "../../common/controller/channels" + ); + const paths = getPaths(); + console.log(""); + console.log("# channel config.json"); + const slugs: string[] = []; + for (const slug of sortedDir(path.join(LIVE, "channels"))) { + const live = path.join(LIVE, "channels", slug, "config.json"); + if (!fs.existsSync(live)) continue; + slugs.push(slug); + const file = path.join(paths.channelsDir, slug, "config.json"); + fs.mkdirSync(path.dirname(file), { recursive: true }); + fs.copyFileSync(live, file); + console.log(`## ${slug}`); + const unknown = rawKeys(rawJson(file)).filter((k) => !declared.includes(k)); + console.log(`unknown keys: ${JSON.stringify(unknown.sort())}`); + const config = await readChannelConfig(paths, slug); + console.log(canonical(config)); + if (!config) continue; + try { + await writeChannelConfig(paths, slug, config); + console.log(`written md5: ${md5(fs.readFileSync(file))}`); + } catch (e) { + console.log(`WRITE THREW: ${(e as Error).message}`); + } + } + return slugs; +} + +type Pair = { + filename: string; + load: (dir: string) => Promise<unknown>; + // Absent for the two markers whose only writer stamps the clock. + write?: (dir: string, value: never) => Promise<void>; +}; + +async function sidecarPairs(): Promise<Pair[]> { + const attribution = await import("../../common/lib/attribution-server"); + const diarization = await import("../../common/lib/diarization-server"); + const availability = await import("../../common/lib/availability-server"); + const dlo = await import("../../common/lib/downloadOutcome-server"); + const tro = await import("../../common/lib/transcribeOutcome-server"); + const dnc = await import("../../common/lib/doNotClean-server"); + const etc = await import("../../common/lib/excludeTruncatedCheck-server"); + const digest = await import("../../common/lib/digest-server"); + const { ATTRIBUTION_FILENAME } = await import("../../common/lib/attribution"); + const { DIARIZATION_FILENAME } = await import("../../common/lib/diarization"); + const { AVAILABILITY_FILENAME } = await import("../../common/lib/availability"); + const { DOWNLOAD_OUTCOME_FILENAME } = await import("../../common/lib/downloadOutcome"); + const { TRANSCRIBE_OUTCOME_FILENAME } = await import("../../common/lib/transcribeOutcome"); + const { DO_NOT_CLEAN_FILENAME } = await import("../../common/lib/doNotClean"); + const { EXCLUDE_TRUNCATED_CHECK_FILENAME } = await import( + "../../common/lib/excludeTruncatedCheck" + ); + const { DIGEST_FILENAME, DIGEST_OVERRIDES_FILENAME } = await import( + "../../common/lib/digest" + ); + return [ + { filename: ATTRIBUTION_FILENAME, load: attribution.loadAttribution, write: attribution.writeAttribution }, + { filename: DIARIZATION_FILENAME, load: diarization.loadDiarization, write: diarization.writeDiarization }, + { filename: AVAILABILITY_FILENAME, load: availability.loadAvailability, write: availability.writeAvailability }, + { filename: DOWNLOAD_OUTCOME_FILENAME, load: dlo.loadDownloadOutcome, write: dlo.writeDownloadOutcome }, + { filename: TRANSCRIBE_OUTCOME_FILENAME, load: tro.loadTranscribeOutcome, write: tro.writeTranscribeOutcome }, + { filename: DO_NOT_CLEAN_FILENAME, load: dnc.loadDoNotClean }, + { filename: EXCLUDE_TRUNCATED_CHECK_FILENAME, load: etc.loadExcludeTruncatedCheck }, + { filename: DIGEST_FILENAME, load: digest.loadDigest, write: digest.writeDigest }, + { filename: DIGEST_OVERRIDES_FILENAME, load: digest.loadDigestOverrides, write: digest.writeDigestOverrides }, + ] as Pair[]; +} + +// A deterministic sample: channels in sorted order, video ids in sorted order, +// the first SAMPLE_N dirs holding each filename. +function sample(slugs: string[], filenames: string[]): Map<string, string[]> { + const out = new Map<string, string[]>(filenames.map((f) => [f, []])); + for (const slug of slugs) { + const data = path.join(LIVE, "channels", slug, "data"); + for (const id of sortedDir(data)) { + const dir = path.join(data, id); + const present = new Set(sortedDir(dir)); + for (const f of filenames) { + const list = out.get(f)!; + if (list.length < SAMPLE_N && present.has(f)) list.push(dir); + } + } + if ([...out.values()].every((l) => l.length >= SAMPLE_N)) break; + } + return out; +} + +async function sidecars(slugs: string[]): Promise<void> { + const pairs = await sidecarPairs(); + const picked = sample(slugs, pairs.map((p) => p.filename)); + console.log(""); + console.log(`# sidecars (first ${SAMPLE_N} per filename, sorted slugs × sorted ids)`); + let n = 0; + for (const pair of pairs) { + const dirs = picked.get(pair.filename)!; + console.log(`## ${pair.filename} — ${dirs.length} dirs`); + let nulls = 0; + for (const live of dirs) { + const rel = path.relative(path.join(LIVE, "channels"), live); + const dir = path.join(SCRATCH, "sidecars", String(n++)); + fs.mkdirSync(dir, { recursive: true }); + fs.copyFileSync(path.join(live, pair.filename), path.join(dir, pair.filename)); + const loaded = await pair.load(dir); + if (loaded === null) nulls++; + let written = "-"; + if (loaded !== null && pair.write) { + await pair.write(dir, loaded as never); + const file = path.join(dir, pair.filename); + written = fs.existsSync(file) ? md5(fs.readFileSync(file)) : "removed"; + } + console.log(`${rel} load=${md5(canonical(loaded))} write=${written}`); + } + console.log(`null loads: ${nulls}`); + } +} + +async function freeze(to: string): Promise<void> { + const copy = (rel: string) => { + const dest = path.join(to, rel); + fs.mkdirSync(path.dirname(dest), { recursive: true }); + fs.copyFileSync(path.join(LIVE, rel), dest); + }; + for (const id of sortedDir(path.join(LIVE, "sites"))) { + const rel = path.join("sites", id, "site.json"); + if (fs.existsSync(path.join(LIVE, rel))) copy(rel); + } + const slugs: string[] = []; + for (const slug of sortedDir(path.join(LIVE, "channels"))) { + const rel = path.join("channels", slug, "config.json"); + if (!fs.existsSync(path.join(LIVE, rel))) continue; + slugs.push(slug); + copy(rel); + } + const pairs = await sidecarPairs(); + const picked = sample(slugs, pairs.map((p) => p.filename)); + let files = 0; + for (const [filename, dirs] of picked) { + for (const dir of dirs) { + copy(path.relative(LIVE, path.join(dir, filename))); + files++; + } + } + console.error(`froze ${slugs.length} configs and ${files} sidecar files into ${to}`); +} + +try { + if (process.env.FREEZE_TO) { + await freeze(path.resolve(process.env.FREEZE_TO)); + process.exit(0); + } + const declared = await declaredKeys(); + await sites(declared.site); + const slugs = await channels(declared.channel); + await sidecars(slugs); +} finally { + fs.rmSync(SCRATCH, { recursive: true, force: true }); +}