Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 0856b35b90bd7e9864517ec1049df99731d131b5
parent 541c9fb6e0c322ab5253ffc39f53dcba7fc17f11
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Thu, 24 Sep 2026 19:14:02 -0400

plans: slice W record, writers numbers tool, changelog

phase3-writers-numbers.ts measures every folded JSON writer's bytes over
frozen live samples on both sides (clock frozen, never writes the corpus,
never boots a server): 323 lines each, 315 old-idiom=same, diff empty.
The record carries the commit table, the post-fold grep, the gates and
the 188 orphan temp files found in the live corpus (left for the
operator).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Diffstat:
Meditor/CHANGELOG.md | 1+
Mplans/one-core-phase-3.md | 87+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aplans/tools/phase3-writers-numbers.ts | 439+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
3 files changed, 527 insertions(+), 0 deletions(-)

diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,6 +1,7 @@ # Changelog ## [Unreleased] +- **Every file the editor replaces atomically is now written one way, and a failed write no longer leaves a temp file behind.** Twenty-six places wrote a file by writing `<file>.tmp-<pid>` and renaming it over the original — the channel roster, maybe-missing and metadata-scan records, the scheduler and auto-queue state, worker defaults, widget presets, the homepage config, relocation markers, shard configs, the duplicate and media-scan reports and their review decisions, the saved-video backup manifest, both cue normalizers, the playlist, the failed-transcriptions list, the X cookie jar, a site's CHANGELOG cut, a saved video copied into its store across drives, and the video page's VTT promote and remark. They now all go through one writer (`common/lib/jsonFile-server.ts`), which gives every write its own temp name and queues writes to the same file one behind another, so two jobs touching one channel's roster at once cannot trip over each other's temp file. A write that fails now removes its temp: the live `.auto-queue/` holds 175 `state.json.tmp-…` files (173 of them empty) from the day `/home` filled up (2026-09-11), each one a failed write the old code left behind; nothing deletes those old ones for you — `find transcripts -name '*.tmp-*'` lists them. No file's contents change — every writer puts the same bytes on disk it did before, measured over the live corpus. The cookie jar is still created readable only by you. - **Every channel table and every job-in-flight line is now drawn one way.** The /channels rack, the dashboard's Channels table and the work tables on the operation pages and /cleanup are one table with a column set per page, over one channel row built on the server (which no longer ships a channel's config to the browser); the dashboard's "Needs work" seed is computed by the same code the widget endpoint serves. On the jobs side, /jobs rows, the "Active jobs" cards on channel/video/operation pages, the monitor widget's Active jobs strip and the operations board's "In flight" list are one job row in three sizes, with one rule for which buttons (Retry / Reorder / Drain / Cancel / Force-release) a job gets. **What you might notice:** a work table's report column reads "stale"/"missing" like the rack's instead of a date; the dashboard's Sync button is the rack's; a lane line on /jobs offers Force-release while its runner is running; widget job lines show who asked for the job; an in-flight download on the operations board links to its job page. Nothing a count says moved. - **`site.json`, each channel's `config.json` and the per-video sidecars now have one schema each, and the two config files have generated key tables.** **`SITE.md`** and **`CHANNEL.md`** (new, repo root) list every key with its default and meaning, generated by `common/bin/file-schemas-docs.ts` and checked by a test. Nothing an operator has configured reads or saves differently: every live `site.json` and `config.json`, and a 1,763-file sample of sidecars, read and write back byte-for-byte as before. **Fixed:** a social-channel fetch no longer undoes Configure-form edits made while it was running (it used to write back the whole config it read when it started). Every change to a channel's config now re-reads the file at the moment it saves and changes only its own fields, so a sync stamping its time and a form save made at the same moment both land. Two writes to the same file from the editor no longer share one temporary file. - **`settings.json` has one schema and one writer, and its key table is generated.** Every key, its default, its clamp and its documentation is now one zod schema (`common/lib/settingsSchema.ts`); `getSettings`/`writeSettings` both parse through it, and every settings form saves through one helper (`editor/app/settings/saveSettings.ts`) that merges only what the form changed. **`SETTINGS.md`** (new, repo root) lists every key with its default and what it does, and `settings.json.example` is now the full default object — both generated by `common/bin/settings-example.ts` and checked by a test, so neither can drift. Nothing an operator has configured reads differently. **Fixed:** adding or editing a storage location on `/storage` no longer erases the record of which location the saved-video store is on (`storage.savedVideosLocationId`). diff --git a/plans/one-core-phase-3.md b/plans/one-core-phase-3.md @@ -703,6 +703,93 @@ undated candidates`, twice — once per Next module graph). IO pressure at boot `full avg10 7.26` (58–66 % at the last rollout, when a remux was saturating the platter), which is the difference from the 25 minutes release 2 paid. +### Slice W, as shipped — one write idiom (2026-09-24) + +Branch `one-core/phase-3-w` off `main` `4130aca1`, five commits, unmerged. Slice 4b left +**16 JSON write sites in 13 files** on the per-pid temp name `${file}.tmp-${process.pid}`, +plus the text and binary tmp + rename writers, plus two modules (`metadataScanStore`, +`autoQueueState`) that had grown their own per-module-copy write counters to dodge the +collision. All of them now go through `common/lib/jsonFile-server.ts`, byte-for-byte. + +| sha | what | +|---|---| +| `db9e3f44` | `writeFileAtomic(file, data: string \| Buffer, {mkdir, mode})` under `writeJsonAtomic` (now `jsonText` → `writeFileAtomic`): the same per-path chain on `globalThis`, the same `${file}.tmp-${pid}-${seq}-${random}` name. `mode` goes to `writeFile(tmp, data, {mode})`, i.e. onto the temp at creation, before the rename — the ordering `xSessionBroker` had. `copyFileAtomic(src, dest, {mkdir})` on the same chain (decided: the saved-video move folds rather than stays). No new sync twin. Tests: string / Buffer / empty bytes, mode 0o600 on the file and on every temp seen, one chain shared by the two writers on one path (40 interleaved writes, last issued lands), copy | +| `9bfd15cd` | Race class: `rosterStore.writeRoster` (mkdir), `maybeMissingStore.writeMaybeMissing` (no mkdir), `metadataScanStore.writeMetadataScan` (no mkdir) on `writeJsonAtomic`. The counters `metadataScanStore.ts:188-191` and `autoQueueState.ts:141` deleted; `autoQueueState` on `writeJsonAtomic` (mkdir). New `autoQueueState.test.ts` "overlapping writes do not collide on the tmp file" (12 concurrent writes, last issued lands, no temp left); `metadataScanStore.test.ts:253` and `rosterStore.test.ts:184-186` unchanged and green | +| `ae6ed9d2` | Process-global + per-video JSON: `syncSchedulerState`, `workerDefaults`, `widgetPresets`, `homepage` (mkdir `homepageDir` = the file's parent), `migrate-channel-priority`, `relocateDir.writeDirMarker`, `shard.saveShardConfig`, `duplicateShorts` ×2, `scanCorruptMedia` ×2, `backupSavedVideos` manifest, `normalizeLiveChat`, `normalizeTranscript`, `videoActions.ts` remark (`'{"transcription":[]}\n'` → `writeJsonAtomic(file, {transcription: []}, {indent: 0})`). Compact without newline (`{indent: 0, newline: false}`) for the two reports and the two cue files. mkdir exactly where a site had one. New `controller/compactJsonWriters.test.ts` (the four compact writers' bytes = `JSON.stringify` of their parse, no newline; passes on the parent too) and the literal pinned in `jsonFile-server.test.ts` | +| `588fd7e9` | Text / binary: `failedTranscriptions` prune + clear, `runYtdlp.writePlaylistFile`, `xSessionBroker.writeCookieJar` (`{mkdir: true, mode: 0o600}`), `savedVideo-server.moveFileCrossDevice` (`copyFileAtomic`), `sites/lib/cutReleaseAction.ts`, `videoActions.ts` VTT promote. One comment on `storageWatch.ts`'s `let timer` (a per-copy singleton, not a temp name, one caller) | +| (this) | `plans/tools/phase3-writers-numbers.ts`, this record, the changelog bullet | + +`videoActions.ts` changed at its two write sites and one added import line only — no export +renamed, no signature changed (slice 3b owns its import list). + +**After commit 4,** `git grep -n 'tmp-${process.pid}' -- common editor`: + +``` +common/controller/buildIndex.ts:971: tmpPath = `${outPath}.tmp-${process.pid}`; +common/controller/buildStats.ts:220: const tmp = `${outPath}.tmp-${process.pid}`; +common/controller/transcode.ts:31: `audio.tmp-${process.pid}.${opts.targetFormat}`, +common/controller/transcribeOne.ts:142: const tmpBase = `transcript.tmp-${process.pid}`; +common/lib/jsonFile-server.test.ts:65: assert.ok(a.startsWith(`/x/config.json.tmp-${process.pid}-`)); +common/lib/jsonFile-server.ts:9:// of them spelling the temp file `${file}.tmp-${process.pid}`. That name is the +common/lib/jsonFile-server.ts:136: return `${file}.tmp-${process.pid}-${seq}-${randomBytes(4).toString("hex")}`; +``` + +The two export page writers, as planned, **plus two the plan did not list**: `transcode.ts:31` +and `transcribeOne.ts:142` are not writers of ours — they name the output file an external +process (ffmpeg; whisper / chough) writes, which the code then renames. Nothing to fold; left +and named. The last three hits are the shared writer itself, its history comment and its test. +`git grep 'rename(tmp'` over `common editor` finds only `buildIndex`, `buildStats`, `transcode`. + +**Out of scope by name:** `buildIndex.ts:971` (the streaming `createWriteStream` page writer) +and `buildStats.ts:220` (a hand-joined array) — restructuring, not a fold; +`scripts/diarize.mjs`; everything under `umtool/`. The module-level `storageWatch` timer and +the TTL caches in `autoRunner.ts` / `recencyIndex.ts` are not temp names. + +**Behaviour changes (intended).** (1) Every folded write is chained per absolute path with every +other `writeFileAtomic`/`writeJsonAtomic`/`copyFileAtomic` in the process, across module +copies — the roster's writers in `runYtdlp`, `quickAvailabilityCheck` and the +`pipelineActions` server action now serialise. (2) A failed write removes its temp; the old +code left it. (3) Temp names changed shape (`<file>.tmp-<pid>-<seq>-<hex>`); the remark's temp +was `transcript.tmp-<pid>.json` and is now `transcript.json.tmp-…`. Nothing reads temp names. + +**Found and left: 173 orphan temps in the live corpus.** `transcripts/.auto-queue/` holds +**175** `state.json.tmp-2514131-NNNN` files, all dated 2026-09-11 — the day `/home` hit 100 % — +**173 of them 0 bytes** (two are partial, 16–20 KB): each a failed `writeFile` (ENOSPC) the old +code never cleaned. Thirteen more elsewhere: seven `snapshot.json.tmp-<pid>`, three sidecar +temps (`availability.json`, `download-outcome.json`) — all from pre-4b writers — and three +`transcript.tmp-<pid>` whisper output bases (`transcribeOne`, an external process's file). The +new writer cannot leave more of the first kind; the existing ones are corpus files and were +**not touched** — the operator's to delete (`find transcripts -name '*.tmp-*'` lists all 188). + +**Numbers** (`plans/tools/phase3-writers-numbers.ts`, one process, never writes the corpus, +never boots a server, clock frozen, only names present on both sides). Inputs frozen once +(`FREEZE_TO`, 1,787 files) and both runs read the frozen tree: the parent `4130aca1` from a +detached scratch worktree, the branch from `588fd7e9` + the tool. Per writer and sample it +loads through the module's reader, writes through the module's writer into scratch, and prints +the md5 of the bytes plus whether they equal the OLD idiom's bytes (`JSON.stringify(v, null, +2) + "\n"`, or `JSON.stringify(v)` for the compact four). Coverage: 53 rosters, 52 +maybe-missing, 1 metadata-scan, scheduler, auto-queue (33,004 B), worker defaults, widget +presets, homepage, the duplicates report (empty scan), the media-scan report (a merge of the +live 27,756 B report), a relocation marker (synthetic — none live), the backup manifest, 100 +`transcript.cues.json` and 100 `live_chat.cues.json` re-normalized. 323 lines each side, +**315 `old-idiom=same`, 0 DIFF, 0 THREW, 0 leftover temps; `diff` before/after empty** +(`w-numbers-{before,after}.txt`). Not measurable without a live sample: shard configs (none +live), duplicate overrides (the live file has no clusters), media-scan overrides (no file), +the priority migration (a CLI whose only write is the default-options `writeJsonAtomic`, +pinned by `jsonText`'s tests) — and the remark literal, pinned by a unit test. + +**Gates** (worktree `one-core-phase-3-w`, ports 3301/3311/3310/3320): +- `pnpm -r --no-bail --workspace-concurrency=1 exec tsc --noEmit` — clean before every commit. +- common **1733/1733** (1723 + 4 `writeFileAtomic`/`copyFileAtomic` + 1 literal + 1 + `autoQueueState` overlap + 4 compact-writer bytes); editor unit (`tsx --test + "app/**/*.test.ts"`) **67/67**; `test:scripts` **156 + 1 skip** (a first run while this + slice's own e2e held the machine lock failed the queue-lock banner test — the lock was + taken; re-run with it free, green); mcp **219/219**. +- `pnpm --filter editor exec next build` exit 0, `ƒ /api/view/[name]` in the table; + `pnpm --filter export exec next build` exit 0; `ZodError|_zod` over both `.next/static`: 0 + files each. +- e2e, the prompt's 20 specs (all exist): **140 passed, 0 failed, 10.7 min**. + ## Next release — slice 3b and Phase 4 (inventory kept from 2026-09-23) Slices 3a and 4b shipped in the release above (2026-09-24); the slice 3b bullets and the diff --git a/plans/tools/phase3-writers-numbers.ts b/plans/tools/phase3-writers-numbers.ts @@ -0,0 +1,439 @@ +#!/usr/bin/env tsx +// The one-core Phase 3 release 4 slice W measurement: the bytes every folded +// tmp + rename JSON writer puts on disk, over live samples — printed +// deterministically so a run on the parent commit and a run on the branch can +// be diffed. Slice W moved these writers onto lib/jsonFile-server.ts and +// claims it changed no byte; an empty diff is that claim. +// +// Model: phase3-files-numbers.ts next door, and the same rules. +// +// NEVER WRITES THE CORPUS. Every input is COPIED into a scratch tree under +// os.tmpdir() first; every read-for-measurement and every write happens on the +// copy, and the scratch tree is deleted at the end. The live corpus is only +// ever read (readdir, stat, readFile, copyFile source). +// +// NEVER BOOTS A SERVER. The writers are called in-process; instrumentation.ts +// is not loaded, so nothing is armed. +// +// ONE PROCESS. `TRANSCRIPTS_DIR` is pointed at the scratch tree BEFORE any +// module that memoises `getPaths()` is imported. +// +// THE CLOCK IS FROZEN. Several writers stamp `new Date()` into what they write +// (a decision's `decidedAt`, a report's `generatedAt`); `Date` is replaced with +// a subclass whose "now" is fixed, so two runs write the same bytes. +// +// USES ONLY NAMES PRESENT ON BOTH SIDES of slice W (the modules' exported +// readers and writers, unchanged by the slice), so the SAME file runs on the +// parent and on the branch. +// +// Per writer and sample it prints the md5 of the bytes written and whether they +// equal the OLD idiom's bytes for the same value — `JSON.stringify(v, null, 2) +// + "\n"` or, for the compact four, `JSON.stringify(v)` — where v is the parse +// of what was written. The old-idiom column pins the FORMAT on each side; the +// before/after diff of the md5 column pins the CONTENT. +// +// Usage, from the repo root: +// LIVE_TRANSCRIPTS_DIR=/abs/transcripts node_modules/.bin/tsx plans/tools/phase3-writers-numbers.ts > out.txt +// Default LIVE_TRANSCRIPTS_DIR: the primary checkout's, a sibling of this repo. +// SAMPLE_N (default 100): video dirs sampled per cue file. +// +// FREEZING THE INPUTS. The live editor rewrites the scheduler and auto-queue +// state, rosters and maybe-missing records all day, so a parent run and a +// branch run minutes apart would differ for reasons that are not this code. +// `FREEZE_TO=/abs/dir` copies exactly what a measurement reads, in the same +// layout, into that directory and exits; both runs then take +// `LIVE_TRANSCRIPTS_DIR=/abs/dir`, and a frozen tree samples itself. + +import { createHash } from "node:crypto"; +import fs from "node:fs"; +import os from "node:os"; +import path from "node:path"; +import { fileURLToPath } from "node:url"; + +const HERE = path.dirname(fileURLToPath(import.meta.url)); +const REPO = path.resolve(HERE, "..", ".."); +const LIVE = + process.env.LIVE_TRANSCRIPTS_DIR ?? + path.join(path.dirname(REPO), "yt-dlp-transcript-browser", "transcripts"); +const SAMPLE_N = Number(process.env.SAMPLE_N ?? 100); +// A live-chat file can be hundreds of MB; the sample takes small ones. +const LIVE_CHAT_MAX_BYTES = 4 * 1024 * 1024; + +const SCRATCH = fs.mkdtempSync(path.join(os.tmpdir(), "phase3-writers-")); +process.env.TRANSCRIPTS_DIR = SCRATCH; +process.env.TZ = "UTC"; + +const FIXED_NOW = Date.UTC(2026, 8, 24, 12, 0, 0); +const RealDate = Date; +class FixedDate extends RealDate { + constructor(...args: unknown[]) { + if (args.length === 0) super(FIXED_NOW); + else super(...(args as [number])); + } + static now(): number { + return FIXED_NOW; + } +} +globalThis.Date = FixedDate as DateConstructor; + +type Format = "pretty" | "compact"; + +function md5(b: string | Buffer): string { + return createHash("md5").update(b).digest("hex"); +} + +function oldIdiom(v: unknown, format: Format): string { + return format === "compact" ? JSON.stringify(v) : JSON.stringify(v, null, 2) + "\n"; +} + +function sortedDir(dir: string): string[] { + try { + return fs.readdirSync(dir).sort(); + } catch { + return []; + } +} + +// Copy a live file (path relative to the corpus root) into the scratch tree. +function stage(rel: string): boolean { + const src = path.join(LIVE, rel); + if (!fs.existsSync(src)) return false; + const dest = path.join(SCRATCH, rel); + fs.mkdirSync(path.dirname(dest), { recursive: true }); + fs.copyFileSync(src, dest); + return true; +} + +// Report the file a writer just wrote. +function report(label: string, file: string, format: Format): void { + if (!fs.existsSync(file)) { + console.log(`${label} ABSENT`); + return; + } + const bytes = fs.readFileSync(file); + let idiom: string; + try { + idiom = oldIdiom(JSON.parse(bytes.toString("utf8")), format) === bytes.toString("utf8") + ? "old-idiom=same" + : "old-idiom=DIFF"; + } catch { + idiom = "old-idiom=UNPARSEABLE"; + } + const leftovers = sortedDir(path.dirname(file)).filter((n) => + n.startsWith(`${path.basename(file)}.tmp-`), + ); + console.log( + `${label} ${bytes.length}B md5=${md5(bytes)} ${idiom}` + + (leftovers.length ? ` LEFT ${leftovers.length} temp(s)` : ""), + ); +} + +async function run(label: string, fn: () => Promise<void>): Promise<void> { + try { + await fn(); + } catch (e) { + console.log(`${label} THREW: ${(e as Error).message}`); + } +} + +async function channelStores(slugs: string[]): Promise<void> { + const { getPaths } = await import("../../common/lib/paths"); + const roster = await import("../../common/controller/rosterStore"); + const mm = await import("../../common/controller/maybeMissingStore"); + const ms = await import("../../common/controller/metadataScanStore"); + const shard = await import("../../common/controller/shard"); + const paths = getPaths(); + + console.log("# per-channel stores (every live file)"); + for (const slug of slugs) { + const ch = path.join("channels", slug); + if (stage(path.join(ch, roster.ROSTER_FILENAME))) { + await run(`roster ${slug}`, async () => { + await roster.writeRoster(paths, slug, await roster.loadRoster(paths, slug)); + report(`roster ${slug}`, roster.rosterPath(paths, slug), "pretty"); + }); + } + if (stage(path.join(ch, mm.MAYBE_MISSING_FILENAME))) { + await run(`maybe-missing ${slug}`, async () => { + const rec = await mm.loadMaybeMissing(paths, slug); + if (!rec) return void console.log(`maybe-missing ${slug} load=null`); + await mm.writeMaybeMissing(paths, slug, rec); + report( + `maybe-missing ${slug}`, + path.join(paths.channelsDir, slug, mm.MAYBE_MISSING_FILENAME), + "pretty", + ); + }); + } + if (stage(path.join(ch, ms.METADATA_SCAN_FILENAME))) { + await run(`metadata-scan ${slug}`, async () => { + // The writer is private; an upsert that changes one entry's title by a + // fixed suffix forces exactly one write of the whole store. + const scan = await ms.loadMetadataScan(paths, slug); + const id = Object.keys(scan.entries).sort()[0]; + const upsert = id + ? { entries: { [id]: { ...scan.entries[id], title: `${scan.entries[id].title}·` } } } + : {}; + await ms.upsertMetadataScan(paths, slug, upsert, "2026-09-24T12:00:00.000Z"); + report(`metadata-scan ${slug}`, ms.metadataScanPath(paths, slug), "pretty"); + }); + } + for (const op of shard.SHARD_OPS) { + const rel = path.relative(SCRATCH, shard.shardFile(paths, slug, op)); + if (!stage(rel)) continue; + await run(`shard-${op} ${slug}`, async () => { + const cfg = await shard.loadShardConfig(paths, slug, op); + if (!cfg) return void console.log(`shard-${op} ${slug} load=null`); + await shard.saveShardConfig(paths, slug, op, cfg); + report(`shard-${op} ${slug}`, shard.shardFile(paths, slug, op), "pretty"); + }); + } + } +} + +async function globalStores(): Promise<void> { + const { getPaths } = await import("../../common/lib/paths"); + const paths = getPaths(); + const rel = (abs: string) => path.relative(SCRATCH, abs); + console.log(""); + console.log("# process-global stores"); + + const sched = await import("../../common/jobs/syncSchedulerState"); + if (stage(rel(paths.schedulerStateFile))) { + await run("scheduler", async () => { + await sched.writeSchedulerState(paths, await sched.readSchedulerState(paths)); + report("scheduler", paths.schedulerStateFile, "pretty"); + }); + } + const aq = await import("../../common/jobs/autoQueueState"); + if (stage(rel(paths.autoQueueStateFile))) { + await run("auto-queue", async () => { + await aq.writeAutoQueueState(paths, await aq.readAutoQueueState(paths)); + report("auto-queue", paths.autoQueueStateFile, "pretty"); + }); + } + const wd = await import("../../common/jobs/workerDefaults"); + if (stage(rel(paths.workerDefaultsFile))) { + await run("worker-defaults", async () => { + const d = wd.readWorkerDefaults(paths); + if (!d) return void console.log("worker-defaults load=null"); + await wd.writeWorkerDefaults(paths, d.enabledWorkerIds); + report("worker-defaults", paths.workerDefaultsFile, "pretty"); + }); + } + const wp = await import("../../common/lib/widgetPresets"); + if (stage(rel(paths.widgetPresetsFile))) { + await run("widget-presets", async () => { + const list = await wp.readWidgetPresets(paths); + if (list.length === 0) return void console.log("widget-presets empty"); + // Saving over an existing name rewrites the whole list, unchanged. + await wp.addWidgetPreset(paths, list[0].name, list[0].query); + report("widget-presets", paths.widgetPresetsFile, "pretty"); + }); + } + const hp = await import("../../common/lib/homepage"); + if (stage(rel(paths.homepageConfigFile))) { + await run("homepage", async () => { + await hp.writeHomepageConfig(hp.getHomepageConfig(paths), paths); + report("homepage", paths.homepageConfigFile, "pretty"); + }); + } + + const dup = await import("../../common/controller/duplicateShorts"); + if (stage(rel(dup.duplicateOverridesPath(paths)))) { + await run("duplicate-overrides", async () => { + const cur = await dup.readDuplicateOverrides(paths); + const id = Object.keys(cur.clusters).sort()[0]; + if (!id) return void console.log("duplicate-overrides empty"); + const c = cur.clusters[id] as Record<string, unknown>; + await dup.updateDuplicateOverride(paths, id, { + ...(typeof c.canonicalSlug === "string" ? { canonicalSlug: c.canonicalSlug } : {}), + ...(typeof c.notDuplicate === "boolean" ? { notDuplicate: c.notDuplicate } : {}), + ...(typeof c.confirmed === "boolean" ? { confirmed: c.confirmed } : {}), + ...(typeof c.note === "string" ? { note: c.note } : {}), + }); + report("duplicate-overrides", dup.duplicateOverridesPath(paths), "pretty"); + }); + } + // The report writer runs at the end of a detection pass; over the scratch + // tree (configs and stores, no data/) that pass sees no videos. + await run("duplicates-report", async () => { + await dup.detectDuplicateShorts({ paths, onLog: () => {} }); + report("duplicates-report (no videos)", path.join(paths.transcriptsDir, "duplicates.json"), "compact"); + }); + + const scm = await import("../../common/controller/scanCorruptMedia"); + if (stage(rel(scm.mediaScanOverridesPath(paths)))) { + await run("media-scan-overrides", async () => { + const cur = await scm.readMediaScanOverrides(paths); + const key = Object.keys(cur.reviewed).sort()[0]; + if (!key) return void console.log("media-scan-overrides empty"); + const note = (cur.reviewed[key] as { note?: string }).note; + await scm.updateMediaScanOverride(paths, key, { reviewed: true, ...(note ? { note } : {}) }); + report("media-scan-overrides", scm.mediaScanOverridesPath(paths), "pretty"); + }); + } + // A scan of a channel that does not exist merges the live report's every + // finding back unchanged, through the report writer. + if (stage("media-scan.json")) { + await run("media-scan-report", async () => { + await scm.scanCorruptMedia({ paths, channels: ["__phase3-writers-none__"], onLog: () => {} }); + report("media-scan-report (merge of live)", path.join(paths.transcriptsDir, "media-scan.json"), "compact"); + }); + } + + const rd = await import("../../common/controller/relocateDir"); + // No live marker exists outside a move; a fixed one exercises the writer. + await run("relocation-marker", async () => { + const file = path.join(SCRATCH, "channels", ".phase3-writers-marker.json"); + fs.mkdirSync(path.dirname(file), { recursive: true }); + await rd.writeDirMarker(file, { + target: "/mnt/x/slug/data", + direction: "out", + startedAt: "2026-09-24T12:00:00.000Z", + phase: "copy", + }); + report("relocation-marker (synthetic)", file, "pretty"); + }); + + const bk = await import("../../common/controller/backupSavedVideos"); + await run("backup-manifest", async () => { + const dest = path.join(SCRATCH, ".phase3-writers-backup"); + const r = await bk.backupSavedVideos({ paths, dest, onLog: () => {} }); + report(`backup-manifest (${r.entries} saved)`, r.manifestPath, "pretty"); + }); +} + +// The first SAMPLE_N video dirs (sorted slugs × sorted ids) holding `filename`. +function sampleDirs(slugs: string[], filename: string, maxBytes?: number): string[] { + const out: string[] = []; + for (const slug of slugs) { + const data = path.join(LIVE, "channels", slug, "data"); + for (const id of sortedDir(data)) { + if (out.length >= SAMPLE_N) return out; + const f = path.join(data, id, filename); + try { + const st = fs.statSync(f); + if (maxBytes !== undefined && st.size > maxBytes) continue; + out.push(path.join(data, id)); + } catch { + // not in this dir + } + } + } + return out; +} + +const MEDIA = /\.(opus|m4a|mp3|webm|mp4|mkv|wav|ogg|aac|flac|part|ytdl|jpg|jpeg|png|webp)$/i; + +// Copy a video dir's non-media files (metadata, raw transcripts, the prior cue +// file whose format tag a re-normalize reads) into the scratch tree. +function stageVideoDir(live: string, n: number): string { + const dir = path.join(SCRATCH, "videos", String(n)); + fs.mkdirSync(dir, { recursive: true }); + for (const name of sortedDir(live)) { + if (MEDIA.test(name)) continue; + const src = path.join(live, name); + if (!fs.statSync(src).isFile()) continue; + fs.copyFileSync(src, path.join(dir, name)); + } + return dir; +} + +async function cueWriters(slugs: string[]): Promise<void> { + const nt = await import("../../common/controller/normalizeTranscript"); + const nl = await import("../../common/controller/normalizeLiveChat"); + const vs = await import("../../common/lib/videoStatus"); + let n = 0; + console.log(""); + console.log(`# ${vs.CUES_JSON_FILENAME} (first ${SAMPLE_N} dirs holding one, force re-normalize)`); + for (const live of sampleDirs(slugs, vs.CUES_JSON_FILENAME)) { + const rel = path.relative(path.join(LIVE, "channels"), live); + const dir = stageVideoDir(live, n++); + await run(rel, async () => { + const out = await nt.normalizeTranscript({ videoDir: dir, channelSlug: "x", force: true } as never); + if ((out as { status: string }).status !== "wrote") { + return void console.log(`${rel} ${(out as { status: string }).status}`); + } + report(rel, path.join(dir, vs.CUES_JSON_FILENAME), "compact"); + }); + } + console.log(""); + console.log(`# ${vs.LIVE_CHAT_CUES_FILENAME} (first ${SAMPLE_N} dirs with a live chat ≤ 4 MB)`); + for (const live of sampleDirs(slugs, vs.LIVE_CHAT_FILENAME, LIVE_CHAT_MAX_BYTES)) { + const rel = path.relative(path.join(LIVE, "channels"), live); + const dir = stageVideoDir(live, n++); + await run(rel, async () => { + const out = await nl.normalizeLiveChat({ videoDir: dir, channelSlug: "x", force: true }); + if (out.status !== "wrote") return void console.log(`${rel} ${out.status}`); + report(rel, path.join(dir, vs.LIVE_CHAT_CUES_FILENAME), "compact"); + }); + } +} + +// The files the measurement reads, relative to the corpus root. +async function inputs(slugs: string[]): Promise<string[]> { + const { getPaths } = await import("../../common/lib/paths"); + const shard = await import("../../common/controller/shard"); + const dup = await import("../../common/controller/duplicateShorts"); + const scm = await import("../../common/controller/scanCorruptMedia"); + const vs = await import("../../common/lib/videoStatus"); + const paths = getPaths(); + const rel = (abs: string) => path.relative(SCRATCH, abs); + const out: string[] = []; + for (const s of slugs) { + const ch = path.join("channels", s); + out.push( + path.join(ch, "config.json"), + path.join(ch, "roster.json"), + path.join(ch, "maybe-missing.json"), + path.join(ch, "metadata-scan.json"), + ...shard.SHARD_OPS.map((op) => rel(shard.shardFile(paths, s, op))), + ); + } + out.push( + rel(paths.schedulerStateFile), + rel(paths.autoQueueStateFile), + rel(paths.workerDefaultsFile), + rel(paths.widgetPresetsFile), + rel(paths.homepageConfigFile), + rel(dup.duplicateOverridesPath(paths)), + rel(scm.mediaScanOverridesPath(paths)), + "media-scan.json", + ); + const dirs = [ + ...sampleDirs(slugs, vs.CUES_JSON_FILENAME), + ...sampleDirs(slugs, vs.LIVE_CHAT_FILENAME, LIVE_CHAT_MAX_BYTES), + ]; + for (const d of dirs) { + for (const name of sortedDir(d)) { + if (MEDIA.test(name) || !fs.statSync(path.join(d, name)).isFile()) continue; + out.push(path.relative(LIVE, path.join(d, name))); + } + } + return [...new Set(out)].filter((r) => fs.existsSync(path.join(LIVE, r))); +} + +try { + const slugs = sortedDir(path.join(LIVE, "channels")).filter((s) => + fs.existsSync(path.join(LIVE, "channels", s, "config.json")), + ); + if (process.env.FREEZE_TO) { + const to = path.resolve(process.env.FREEZE_TO); + const files = await inputs(slugs); + for (const r of files) { + fs.mkdirSync(path.dirname(path.join(to, r)), { recursive: true }); + fs.copyFileSync(path.join(LIVE, r), path.join(to, r)); + } + console.error(`froze ${files.length} files into ${to}`); + fs.rmSync(SCRATCH, { recursive: true, force: true }); + process.exit(0); + } + // Configs first: several writers read the channel list. + for (const s of slugs) stage(path.join("channels", s, "config.json")); + await channelStores(slugs); + await globalStores(); + await cueWriters(slugs); +} finally { + fs.rmSync(SCRATCH, { recursive: true, force: true }); +}