Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit b741e41b251b98d326a8f2154b95db35b99727ab
parent f130084a614cf3757349bb6ec8f2bdfac7d8f58d
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Thu, 24 Sep 2026 22:16:58 -0400

plans: release 5 slice X record — transcriptDownloads shipped on one-core/r5-exports

The record (commits, gates, numbers, the hub finding), the superseded note
on site-exports-off.md, the numbers tool's key literal, and the
[Unreleased] changelog bullet.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
Meditor/CHANGELOG.md | 1+
Mplans/release-5.md | 72++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mplans/site-exports-off.md | 5+++++
Mplans/tools/phase3-files-numbers.ts | 5+++--
4 files changed, 81 insertions(+), 2 deletions(-)

diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,6 +1,7 @@ # Changelog ## [Unreleased] +- **A site can turn off its visitors' per-video transcript downloads.** The transcript viewer on a published site has always offered three ways to take a video's text away: a **Download** menu (txt, srt, json), **Copy MD**, and **Copy download command** (a `yt-dlp` line for a marked clip). A site's settings form now has a checkbox for them, *Per-video transcript downloads*, beside the archive zips one. Unticked, the site's next build shows none of the three; **Share** and the clip marks stay. It is on by default, so a site nobody touches is unchanged, and the file stores `"transcriptDownloads": false` only when it is off (`SITE.md` has the key). The site's machine contract (`/corpus.json`, `llms.txt`, the manifests and shards the MCP server and report-to-video read) is published either way. The hub has no site settings of its own and keeps the controls. The editor's own video pages are unaffected. - **Channel rows no longer scroll over a group's controls on `/channels`.** Scrolled down and to the right, the pinned Slug column of every row painted over the pinned group header and its five station buttons (Sync, Download, Transcribe, Digest and the speaker lane), and took the clicks. The pinned Slug cell and the group header sat at the same stacking level, and the later rows won. The rack now has one named layer order, kept in one file: the Advanced panel, then the column header, then the group header, then the pinned checkbox and Slug cells. Nothing ties any more. The screenshot audit found four more problems, fixed as well. A group header's name and buttons now stay on screen however far the columns scroll across (they used to scroll off to the left). An Advanced panel opened near the bottom or the right edge scrolls itself into view instead of being cut off. The rule above a pinned group header moves with it instead of leaving a gap the rows showed through. On a phone, the column header no longer paints over the selection bar pinned to the bottom of the screen. - **A group's Transcribe works for YouTube channels, and it counts what it queues.** The station used to be disabled for every `youtube`-handling channel with the message "a youtube-handling channel never runs whisper". That was wrong. A YouTube video that came down with no captions is transcription work like any other, and the automatic runner already treats it that way. Transcribe now counts two kinds of video, after the usual members-only, deleted and private exclusions: downloaded videos with no transcript at all, and downloaded videos whose only transcript is YouTube's auto-captions. Pressing it queues exactly those videos, by id, as the channel page does: up to two jobs per channel on the transcription queue. A video downloaded before it went private, members-only or deleted is no longer transcribed by the group button, because it was never in the figure. Pressing it again while either job runs says *already running*. The wording names no method ("…has downloaded audio to transcribe", "…each takes minutes"). **This figure can now be higher than the Transcription band in the same rack on channels with many auto-caption-only videos.** The band counts videos with no transcript at all, while the station counts everything its button would queue. That is intended. - **The editor's atomic JSON, text and binary writes now go one way, and a failed write no longer leaves a temp file behind.** Nineteen JSON write sites and seven text and binary ones each wrote `<file>.tmp-<pid>` and renamed it over the original — the channel roster, maybe-missing and metadata-scan records, the scheduler and auto-queue state, worker defaults, widget presets, the homepage config, relocation markers, shard configs, the duplicate and media-scan reports and their review decisions, the saved-video backup manifest, both cue normalizers, the playlist, the failed-transcriptions list, the X cookie jar, a site's CHANGELOG cut, a saved video copied into its store across drives, and the video page's VTT promote and remark. They now all go through one writer (`common/lib/jsonFile-server.ts`), which gives every write its own temp name and queues writes to the same file one behind another, so two jobs touching one channel's roster at once cannot trip over each other's temp file. A write that fails now removes its temp: the live `.auto-queue/` holds 175 `state.json.tmp-…` files (173 of them empty) from the day `/home` filled up (2026-09-11), each one a failed write the old code left behind; nothing deletes those old ones for you — `find transcripts -name '*.tmp-*'` lists them. Four temp names stay, on purpose: the export build's two page writers `buildIndex.ts` (a streaming page writer) and `buildStats.ts` (a hand-joined array) — folding them is a restructuring, not a swap — and `transcode.ts` / `transcribeOne.ts` name the output file ffmpeg or the transcription app writes, which is not our write to fold. No file's contents change — every writer puts the same bytes on disk it did before, measured over the live corpus. The cookie jar is still created readable only by you. diff --git a/plans/release-5.md b/plans/release-5.md @@ -141,3 +141,75 @@ STATE. umtool: rebuild only if `git diff --stat <live>..<new> -- umtool` is non- `one-core-phase-3.md` "Next — Phase 4". ## Record + +### Slice X, as shipped — visitor exports off, per site (2026-09-24) + +Branch `one-core/r5-exports` off `main` `f4da04a9`. One new `site.json` key, `transcriptDownloads` +(boolean, absent = on), gating the transcript modal's three per-video export controls on the export +site; the site form writes it. `archives: false` needed no code — it gained test coverage. + +| sha | what | +|---|---| +| `acf0c676` | `common/lib/siteSchema.ts`: `transcriptDownloads` in the `Site` type, `SITE_FIELD_DOCS` (after `duplicates`), a zod field that is off only on an explicit `false`, `siteToDisk` writes it only when `false` — the `archives` pattern, four places. `siteSchema.test.ts` extends the existing default / opt-out / siteToDisk / round-trip / writeSite tests. `SITE.md` regenerated (`--check` green) | +| `d301d757` | `PlayerProvider({children, features?})`: `PlayerFeatures = {transcriptDownloads}`, `DEFAULT_PLAYER_FEATURES` all on, exposed as `usePlayer().features` (memoised on the flag, so an inline `features={{…}}` does not rebuild the context). `TranscriptModal` renders none of Copy download command, the Download dropdown and Copy MD when off; Share, the clip marks and every accessible name are unchanged when on. Mounts: `(workspace)/layout.tsx` → `SiteWorkspace` prop, `duplicates/page.tsx` directly, `HubHome`/`AskHub` via a prop from their server pages (`(workspace)/page.tsx`, `(workspace)/ask/page.tsx`) — all `currentSite().transcriptDownloads !== false` | +| `a42b1737` | `SiteForm.tsx` checkbox after the archive size cap, label "Per-video transcript downloads (Download menu, Copy Markdown, Copy download command)"; `actions.ts` reads it like `archives`. `editor/e2e/sites-crud.spec.ts`: one new test round-trips BOTH opt-outs (no spec covered the `archives` checkbox before) | +| `9958b42b` | export e2e: `transcript-downloads.spec.ts` (2), `archives-off.spec.ts` (2) | +| *(this commit)* | `plans/tools/phase3-files-numbers.ts` literal gains the key; this record; `site-exports-off.md` superseded note; changelog | + +**Deviations / findings.** +1. *The hub has no `site.json`.* `sites/_homepage/` holds `homepage.json`; in hub mode + `currentSite()` is `hubSite()`, a `Site` synthesised from the homepage config that never sets + the key. The hub pages are wired through `currentSite()` like every other mount, so they follow + the key IF the hub ever gets one — today the hub always shows the three controls. Giving the + hub a switch means a `homepage.json` key and `hubSite()` passing it through + (`common/lib/homepage.ts`, not this slice's file). **Open for the operator:** is the hub + (`archilyzer.pages.dev`'s Browse/Ask) a surface the exports-off decision covers? +2. *Client mounts cannot call `currentSite()`* (it reads `site.json` off disk). `SiteWorkspace`, + `HubHome` and `AskHub` are `"use client"`, so each takes a boolean prop from its server parent; + `duplicates/page.tsx` is a server component and passes the features itself. +3. *Export e2e varies per-site config by flipping the fixture.* The suite serves ONE site + (`SITE_ID=testsite`) from one `next dev`; no spec varied `site.json` before (the per-site + variants in the suite are route mocks of client data). `currentSite()` is uncached and `next + dev` re-renders the layout per request, so the off test writes `transcriptDownloads: false` + into `e2e/fixtures/sites/testsite/site.json`, asserts, and `afterEach` restores the bytes; a + run killed mid-test leaves the key, which `beforeAll` strips on the next run. A second fixture + site would need a second `next dev` of the same app dir. +4. *`archives: false` is a build-time switch, so its coverage is a compose run, not a page.* The + Downloads page and its Header/Footer links key off `public/archives/manifest.json` + (`hasArchives()`), which `compose-site` writes; the dev server reads the checked-out + `export/public` (symlinked from the primary). `archives-off.spec.ts` runs `compose-site` into a + scratch `EXPORT_PUBLIC_DIR` under the test's output dir over a scratch channel-less site: + `archives: false` → no manifest, a pre-seeded stale one removed; key absent → a fresh manifest. + The "no `/downloads` link" half follows from `hasArchives()` and is not rendered in a test. +5. *The actions themselves are not gated.* `copyDownloadCommand` / `downloadTranscriptFile` / + `copyTranscriptMarkdown` stay on the context; the three buttons are their only callers (grep), + and the data is in the visitor's browser either way — this removes the affordance, it is not + access control. + +**Gates** (worktree root). +- tsc `pnpm -r --no-bail --workspace-concurrency=1 exec tsc --noEmit`: clean. +- common `pnpm --filter yt-dlp-transcript-common test`: **1738/1738** (assertions added inside + existing tests; count unchanged). `file-schemas-docs.ts --check`: green. +- editor unit (`editor/`, `tsx --test "app/**/*.test.ts"`): **72/72**. `pnpm run test:scripts`: + **156 pass + 1 skip**. mcp: **219/219**. +- `pnpm --filter editor exec next build`: ok (47 s). `pnpm --filter export exec next build`: ok + (32 s); no dangling `export/public` links before either. +- EXPORT e2e, full suite (`node scripts/worktree.mjs run -- pnpm --filter export run e2e`): + **192 passed, 0 failed, 8.5 min** (188 + the 4 new). The fixture was back to its committed + bytes afterwards (`git status` clean). +- EDITOR e2e (`pnpm e2e` with `sites-crud site-publish-preview site-scope cut-release view-route + settings`, all six exist; `sites-crud` is now the spec covering the `archives` checkbox): + **50 passed, 0 failed, 1.7 min** (after ~7.5 min behind slice R's run on the queue). + +**Numbers** (`plans/tools/phase3-files-numbers.ts`, `TMPDIR` = the job's scratch). Inputs frozen +once (`FREEZE_TO`) from the live corpus: 6 `site.json`, 71 `config.json`, 1,763 sidecars. Before = +the tool on `f4da04a9` code (branch changes stashed); after = the branch with the literal updated +(it asserts the literal equals `SITE_FIELD_DOCS`). **Unknown-key report: empty on both sides**, 77 +files each. **Diff: 18 lines, exactly one per site** — each parsed site gains +`"transcriptDownloads": true`; every `writeSite(getSite())` file is byte-identical (no site gains +the key on disk). A third run over a scratch copy with jeralyzer's `site.json` carrying +`transcriptDownloads: false` + `archives: false`: unknown keys still empty, both parse `false`, +and the written file keeps both keys. + +**Left.** Rollout item 5 (untick both on the five published sites, build + deploy each) is the +operator's, after merge. The hub question in finding 1. diff --git a/plans/site-exports-off.md b/plans/site-exports-off.md @@ -1,5 +1,10 @@ # Plan — visitor exports off on the published sites, kept as an option for the OSS release +> **Superseded by [`release-5.md`](release-5.md) slice X** (shipped as "### Slice X, as shipped" +> in its Record). Kept as written for the inventory; where the two differ, release-5.md wins — +> notably there is no `pnpm ops site-config` (the site form is the one writer), and the hub has no +> `site.json` of its own. + **Asked 2026-09-24** (operator, mid release 4): turn off "exports" on every existing published site to strengthen the legal footing of the operator's own instances, while keeping the capability as a per-site option so anyone running the OSS release can leave it on. Scheduled as diff --git a/plans/tools/phase3-files-numbers.ts b/plans/tools/phase3-files-numbers.ts @@ -54,13 +54,14 @@ const SCRATCH = fs.mkdtempSync(path.join(os.tmpdir(), "phase3-files-")); process.env.TRANSCRIPTS_DIR = SCRATCH; process.env.TZ = "UTC"; -// The keys each file may carry, as of 2026-09-24. On the branch these must +// The keys each file may carry, as of 2026-09-24 (+ transcriptDownloads, release +// 5). On the branch these must // equal the docs records (asserted below). const SITE_KEYS_LITERAL = [ "siteId", "siteTitle", "siteDescription", "headerTitle", "homeTagline", "accent", "socialLinks", "groups", "defaultGroupId", "channels", "cloudflareProject", "siteUrl", "relatedSites", "pwa", "archives", - "archiveMaxBytes", "duplicates", "hubUrl", + "archiveMaxBytes", "duplicates", "transcriptDownloads", "hubUrl", ]; const CHANNEL_KEYS_LITERAL = [ "handling", "sourceKind", "postFetcher", "socialHandle", "platform", "name",