Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit bfbdcbb85bdcf3fa2a1cfdce260456b3ed29b3a8
parent 161912b93d84f9077b435b2f5e8a130e1f7344bf
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Tue,  6 Oct 2026 10:00:40 -0400

docs + export e2e: alternate tracks

Changelog [Unreleased] (editor, export), README and PUBLISH say what an
alternate track is and what the first index build does with them. The export
fixture's transcript-only video carries an uploaded `en` that says a word its
primary never does; transcript-tracks.spec.ts switches to it in the reader,
checks the share link carries it, and finds the word by search, named and
opened on its track.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
MPUBLISH.md | 10++++++++++
MREADME.md | 5++++-
Meditor/CHANGELOG.md | 5++++-
Mexport/CHANGELOG.md | 1+
Mexport/e2e/fixtures/data.ts | 10++++++++++
Aexport/e2e/transcript-tracks.spec.ts | 72++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
6 files changed, 101 insertions(+), 2 deletions(-)

diff --git a/PUBLISH.md b/PUBLISH.md @@ -38,6 +38,16 @@ the stats datasets and the chart templates — `archilyzer index`, `build stats` `export/public`, plus its download archives — `archilyzer compose site <id>`) and `next build`. `--nodata` skips the data phase and reuses the last one's staging. +A transcript record in the shared pages carries its other English caption tracks +(`altTracks`, with the transcript's own `track`) only where one's words differ from +the transcript's — the served `en` beside `en-orig`, a regional or auto-translated +track, the captions a local transcription replaced (`common/lib/captionTracks.ts`). +Identical tracks add nothing, so most records' bytes are what they were. Search reads +those tracks with the transcript and names the track of a hit only one holds; the +reader switches to them. English VTTs are not published as subtitle tracks. The first +index build after this reads them once (`Alternate tracks v1: N record(s) re-read.`), +for exactly the records that can hold one. + ## The three ways to drive it The editor's **/sites** page, `pnpm ops` (HTTP to a running editor, with its diff --git a/README.md b/README.md @@ -385,7 +385,10 @@ differs is emphasis: where you would rather have the better text. Where YouTube serves both, the original-audio captions (`en-orig`) are read before the served `en` track, which can reword what was said; a track with no text falls through to the next, and a video's - page can pin another (`common/lib/videoStatus.ts`, "the caption-track rule"). + page can pin another (`common/lib/videoStatus.ts`, "the caption-track rule"). The + other English tracks are kept where their words differ — uploaded captions are not + always what was said: search reads them too and says which track a hit is in, and + the transcript reader switches to them (`common/lib/captionTracks.ts`). - **The output is yours to brand.** Site title, header, description, tagline, social links and channel grouping are all per-site configuration. - **One corpus can publish several sites.** Channels are grouped into sites, so a diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,9 +1,12 @@ # Changelog ## [Unreleased] +- **A video's other English tracks are readable and searchable where their words differ.** Uploaded captions are not always a transcript of what was said, so the tracks beside the transcript stay: the served `en` beside `en-orig`, a regional or auto-translated track, and the captions a local transcription replaced. One is kept where its words differ from the transcript's and from every track kept before it; identical tracks, most of them, add nothing. The index keeps them in an `alts` sub-DB and writes `track` and `altTracks` onto the transcript record only then, so every other record's page is what it was. A search hit in a word only an alternate holds names the track; one every track says is found once, in the transcript. The video page's **Transcript** card reads the transcript and switches tracks ("Track: original audio captions ▾"); switching changes nothing on disk, and **Set as transcript** stays the way the transcript itself changes. English VTTs are no longer shipped as subtitle tracks. One notion of a track — ids, plain labels, which are kept, how a hit across them is found — lives in `common/lib/captionTracks.ts`. +- **The next index build reads the alternate tracks once.** Every record that can hold one — two or more English VTTs, or a transcription beside captions — is re-read from disk, and nothing else; the log says `Alternate tracks v1: N record(s) re-read.` and how many hold a track whose words differ. The version is recorded only when no channel is held. A transcribed video's captions now count toward its change time, so a later caption fetch reaches the index. +- **The MCP reads every English track.** `search_transcripts` and the query-tree tools match a record's alternate tracks and tag a snippet from one (`[in uploaded captions 1:30]`); `get_transcript` names a video's tracks in its header and reads another with `track`; `get_transcripts` windows a match only an alternate holds, under its name; `get_video_metadata` lists the other tracks without their cues. The sweep plan says what such a hit is before it is quoted. - **A Wayback Machine capture is a copy, and says of what.** A capture URL (`web.archive.org/web/<timestamp>[id_|im_|…]/<original>`) names its record by what it is a capture of: an archived YouTube page by its YouTube id (no longer `watch`), a JW Player file by its media id (no longer `<id>-<rendition>.mp4`). Every download of a capture writes `wayback.json` (the original URL, the capture's timestamp, the capture page and its raw bytes); the video page says "Archived copy (Wayback Machine, <date>) of <original>"; a citation links the original, marked as possibly gone, and the Wayback copy, and its moment link is the capture, which plays (a capture URL never takes a time param). An existing record whose page is a capture is renamed to its id by the next snapshot. - **`archilyzer wayback refresh <slug> [--titles <file>] [--dry-run]`** brings a channel's Wayback copies up to that offline: `wayback.json`, the dir renamed through the snapshot's own reconcile pass with its roster entry moved, and with `--titles` (`id → {title, upload_date}`) the title and date of a raw file that has none, recorded in the metadata history as `wayback-provenance`. A record a live job holds is skipped and named; a second run changes nothing. -- **A video's captions are read from its original-audio track first, and a track with no text never hides one that has it.** Where YouTube serves both, `transcript.en-orig.vtt` (the captions of the original audio) is read before `transcript.en.vtt`, whose text can be a rewrite of what was said; then regional tracks (`en-US`, `en-GB`, …), then auto-translated `en-en-*` ones. The transcript is the first track in that order that has cues. One rule (`englishVttsByPreference` / `readEnglishVttCues` in `common/lib/videoStatus.ts`) serves the index, normalize and report compose, so search, the export, the MCP and report videos read the same words. **Set as transcript** on a video's page copies the chosen track to `transcript.en.vtt` and pins it there with `transcript-pin.json`, which ranks it first; deleting `transcript-pin.json` returns the video to the automatic pick. The served `en` track is listed as an alternate subtitle track where `en-orig` is the transcript. Videos with a human-made `en` track are not counted as auto-captions-only, so the replace-auto-captions lane still leaves them alone. +- **A video's captions are read from its original-audio track first, and a track with no text never hides one that has it.** Where YouTube serves both, `transcript.en-orig.vtt` (the captions of the original audio) is read before `transcript.en.vtt`, whose text can be a rewrite of what was said; then regional tracks (`en-US`, `en-GB`, …), then auto-translated `en-en-*` ones. The transcript is the first track in that order that has cues. One rule (`englishVttsByPreference` / `readEnglishVttCues` in `common/lib/videoStatus.ts`) serves the index, normalize and report compose, so search, the export, the MCP and report videos read the same words. **Set as transcript** on a video's page copies the chosen track to `transcript.en.vtt` and pins it there with `transcript-pin.json`, which ranks it first; deleting `transcript-pin.json` returns the video to the automatic pick. The served `en` track stays readable and searchable as an alternate track where its words differ (below). Videos with a human-made `en` track are not counted as auto-captions-only, so the replace-auto-captions lane still leaves them alone. - **Captions in cue blocks are read.** A VTT with no inline word timing (uploaded captions, and the `en` track YouTube serves for some livestream recordings: two lines a cue, `&nbsp;` at each line end) is read cue by cue; it used to parse to no cues, which left those videos with no text in the index. - **The next index build re-reads the caption records the new rule reaches, once.** A record with more than one English track, or with no cues stored, is re-read from disk; nothing else is. The build log says `Caption track v1: N record(s) re-read.` and, per channel, how many now read different text and how many had none and now do. The version is recorded only when no channel is held. A `transcript.cues.json` written before the rule is stale only where the rule reads something else — its `transcript.en.vtt` and `transcript.en-orig.vtt` differ, or it holds no cues and a cue-block parse or the next track may have them — so a **Normalize** run rewrites exactly those, and the digest and attribution lanes hold them until it does; a cues file now records the track it was read from (`vttFile`) and the rule (`captionTrackRule`). umtool's report-to-video refuses such a stale local record by name rather than cut from it. - **A cited moment at the very end of a recording prepares.** Prepare evidence media cuts a clip whose padding runs past the recording's end at the end (the recording's duration from its metadata), where it found no media for the padded span; a span that starts past the end is still refused. report-to-video keeps its strict rule. diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md @@ -1,6 +1,7 @@ # Changelog ## [Unreleased] +- **Search reads every English track of a video, and the transcript switches tracks.** Where a video has another English caption track whose words differ from its transcript — the uploaded captions beside the original audio's, a regional or auto-translated track — a query matches it too: a hit only that track holds says so ("in uploaded captions") and opens the transcript on that track at that moment, and a word both say is found once, in the transcript. The transcript reader shows a small "Track:" switcher beside the mode buttons on such a video; the transcript stays the default, and the choice rides on the share link (`vt`). Downloads and Copy MD take the track on show. Needs an index build and a rebuild and deploy of each site. - **A citation of a Wayback Machine copy links its original and the copy.** A cited record downloaded from a Wayback capture shows "Original (may be gone)", the original at the cited second where its platform takes one, and "Wayback Machine copy, <capture date>", the capture page, which plays. Its moment link is the capture: a capture URL never takes a time param. - **Transcripts read the original-audio captions.** Where a video has both, its transcript is YouTube's `en-orig` track (the captions of what was said) rather than the served `en`, which can reword it; a track with no text falls through to the next. Videos whose only captions are in cue blocks (some livestream recordings) have their text. - **A report shows its revision, and every edit to it can be checked.** A report's date line ends with "revision N", linking to its history ("edited since revision N" when the report has changed since). The history page, `/reports/<id>/history/`, lists every revision, newest first: its number, date (UTC), commit hash and the sha256 of its `report.json`, what changed (claims added or removed, verdicts changed, claims edited, citations added or removed, quotes edited, title, series or subtitle changed), and each changed claim's title, text, verdict and findings with the words removed struck through and the words added marked. The same data is in `history.json` beside the page. Each report's history is its own git repository, published for cloning: `git clone <site>/reports/<id>/history/repo`. Each commit names the site as its author, with its date in UTC. The footer of the report's HTML, PDF and Markdown downloads begins with the revision number, and its sha256 can be looked up on the history page. Needs `reports export` and a rebuild and deploy of each site with reports. diff --git a/export/e2e/fixtures/data.ts b/export/e2e/fixtures/data.ts @@ -292,6 +292,16 @@ export function transcriptPage() { description: "Filmed on location with a zebra in the background.", tags: ["news"], cues: transcriptCues("transcript-only video"), + // An ALTERNATE English track (lib/captionTracks.ts): the uploaded + // captions, whose words differ from the primary's. "zeppelin" is said + // only here, at 200 s — what transcript-tracks.spec.ts searches for. + track: "en-orig", + altTracks: [ + { + track: "en", + cues: [{ start: 200, end: 204, text: "uploaded words about a zeppelin" }], + }, + ], }, { ...makeSummary(VIDEO_CHAT_SMALL, "Small live chat"), diff --git a/export/e2e/transcript-tracks.spec.ts b/export/e2e/transcript-tracks.spec.ts @@ -0,0 +1,72 @@ +import { expect, test, type Page } from "@playwright/test"; +import { CHANNEL_SLUG, VIDEO_TRANSCRIPT_ONLY } from "./fixtures/data"; +import { expectModalOpen, installRoutes } from "./helpers"; + +// A record's other English tracks (lib/captionTracks.ts). VIDEO_TRANSCRIPT_ONLY +// carries an en-orig primary and an uploaded `en` that says "zeppelin" at +// 200 s, which its primary never does; the other fixture videos have no +// alternate. + +test.use({ + permissions: ["clipboard-read", "clipboard-write"], +}); + +const SLUG = `${CHANNEL_SLUG}/${VIDEO_TRANSCRIPT_ONLY}`; +const leafInput = (page: Page) => + page.locator('input[data-testid^="leaf-query-"]').first(); + +test.describe("transcript tracks", () => { + test.beforeEach(async ({ page }) => { + await installRoutes(page); + }); + + test("the reader shows the primary and switches to the uploaded captions", async ({ page }) => { + await page.goto(`/?v=${SLUG}`); + await expectModalOpen(page); + const cues = page.getByTestId("cue-list"); + await expect(cues).toContainText("transcript-only video — alpha line"); + + const switcher = page.getByTestId("track-switcher"); + await expect(switcher).toHaveValue("en-orig"); + await expect(switcher.locator("option")).toHaveText([ + "original audio captions (default)", + "uploaded captions", + ]); + await switcher.selectOption("en"); + await expect(cues).toContainText("uploaded words about a zeppelin"); + await expect(cues).not.toContainText("alpha line"); + expect(new URL(page.url()).searchParams.get("vt")).toBe("en"); + + // The share link reopens this track; back on the primary, it is off the URL. + await page.getByRole("button", { name: "Copy share link at current time" }).click(); + const clip = new URL(await page.evaluate(() => navigator.clipboard.readText())); + expect(clip.searchParams.get("vt")).toBe("en"); + await switcher.selectOption("en-orig"); + await expect(cues).toContainText("alpha line"); + expect(new URL(page.url()).searchParams.get("vt")).toBeNull(); + }); + + test("a record with no alternate shows no switcher", async ({ page }) => { + await page.goto(`/?v=${CHANNEL_SLUG}/vid-chat-small`); + await expectModalOpen(page); + await expect(page.getByTestId("cue-list")).toContainText("small chat video"); + await expect(page.getByTestId("track-switcher")).toHaveCount(0); + }); + + test("search finds a word only the uploaded captions hold, says so, and opens that track", async ({ page }) => { + await page.goto("/"); + await leafInput(page).fill("zeppelin"); + await page.getByTestId("search-submit").click(); + + const card = page.locator(`[data-result-slug="${SLUG}"]`); + await expect(card).toBeVisible({ timeout: 15_000 }); + await expect(page.locator("[data-card-header]")).toHaveCount(1); + const hit = page.getByRole("button", { name: /in uploaded captions/i }).first(); + await expect(hit).toContainText("zeppelin"); + await hit.click(); + + await expectModalOpen(page); + await expect(page.getByTestId("track-switcher")).toHaveValue("en"); + await expect(page.getByTestId("cue-list")).toContainText("uploaded words about a zeppelin"); + }); +});