Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 0d33a251c076ed4338fa98e4c539a3fc855e95ad
parent c5dcbf938e801884375de2f90ab117cc3ad89fce
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Thu, 20 Aug 2026 01:20:40 -0400

docs: the video path no longer needs a corpus, so stop saying it does

Both files described fetching cue windows over HTTP as the missing piece and
warned that report-to-video was on an unmerged branch. Both are now shipped and
merged, so the docs describe what the flags actually are (--site-origin,
--cue-source, --resolve-site-ids) instead of what someone would have to build.

Adds the two things that will otherwise bite: a published archive and a live
corpus can hold different cues for the same video once the corpus moves past its
last publish (measured: 65 of 84 cue texts rewritten, up to 2.24 s of drift), and
Rumble's two ids mean a locally-authored manifest misses on every Rumble clip
unless it says which id the archive uses.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Diffstat:
MAGENTS.md | 38+++++++++++++++++++++-----------------
MREADME.md | 47++++++++++++++++++++++++++++++-----------------
2 files changed, 51 insertions(+), 34 deletions(-)

diff --git a/AGENTS.md b/AGENTS.md @@ -65,23 +65,27 @@ seconds** rather than whole videos. Searching text first is what makes fetching boundary is where the caption line wrapped, so cutting there ends mid-thought. Neither a report nor an MCP snippet carries an end. -The scripts currently read `transcript.cues.json` from local disk, so point -`CHANNELS_DIR` at a channels directory holding the cited videos; both otherwise default -it to an absolute path on the original author's machine. - -**That is a tooling limit, not a data one.** Published archives already serve -`cues: [{ start, end, text }]` on their transcript pages — documented under -`shardScheme` in `/corpus.json`, and verified 2026-08-20 against -`https://jeralyzer.pages.dev/transcripts/chrissie-mayr/page-0000.json` -(`{"start":8.12,"end":10.31,…}`). Teaching `resolve-windows.mjs` and `build-video.mjs` -to fetch windows over HTTP — reusing the shard walk in `mcp/` — is what would let the -whole video path run with **no local corpus at all**. - -Fetched clips are cached on disk under the report's `out/` (`clips-raw/`, `segments/`, -`cards/`, `<slug>.mp4`) and reused on rebuild. The per-report `video.manifest.json` is -the regeneration source of truth, not the report. - -As of 2026-08-20 this tooling is on the `feat/umtool-projects` branch, not on `main`. +Cues resolve through `scripts/report-to-video/cues.mjs`: a local corpus when there is +one (`CHANNELS_DIR`, now resolved relative to the repo), else the published archive the +manifest names in `provenance.siteOrigin`. The shard walk is the contract in +`/corpus.json`: corpus → channel transcripts manifest → `slugToPage` → `page-<NNNN>.json` +(zero-padded to four — `page-0.json` is a 404) → the record whose `id` matches. Manifests +and shards are cached in memory and on disk. + +**The two sources can disagree, and not by rounding.** An archive is a snapshot; a corpus +keeps moving. Measured 2026-08-20 (local 2026-08-13 against a 2026-08-07 publish): three +of four videos byte-identical, the fourth with 65 of 84 cue texts rewritten and timings +shifted by up to **2.24 s**. `--cue-source auto|local|http` makes the choice explicit; +`local` refuses to fall back rather than silently cut from other cues. + +**Rumble videos have two ids** — the archive keys by the EMBED id, a local cue dir is +named for the URL SLUG — so a locally-authored manifest misses on every Rumble clip. It +fails loudly; fix with a per-clip `siteVideo`/`siteChannel`, or `--resolve-site-ids` to +scan the channel's shards (opt-in: a shard is up to 8 MB). Do **not** derive the id from +`citeUrl`: a citeUrl may deliberately cite a different recording (a mirror that reads +better), whose clock is not the same. + +This tooling is on `main` as of 2026-08-20. # The corpus, and where the live sites are configured diff --git a/README.md b/README.md @@ -349,31 +349,44 @@ node scripts/report-to-video/resolve-windows.mjs <report>/video.manifest.json -- node scripts/report-to-video/build-video.mjs <report>/video.manifest.json ``` -**A note on clip boundaries, and what it means for running without a corpus.** Cutting -on the raw cue span cuts mid-thought, because a cue boundary is just where the caption -line wrapped. `resolve-windows.mjs` widens each clip outward to a whole sentence, and -that needs cue **end** times — which neither a report nor an MCP snippet carries. +**Clip boundaries, and why this runs without a corpus.** Cutting on the raw cue span +cuts mid-thought, because a cue boundary is just where the caption line wrapped. +`resolve-windows.mjs` widens each clip outward to a whole sentence, and that needs cue +**end** times — which neither a report nor an MCP snippet carries. -Today those come from `transcript.cues.json` on local disk (`CHANNELS_DIR`), so the -polished path currently wants a corpus. **That is a tooling limit, not a data one:** -published archives already serve end times, documented under `shardScheme` in -`/corpus.json` and verifiably present — +Those come from the archive itself. A published instance serves the same record a local +corpus holds, documented under `shardScheme` in `/corpus.json`: ```bash curl -s https://jeralyzer.pages.dev/transcripts/chrissie-mayr/page-0000.json | head -c 220 # [{"slug":…,"cues":[{"start":8.12,"end":10.31,"text":"squirrels move fast but that's just the"}… ``` -So teaching those two scripts to fetch windows over HTTP — reusing the shard walk the -MCP server already implements — is what would make the whole video path run with **no -local corpus at all**. Until then: crude clips (a cited second plus padding) work -against a public instance right now with nothing but `yt-dlp`; sentence-accurate ones -want `CHANNELS_DIR` pointed at the cited channels. +So the scripts read cues from a local corpus when there is one and from the published +archive when there is not — **no corpus required, and no configuration either**, since a +manifest already records the archive it was built against (`provenance.siteOrigin`). -> **Status: the video pipeline is not on `main`.** `scripts/report-to-video/` and the -> `umtool/` bench (port 3050) that drives it live on the `feat/umtool-projects` branch -> and have not been merged. Both scripts also default `CHANNELS_DIR` to an absolute -> path on the original author's machine, so set it explicitly until that is fixed. +| Flag | For | +|---|---| +| `--site-origin <url>` | Read cues from a specific archive, overriding the manifest. | +| `--cue-source auto\|local\|http` | Which source to trust. `auto` is local-first. | +| `--resolve-site-ids` | Recover from an id mismatch by scanning a channel's shards. Slow; see below. | + +> **The two sources can disagree, and not by rounding.** An archive is a snapshot; a +> corpus keeps moving. Measured on this corpus — local six days newer than the publish — +> three of four videos were byte-identical and the fourth had **65 of its 84 cue texts +> rewritten, with timings shifted by up to 2.24 s**. That is enough to cut in the wrong +> place, which is why the source is a flag rather than an implementation detail. Use +> `--cue-source http` when you want the clip to match what a reader following the +> citation will actually see, and `local` to refuse to fall back at all. + +> **Rumble videos have two ids.** The archive keys a recording by its **embed** id while +> a local cue directory is named for the **URL slug**, so a manifest authored against +> local directories misses on every Rumble clip. That fails with the diagnosis rather +> than silently; fix it by adding `siteVideo` to the clip, or pass `--resolve-site-ids` +> to find the record by scanning the channel's shards (each up to 8 MB, which is why it +> is opt-in). A clip's `citeUrl` is deliberately *not* used for this — it may point at a +> different recording on purpose, and that recording's clock is not the same one. ## How it fits together