commit 0d33a251c076ed4338fa98e4c539a3fc855e95ad
parent c5dcbf938e801884375de2f90ab117cc3ad89fce
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Thu, 20 Aug 2026 01:20:40 -0400
docs: the video path no longer needs a corpus, so stop saying it does
Both files described fetching cue windows over HTTP as the missing piece and
warned that report-to-video was on an unmerged branch. Both are now shipped and
merged, so the docs describe what the flags actually are (--site-origin,
--cue-source, --resolve-site-ids) instead of what someone would have to build.
Adds the two things that will otherwise bite: a published archive and a live
corpus can hold different cues for the same video once the corpus moves past its
last publish (measured: 65 of 84 cue texts rewritten, up to 2.24 s of drift), and
Rumble's two ids mean a locally-authored manifest misses on every Rumble clip
unless it says which id the archive uses.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Diffstat:
| M | AGENTS.md | | | 38 | +++++++++++++++++++++----------------- |
| M | README.md | | | 47 | ++++++++++++++++++++++++++++++----------------- |
2 files changed, 51 insertions(+), 34 deletions(-)
diff --git a/AGENTS.md b/AGENTS.md
@@ -65,23 +65,27 @@ seconds** rather than whole videos. Searching text first is what makes fetching
boundary is where the caption line wrapped, so cutting there ends mid-thought.
Neither a report nor an MCP snippet carries an end.
-The scripts currently read `transcript.cues.json` from local disk, so point
-`CHANNELS_DIR` at a channels directory holding the cited videos; both otherwise default
-it to an absolute path on the original author's machine.
-
-**That is a tooling limit, not a data one.** Published archives already serve
-`cues: [{ start, end, text }]` on their transcript pages — documented under
-`shardScheme` in `/corpus.json`, and verified 2026-08-20 against
-`https://jeralyzer.pages.dev/transcripts/chrissie-mayr/page-0000.json`
-(`{"start":8.12,"end":10.31,…}`). Teaching `resolve-windows.mjs` and `build-video.mjs`
-to fetch windows over HTTP — reusing the shard walk in `mcp/` — is what would let the
-whole video path run with **no local corpus at all**.
-
-Fetched clips are cached on disk under the report's `out/` (`clips-raw/`, `segments/`,
-`cards/`, `<slug>.mp4`) and reused on rebuild. The per-report `video.manifest.json` is
-the regeneration source of truth, not the report.
-
-As of 2026-08-20 this tooling is on the `feat/umtool-projects` branch, not on `main`.
+Cues resolve through `scripts/report-to-video/cues.mjs`: a local corpus when there is
+one (`CHANNELS_DIR`, now resolved relative to the repo), else the published archive the
+manifest names in `provenance.siteOrigin`. The shard walk is the contract in
+`/corpus.json`: corpus → channel transcripts manifest → `slugToPage` → `page-<NNNN>.json`
+(zero-padded to four — `page-0.json` is a 404) → the record whose `id` matches. Manifests
+and shards are cached in memory and on disk.
+
+**The two sources can disagree, and not by rounding.** An archive is a snapshot; a corpus
+keeps moving. Measured 2026-08-20 (local 2026-08-13 against a 2026-08-07 publish): three
+of four videos byte-identical, the fourth with 65 of 84 cue texts rewritten and timings
+shifted by up to **2.24 s**. `--cue-source auto|local|http` makes the choice explicit;
+`local` refuses to fall back rather than silently cut from other cues.
+
+**Rumble videos have two ids** — the archive keys by the EMBED id, a local cue dir is
+named for the URL SLUG — so a locally-authored manifest misses on every Rumble clip. It
+fails loudly; fix with a per-clip `siteVideo`/`siteChannel`, or `--resolve-site-ids` to
+scan the channel's shards (opt-in: a shard is up to 8 MB). Do **not** derive the id from
+`citeUrl`: a citeUrl may deliberately cite a different recording (a mirror that reads
+better), whose clock is not the same.
+
+This tooling is on `main` as of 2026-08-20.
# The corpus, and where the live sites are configured
diff --git a/README.md b/README.md
@@ -349,31 +349,44 @@ node scripts/report-to-video/resolve-windows.mjs <report>/video.manifest.json --
node scripts/report-to-video/build-video.mjs <report>/video.manifest.json
```
-**A note on clip boundaries, and what it means for running without a corpus.** Cutting
-on the raw cue span cuts mid-thought, because a cue boundary is just where the caption
-line wrapped. `resolve-windows.mjs` widens each clip outward to a whole sentence, and
-that needs cue **end** times — which neither a report nor an MCP snippet carries.
+**Clip boundaries, and why this runs without a corpus.** Cutting on the raw cue span
+cuts mid-thought, because a cue boundary is just where the caption line wrapped.
+`resolve-windows.mjs` widens each clip outward to a whole sentence, and that needs cue
+**end** times — which neither a report nor an MCP snippet carries.
-Today those come from `transcript.cues.json` on local disk (`CHANNELS_DIR`), so the
-polished path currently wants a corpus. **That is a tooling limit, not a data one:**
-published archives already serve end times, documented under `shardScheme` in
-`/corpus.json` and verifiably present —
+Those come from the archive itself. A published instance serves the same record a local
+corpus holds, documented under `shardScheme` in `/corpus.json`:
```bash
curl -s https://jeralyzer.pages.dev/transcripts/chrissie-mayr/page-0000.json | head -c 220
# [{"slug":…,"cues":[{"start":8.12,"end":10.31,"text":"squirrels move fast but that's just the"}…
```
-So teaching those two scripts to fetch windows over HTTP — reusing the shard walk the
-MCP server already implements — is what would make the whole video path run with **no
-local corpus at all**. Until then: crude clips (a cited second plus padding) work
-against a public instance right now with nothing but `yt-dlp`; sentence-accurate ones
-want `CHANNELS_DIR` pointed at the cited channels.
+So the scripts read cues from a local corpus when there is one and from the published
+archive when there is not — **no corpus required, and no configuration either**, since a
+manifest already records the archive it was built against (`provenance.siteOrigin`).
-> **Status: the video pipeline is not on `main`.** `scripts/report-to-video/` and the
-> `umtool/` bench (port 3050) that drives it live on the `feat/umtool-projects` branch
-> and have not been merged. Both scripts also default `CHANNELS_DIR` to an absolute
-> path on the original author's machine, so set it explicitly until that is fixed.
+| Flag | For |
+|---|---|
+| `--site-origin <url>` | Read cues from a specific archive, overriding the manifest. |
+| `--cue-source auto\|local\|http` | Which source to trust. `auto` is local-first. |
+| `--resolve-site-ids` | Recover from an id mismatch by scanning a channel's shards. Slow; see below. |
+
+> **The two sources can disagree, and not by rounding.** An archive is a snapshot; a
+> corpus keeps moving. Measured on this corpus — local six days newer than the publish —
+> three of four videos were byte-identical and the fourth had **65 of its 84 cue texts
+> rewritten, with timings shifted by up to 2.24 s**. That is enough to cut in the wrong
+> place, which is why the source is a flag rather than an implementation detail. Use
+> `--cue-source http` when you want the clip to match what a reader following the
+> citation will actually see, and `local` to refuse to fall back at all.
+
+> **Rumble videos have two ids.** The archive keys a recording by its **embed** id while
+> a local cue directory is named for the **URL slug**, so a manifest authored against
+> local directories misses on every Rumble clip. That fails with the diagnosis rather
+> than silently; fix it by adding `siteVideo` to the clip, or pass `--resolve-site-ids`
+> to find the record by scanning the channel's shards (each up to 8 MB, which is why it
+> is opt-in). A clip's `citeUrl` is deliberately *not* used for this — it may point at a
+> different recording on purpose, and that recording's clock is not the same one.
## How it fits together