commit add512464a999f38383331c608f2d7b79f0e5e6b
parent 1de3212e41b9c81ff6089008f51156e39c813f10
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Tue, 4 Aug 2026 12:42:27 -0400
Release export 0.8.3
Diffstat:
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md
@@ -1,6 +1,6 @@
# Changelog
-## [Unreleased]
+## [0.8.3] - 2026-08-04
- **The archive now recognises far more cross-platform re-uploads as the same video.** Duplicate detection compared two transcripts and called them the same only above a similarity of 0.6 — a threshold tuned for two transcripts of the same *text*, which quietly failed the case detection exists for. The two sides of a YouTube↔Rumble mirror are transcribed by **different speech-recognition engines**, and word-level disagreement between them lands a five-word-window comparison at roughly 0.35–0.60, i.e. just *under* the old cutoff. The threshold is now **0.35**, which takes the archive from 2,846 duplicate clusters over 5,746 videos to **7,434 clusters over 14,997 videos** — so a search result is far more likely to tell you the same recording exists elsewhere, and to offer you the jump. The change was bracketed at four values on the full archive before being made, and every step is a **strict superset**: no video that was previously flagged as a duplicate stopped being one. What it newly admits was inspected rather than counted — 96–97% have byte-identical titles, 99% span two platforms, and the handful of same-channel cases were read individually. Nothing about *how* a duplicate is decided changed: pairs are still confirmed by comparing real transcripts, a pair that merely shares a title and a runtime is still an internal review item that never reaches you, and the "jump to this moment in the other copy" button still only carries your timestamp when the two were *measured* as aligned.
- **AI chapters: jump straight to the part of a video you want.** Videos that have been through the local-AI digest pass now ship their derived **chapters and topic tags** to the site, and the player gains a third panel beside Transcript and Live chat. Open it and you get a titled list of moments — click one and the player **seeks there**; the chapter you're currently inside stays marked as the video plays. The layer is **sparse on purpose and honest about it**: only a small, growing fraction of the archive has been digested (generation is a multi-week GPU pass), so the control simply **isn't shown** on a video that has no digest, rather than offering a button that opens an empty panel — and if a digest can't be loaded, the panel snaps back to the transcript with a notice instead of stranding you. What ships is the **composed** digest: any human correction is applied and any chapter a human rejected is dropped, so you see what a person approved rather than raw model output, with hand-edited chapters marked **edited** and a provenance line naming the model that wrote the rest. **A digest borrowed from a duplicate upload says so, prominently** — when the same recording exists twice in the archive, one copy's chapters can be shared onto the other, and the panel names the source video and the measured timing offset rather than passing them off as native (plausible chapters describing a *different* upload is the failure that looks like success). Digests live at `/digests/<channel>/` under the same paginated-shard scheme as transcripts and posts, are offline-cached by the service worker, and are described in `corpus.json` — which bumps to **spec 3** with a `digestScheme` and per-channel `digests` manifest pointers, so an AI tool reading the corpus can navigate them and knows that an absent video means "not yet generated" rather than "nothing to say". Deep links carry the panel (`?vm=digest`), and share links reopen on it. Operator telemetry (why the model's proposals were rejected, the regeneration history) is deliberately **not** shipped — that stays in the editor. See `common/lib/digests.ts`, `common/components/{digestCache,digestStore}.ts`, `common/components/{PlayerProvider,TranscriptModal,urlState}.tsx/ts`, and `export/e2e/modal-digest.spec.ts`.
- **Search results tell you when a video exists elsewhere in the archive, and take you there.** A result that belongs to a duplicate cluster now carries a **Dupe** badge, and a strip under the card header offers one button per other copy — the same recording mirrored to another platform, or re-uploaded on another channel. Clicking one opens that copy in the player. **The jump is honest about what it knows:** matching content does not imply matching timings (a mirror with a longer intro carries the same words at shifted times), so a button only carries your current timestamp when detection *measured* the two as aligned; otherwise it says so and opens the other copy from the start. A copy whose alignment was never measured is treated as not aligned. Only clusters whose transcripts were actually compared reach the site — a pair that merely shares a title and a runtime stays an internal review item and is never asserted to you. In hub mode the badge is limited to same-origin results, since the duplicate index is per-site. Sites with no duplicate report are entirely unaffected.