commit a67e817dcdcf2ecd5a37f986424f14c750db53d7
parent 2a72cda40dce8339a22f2a0084e86209a3df3893
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Wed, 29 Jul 2026 11:50:33 -0400
Changelogs for the sweep controls, the threshold change, and the review path
Three user-visible changes this session: readers get many more duplicate
cross-uploads recognised (the 0.6 → 0.35 threshold change, which is a
publishing decision and belongs in the export log); operators get a corpus-wide
sweep they can start, leave, pause and watch; and the review path gets the
confirm buttons plus the failure evidence that previously existed nowhere.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Diffstat:
2 files changed, 4 insertions(+), 0 deletions(-)
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -1,6 +1,9 @@
# Changelog
## [Unreleased]
+- **A corpus-wide digest backfill can now be started, left alone, and watched.** The digest layer could generate, but only one channel at a time from that channel's own page — a full-archive pass meant 63 manual launches, and a server restart silently ended it with nothing to say so. There is now a **Start Digest Sweep** control on the dashboard that walks every channel in turn, **heaviest first by remaining audio-hours** (cost is audio, not videos: one VOD channel outweighs every duplicate mirror in the archive combined), and it **survives a restart** — the sweep is re-armed at boot the way the auto-download and auto-transcribe runners already were. It stores no work-list, so a resumed sweep re-does nothing: what still needs generating is re-derived from disk every time, which also means a transcript that finishes mid-sweep, or a duplicate cluster you confirm, is simply picked up on the next pass. A separate **Pause Digests** control holds a running sweep at zero without ending it (the pause flag has existed since the digest layer shipped and nothing could set it). **The digest lane now yields the GPU to transcription**: the two were deliberately on separate queues so they wouldn't serialise, whose unintended consequence was that the model and whisper competed for the same card — measured at 90 seconds per audio-hour against the 27 an idle machine managed. While transcription is working the digest lane steps aside and resumes when the card is free; it can be turned off in Settings. Coverage is now visible — a digest instrument on the dashboard with the corpus percentage (to two decimals, because rounding 0.13% up to 1% flatters an 80-day job), a **No digest** column on the channels table, and a per-channel count on the needs-work rows. Progress bars also **work during a regeneration** for the first time: they re-counted digest files from disk, and a regenerated digest is rewritten in place, so a job that was working sat at 0% for its whole run. Time-remaining estimates for digest work are now computed in **seconds per audio-hour** rather than by averaging videos, which for this archive is wrong by more than an order of magnitude between a VOD channel and a shorts channel. New `common/bin/digest-plan.ts` prices the whole backfill in audio-hours before you commit hardware to it.
+- **Duplicate clusters awaiting review can finally be confirmed or rejected.** A cluster whose evidence is only a shared title and runtime shares no AI digests and reaches no built site until a human confirms it — and the code to record that confirmation existed, complete, with **no way to reach it from anywhere in the app**. Roughly 166 clusters were therefore stuck permanently. `/actionable` duplicate cards now carry **Confirm** / **Not a duplicate** / **Undo**, badges that show what has already been decided, and confirming immediately shares the canonical copy's digest to any aligned member rather than making you wait for its next sweep.
+- **Digest passes that produced nothing now leave evidence.** When every chapter a model proposed was rejected by a guard — or when it proposed nothing at all — the generator deliberately wrote no digest, so the video would be retried. The side effect was that the **worst** outputs recorded their warnings nowhere but a job log that rotates after 30 days, exactly the videos a review pass most needs to find. Those failures are now recorded on the artifact, and they distinguish *"the model proposed nothing"* from *"the model proposed chapters and every one was rejected"* — which look identical from outside and need opposite fixes. A new **digest warnings** section on `/actionable` and a matching **Digest warnings** filter on the channel video list surface them.
- **You can now read what the digest layer produced, and correct it, from the video page.** The digest generator shipped with nowhere for a human to look at its output — the channel Digest stage is a queue-and-count card, and the per-video page had no digest reference at all. Each video page now carries an **AI digest** panel that leads with what might be *wrong*: the `warnings` recorded during generation (grouped by guard, with the first offending value verbatim), then the **provenance** of each section — engine, requested vs actual model, lane, prompt version, prompt variant, context hash, and how many chunks of how many came back usable — then the chapters and tags themselves. A **freshness badge** says whether the digest still matches what a regeneration would produce right now, and a stale section spells out what it *would* be replaced with; changing an engine, model, prompt version or prompt shape is exactly what makes it stale, so this is where you see that a config change has invalidated your corpus. A digest that was **shared from a duplicate cluster's canonical member** says so plainly, links to the video it came from, and shows the measured cue offset that made placing it here safe — a borrowed digest is never presented as native. From the panel you can regenerate this one video on either lane (live log, cancellable, the same job machinery as a channel sweep) and hand-correct chapters: retitle one, or untick it to suppress it without deleting it, so a regeneration that re-emits the same item cannot resurrect something you rejected. Corrections are written **only** to `ai-digest.overrides.json` and never to the generated file. Settings also gains the **Digest** section the rest of the app has been pointing at (local engine, sections to generate, timestamp mode, prompt-variant label, per-engine model / context window / temperature, and the metered lane with its spend cap and long-tail cutoff), and **Digest** now appears in the channel status-header badges. See `editor/app/channels/[slug]/videos/[id]/components/DigestPanel.tsx`, `editor/app/settings/components/{SettingsForm,DigestAppsField}.tsx`, and `editor/e2e/digest.spec.ts`.
- **Duplicate detection now runs over the whole archive, not just shorts.** It previously ran out of memory on a full-corpus pass and was left off. Two things were actually wrong, and both are fixed: candidate pairs were generated by pairing every short with every longer video (half a billion pairs), and transcript fingerprints for the entire corpus were held in memory at once. Detection now works one block of similar videos at a time — fingerprint, compare, discard — so a corpus-wide run finishes in minutes at ordinary memory. A new **`--blocking`** flag on `duplicate-shorts` chooses how candidates are proposed: `title` (default for a full-archive run — fast, finds cross-platform re-uploads that kept their name), `duration` (slower, but the only one that finds a *re-titled* mirror), or `both`. The choice only affects which pairs get *considered*; what counts as a duplicate is still decided by comparing the actual transcripts.
- **Videos that share a title and a runtime are now surfaced for review instead of being dropped.** When one side has no transcript there is nothing to compare, so detection can't rule either way. Rather than discarding the pair, it is reported as a cluster badged **needs review** — visible in the archive's own tooling, excluded from every built site, and blocked from sharing AI digests until you confirm it. Clusters whose transcripts *were* compared are unaffected and behave exactly as before. Note that this test is doing real work: on the full archive, **63% of same-title, same-length pairs turned out not to be the same video.**
diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md
@@ -1,6 +1,7 @@
# Changelog
## [Unreleased]
+- **The archive now recognises far more cross-platform re-uploads as the same video.** Duplicate detection compared two transcripts and called them the same only above a similarity of 0.6 — a threshold tuned for two transcripts of the same *text*, which quietly failed the case detection exists for. The two sides of a YouTube↔Rumble mirror are transcribed by **different speech-recognition engines**, and word-level disagreement between them lands a five-word-window comparison at roughly 0.35–0.60, i.e. just *under* the old cutoff. The threshold is now **0.35**, which takes the archive from 2,846 duplicate clusters over 5,746 videos to **7,434 clusters over 14,997 videos** — so a search result is far more likely to tell you the same recording exists elsewhere, and to offer you the jump. The change was bracketed at four values on the full archive before being made, and every step is a **strict superset**: no video that was previously flagged as a duplicate stopped being one. What it newly admits was inspected rather than counted — 96–97% have byte-identical titles, 99% span two platforms, and the handful of same-channel cases were read individually. Nothing about *how* a duplicate is decided changed: pairs are still confirmed by comparing real transcripts, a pair that merely shares a title and a runtime is still an internal review item that never reaches you, and the "jump to this moment in the other copy" button still only carries your timestamp when the two were *measured* as aligned.
- **AI chapters: jump straight to the part of a video you want.** Videos that have been through the local-AI digest pass now ship their derived **chapters and topic tags** to the site, and the player gains a third panel beside Transcript and Live chat. Open it and you get a titled list of moments — click one and the player **seeks there**; the chapter you're currently inside stays marked as the video plays. The layer is **sparse on purpose and honest about it**: only a small, growing fraction of the archive has been digested (generation is a multi-week GPU pass), so the control simply **isn't shown** on a video that has no digest, rather than offering a button that opens an empty panel — and if a digest can't be loaded, the panel snaps back to the transcript with a notice instead of stranding you. What ships is the **composed** digest: any human correction is applied and any chapter a human rejected is dropped, so you see what a person approved rather than raw model output, with hand-edited chapters marked **edited** and a provenance line naming the model that wrote the rest. **A digest borrowed from a duplicate upload says so, prominently** — when the same recording exists twice in the archive, one copy's chapters can be shared onto the other, and the panel names the source video and the measured timing offset rather than passing them off as native (plausible chapters describing a *different* upload is the failure that looks like success). Digests live at `/digests/<channel>/` under the same paginated-shard scheme as transcripts and posts, are offline-cached by the service worker, and are described in `corpus.json` — which bumps to **spec 3** with a `digestScheme` and per-channel `digests` manifest pointers, so an AI tool reading the corpus can navigate them and knows that an absent video means "not yet generated" rather than "nothing to say". Deep links carry the panel (`?vm=digest`), and share links reopen on it. Operator telemetry (why the model's proposals were rejected, the regeneration history) is deliberately **not** shipped — that stays in the editor. See `common/lib/digests.ts`, `common/components/{digestCache,digestStore}.ts`, `common/components/{PlayerProvider,TranscriptModal,urlState}.tsx/ts`, and `export/e2e/modal-digest.spec.ts`.
- **Search results tell you when a video exists elsewhere in the archive, and take you there.** A result that belongs to a duplicate cluster now carries a **Dupe** badge, and a strip under the card header offers one button per other copy — the same recording mirrored to another platform, or re-uploaded on another channel. Clicking one opens that copy in the player. **The jump is honest about what it knows:** matching content does not imply matching timings (a mirror with a longer intro carries the same words at shifted times), so a button only carries your current timestamp when detection *measured* the two as aligned; otherwise it says so and opens the other copy from the start. A copy whose alignment was never measured is treated as not aligned. Only clusters whose transcripts were actually compared reach the site — a pair that merely shares a title and a runtime stays an internal review item and is never asserted to you. In hub mode the badge is limited to same-origin results, since the duplicate index is per-site. Sites with no duplicate report are entirely unaffected.
- **One search now covers video transcripts *and* social posts.** Archived X/Twitter and Bluesky posts ship as a parallel corpus beside transcripts and live chat, and compose into the same boolean query tree — so `(transcripts:"foo" OR posts:"foo")` returns both kinds in one ranked, newest-first result set. `LayerScope` gains `"posts"` (whitelisted in `qt=` deserialization, so a shared link round-trips a posts leaf), the leaf scope selector gains **Posts**, and the filter row gains a **Posts** media kind beside Videos and Livestreams — a third kind, because a post is neither, and folding it into the video toggle would silently drop the whole corpus. Post and video slugs live in disjoint namespaces, partitioned per-leaf by the eval engine so a transcripts leaf never fetches a post and an AND across the two can't collapse to nothing. Post result cards drop what doesn't apply (no seek gutter, no livestream/age badges, no VOD expiry) and lead with the post body; opening one shows a new **PostModal** — a sibling of the transcript reader, not a generalization of it — with the post, its archived thread, its outbound links and its engagement counts. Date filters work unchanged: every post carries a derived `uploadDate`. Posts are cached and served under `/posts/`, offline-cached by the service worker, and CORS-readable so a federating hub merges them across origins.