Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 1476904975d1d186034e0f0c6a84bce9279d3adb
parent 01c3073598b4a553981ac11ad8c6daf0c6986789
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Mon, 27 Jul 2026 00:09:25 -0400

Corpus-wide duplicate detection: blocking, review queue, alignment

Detection previously OOMed on a full-corpus pass and was left off. Two
things were actually wrong and both are fixed: candidates were generated
by pairing every short with every longer video (half a billion pairs), and
transcript fingerprints for the whole corpus were held in memory at once.
Detection now works one block at a time — fingerprint, compare, discard —
so a corpus-wide run finishes in minutes at ordinary memory.

- --blocking {title,duration,both} on duplicate-shorts chooses how pairs
  are PROPOSED. It never decides what counts as a duplicate; that is still
  the transcript comparison. title is the default for a full-archive run
  (fast, catches cross-platform re-uploads that kept their name);
  duration is the only one that finds a re-titled mirror.
- A pair whose sides share a title and a runtime but where one side has no
  transcript can't be ruled either way, so it is no longer discarded: it
  becomes a needsReview cluster, visible in the editor, excluded from
  every built site, and blocked from sharing a digest until confirmed.
  This test is doing real work — on the full archive 63% of same-title,
  same-length pairs turned out NOT to be the same video.
- Confirmed clusters now persist measured per-member alignment against the
  canonical video. Matching content does not imply matching timings, so
  this is what lets the viewer carry a timestamp into another copy only
  when the two were actually measured as aligned, and what lets a shared
  digest be placed correctly rather than plausibly (digestSharing.ts).
- Viewer: a Dupe badge on search results plus a per-copy jump strip that
  says so when it cannot carry your timestamp. Hub mode limits the badge
  to same-origin results, since the duplicate index is per-site.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Diffstat:
MPLAN.md | 52++++++++++++++++++++++++++++++++++++++++++++++++----
Mcommon/bin/compose-site.ts | 45+++++++++++++++++++++++++++++++++++++++++----
Mcommon/bin/duplicate-shorts.ts | 27+++++++++++++++++++++++----
Mcommon/components/SearchResults.tsx | 114+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/components/badges.tsx | 20++++++++++++++++++++
Acommon/components/duplicatesCache.ts | 83+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/controller/digestSharing.ts | 10++++++++--
Mcommon/controller/duplicateShorts.ts | 781+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++------------
Mcommon/lib/duplicates.test.ts | 291++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/lib/duplicates.ts | 165+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++----
Meditor/CHANGELOG.md | 3+++
Meditor/app/actionable/page.tsx | 5+++++
Meditor/e2e/duplicate-shorts.spec.ts | 129++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-----
Mexport/CHANGELOG.md | 1+
Mexport/app/duplicates/DuplicatesClient.tsx | 9+++++++++
Mexport/e2e/helpers.ts | 11+++++++++++
Aexport/e2e/search-duplicates.spec.ts | 199+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mplans/FACTS.md | 279+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++--
Mplans/STATE.md | 264+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++--------------------
19 files changed, 2271 insertions(+), 217 deletions(-)

diff --git a/PLAN.md b/PLAN.md @@ -274,6 +274,14 @@ bundled inside a payload that policy may withhold. Digests get their own page tr bump and existing transcript caches survive. - `corpus.ts` — `CORPUS_SPEC_VERSION` 2 → 3, a `DIGEST_SCHEME` modelled on `POST_SCHEME`, and a digest manifest pointer on `CorpusChannel.manifests`. Update `corpus.test.ts`. +- **Carry `derivedFrom` through to the client.** *(Added by the duplicate-detection work.)* + Once digests are shared across a duplicate cluster, a video's digest may have been + generated for a *different* video. `DigestRecord.derivedFrom` already records that + (`{slug, clusterId, sharedAt, offsetSeconds}`, written by `writeSharedDigest`), but it + stops at the record — the page-tree schema above must include it, or the viewer cannot + tell a native digest from a borrowed one. Sharing that is invisible is sharing that is + indistinguishable from a claim, which Phase 3 then has no way to be honest about. This is + a dependency Phase 2 creates, not an optional extra. ### Phase 2.5 — Observability for the backfill @@ -314,6 +322,14 @@ tune or safely interrupt. use hard-coded `bg-zinc-900/80`, `ring-white/10`, `text-white` throughout. Match that local palette; do not introduce `bg-card`/`text-foreground` here. Migrating the modal chrome is out of scope. +- **Show a borrowed digest as borrowed.** *(Added by the duplicate-detection work.)* When + `derivedFrom` is set (Phase 2), the provenance line must name the member the digest was + generated *for* and the measured offset — not present it as native to the video being + watched. The chapters are placed by the canonical member's timeline; the detector's + `aligned` verdict is why sharing was allowed at all, and its `offsetSeconds` is the + residual error the viewer is looking at. A shared digest presented as native is the + failure mode that looks like success: every chapter plausible, all of them describing a + different upload. ### Phase 4 — Search indexing @@ -441,7 +457,20 @@ fourth meaning of "report". rather than building a parallel one: its `SectionConfig` array with counters in `loadActionable.ts` is the designed extension point. New sections: digests needing review (driven by the `warnings` array Phase 1 persists), proposed channel context awaiting -promotion, uncertain attribution, viewer feedback. +promotion, uncertain attribution, viewer feedback, and **duplicate clusters awaiting +confirmation**. + +**Duplicate clusters awaiting confirmation** *(added by the duplicate-detection work — UI +only, the data and the write path already exist).* A `title-duration` cluster is a suspect: +two videos share a title and a near-identical runtime, and nothing compared their content +because at least one side has no transcript. It ships nowhere (`clusterIsPublishable`) and +shares no derived work (`clusterMaySharePartial`) until a human records `confirmed: true`. +Corpus-wide that is currently **335 clusters**. Everything needed is in place: +- the queue is `report.clusters.filter(c => c.needsReview)` minus `isClusterReviewed(...)`; +- `updateDuplicateOverride(paths, clusterId, {confirmed})` already accepts and persists it, + and its clear-heuristic already refuses to delete a confirmation; +- `/actionable` already renders cluster cards with a `needs review` badge. +The missing piece is two buttons — *these are the same* / *these are not* — on that card. Actions per item: approve · edit inline · regenerate · dismiss · **"add note to channel context"**. The last is the one that compounds — a correction applied to one video fixes one @@ -452,9 +481,24 @@ when one engine is systematically weaker on a given channel. **11b — viewer feedback (can follow the corpus going public).** A "flag this" control per chapter, with categories including **"misrepresents what was said"**. Transport in preference -order: an optional per-site `feedbackUrl` POST, else copy-to-clipboard / download-JSON. The -existing `r2-proxy/` package is a plausible minimal collector — evaluate it before standing -up new infrastructure. Ingestion is an editor action accepting pasted JSON. Treat submissions +order: an optional per-site `feedbackUrl` POST, else copy-to-clipboard / download-JSON. + +**Collector evaluated — the answer is no, not `r2-proxy/`.** *(Recorded by the +duplicate-detection work so the question is not re-opened.)* The proxy is **read-only by +explicit design**: `src/index.ts:35-41` rejects every non-GET/HEAD method, and `KEY_RE` at +`:27` carries a comment stating the proxy must never become a general oracle over the bucket. +Turning it into a write endpoint works directly against that intent, for a feature that does +not need a server at all. Its one genuinely reusable piece is the `RATE_LIMITER` binding. +**Use the zero-infrastructure fallback this phase already names** — copy-to-clipboard / +download-JSON, pasted into an editor action. It needs no deployment and matches the existing +idiom at `PlayerProvider.tsx:131-136`. + +**New trigger for this phase:** sharing one digest across a duplicate pair is precisely the +case a viewer needs to be able to flag. A borrowed digest can be misplaced (the mirror drifts +after the anchors that were measured) or simply wrong for that upload, and the viewer is the +only party who will ever notice — the detector already believes the two are the same video. + +Ingestion is an editor action accepting pasted JSON. Treat submissions as **untrusted input rendered in an admin UI**: escape it, never feed it into a prompt unreviewed. Store under `transcripts/.feedback/`, a sibling of `.jobs` and `.bookmarks`, outside the build trees. diff --git a/common/bin/compose-site.ts b/common/bin/compose-site.ts @@ -22,7 +22,10 @@ import { getSite, resolveSocialLinks, resolveHubUrl, type Site } from "../lib/si import { getSettings } from "../lib/settings"; import { DUPLICATES_FILENAME, + DUPLICATE_OVERRIDES_FILENAME, + clusterIsPublishable, filterClusterToChannels, + sanitizeDuplicateOverrides, type DuplicateReport, } from "../lib/duplicates"; import type { Manifest, SubsManifest } from "../lib/manifest"; @@ -743,11 +746,24 @@ async function main(): Promise<void> { // site (per-site `duplicates` opt-out) AND there's at least one in-scope // cluster — so its mere presence is what hasDuplicates() keys off to show the // nav link. No file → the page shows its empty state and the link self-hides. + // + // UNCONFIRMED SUSPECTS ARE NOT SHIPPED. A `needsReview` cluster is evidence + // that two videos share a title and a runtime — nothing compared their + // content — so it is an internal review queue, not something to assert to a + // viewer. It reaches the public site only once a human records `confirmed` in + // duplicates.overrides.json, which is why that file (never read at compose + // time before) is read here. const dupSrc = path.join(paths.transcriptsDir, DUPLICATES_FILENAME); const dupDest = path.join(paths.exportPublicDir, DUPLICATES_FILENAME); - // Gate the re-filter/re-serialize on the source's mtime+size, the member set, - // and the feature flag. The cache value is prefixed written|/empty| so a skip - // can self-heal if public/ was wiped out of band (presence must match). + const dupOverridesSrc = path.join( + paths.transcriptsDir, + DUPLICATE_OVERRIDES_FILENAME, + ); + // Gate the re-filter/re-serialize on the source's mtime+size, the OVERRIDES' + // mtime+size (a confirmation changes what ships without touching the report), + // the member set, and the feature flag. The cache value is prefixed + // written|/empty| so a skip can self-heal if public/ was wiped out of band + // (presence must match). let dupStat: { mtimeMs: number; size: number } | null = null; try { const s = await stat(dupSrc); @@ -755,8 +771,18 @@ async function main(): Promise<void> { } catch { dupStat = null; } + let dupOvStat: { mtimeMs: number; size: number } | null = null; + try { + const s = await stat(dupOverridesSrc); + dupOvStat = { mtimeMs: s.mtimeMs, size: s.size }; + } catch { + dupOvStat = null; + } const dupEnabled = site.duplicates !== false && dupStat !== null; - const dupKey = `${dupEnabled}|${dupStat?.mtimeMs ?? ""}|${dupStat?.size ?? ""}|${[...memberSlugs].sort().join(",")}`; + const dupKey = + `${dupEnabled}|${dupStat?.mtimeMs ?? ""}|${dupStat?.size ?? ""}` + + `|${dupOvStat?.mtimeMs ?? ""}|${dupOvStat?.size ?? ""}` + + `|${[...memberSlugs].sort().join(",")}`; const dupDestPresent = await exists(dupDest); const dupCachedWritten = cache.duplicates?.startsWith("written|") ?? false; const dupCacheHit = @@ -772,7 +798,18 @@ async function main(): Promise<void> { await readFile(dupSrc, "utf8"), ) as DuplicateReport; const memberSet = new Set(memberSlugs); + // Never throws: an unreadable or malformed overrides file reads as "no + // decisions recorded", which fails CLOSED — every suspect stays internal. + let dupOverrides = sanitizeDuplicateOverrides(null); + try { + dupOverrides = sanitizeDuplicateOverrides( + JSON.parse(await readFile(dupOverridesSrc, "utf8")), + ); + } catch { + // no decisions recorded + } const clusters = report.clusters + .filter((c) => clusterIsPublishable(c, dupOverrides)) .map((c) => filterClusterToChannels(c, memberSet)) .filter((c): c is NonNullable<typeof c> => c !== null); if (clusters.length > 0) { diff --git a/common/bin/duplicate-shorts.ts b/common/bin/duplicate-shorts.ts @@ -1,15 +1,21 @@ #!/usr/bin/env tsx -// On-demand duplicate-shorts detection. Runs AFTER build:index + build:stats -// (it reads the cues + statsByPath those populate), so it is deliberately not -// chained into build:data. +// On-demand duplicate detection. Runs AFTER build:index + build:stats (it reads +// the cues + statsByPath those populate), so it is deliberately not chained into +// build:data. // // Flags: // --threshold N duration cutoff in seconds (default 180) -// --all-durations no length filter; enables containment matching +// --all-durations no length filter; enables the containment sweep +// --blocking S how candidate pairs are nominated: title | duration | +// both. Defaults to `title` corpus-wide and `duration` in +// shorts mode. Nomination is never evidence — every pair is +// then judged on transcript content — so this trades recall +// against runtime, not correctness. // --near F Jaccard near-duplicate threshold (default 0.6) // --tolerance N duration bucketing window in seconds (default 2) import { getPaths } from "../lib/paths"; import { detectDuplicateShorts } from "../controller/duplicateShorts"; +import type { DuplicateBlocking } from "../lib/duplicates"; import { parseFlags } from "./_parseFlags"; const flags = parseFlags(process.argv.slice(2)); @@ -21,9 +27,22 @@ const thresholdSeconds = ? Number(flags.threshold) : undefined; +const BLOCKING: ReadonlyArray<DuplicateBlocking> = ["title", "duration", "both"]; +const blockingFlag = flags.blocking; +if ( + blockingFlag !== undefined && + !BLOCKING.includes(blockingFlag as DuplicateBlocking) +) { + console.error( + `--blocking must be one of: ${BLOCKING.join(", ")} (got "${blockingFlag}")`, + ); + process.exit(1); +} + detectDuplicateShorts({ paths: getPaths(), thresholdSeconds, + blocking: blockingFlag as DuplicateBlocking | undefined, nearThreshold: flags.near !== undefined ? Number(flags.near) : undefined, durationToleranceSeconds: flags.tolerance !== undefined ? Number(flags.tolerance) : undefined, diff --git a/common/components/SearchResults.tsx b/common/components/SearchResults.tsx @@ -14,9 +14,20 @@ import { Checkbox } from "./ui/checkbox"; import type { LayerHit } from "./searchPipeline"; import { AgeRestrictedBadge, + DuplicateBadge, LivestreamBadge, VodExpiredBadge, } from "./badges"; +import { + duplicateJumpSeconds, + duplicateSiblings, + useDuplicates, + type DuplicateLookup, +} from "./duplicatesCache"; +import type { DuplicateVideoRef } from "../lib/duplicates"; +import { splitId } from "./originId"; +import { usePlayer } from "./PlayerProvider"; +import { formatTimestamp } from "../lib/vtt"; import { vodExpiry } from "../lib/vodExpiry"; import { Button } from "./ui/button"; import { LayerSwatch } from "./LayerSwatch"; @@ -298,6 +309,20 @@ function VirtualResultList({ }) { const flatRows = useMemo(() => buildResultRows(resultGroups), [resultGroups]); + // "This exists elsewhere in the archive." One root-level JSON, fetched once; + // absent on most sites, which is why every use of it is additive. + const duplicates = useDuplicates(); + const { openTranscript } = usePlayer(); + // Opening a sibling is not opening a hit, so it goes straight to the player + // rather than through openWithMode: there is no LayerHit to carry a scope, and + // the target is always a video's transcript view. + const openDuplicate = useCallback( + (slug: string, seconds: number) => { + openTranscript(slug, seconds, { mode: "transcript" }); + }, + [openTranscript], + ); + const listRef = useRef<HTMLDivElement | null>(null); const [scrollMargin, setScrollMargin] = useState(0); @@ -361,6 +386,8 @@ function VirtualResultList({ selected={selectedSlugs.has(row.group.slug)} onToggleSelect={onToggleSelect} onAsk={onAsk} + duplicates={duplicates} + openDuplicate={openDuplicate} /> )} </VirtualRow> @@ -380,6 +407,8 @@ const ResultCard = memo(function ResultCard({ selected, onToggleSelect, onAsk, + duplicates, + openDuplicate, }: { group: ResultGroup; leavesById: ReadonlyMap<string, LeafInfo>; @@ -390,7 +419,29 @@ const ResultCard = memo(function ResultCard({ selected: boolean; onToggleSelect: (slug: string) => void; onAsk: (slug: string) => void; + duplicates: DuplicateLookup; + openDuplicate: (slug: string, seconds: number) => void; }) { + // Other copies of this video elsewhere in the archive. + // + // FEDERATION GUARD: in hub mode a group's slug may be an origin-qualified id, + // and /duplicates.json is same-origin only — a cross-origin row's bare slug + // could collide with a local one and point the viewer at the wrong video. A + // post has no duplicate story at all. Both are simply excluded. + const siblings = useMemo(() => { + if (group.post) return []; + if (splitId(group.slug).origin !== "") return []; + return duplicateSiblings(duplicates, group.slug); + }, [duplicates, group.slug, group.post]); + + // Where the viewer currently is in THIS copy: the moment the card is showing + // them (the player's playhead when this card is what's open, otherwise the top + // hit). Only ever honoured for a sibling the detector measured as aligned. + const hereSeconds = + activeVideo === group.slug && activeTime !== null + ? activeTime + : group.hits[0]?.start; + // Bucket hits per contributing leaf so the user sees one section per // layer rather than an interleaved mishmash. const buckets = useMemo(() => { @@ -451,6 +502,9 @@ const ResultCard = memo(function ResultCard({ </span> {group.isLivestream && <LivestreamBadge />} {group.ageRestricted && <AgeRestrictedBadge />} + {siblings.length > 0 && ( + <DuplicateBadge count={siblings.length} /> + )} {(() => { const e = vodExpiry(group.platform, group.uploadDate); return e?.likelyExpired ? ( @@ -477,6 +531,13 @@ const ResultCard = memo(function ResultCard({ <MessageSquareIcon className="size-3.5" /> Ask </button> </div> + {siblings.length > 0 && ( + <DuplicateStrip + siblings={siblings} + hereSeconds={hereSeconds} + openDuplicate={openDuplicate} + /> + )} <ul className="flex flex-col divide-y divide-border"> {Array.from(buckets.entries()).map(([leafId, hits]) => { if (hits.length === 0) return null; @@ -531,6 +592,59 @@ const ResultCard = memo(function ResultCard({ }); ResultCard.displayName = "ResultCard"; +// The jump controls for a card's other copies. Lives outside the card's +// open-the-video <Button> so these can be real buttons, and handles more than +// one sibling, which a single control in the header could not. +// +// Each button STATES where it will land, because the two cases are genuinely +// different promises: an aligned sibling was measured to carry the same words at +// the same times, so the same timestamp is meaningful there; an unmeasured or +// misaligned one is opened from the start, because a mirror with a longer intro +// would otherwise drop the viewer mid-sentence in the wrong place and look like +// it worked. Never claim the moment we did not measure. +function DuplicateStrip({ + siblings, + hereSeconds, + openDuplicate, +}: { + siblings: ReadonlyArray<DuplicateVideoRef>; + hereSeconds: number | undefined; + openDuplicate: (slug: string, seconds: number) => void; +}) { + return ( + <div + data-duplicate-strip="" + className="flex flex-wrap items-center gap-2 border-b border-border bg-muted/40 px-3 py-1.5 text-xs" + > + <span className="text-muted-foreground">Also in this archive:</span> + {siblings.map((ref) => { + const to = duplicateJumpSeconds(ref, hereSeconds); + const label = ref.channel || ref.channelSlug; + return ( + <button + key={ref.slug} + type="button" + data-duplicate-slug={ref.slug} + data-duplicate-aligned={ref.aligned === true ? "true" : "false"} + onClick={() => openDuplicate(ref.slug, to)} + title={ + to > 0 + ? `Open the ${ref.platform} copy on ${label} at the same moment — its timings were measured to line up with this one` + : ref.aligned === true + ? `Open the ${ref.platform} copy on ${label}` + : `Open the ${ref.platform} copy on ${label} from the start — its timings were not verified to line up with this one` + } + className="rounded border border-border bg-background px-1.5 py-0.5 transition-colors hover:bg-accent hover:text-accent-foreground" + > + {label} · {ref.platform} + {to > 0 ? ` @ ${formatTimestamp(to)}` : ""} + </button> + ); + })} + </div> + ); +} + // One hit row inside a card. Memoized so that (a) scroll-frame re-renders of // the parent list don't re-run highlight() for every hit, and (b) when the // modal target changes only the previously-active and newly-active rows diff --git a/common/components/badges.tsx b/common/components/badges.tsx @@ -23,6 +23,26 @@ export function AgeRestrictedBadge() { ); } +// Flags a search result that also exists elsewhere in the archive — a mirror on +// another platform, or a re-upload on another channel. Non-interactive on +// purpose: it sits inside the card's open-the-video button, and the jump +// controls live in their own strip below the header where they can be real +// buttons (and where there is room for more than one sibling). +export function DuplicateBadge({ count }: { count: number }) { + return ( + <span + title={ + count === 1 + ? "This video also exists elsewhere in the archive" + : `This video also exists in ${count} other places in the archive` + } + className={`${badgeBase} bg-sky-100 text-sky-800 dark:bg-sky-900/40 dark:text-sky-200`} + > + {count === 1 ? "Dupe" : `${count} dupes`} + </span> + ); +} + // Flags a stream VOD that is likely past its platform's retention window (Kick // ~30d, Twitch 7–60d), so playback probably fails. `tooltip` explains the // platform-specific retention on hover (see lib/vodExpiry). diff --git a/common/components/duplicatesCache.ts b/common/components/duplicatesCache.ts @@ -0,0 +1,83 @@ +"use client"; + +// Client fetch for a site's shipped /duplicates.json, indexed by member slug so +// a search result can answer "does this video exist elsewhere?" in O(1). Mirrors +// aliasesCache / summariesCache: one root-level JSON, fetched once, cached +// forever (it is static per export build). +// +// A missing file resolves to an empty map rather than an error. That is the +// common case, not an edge one — compose-site writes the file only when the site +// has at least one shippable cluster, and every duplicate affordance here is +// purely additive, so its absence must never break search. +// +// WHAT SHIPS HERE IS ALREADY FILTERED. compose-site drops `needsReview` clusters +// unless a human confirmed them, so anything in this map has had its CONTENT +// compared, not just its title and runtime. The one thing still worth checking +// per member is `aligned` — see duplicateSiblings below. + +import { useQuery } from "@tanstack/react-query"; +import { idBaseUrl } from "./originId"; +import type { + DuplicateCluster, + DuplicateReport, + DuplicateVideoRef, +} from "../lib/duplicates"; + +export type DuplicateLookup = ReadonlyMap<string, DuplicateCluster>; + +const EMPTY: DuplicateLookup = new Map(); + +export async function fetchDuplicates(origin = ""): Promise<DuplicateLookup> { + let report: DuplicateReport | null = null; + try { + const r = await fetch(`${idBaseUrl(origin)}/duplicates.json`); + if (!r.ok) return EMPTY; + report = (await r.json()) as DuplicateReport; + } catch { + return EMPTY; + } + const out = new Map<string, DuplicateCluster>(); + for (const cluster of report?.clusters ?? []) { + for (const ref of cluster?.videoRefs ?? []) { + if (ref?.slug) out.set(ref.slug, cluster); + } + } + return out; +} + +export function useDuplicates(origin = ""): DuplicateLookup { + const { data } = useQuery<DuplicateLookup>({ + queryKey: ["duplicates", origin], + queryFn: () => fetchDuplicates(origin), + staleTime: Infinity, // static per export build + }); + return data ?? EMPTY; +} + +// The other members of `slug`'s cluster, or [] when it is in none. +export function duplicateSiblings( + lookup: DuplicateLookup, + slug: string, +): DuplicateVideoRef[] { + const cluster = lookup.get(slug); + if (!cluster) return []; + return cluster.videoRefs.filter((r) => r.slug !== slug); +} + +// Where to land when opening a sibling, given where the viewer is in THIS copy. +// +// The honesty rule. Matching content does NOT imply matching timings: a mirror +// with a longer intro or an extra ad break carries the same words at shifted +// times, so seeking to the same timestamp lands in the wrong place while looking +// perfectly plausible. `aligned` is the detector's measured verdict on exactly +// that, and it is only ever true when several anchors were located within the +// tolerance. Absent (an older report, or no timed cues to measure) reads as NOT +// aligned, so the fallback is the start of the video — a jump the viewer can +// always make sense of. +export function duplicateJumpSeconds( + ref: DuplicateVideoRef, + seconds: number | undefined, +): number { + if (ref.aligned !== true) return 0; + return typeof seconds === "number" && seconds > 0 ? seconds : 0; +} diff --git a/common/controller/digestSharing.ts b/common/controller/digestSharing.ts @@ -73,7 +73,10 @@ export async function buildDigestClusterPlan( for (const cluster of report.clusters) { const canonicalSlug = resolveCanonicalSlug(cluster, overrides); // null → a human said "not a duplicate": every member stands alone. - if (!canonicalSlug || !clusterMaySharePartial(cluster)) { + // The overrides also carry the `confirmed` flag that is the ONLY thing + // letting a needsReview cluster share, so they must be passed here — without + // them every confirmed suspect would silently keep generating twice. + if (!canonicalSlug || !clusterMaySharePartial(cluster, overrides)) { plan.independentClusters++; continue; } @@ -220,8 +223,11 @@ export async function shareClusterFromCanonical( onLog?: (msg: string) => void; } = {}, ): Promise<ShareOutcome[]> { - if (!clusterMaySharePartial(cluster)) return []; + // Load the overrides BEFORE the gate, not after: `confirmed` lives in them and + // is what unblocks a needsReview cluster, so testing the gate first would + // refuse to share from every cluster a human had just approved. const overrides = opts.overrides ?? (await readDuplicateOverrides(paths)); + if (!clusterMaySharePartial(cluster, overrides)) return []; const canonicalSlug = resolveCanonicalSlug(cluster, overrides); if (!canonicalSlug) return []; return shareDigestToCluster({ diff --git a/common/controller/duplicateShorts.ts b/common/controller/duplicateShorts.ts @@ -1,22 +1,41 @@ -// Cross-platform duplicate "shorts" detection — a global pass over every -// channel's videos (NOT per-channel, since duplicates are cross-channel and -// cross-platform by definition). +// Cross-platform duplicate detection — a global pass over every channel's videos +// (NOT per-channel, since duplicates are cross-channel and cross-platform by +// definition). // -// Hybrid cascade: -// Phase 1 (cheap) — pre-cluster candidates by rounded duration, the only -// signal stable across platforms and re-titles. Consumes -// the already-aggregated VideoStat index (LMDB statsByPath, -// populated by buildStats) rather than re-walking dirs. -// Phase 2 (confirm)— for each candidate pair, compare transcript content -// (read from the LMDB `cues` sub-db that buildIndex -// populates): exact text hash → 5-gram Jaccard → -// (all-durations only) containment for "short is a clip of -// a longer video". Falls back to a metadata-only match when -// a transcript is missing/empty. +// THE PRE-FILTER PROPOSES; THE TRANSCRIPT DISPOSES. // -// Output is a single global transcripts/duplicates.json. Flag-only: nothing is -// merged or deleted. Requires build:index (cues) + build:stats (statsByPath) to -// have run first. +// A blocking strategy (§ "Phase 1" below) only NOMINATES pairs. It is never +// evidence on its own. Every nominated pair then goes through the transcript +// cascade, which has three outcomes: +// +// both sides have a transcript, content matches → transcript-exact/near. +// Confirmed; may share. +// both sides have a transcript, content does NOT → REJECTED, no cluster. +// This is the whole value of +// testing the suspects. +// one side has no transcript → untestable. Kept only when +// the pre-filter's own claim +// is strong enough to be +// worth a human's attention +// (title + near-identical +// runtime), as a needsReview +// `title-duration` suspect +// that shares nothing. +// +// STREAMING, NOT BATCHING. The first implementation built one global candidate +// array, then fingerprinted every video it referenced, then evaluated. Both +// halves fail corpus-wide: the containment pass alone was an unblocked cartesian +// product (498 M pairs — V8 throws RangeError building the array), and holding +// 5-word shingle sets for the whole corpus at once (a 3.5 h video is ~40 k +// strings) exhausted a 4 GB heap. Neither is inherent to corpus-wide detection. +// Blocking fixes the first; this file's block-at-a-time iteration fixes the +// second: fingerprint one block, evaluate its pairs, keep only the confirmed +// matches, release. Peak memory is O(largest block), not O(corpus), and no +// multi-million-element pair array is ever materialised. +// +// Inputs are the already-aggregated VideoStat index (LMDB statsByPath, from +// buildStats) plus the `cues` sub-db (from buildIndex). Output is a single +// global transcripts/duplicates.json. Flag-only: nothing is merged or deleted. import path from "node:path"; import { mkdir, rename, writeFile, readFile } from "node:fs/promises"; @@ -30,8 +49,15 @@ import type { Cue } from "../lib/vtt"; import { parseTranscriptJson } from "../lib/whisper"; import { readVideoFiles, pickIndexTranscript } from "../lib/videoStatus"; import { + DEFAULT_ALIGNMENT_TOLERANCE_SECONDS, DEFAULT_CONTAINMENT_THRESHOLD, DEFAULT_DURATION_TOLERANCE_SECONDS, + DEFAULT_TITLE_DURATION_RATIO, + MAX_DURATION_BLOCK_SIZE, + MAX_TITLE_GROUP_SIZE, + durationsCompatible, + measureAlignment, + normalizeTitleKey, DEFAULT_NEAR_THRESHOLD, DEFAULT_SHINGLE_SIZE, DEFAULT_SHORT_THRESHOLD_SECONDS, @@ -46,6 +72,7 @@ import { sanitizeDuplicateOverrides, shingles, strongerMatch, + type DuplicateBlocking, type DuplicateOverrides, type DuplicateCluster, type DuplicateMatchKind, @@ -54,6 +81,17 @@ import { type DuplicateVideoRef, } from "../lib/duplicates"; +// A normalized title shorter than this carries no blocking information — +// pairing on it would rebuild the cartesian product the pass exists to avoid. +const MIN_TITLE_KEY_LENGTH = 8; + +// Hard budget for the shorts-vs-longer containment sweep, which is a genuine +// cartesian product and the one pass blocking cannot rescue. Under the budget it +// runs (and finds clip-of-longer matches, which nothing else can); over it, it +// is SKIPPED AND REPORTED rather than silently truncated or left to die. Scaling +// containment properly wants MinHash/LSH signatures — see STATE.md. +const MAX_CONTAINMENT_PAIRS = 5_000_000; + type PathKey = [string, string]; type IndexKey = [string, string, string]; type StatsRecord = { metaMs: number; stat: VideoStat }; @@ -67,11 +105,21 @@ export type DetectDuplicateShortsOptions = { nearThreshold?: number; containmentThreshold?: number; shingleSize?: number; + // Candidate-nomination strategy. Defaults to "title" corpus-wide (where a + // duration sweep is affordable but far slower) and "duration" in shorts mode, + // which preserves today's proven behaviour at that scale. + blocking?: DuplicateBlocking; + titleDurationRatio?: number; + // Tolerance for the per-cluster timing measurement written onto each ref. + alignmentToleranceSeconds?: number; onLog?: (msg: string) => void; signal?: AbortSignal; }; -// Per-participant transcript fingerprint, computed once and reused across pairs. +// Per-participant transcript fingerprint, computed once per BLOCK and released +// when the block is done. Holding these for the whole corpus is what exhausted +// the heap; the shingle set is the expensive part (~40 k strings for a 3.5 h +// video), which is why it never outlives its block. type Fingerprint = { stat: VideoStat; hasTranscript: boolean; @@ -81,6 +129,24 @@ type Fingerprint = { type Entry = { stat: VideoStat; videoDir: string }; +// One nominated pair. `suspectKind` is what this pair's evidence is worth when +// the content CANNOT be compared because a side has no transcript: a shared +// title AND a near-identical runtime is a claim worth a human's attention, so it +// survives as a `title-duration` suspect; a shared duration alone is not, and +// treating it as one is exactly what produced enormous false clusters of +// unrelated same-length videos before, so it is null and the pair is dropped. +type Nomination = { + a: Entry; + b: Entry; + // Skip the similarity test and go straight to containment — a clip can never + // pass Jaccard against the full recording it was cut from. + containmentOnly: boolean; + suspectKind: DuplicateMatchKind | null; +}; + +// A unit of work: fingerprint these members, evaluate these pairs, release. +type Block = { members: Entry[]; nominations: Nomination[] }; + type PairResult = { a: string; // slug b: string; // slug @@ -89,6 +155,21 @@ type PairResult = { contained: boolean; }; +// Ordered slug-pair key, so "both" can union two nomination streams without +// evaluating the overlap twice. +function pairKey(a: string, b: string): string { + return a < b ? `${a}\n${b}` : `${b}\n${a}`; +} + +function membersOf(nominations: Nomination[]): Entry[] { + const m = new Map<string, Entry>(); + for (const n of nominations) { + m.set(n.a.stat.slug, n.a); + m.set(n.b.stat.slug, n.b); + } + return [...m.values()]; +} + export async function detectDuplicateShorts( opts: DetectDuplicateShortsOptions, ): Promise<DuplicateReport> { @@ -107,12 +188,24 @@ export async function detectDuplicateShorts( opts.containmentThreshold ?? DEFAULT_CONTAINMENT_THRESHOLD; const shingleSize = opts.shingleSize ?? DEFAULT_SHINGLE_SIZE; + // Corpus-wide defaults to TITLE blocking: it is effectively linear and it is + // the signal that finds cross-platform re-uploads, which keep their name. + // Duration blocking is viable corpus-wide too since the streaming rewrite (it + // is the only strategy that catches a RE-TITLED mirror) but it nominates + // millions of pairs where title nominates thousands, so it is opt-in. + const blocking = + opts.blocking ?? (thresholdSeconds === null ? "title" : "duration"); + const titleDurationRatio = + opts.titleDurationRatio ?? DEFAULT_TITLE_DURATION_RATIO; + const alignmentTolerance = + opts.alignmentToleranceSeconds ?? DEFAULT_ALIGNMENT_TOLERANCE_SECONDS; const runConfig: DuplicateRunConfig = { thresholdSeconds, durationToleranceSeconds: W, nearThreshold, containmentThreshold, shingleSize, + blocking, }; // ---- Open the index: statsByPath (Phase-1 input) + cues (Phase-2 input) --- @@ -159,113 +252,248 @@ export async function detectDuplicateShorts( `(threshold=${thresholdSeconds === null ? "all" : `${thresholdSeconds}s`}).`, ); - // ---- Phase 1: bucket by rounded duration --------------------------------- - const bucketOf = (d: number) => Math.round(d / W); - const buckets = new Map<number, Entry[]>(); - for (const e of participants) { - const b = bucketOf(e.stat.duration); - const arr = buckets.get(b); - if (arr) arr.push(e); - else buckets.set(b, [e]); - } + // ---- Phase 1 + 2, interleaved: nominate a block, judge it, release it ----- + // + // The two phases are no longer separate passes. Each block is fingerprinted, + // evaluated and dropped before the next is built, so peak memory is O(largest + // block) and only the (few, by construction) confirmed matches survive. + const stats = { + nominated: 0, + evaluated: 0, + confirmed: 0, + rejected: 0, + suspects: 0, + titleGroups: 0, + oversizedGroups: 0, + oversizedGroupVideos: 0, + oversizedBlocks: 0, + oversizedBlockVideos: 0, + untitled: 0, + fingerprinted: 0, + rawFallbacks: 0, + rawHits: 0, + peakRssMb: 0, + }; + const noteRss = () => { + const mb = Math.round(process.memoryUsage().rss / 1_048_576); + if (mb > stats.peakRssMb) stats.peakRssMb = mb; + }; - // Candidate pairs. Same or adjacent duration bucket + relative-duration guard, - // or (all-durations mode) short/longer pairs eligible for containment. - const maxDelta = (d: number) => Math.max(W, Math.ceil(0.02 * d)); - type Candidate = { a: Entry; b: Entry; containmentOnly: boolean }; - const candidates: Candidate[] = []; - const sortedBuckets = [...buckets.keys()].sort((x, y) => x - y); - for (const b of sortedBuckets) { - const here = buckets.get(b) as Entry[]; - const next = buckets.get(b + 1) ?? []; - for (let i = 0; i < here.length; i++) { - for (let j = i + 1; j < here.length; j++) { - if (durationsClose(here[i], here[j], maxDelta)) - candidates.push({ a: here[i], b: here[j], containmentOnly: false }); + // Fingerprint one block's members, re-using anything the previous block + // already computed. Primary source is the LMDB `cues` sub-db — but parseVtt + // only extracts text from YouTube karaoke-tagged cues, so plain-VTT + // transcripts (most non-YouTube captions, whisper-as-vtt, manual subs) store + // as zero cues and would look transcript-less. For those, fall back to reading + // the raw transcript file off disk and extracting plain text, for comparison + // only — nothing is persisted. + const fingerprintBlock = async ( + members: Entry[], + reuse: ReadonlyMap<string, Fingerprint>, + ): Promise<Map<string, Fingerprint>> => { + const out = new Map<string, Fingerprint>(); + const rawFallbacks: Entry[] = []; + for (const e of members) { + const s = e.stat; + if (out.has(s.slug)) continue; + const cached = reuse.get(s.slug); + if (cached) { + out.set(s.slug, cached); + continue; + } + const cues = cuesDb.get([s.uploadDate, s.channelSlug, s.id]); + if (cues && cues.length > 0) { + stats.fingerprinted++; + out.set(s.slug, fingerprintFrom(s, cues, shingleSize)); + } else { + rawFallbacks.push(e); } } - for (const x of here) { - for (const y of next) { - if (durationsClose(x, y, maxDelta)) - candidates.push({ a: x, b: y, containmentOnly: false }); + if (rawFallbacks.length > 0) { + stats.rawFallbacks += rawFallbacks.length; + const limit = pLimit(16); + await Promise.all( + rawFallbacks.map((e) => + limit(async () => { + opts.signal?.throwIfAborted(); + const cues = await readRawCues(opts.paths.channelsDir, e); + if (cues && cues.length > 0) stats.rawHits++; + stats.fingerprinted++; + out.set(e.stat.slug, fingerprintFrom(e.stat, cues, shingleSize)); + }), + ), + ); + } + return out; + }; + + // Cues WITH real start times, for the alignment measurement only. Deliberately + // LMDB-only: the raw-VTT fallback stamps every cue at start 0, which is fine + // for a set-based content fingerprint and worthless — actively misleading — + // for measuring an offset. No timed cues → no measurement → `aligned` stays + // undefined, which every consumer must read as "not aligned". + const timedCues = (e: Entry | undefined): Cue[] | null => { + if (!e) return null; + const s = e.stat; + const cues = cuesDb.get([s.uploadDate, s.channelSlug, s.id]); + return cues && cues.length > 0 ? cues : null; + }; + + const streams: Generator<Block>[] = []; + if (blocking === "title" || blocking === "both") { + streams.push(titleBlocks(participants, titleDurationRatio, stats, log)); + } + if (blocking === "duration" || blocking === "both") { + streams.push(durationBlocks(participants, W, stats, log)); + } + // Only "both" needs cross-stream dedup. It is deliberately one-directional: + // the FIRST stream records its nominations and later streams consult them. + // Title blocking is pushed first precisely so the remembered set is the small + // one — recording the duration stream's millions of pair keys as well would + // reintroduce, in the dedup set, exactly the whole-corpus retention the + // streaming rewrite exists to remove. + const seenPairs = blocking === "both" ? new Set<string>() : null; + + const matches: PairResult[] = []; + const hasTranscriptBySlug = new Map<string, boolean>(); + + // 2-block sliding fingerprint cache: a duration block is bucket[b] ∪ + // bucket[b+1], so every member except the last bucket is re-used by the very + // next block. Without this, each bucket would be read and shingled twice. + let prevFps = new Map<string, Fingerprint>(); + let blocksDone = 0; + + for (const [streamIndex, stream] of streams.entries()) { + const recordPairs = seenPairs !== null && streamIndex === 0; + const skipSeenPairs = seenPairs !== null && streamIndex > 0; + for (const block of stream) { + opts.signal?.throwIfAborted(); + let nominations = block.nominations; + if (recordPairs) { + for (const n of nominations) { + seenPairs.add(pairKey(n.a.stat.slug, n.b.stat.slug)); + } + } else if (skipSeenPairs) { + nominations = nominations.filter( + (n) => !seenPairs.has(pairKey(n.a.stat.slug, n.b.stat.slug)), + ); + } + stats.nominated += nominations.length; + if (nominations.length === 0) continue; + + // When dedup narrowed the nominations, re-derive the members from what is + // actually left rather than fingerprinting videos no surviving pair needs. + const fps = await fingerprintBlock( + skipSeenPairs ? membersOf(nominations) : block.members, + prevFps, + ); + for (const [slug, fp] of fps) hasTranscriptBySlug.set(slug, fp.hasTranscript); + + for (const n of nominations) { + const fa = fps.get(n.a.stat.slug); + const fb = fps.get(n.b.stat.slug); + if (!fa || !fb) continue; + stats.evaluated++; + const res = n.containmentOnly + ? evalContainment(fa, fb, containmentThreshold) + : evalBlocked(fa, fb, { + nearThreshold, + containmentThreshold, + allowContainment: includeContainment, + suspectKind: n.suspectKind, + }); + if (!res) { + stats.rejected++; + continue; + } + if (res.kind === "title-duration") stats.suspects++; + else stats.confirmed++; + matches.push({ a: n.a.stat.slug, b: n.b.stat.slug, ...res }); + } + prevFps = fps; + noteRss(); + // A corpus-wide duration run is long enough that silence is + // indistinguishable from a hang. Heartbeat with the numbers that matter. + if (++blocksDone % 100 === 0) { + log( + ` …${blocksDone} block(s), ${stats.evaluated} pair(s) tested, ` + + `${stats.confirmed} confirmed, RSS ${Math.round(process.memoryUsage().rss / 1_048_576)} MB`, + ); } } + prevFps = new Map(); } - // Containment candidates: a short paired with any meaningfully-longer video. + // ---- The containment sweep: shorts vs every meaningfully-longer video ----- + // + // The one pass blocking cannot rescue — a clip and its parent share neither a + // duration nor, usually, a title, so nothing nominates them but a cartesian + // product. Bounded by an explicit pair budget and skipped-with-a-log when it + // does not fit, rather than silently truncated. Memory stays bounded because + // the budget implicitly bounds the short side: at 76 k longer videos, fitting + // under the budget means only a handful of shorts. if (includeContainment) { - const shortsForContainment = participants.filter( + const shorts = participants.filter( (e) => - e.stat.duration > 0 && - e.stat.duration <= DEFAULT_SHORT_THRESHOLD_SECONDS, + e.stat.duration > 0 && e.stat.duration <= DEFAULT_SHORT_THRESHOLD_SECONDS, ); - for (const s of shortsForContainment) { + const budget = shorts.length * participants.length; + if (shorts.length === 0) { + // nothing to sweep + } else if (budget > MAX_CONTAINMENT_PAIRS) { + log( + `Containment sweep SKIPPED: ${shorts.length} short(s) × ${participants.length} video(s) ` + + `= ~${budget.toLocaleString("en-US")} pairs, over the ${MAX_CONTAINMENT_PAIRS.toLocaleString("en-US")} budget. ` + + `Clip-of-longer duplicates are NOT covered by this run.`, + ); + } else { + const shortFps = await fingerprintBlock(shorts, new Map()); + for (const [slug, fp] of shortFps) { + hasTranscriptBySlug.set(slug, fp.hasTranscript); + } + const shortSlugs = new Set(shorts.map((e) => e.stat.slug)); + let swept = 0; for (const v of participants) { - if (v === s) continue; - if (v.stat.duration >= s.stat.duration * 1.5) - candidates.push({ a: s, b: v, containmentOnly: true }); + opts.signal?.throwIfAborted(); + // One long video at a time: fingerprint, compare against every short, + // release. O(shorts + 1) fingerprints held. + const relevant = shorts.filter( + (s) => + s.stat.slug !== v.stat.slug && + v.stat.duration >= s.stat.duration * 1.5, + ); + if (relevant.length === 0) continue; + const fv = shortSlugs.has(v.stat.slug) + ? shortFps.get(v.stat.slug) + : (await fingerprintBlock([v], new Map())).get(v.stat.slug); + if (!fv) continue; + hasTranscriptBySlug.set(v.stat.slug, fv.hasTranscript); + for (const s of relevant) { + const fs = shortFps.get(s.stat.slug); + if (!fs) continue; + swept++; + stats.evaluated++; + const res = evalContainment(fs, fv, containmentThreshold); + if (!res) { + stats.rejected++; + continue; + } + stats.confirmed++; + matches.push({ a: s.stat.slug, b: v.stat.slug, ...res }); + } + noteRss(); } + stats.nominated += swept; + log(`Containment sweep: ${swept} pair(s) evaluated.`); } } - log( - `Phase 1: ${buckets.size} duration buckets → ${candidates.length} candidate pairs.`, - ); - // ---- Build fingerprints for every video in a candidate ------------------ - // Primary source is the LMDB `cues` sub-db. But parseVtt only extracts text - // from YouTube karaoke-tagged cues, so plain-VTT transcripts (most non-YouTube - // captions, whisper-as-vtt, manual subs) store as zero cues and would look - // transcript-less. For those, fall back to reading the raw transcript file off - // disk and extracting plain text — for comparison only, not stored anywhere. - const needed = new Map<string, Entry>(); - for (const c of candidates) { - needed.set(c.a.stat.slug, c.a); - needed.set(c.b.stat.slug, c.b); - } - const fingerprints = new Map<string, Fingerprint>(); - const rawFallbacks: Entry[] = []; - for (const e of needed.values()) { - opts.signal?.throwIfAborted(); - const s = e.stat; - const cues = cuesDb.get([s.uploadDate, s.channelSlug, s.id]); - if (cues && cues.length > 0) { - fingerprints.set(s.slug, fingerprintFrom(s, cues, shingleSize)); - } else { - rawFallbacks.push(e); - } - } - await root.close(); - - let rawHits = 0; - if (rawFallbacks.length > 0) { - const limit = pLimit(16); - await Promise.all( - rawFallbacks.map((e) => - limit(async () => { - opts.signal?.throwIfAborted(); - const cues = await readRawCues(opts.paths.channelsDir, e); - if (cues && cues.length > 0) rawHits++; - fingerprints.set(e.stat.slug, fingerprintFrom(e.stat, cues, shingleSize)); - }), - ), - ); - } log( - `Phase 2: fingerprinted ${fingerprints.size} candidate videos ` + - `(${rawFallbacks.length} missing LMDB cues; ${rawHits} recovered from raw transcripts).`, + `Phase 2: ${stats.evaluated} nominated pair(s) tested → ${stats.confirmed} confirmed, ` + + `${stats.rejected} rejected by content, ${stats.suspects} untestable suspect(s). ` + + `Fingerprinted ${stats.fingerprinted} video(s) (${stats.rawFallbacks} missing LMDB cues; ` + + `${stats.rawHits} recovered from raw transcripts).`, ); - - // ---- Phase 2: evaluate each candidate pair ------------------------------- - const matches: PairResult[] = []; - for (const c of candidates) { - opts.signal?.throwIfAborted(); - const fa = fingerprints.get(c.a.stat.slug) as Fingerprint; - const fb = fingerprints.get(c.b.stat.slug) as Fingerprint; - const res = c.containmentOnly - ? evalContainment(fa, fb, containmentThreshold) - : evalSimilar(fa, fb, nearThreshold); - if (res) matches.push({ a: c.a.stat.slug, b: c.b.stat.slug, ...res }); - } + prevFps = new Map(); // ---- Cluster matched pairs (union-find) ---------------------------------- const uf = new UnionFind(); @@ -297,10 +525,21 @@ export async function detectDuplicateShorts( } } + // Refs are rebuilt from the participant index, NOT from fingerprints — those + // were released with their block. VideoStat is small metadata and the index is + // already held for the whole run, so this costs nothing. + const bySlug = new Map<string, Entry>(); + for (const e of participants) bySlug.set(e.stat.slug, e); + const clusters: DuplicateCluster[] = components.map((slugs) => { const agg = aggByRoot.get(slugs[0]) as Agg; const refs = slugs - .map((slug) => toRef(fingerprints.get(slug) as Fingerprint)) + .map((slug) => + toRef( + (bySlug.get(slug) as Entry).stat, + hasTranscriptBySlug.get(slug) ?? false, + ), + ) .sort((x, y) => x.slug.localeCompare(y.slug)); const platforms = new Set(refs.map((r) => r.platform)); const channels = new Set(refs.map((r) => r.channelSlug)); @@ -313,6 +552,11 @@ export async function detectDuplicateShorts( durationBucket: Math.round(minDuration), crossPlatform: platforms.size > 1, crossChannel: channels.size > 1, + // Nothing compared this cluster's CONTENT — every pair in it was + // untestable. It is a suspect for human review and shares nothing until + // someone records `confirmed`. Written only when true so that a + // content-confirmed cluster serialises exactly as it did before. + ...(agg.matchKind === "title-duration" ? { needsReview: true } : {}), videoRefs: refs, }; // Record the rule's choice of canonical member: the one derived work (an AI @@ -322,9 +566,57 @@ export async function detectDuplicateShorts( return { ...cluster, canonicalSlug: pickCanonicalSlug(cluster) }; }); + // ---- Measure timing alignment against each cluster's canonical member ----- + // + // Content similarity says NOTHING about timing: a mirror with a longer intro + // matches on text at shifted times. Anything that seeks into a sibling — a + // shared digest's chapters, the search-result "jump to this moment" — is wrong + // without this, and wrong in the way that looks right. digestSharing measured + // it already but threw the result away; persisting it is what lets the viewer + // be honest about when a jump is trustworthy. + // + // Confirmed clusters only, and matches are few by construction, so re-reading + // those cue lists is cheap. Suspects are skipped: they share nothing anyway. + let aligned = 0; + let alignmentPairs = 0; + for (const cluster of clusters) { + opts.signal?.throwIfAborted(); + if (cluster.needsReview || cluster.contained) continue; + const canonicalSlug = cluster.canonicalSlug; + if (!canonicalSlug) continue; + const canonicalCues = timedCues(bySlug.get(canonicalSlug)); + if (!canonicalCues) continue; + for (const ref of cluster.videoRefs) { + if (ref.slug === canonicalSlug) { + ref.offsetSeconds = 0; + ref.aligned = true; + continue; + } + const cues = timedCues(bySlug.get(ref.slug)); + if (!cues) continue; + alignmentPairs++; + const a = measureAlignment(canonicalCues, cues, { + toleranceSeconds: alignmentTolerance, + }); + ref.aligned = a.aligned; + ref.offsetSeconds = Number.isFinite(a.maxOffsetSeconds) + ? Math.round(a.maxOffsetSeconds * 100) / 100 + : null; + if (a.aligned) aligned++; + } + } + if (alignmentPairs > 0) { + log( + `Alignment: ${aligned}/${alignmentPairs} mirror(s) within ${alignmentTolerance}s of their canonical member.`, + ); + } + await root.close(); + const rank: Record<DuplicateMatchKind, number> = { "transcript-exact": 2, "transcript-near": 1, + // Suspects sort last: they are a review queue, not a result. + "title-duration": 0, }; clusters.sort( (a, b) => @@ -348,10 +640,15 @@ export async function detectDuplicateShorts( const tmp = `${outPath}.tmp-${process.pid}`; await writeFile(tmp, JSON.stringify(report)); await rename(tmp, outPath); + noteRss(); + const suspectClusters = clusters.filter((c) => c.needsReview).length; log( - `Done: ${clusters.length} duplicate cluster(s) over ${videosInClusters} video(s) → ${DUPLICATES_FILENAME} ` + + `Done: ${clusters.length} duplicate cluster(s) over ${videosInClusters} video(s) ` + + `(${suspectClusters} awaiting review) → ${DUPLICATES_FILENAME} ` + `in ${Math.round((Date.now() - startedAt) / 1000)}s ` + - `(${videosScanned} scanned, ${candidates.length} candidate pairs, ${matches.length} confirmed).`, + `(blocking=${blocking}, ${videosScanned} scanned, ${stats.nominated} nominated pair(s), ` + + `${stats.confirmed} confirmed, ${stats.rejected} rejected by content, ${stats.suspects} suspect(s), ` + + `peak RSS ${stats.peakRssMb} MB).`, ); return report; } @@ -398,7 +695,14 @@ export async function readDuplicateOverrides( export async function updateDuplicateOverride( paths: Paths, clusterId: string, - patch: { canonicalSlug?: string; notDuplicate?: boolean; note?: string }, + patch: { + canonicalSlug?: string; + notDuplicate?: boolean; + // "I looked, and these really are the same video." The ONLY thing that lets + // a needsReview (title+duration) cluster share derived work. + confirmed?: boolean; + note?: string; + }, ): Promise<DuplicateOverrides> { const current = await readDuplicateOverrides(paths); const existing = current.clusters[clusterId] ?? {}; @@ -410,12 +714,16 @@ export async function updateDuplicateOverride( ...(patch.notDuplicate !== undefined ? { notDuplicate: patch.notDuplicate } : {}), + ...(patch.confirmed !== undefined ? { confirmed: patch.confirmed } : {}), ...(patch.note !== undefined ? { note: patch.note } : {}), decidedAt: new Date().toISOString(), }; // An empty patch clears the decision (back to "awaiting review") rather than - // leaving a decidedAt-only stub that would read as reviewed. - if (!next.canonicalSlug && next.notDuplicate !== true) { + // leaving a decidedAt-only stub that would read as reviewed. `confirmed` has + // to be part of that test: without it, clearing a canonical choice on a + // confirmed suspect would DELETE the confirmation and silently un-share the + // cluster's derived work. + if (!next.canonicalSlug && next.notDuplicate !== true && next.confirmed !== true) { delete current.clusters[clusterId]; } else { current.clusters[clusterId] = next; @@ -429,6 +737,193 @@ export async function updateDuplicateOverride( return out; } +// --------------------------------------------------------------------------- +// Blocking strategies — they NOMINATE, they never decide +// --------------------------------------------------------------------------- + +// Counters the generators fill in as they run, so the caller can report what was +// covered AND what was skipped. Skipping without reporting reads as "covered +// everything" when it did not. +type BlockStats = { + titleGroups: number; + oversizedGroups: number; + oversizedGroupVideos: number; + oversizedBlocks: number; + oversizedBlockVideos: number; + untitled: number; +}; + +// TITLE BLOCKING. One pass to group by exact normalized title, then pair within +// each group when the runtimes agree. Quadratic only INSIDE a group, and a group +// is a handful of videos, so the whole pass is effectively linear. +// +// The key is deliberately EXACT rather than fuzzy: a looser key merges "Episode +// 12" with "Episode 13", and every such merge is a false cluster a human then +// has to reject. The cost is recall on re-titled mirrors, which duration +// blocking covers instead — see STATE.md for the deferred middle grounds. +function* titleBlocks( + participants: Entry[], + ratio: number, + stats: BlockStats, + log: (msg: string) => void, +): Generator<Block> { + const byTitle = new Map<string, Entry[]>(); + for (const e of participants) { + const key = normalizeTitleKey(e.stat.title ?? ""); + if (key.length < MIN_TITLE_KEY_LENGTH) { + stats.untitled++; + continue; + } + const arr = byTitle.get(key); + if (arr) arr.push(e); + else byTitle.set(key, [e]); + } + + // Decide the whole work list BEFORE yielding any of it, so the summary — and + // in particular what was SKIPPED — is reported up front rather than after the + // long evaluation it describes. + const accepted: Entry[][] = []; + for (const [key, group] of byTitle) { + if (group.length < 2) continue; + if (group.length > MAX_TITLE_GROUP_SIZE) { + // A FORMAT, not a title — "live stream", "untitled", a daily show's + // date-less name. Reported, never silently truncated. + stats.oversizedGroups++; + stats.oversizedGroupVideos += group.length; + log( + ` skipping title group of ${group.length} (over ${MAX_TITLE_GROUP_SIZE}): "${key.slice(0, 60)}"`, + ); + continue; + } + accepted.push(group); + } + log( + `Title blocking: ${byTitle.size} distinct title(s), ${accepted.length} group(s) with 2+ members ` + + `(${stats.untitled} video(s) with no usable title; ${stats.oversizedGroups} group(s) over ` + + `${MAX_TITLE_GROUP_SIZE} skipped, covering ${stats.oversizedGroupVideos} video(s)).`, + ); + + for (const group of accepted) { + const nominations: Nomination[] = []; + for (let i = 0; i < group.length; i++) { + for (let j = i + 1; j < group.length; j++) { + if ( + durationsCompatible( + group[i].stat.duration, + group[j].stat.duration, + ratio, + ) + ) { + nominations.push({ + a: group[i], + b: group[j], + containmentOnly: false, + // Same title AND a near-identical runtime: strong enough to be + // worth a human's attention when no transcript can settle it. + suspectKind: "title-duration", + }); + } + } + } + if (nominations.length === 0) continue; + stats.titleGroups++; + yield { members: membersOf(nominations), nominations }; + } +} + +// DURATION BLOCKING. One block per rounded-duration bucket, unioned with the +// next so a pair straddling a bucket edge is still nominated. Emitting +// here×here and here×next (never next×next) makes every pair appear exactly +// once, and yielding bucket b+1's members as part of block b is what lets the +// caller's 2-block fingerprint cache halve the transcript reads. +// +// This is the only strategy that catches a RE-TITLED mirror, and it is viable +// corpus-wide only because blocks are evaluated and released one at a time. +function* durationBlocks( + participants: Entry[], + W: number, + stats: BlockStats, + log: (msg: string) => void, +): Generator<Block> { + const bucketOf = (d: number) => Math.round(d / W); + const buckets = new Map<number, Entry[]>(); + for (const e of participants) { + const b = bucketOf(e.stat.duration); + const arr = buckets.get(b); + if (arr) arr.push(e); + else buckets.set(b, [e]); + } + const maxDelta = (d: number) => Math.max(W, Math.ceil(0.02 * d)); + const sorted = [...buckets.keys()].sort((x, y) => x - y); + + // As with title blocking: decide the work list, report it (skips included), + // then do it. A summary that only arrives after an hour of evaluation is not a + // summary anyone can act on. + const accepted: number[] = []; + for (const b of sorted) { + const size = + (buckets.get(b) as Entry[]).length + (buckets.get(b + 1)?.length ?? 0); + // A bucket can itself be pathological — round numbers attract videos, and a + // corpus can hold thousands that are exactly 60 s. Same cap-and-report + // treatment as an oversized title group. + if (size > MAX_DURATION_BLOCK_SIZE) { + stats.oversizedBlocks++; + stats.oversizedBlockVideos += (buckets.get(b) as Entry[]).length; + log( + ` skipping duration block ~${b * W}s of ${size} video(s) ` + + `(over ${MAX_DURATION_BLOCK_SIZE}).`, + ); + continue; + } + accepted.push(b); + } + log( + `Duration blocking: ${buckets.size} bucket(s) of ${W}s, ${accepted.length} block(s) to evaluate` + + (stats.oversizedBlocks > 0 + ? ` (${stats.oversizedBlocks} block(s) over ${MAX_DURATION_BLOCK_SIZE} skipped, ` + + `covering ${stats.oversizedBlockVideos} video(s))` + : "") + + ".", + ); + + for (const b of accepted) { + const here = buckets.get(b) as Entry[]; + const next = buckets.get(b + 1) ?? []; + const nominations: Nomination[] = []; + for (let i = 0; i < here.length; i++) { + for (let j = i + 1; j < here.length; j++) { + if (durationsClose(here[i], here[j], maxDelta)) { + // A shared duration ALONE is not evidence — it produced enormous + // false clusters of unrelated same-length videos. If the content + // cannot be compared, this pair is worth nothing. + nominations.push({ + a: here[i], + b: here[j], + containmentOnly: false, + suspectKind: null, + }); + } + } + } + for (const x of here) { + for (const y of next) { + if (durationsClose(x, y, maxDelta)) { + nominations.push({ + a: x, + b: y, + containmentOnly: false, + suspectKind: null, + }); + } + } + } + if (nominations.length === 0) continue; + // Members are here ∪ next, so the caller's sliding cache carries bucket b+1 + // straight into the next block. + yield { members: membersOf(nominations), nominations }; + } +} + // --- helpers --------------------------------------------------------------- function durationsClose( @@ -541,6 +1036,16 @@ function evalSimilar( if (a.hash && a.hash === b.hash) { return { kind: "transcript-exact", score: 1, contained: false }; } + // Size prune. Exact, not a heuristic: |A∩B| ≤ min and |A∪B| ≥ max, so + // J ≤ min/max. A pair whose shingle counts are further apart than the + // threshold CANNOT reach it, and skipping it changes no result. Worth having + // because it is O(1) where the Jaccard it replaces is O(min set size), and + // corpus-wide duration blocking evaluates millions of pairs whose members + // share a runtime but not a word count (a music video and a lecture can both + // be ten minutes long). + const small = Math.min(a.shingleSet.size, b.shingleSet.size); + const large = Math.max(a.shingleSet.size, b.shingleSet.size); + if (large === 0 || small / large < nearThreshold) return null; const j = jaccard(a.shingleSet, b.shingleSet); if (j >= nearThreshold) { return { kind: "transcript-near", score: round3(j), contained: false }; @@ -548,6 +1053,46 @@ function evalSimilar( return null; } +// The verdict on one nominated pair — the point where the pre-filter's proposal +// meets the transcript's disposal. Three outcomes, and the middle one is the +// reason nominating aggressively is safe: +// +// both transcripts, content agrees → confirmed (may share) +// both transcripts, content differs → null. REJECTED. Two episodes of a daily +// show can share a title and a runtime and +// be entirely different material; testing +// them is what keeps that out. +// a transcript is missing → nothing can compare the content, so the +// pre-filter's own claim is all there is. +// Kept as a needsReview suspect when that +// claim is strong (title + runtime), and +// dropped when it is not (duration alone). +function evalBlocked( + a: Fingerprint, + b: Fingerprint, + opts: { + nearThreshold: number; + containmentThreshold: number; + allowContainment: boolean; + suspectKind: DuplicateMatchKind | null; + }, +): Omit<PairResult, "a" | "b"> | null { + if (a.hasTranscript && b.hasTranscript) { + const similar = evalSimilar(a, b, opts.nearThreshold); + if (similar) return similar; + // Containment as a FALLBACK inside a pair we already nominated: it costs one + // more set intersection over shingle sets that are already in hand, so it is + // free relative to the nomination that got us here. + if (opts.allowContainment) { + const contained = evalContainment(a, b, opts.containmentThreshold); + if (contained) return contained; + } + return null; + } + if (!opts.suspectKind) return null; + return { kind: opts.suspectKind, score: null, contained: false }; +} + // Containment pair: the shorter transcript's shingles are largely a subset of // the longer one's. Requires both transcripts (no metadata fallback here). function evalContainment( @@ -563,8 +1108,10 @@ function evalContainment( return null; } -function toRef(fp: Fingerprint): DuplicateVideoRef { - const s = fp.stat; +// Built from the participant index rather than a fingerprint: fingerprints are +// released with their block, and VideoStat is the small half of what they held. +// `offsetSeconds` / `aligned` are filled in afterwards by the alignment pass. +function toRef(s: VideoStat, hasTranscript: boolean): DuplicateVideoRef { return { slug: s.slug, channelSlug: s.channelSlug, @@ -574,7 +1121,7 @@ function toRef(fp: Fingerprint): DuplicateVideoRef { title: s.title, duration: s.duration, uploadDate: s.uploadDate, - hasTranscript: fp.hasTranscript, + hasTranscript, }; } diff --git a/common/lib/duplicates.test.ts b/common/lib/duplicates.test.ts @@ -1,12 +1,19 @@ import { test } from "node:test"; import assert from "node:assert/strict"; import { + DEFAULT_TITLE_DURATION_RATIO, + TITLE_DURATION_MIN_TOLERANCE_SECONDS, + clusterIsPublishable, clusterMaySharePartial, + durationsCompatible, + filterClusterToChannels, isClusterReviewed, measureAlignment, + normalizeTitleKey, pickCanonicalSlug, resolveCanonicalSlug, sanitizeDuplicateOverrides, + strongerMatch, type DuplicateCluster, type DuplicateVideoRef, } from "./duplicates"; @@ -162,6 +169,290 @@ test("sanitizeDuplicateOverrides survives a non-object file", () => { assert.deepEqual(sanitizeDuplicateOverrides([1, 2, 3]).clusters, {}); }); +// `confirmed` is the ONLY thing that lets a title+duration suspect ship or share +// derived work. If the sanitizer dropped it, every human confirmation would be +// silently discarded on the next read and the cluster would quietly revert to +// "unreviewed" — a failure that leaves no trace anywhere. +test("sanitizeDuplicateOverrides preserves confirmed", () => { + const o = sanitizeDuplicateOverrides({ + clusters: { c1: { confirmed: true }, c2: { confirmed: "yes" } }, + }); + assert.equal(o.clusters.c1?.confirmed, true); + // Only a real boolean true counts; a truthy string is not a decision. + assert.equal(o.clusters.c2, undefined); +}); + +test("isClusterReviewed accepts a confirmed-only entry", () => { + // A confirmation with no canonical override IS a decision — it is the whole + // point of the review queue — so it must clear the cluster off the worklist. + const c = cluster([ref({ slug: "a/1" }), ref({ slug: "b/2" })]); + assert.equal( + isClusterReviewed( + c, + sanitizeDuplicateOverrides({ clusters: { cluster1: { confirmed: true } } }), + ), + true, + ); +}); + +// --------------------------------------------------------------------------- +// Suspects: needsReview gates sharing until a human confirms +// --------------------------------------------------------------------------- + +test("a needsReview cluster shares nothing without a confirmation", () => { + // Title + near-identical runtime is a suspicion, not a comparison. Two + // episodes of a daily show can share both and be entirely different material, + // and a shared digest would then describe the wrong video convincingly. + const c = cluster([ref({ slug: "a/1" }), ref({ slug: "b/2" })], { + matchKind: "title-duration", + score: null, + needsReview: true, + }); + assert.equal(clusterMaySharePartial(c, null), false); + assert.equal( + clusterMaySharePartial( + c, + sanitizeDuplicateOverrides({ clusters: { cluster1: { canonicalSlug: "a/1" } } }), + ), + false, + "picking a canonical member is not the same as confirming the duplicate", + ); +}); + +test("confirmed unblocks a needsReview cluster", () => { + const c = cluster([ref({ slug: "a/1" }), ref({ slug: "b/2" })], { + matchKind: "title-duration", + score: null, + needsReview: true, + }); + assert.equal( + clusterMaySharePartial( + c, + sanitizeDuplicateOverrides({ clusters: { cluster1: { confirmed: true } } }), + ), + true, + ); +}); + +test("confirmed does NOT unblock a contained (clip-of-longer) cluster", () => { + // A clip is a different artifact from the recording it was cut from, however + // confident a human is that they are related. + const c = cluster([ref({ slug: "a/1" }), ref({ slug: "b/2" })], { + contained: true, + }); + assert.equal( + clusterMaySharePartial( + c, + sanitizeDuplicateOverrides({ clusters: { cluster1: { confirmed: true } } }), + ), + false, + ); +}); + +// --------------------------------------------------------------------------- +// What reaches a built site +// --------------------------------------------------------------------------- + +test("an unconfirmed suspect does NOT ship to a built site", () => { + // compose-site applies this. A suspect asserts a relationship no machine and + // no human has checked, so it stays an internal review queue. + const c = cluster([ref({ slug: "a/1" }), ref({ slug: "b/2" })], { + matchKind: "title-duration", + score: null, + needsReview: true, + }); + assert.equal(clusterIsPublishable(c, null), false); + assert.equal( + clusterIsPublishable( + c, + sanitizeDuplicateOverrides({ clusters: { cluster1: { canonicalSlug: "a/1" } } }), + ), + false, + ); + assert.equal( + clusterIsPublishable( + c, + sanitizeDuplicateOverrides({ clusters: { cluster1: { confirmed: true } } }), + ), + true, + ); +}); + +test("a content-confirmed cluster ships with no overrides at all", () => { + const c = cluster([ref({ slug: "a/1" }), ref({ slug: "b/2" })]); + assert.equal(clusterIsPublishable(c, null), true); +}); + +test("a contained cluster ships even though it shares nothing", () => { + // The two predicates deliberately disagree here: "this is a clip of that" is a + // real relationship worth showing a viewer, and simultaneously a reason never + // to copy the longer video's derived work onto the clip. + const c = cluster([ref({ slug: "a/1" }), ref({ slug: "b/2" })], { + contained: true, + }); + assert.equal(clusterIsPublishable(c, null), true); + assert.equal(clusterMaySharePartial(c, null), false); +}); + +// --------------------------------------------------------------------------- +// Match tiers +// --------------------------------------------------------------------------- + +test("strongerMatch ranks exact > near > title-duration", () => { + assert.equal(strongerMatch("transcript-near", "transcript-exact"), "transcript-exact"); + assert.equal(strongerMatch("transcript-exact", "transcript-near"), "transcript-exact"); + assert.equal(strongerMatch("title-duration", "transcript-near"), "transcript-near"); + assert.equal(strongerMatch("transcript-near", "title-duration"), "transcript-near"); + assert.equal(strongerMatch("title-duration", "title-duration"), "title-duration"); + // Load-bearing for needsReview: a cluster is only a suspect when EVERY pair in + // it was untestable, so one content-confirmed pair must dominate. + assert.equal(strongerMatch("title-duration", "transcript-exact"), "transcript-exact"); +}); + +// --------------------------------------------------------------------------- +// The title blocking key +// --------------------------------------------------------------------------- + +test("normalizeTitleKey folds case, punctuation and diacritics", () => { + assert.equal(normalizeTitleKey("Pokémon: The First Movie!"), "pokemon the first movie"); + assert.equal(normalizeTitleKey(" Multiple spaces "), "multiple spaces"); + assert.equal( + normalizeTitleKey("The Show — Episode 4"), + normalizeTitleKey("the show episode 4"), + ); +}); + +test("normalizeTitleKey drops platform re-upload suffixes", () => { + // These are what a mirror ADDS to an otherwise identical title, so folding + // them is what makes the mirror block with its original. + const base = normalizeTitleKey("Weekly Roundup"); + assert.equal(normalizeTitleKey("Weekly Roundup (reupload)"), base); + assert.equal(normalizeTitleKey("Weekly Roundup [mirror]"), base); + assert.equal(normalizeTitleKey("Weekly Roundup #shorts"), base); +}); + +test("normalizeTitleKey does NOT merge a series", () => { + // The whole reason the key is exact rather than fuzzy: a looser key merges + // consecutive episodes, and every such merge is a false cluster a human then + // has to reject. + assert.notEqual(normalizeTitleKey("Episode 12"), normalizeTitleKey("Episode 13")); + assert.notEqual( + normalizeTitleKey("Morning Show Jan 4"), + normalizeTitleKey("Morning Show Jan 5"), + ); +}); + +test("durationsCompatible allows a re-encode's drift but not a different cut", () => { + // 2% of the longer runtime, floored at 2s. + assert.equal(durationsCompatible(3600, 3620), true, "20s on an hour is 0.6%"); + assert.equal(durationsCompatible(3600, 3800), false, "200s on an hour is 5.6%"); + // The floor keeps sub-second rounding from splitting two short clips: 2% of + // 30s is 0.6s, which nothing survives. + assert.equal(durationsCompatible(30, 31), true); + assert.equal(durationsCompatible(30, 40), false); + assert.equal( + durationsCompatible(1000, 1000 + 1000 * DEFAULT_TITLE_DURATION_RATIO), + true, + "exactly at the ratio is compatible", + ); + assert.equal( + durationsCompatible(100, 100 + TITLE_DURATION_MIN_TOLERANCE_SECONDS), + true, + "exactly at the floor is compatible", + ); +}); + +test("durationsCompatible refuses a missing or zero duration", () => { + // An unknown runtime is not a match — treating 0 as "close to 0" would block + // every metadata-less video together. + assert.equal(durationsCompatible(0, 0), false); + assert.equal(durationsCompatible(600, 0), false); + assert.equal(durationsCompatible(-5, -5), false); +}); + +// --------------------------------------------------------------------------- +// Narrowing a global cluster to one site's channels +// --------------------------------------------------------------------------- + +test("filterClusterToChannels keeps in-site members and recomputes the flags", () => { + const c = cluster( + [ + ref({ slug: "a/1", channelSlug: "a", platform: "youtube" }), + ref({ slug: "b/2", channelSlug: "b", platform: "rumble" }), + ref({ slug: "c/3", channelSlug: "c", platform: "odysee" }), + ], + { crossPlatform: true, crossChannel: true }, + ); + const out = filterClusterToChannels(c, new Set(["a", "b"])); + assert.ok(out); + assert.deepEqual(out.videoRefs.map((r) => r.slug), ["a/1", "b/2"]); + assert.equal(out.crossPlatform, true); + assert.equal(out.crossChannel, true); +}); + +test("filterClusterToChannels drops a cluster that falls below two members", () => { + // One video is not a visible duplicate — there is nothing to switch to. + const c = cluster([ + ref({ slug: "a/1", channelSlug: "a" }), + ref({ slug: "b/2", channelSlug: "b" }), + ]); + assert.equal(filterClusterToChannels(c, new Set(["a"])), null); + assert.equal(filterClusterToChannels(c, new Set(["z"])), null); +}); + +test("filterClusterToChannels clears cross-* flags the survivors no longer earn", () => { + const c = cluster( + [ + ref({ slug: "a/1", channelSlug: "a", platform: "youtube" }), + ref({ slug: "a/2", channelSlug: "a", platform: "youtube" }), + ref({ slug: "b/3", channelSlug: "b", platform: "rumble" }), + ], + { crossPlatform: true, crossChannel: true }, + ); + const out = filterClusterToChannels(c, new Set(["a"])); + assert.ok(out); + assert.equal(out.crossPlatform, false); + assert.equal(out.crossChannel, false); +}); + +test("filterClusterToChannels carries needsReview and the alignment fields through", () => { + // The site-narrowing step must not launder a suspect into a shipped cluster, + // and it must not lose the per-member alignment the viewer's jump depends on. + const c = cluster( + [ + ref({ slug: "a/1", channelSlug: "a", aligned: true, offsetSeconds: 0 }), + ref({ slug: "b/2", channelSlug: "b", aligned: false, offsetSeconds: 41.5 }), + ref({ slug: "c/3", channelSlug: "c" }), + ], + { matchKind: "title-duration", score: null, needsReview: true, canonicalSlug: "a/1" }, + ); + const out = filterClusterToChannels(c, new Set(["a", "b"])); + assert.ok(out); + assert.equal(out.needsReview, true); + assert.equal(out.canonicalSlug, "a/1"); + assert.equal(out.videoRefs[0].aligned, true); + assert.equal(out.videoRefs[1].aligned, false); + assert.equal(out.videoRefs[1].offsetSeconds, 41.5); +}); + +test("filterClusterToChannels leaves matchKind and score as the detector reported them", () => { + // Intentional and documented: they describe the strongest pair in the FULL + // cluster. Recomputing them here would change the meaning of clusters already + // shipped, and there is no similarity data at this point to recompute from. + const c = cluster( + [ + ref({ slug: "a/1", channelSlug: "a" }), + ref({ slug: "b/2", channelSlug: "b" }), + ref({ slug: "c/3", channelSlug: "c" }), + ], + { matchKind: "transcript-exact", score: 1 }, + ); + const out = filterClusterToChannels(c, new Set(["a", "b"])); + assert.ok(out); + assert.equal(out.matchKind, "transcript-exact"); + assert.equal(out.score, 1); +}); + // --------------------------------------------------------------------------- // The alignment gate // --------------------------------------------------------------------------- diff --git a/common/lib/duplicates.ts b/common/lib/duplicates.ts @@ -24,10 +24,41 @@ export const DEFAULT_CONTAINMENT_THRESHOLD = 0.8; // unigrams. export const DEFAULT_SHINGLE_SIZE = 5; -// Only content-confirmed tiers form clusters. Duration coincidence alone is -// never treated as a match (it produced enormous false clusters of unrelated -// same-length videos). -export type DuplicateMatchKind = "transcript-exact" | "transcript-near"; +// Match tiers, weakest first. +// +// `title-duration` is a SUSPECT, not a confirmation. Duration coincidence alone +// was never a match — it produced enormous false clusters of unrelated +// same-length videos — but the same title AND a near-identical runtime is a +// different claim entirely, and it is the only signal available when one side +// has no transcript to compare. Such a cluster is flagged `needsReview` and +// shares nothing until a human confirms it (see clusterMaySharePartial). +export type DuplicateMatchKind = + | "title-duration" + | "transcript-near" + | "transcript-exact"; + +// Two videos with the same normalized title are candidates when their runtimes +// agree within this fraction of the longer one. A mirror re-encode drifts by a +// second or two; 2% also absorbs an ad-break difference on a long video without +// admitting a genuinely different cut. +export const DEFAULT_TITLE_DURATION_RATIO = 0.02; +// ...but never demand tighter than this, so two 30-second clips are not split by +// sub-second rounding. +export const TITLE_DURATION_MIN_TOLERANCE_SECONDS = 2; + +// A title group larger than this is a FORMAT, not a title — "live stream", +// "untitled", a daily show's date-less name. Pairing inside it is quadratic and +// the matches would be noise, so the group is skipped and reported rather than +// silently truncated. +export const MAX_TITLE_GROUP_SIZE = 40; + +// The same treatment for a pathological DURATION block. Round numbers attract +// videos — a corpus can hold thousands of videos that are exactly 60s — and one +// such block is quadratic on its own. Blocks over this are skipped and reported, +// for the same reason: a silent truncation reads as "covered everything". +// Generous, because the proven shorts-mode run has legitimately large buckets +// (34,915 shorts over ~90 two-second buckets) and must keep behaving as it does. +export const MAX_DURATION_BLOCK_SIZE = 2000; export type DuplicateVideoRef = { slug: string; // `${channelSlug}/${id}` @@ -39,6 +70,20 @@ export type DuplicateVideoRef = { duration: number; // seconds uploadDate: string; // YYYYMMDD hasTranscript: boolean; + // Timing alignment against the cluster's canonical member, measured by + // measureAlignment() at detection time (confirmed clusters only — see the + // detector). Both are optional: reports written before these fields existed + // lack them, and absent must be read as "not measured", i.e. NOT aligned. + // + // `offsetSeconds` is the largest |offset| observed across the matched anchors, + // not a signed shift — the same quantity writeSharedDigest already records. + // It is informational; `aligned` is the gate. Anything that seeks INTO a + // sibling (the search-result duplicate badge, a shared digest's chapters) must + // key off `aligned`, because a mirror with a longer intro matches on text at + // shifted times and would otherwise land in the wrong place while looking + // perfectly plausible. + offsetSeconds?: number | null; + aligned?: boolean; }; export type DuplicateCluster = { @@ -49,6 +94,12 @@ export type DuplicateCluster = { durationBucket: number; // representative rounded duration (seconds) crossPlatform: boolean; // members span more than one platform crossChannel: boolean; // members span more than one channelSlug + // True when the cluster's strongest evidence is title+duration only — nothing + // compared the actual content. It is offered for human review and shares no + // derived work until confirmed. Optional: reports written before this field + // existed lack it, and absent means "content-confirmed", which is what those + // reports only ever contained. + needsReview?: boolean; videoRefs: DuplicateVideoRef[]; // The member that OWNS derived work for this cluster: the one an AI digest is // generated for, and the one aligned mirrors copy it from. Chosen by @@ -58,12 +109,29 @@ export type DuplicateCluster = { canonicalSlug?: string; }; +// How candidate pairs are NOMINATED. A blocking strategy never decides anything +// on its own — it only proposes pairs that the transcript cascade then confirms +// or rejects (see the controller's evalBlocked). The choice is therefore about +// recall and cost, not correctness. +// +// "duration" — same/adjacent rounded-duration bucket. The only strategy that +// catches a RE-TITLED mirror. Corpus-viable only since the +// streaming rewrite; still much the more expensive of the two. +// "title" — exact normalized-title groups, paired when the runtimes agree. +// Effectively linear, and it is what actually finds cross-platform +// re-uploads, which keep their name. +// "both" — the union, deduped. +export type DuplicateBlocking = "duration" | "title" | "both"; + export type DuplicateRunConfig = { thresholdSeconds: number | null; // null === all durations durationToleranceSeconds: number; nearThreshold: number; // Jaccard containmentThreshold: number; shingleSize: number; + // Optional: reports written before blocking was configurable lack it, and + // those were all duration-blocked. + blocking?: DuplicateBlocking; }; export type DuplicateReport = { @@ -100,6 +168,10 @@ export type DuplicateClusterOverride = { // keep finding it (content really is similar), which is exactly why the // decision has to be recorded outside the report. notDuplicate?: boolean; + // "I looked, and these really are the same video." The positive counterpart of + // notDuplicate, and the ONLY thing that lets a title-duration cluster share + // derived work. Content-confirmed clusters do not need it. + confirmed?: boolean; decidedAt?: string; note?: string; }; @@ -132,6 +204,7 @@ export function sanitizeDuplicateOverrides(value: unknown): DuplicateOverrides { override.canonicalSlug = e.canonicalSlug.trim(); } if (e.notDuplicate === true) override.notDuplicate = true; + if (e.confirmed === true) override.confirmed = true; if (typeof e.decidedAt === "string") override.decidedAt = e.decidedAt; if (typeof e.note === "string" && e.note.trim()) override.note = e.note.trim(); if (Object.keys(override).length === 0) continue; @@ -226,7 +299,9 @@ export function isClusterReviewed( ): boolean { const o = overrides?.clusters[cluster.clusterId]; if (!o) return false; - return o.notDuplicate === true || Boolean(o.canonicalSlug); + return ( + o.notDuplicate === true || o.confirmed === true || Boolean(o.canonicalSlug) + ); } // --------------------------------------------------------------------------- @@ -407,8 +482,83 @@ export function measureAlignment( // video, not a mirror of it. A clip is a different artifact: the longer video's // chapters describe material the clip does not contain, so sharing wholesale // would be wrong even at a perfect zero offset. -export function clusterMaySharePartial(cluster: DuplicateCluster): boolean { - return !cluster.contained; +// +// A `needsReview` cluster (title+duration only) shares nothing either, until a +// human records `confirmed: true`. The alignment gate would still protect the +// TIMING, but nothing here has compared the CONTENT — two episodes of a daily +// show can share a title and a runtime and be entirely different material, and +// a shared digest would then describe the wrong video convincingly. +export function clusterMaySharePartial( + cluster: DuplicateCluster, + overrides?: DuplicateOverrides | null, +): boolean { + if (cluster.contained) return false; + if (!cluster.needsReview) return true; + return overrides?.clusters[cluster.clusterId]?.confirmed === true; +} + +// Whether a cluster may be SHIPPED to a built site at all. +// +// Distinct from clusterMaySharePartial, which asks whether derived work may flow +// between members. The two disagree in both directions, on purpose: +// +// a `contained` cluster SHIPS (a clip of a longer video is a real, useful +// relationship for a viewer to see) but SHARES NOTHING (the longer video's +// chapters describe material the clip does not contain); +// +// an unconfirmed `needsReview` cluster does NEITHER — nothing compared its +// members' content, so asserting the relationship to a viewer would be +// claiming something no machine and no human has actually checked. It stays an +// internal review queue until someone records `confirmed`. +// +// Fails closed: no overrides means no confirmations means no suspects ship. +export function clusterIsPublishable( + cluster: DuplicateCluster, + overrides?: DuplicateOverrides | null, +): boolean { + if (!cluster.needsReview) return true; + return overrides?.clusters[cluster.clusterId]?.confirmed === true; +} + +// The blocking key for title-based candidate generation. +// +// This is what makes corpus-wide detection tractable. The alternative the first +// implementation used — pairing every short with every longer video — is +// quadratic and measured at 498 MILLION pairs on this corpus. Grouping by an +// exact normalized title instead is one pass and a hash lookup, and it is the +// signal that actually finds cross-platform mirrors, which are re-uploads of the +// same file under the same name. +// +// Deliberately conservative: lowercase, strip punctuation and diacritics, drop a +// few platform suffixes, collapse whitespace. It does NOT stem, fuzzy-match or +// drop stopwords — a looser key merges a series ("Episode 12" vs "Episode 13") +// and every such merge is a false cluster a human then has to reject. +const TITLE_NOISE_RE = + /\s*(?:#shorts?|\(official(?: video| audio)?\)|\[official\]|\|\s*full episode|\(full episode\)|\(reupload\)|\[reupload\]|\(mirror\)|\[mirror\])\s*/gi; + +export function normalizeTitleKey(title: string): string { + return title + .normalize("NFKD") + // Strip combining marks so "Pokémon" and "Pokemon" block together. + .replace(/[\u0300-\u036f]/g, "") + .toLowerCase() + .replace(TITLE_NOISE_RE, " ") + .replace(/[^\p{L}\p{N}]+/gu, " ") + .trim(); +} + +// Are two same-titled videos close enough in runtime to be candidates? +export function durationsCompatible( + a: number, + b: number, + ratio = DEFAULT_TITLE_DURATION_RATIO, +): boolean { + if (!(a > 0) || !(b > 0)) return false; + const tolerance = Math.max( + TITLE_DURATION_MIN_TOLERANCE_SECONDS, + Math.max(a, b) * ratio, + ); + return Math.abs(a - b) <= tolerance; } // --------------------------------------------------------------------------- @@ -512,6 +662,7 @@ export class UnionFind { // Tier ranking so a cluster reports its strongest evidence. const MATCH_RANK: Record<DuplicateMatchKind, number> = { + "title-duration": 0, "transcript-near": 1, "transcript-exact": 2, }; diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,6 +1,9 @@ # Changelog ## [Unreleased] +- **Duplicate detection now runs over the whole archive, not just shorts.** It previously ran out of memory on a full-corpus pass and was left off. Two things were actually wrong, and both are fixed: candidate pairs were generated by pairing every short with every longer video (half a billion pairs), and transcript fingerprints for the entire corpus were held in memory at once. Detection now works one block of similar videos at a time — fingerprint, compare, discard — so a corpus-wide run finishes in minutes at ordinary memory. A new **`--blocking`** flag on `duplicate-shorts` chooses how candidates are proposed: `title` (default for a full-archive run — fast, finds cross-platform re-uploads that kept their name), `duration` (slower, but the only one that finds a *re-titled* mirror), or `both`. The choice only affects which pairs get *considered*; what counts as a duplicate is still decided by comparing the actual transcripts. +- **Videos that share a title and a runtime are now surfaced for review instead of being dropped.** When one side has no transcript there is nothing to compare, so detection can't rule either way. Rather than discarding the pair, it is reported as a cluster badged **needs review** — visible in the archive's own tooling, excluded from every built site, and blocked from sharing AI digests until you confirm it. Clusters whose transcripts *were* compared are unaffected and behave exactly as before. Note that this test is doing real work: on the full archive, **63% of same-title, same-length pairs turned out not to be the same video.** +- **Detection now records how well each copy's timings line up with the original.** Every confirmed cluster stores, per member, whether its transcript timings align with the cluster's canonical video and by how much. This is what lets the viewer offer "jump to this moment in the other copy" only when that moment actually corresponds — and lets a shared AI digest be placed correctly rather than plausibly. - **Every transcript can now be given AI-generated chapters and topic tags, from a model running on your own hardware.** A new **Digest** stage on each channel page sweeps its transcripts through a local model (ollama by default) and writes an `ai-digest.json` next to each one — a list of titled moments with timestamps, plus a short set of topic tags. Nothing leaves the machine unless you opt in: a second, **metered** lane (Claude Code) exists for the long tail of very long videos and is **off by default**, gated behind its own spend cap. The two lanes run on separate job queues on purpose — the local lane is GPU-bound and the metered one is network-bound, so sharing a queue would have halved the throughput of a sweep measured in weeks. **A re-run is cheap by construction.** Each generated section records the exact identity that produced it — engine, requested model, prompt version, prompt shape, and a hash of the channel-context inputs — and regeneration skips any section whose identity already matches. That is what makes "digest this channel again" take minutes instead of restarting a multi-week job, and it is why an alias like `qwen2.5` resolving to `qwen2.5:7b` is deliberately *not* treated as a model change. **Hand corrections are kept in a separate file** (`ai-digest.overrides.json`) that generation never opens, so no merge bug in the generator can destroy work a human did; readers compose the two, an override replaces a generated item by id, and `enabled: false` suppresses one without deleting it, so a regeneration can't resurrect something you rejected. **The output is guarded, not trusted.** Schema-constrained decoding pins every timestamp to a full `HH:MM:SS` and demands English titles, and the parser re-checks each item against the chunk's real time range, monotonicity, seam duplication and empty titles — every rejection recorded as a `warning` on the artifact rather than silently dropped, because a sweep this long is only tunable if its failures are inspectable. Chunk size is **sized to the configured context window**: ollama's default 4096 silently truncates over-long input and the model then summarizes whatever fragment survived, which looks like a bad model and is actually a misconfiguration. Two timestamp numberings are shipped and both are selectable (`absolute`, and `chunk-local`, which re-bases each chunk to `00:00:00` and adds the offset back before any guard runs) because which one is better was a measured question, not a guess — a bake-off harness (`common/bin/digest-bakeoff.ts`, reports under `plans/bakeoff/`) exists to settle it, and the choice is folded into the recorded identity so switching modes correctly invalidates the corpus instead of silently skipping it. Digests are also **shared across duplicate clusters**: a confirmed mirror of an already-digested video borrows its digest rather than paying for it twice, but only when detection *measured* the two as aligned, and the borrowed copy records where it came from and by what offset. Also fixes two wiring bugs found only end-to-end: the job log parser threw on the first line of every digest job (both of its parsers are null for this task, and the code asserted one was not — a whole sweep would have reported "0 generated, N failed" and looked like an engine fault), and saving Settings rebuilt the digest block without carrying `timestampMode`/`promptVariant` through, which would have silently reset the prompt shape on an unrelated save and invalidated every digest generated under it. See `common/lib/{digest,digestPrompt,digestParse,digestApps}.ts`, `common/controller/{digestVideo,digestBatch,digestSharing}.ts`, `editor/app/channels/[slug]/{digestActions.ts,components/stages/DigestStage.tsx}`, and `editor/e2e/digest.spec.ts`. - **Replace YouTube's auto-captions with transcripts of our own.** Most of the corpus rides on YouTube ASR captions, which are noticeably worse than what the transcription workers produce — no punctuation, rolling duplicate cues, `[Music]` filler — and they were *sticky*: `isVideoTranscribed()` counts any English VTT, so a video with only auto-captions was permanently invisible to every transcribe bucket and every transcribe job. There is now an opt-in, strictly-lowest-priority lane that finds those videos, downloads their audio, and transcribes them properly; whisper's `transcript.json` then wins the index pick automatically. Provenance is decided by a 4 KB sniff of the VTT itself (YouTube ASR marks ~96–100% of cues with `align:start position:N%` plus inline word timings; manual tracks mark 0%), with the 490 KB `metadata.info.json` parse kept only as a tie-breaker — so the per-regen cost is one small read per English-VTT-having video. Three snapshot buckets carry it: `autoSubsOnly` (needs audio) → `downloadedAutoSubsOnly` (needs whisper) → `supersededAutoSubs` (done, old VTT kept as a backup). Nothing is automatic by default: the auto-queue gains a per-runner **Replace YouTube auto-captions** switch that appends the bucket to the *tail* of the default union (real work always drains first), and a leaf can target the bucket by name for per-channel opt-in — the default unions are byte-identical to before, so existing setups are untouched. Manually, the Transcribe stage gains a two-step "YouTube auto-captions only" section and a single-video *Replace auto-captions* action, and each subtitle track is now labelled *YouTube auto-captions* / *manual captions*. A caption track whose provenance can't be proven machine-generated is **never** a candidate, and the original VTT is never deleted automatically — the Cleanup stage's purge button is the only thing that removes it (English ASR tracks only, honouring `do-not-clean`), so an AI-vs-YouTube comparison stays possible. See `common/lib/subtitleProvenance.ts`, `common/controller/purgeSupersededAutoSubs.ts`, `common/jobs/autoQueuePolicy.ts`, `editor/e2e/auto-subs-replace.spec.ts`. - **Social posts are archived as a parallel corpus to video transcripts.** The archive can now ingest X/Twitter and Bluesky accounts from the same commentators and search them *together* with video transcripts — one corpus, one set of searches. A channel gains `sourceKind: "social"` (a separate axis from `handling`, so every existing `handling === "transcribe" ? … : …` branch stays binary and can never misroute), plus `postFetcher` and `socialHandle`. Posts are modelled on the live-chat layer, not the video layer: their own month-sharded JSONL on disk (`channels/<slug>/posts/YYYY-MM.jsonl` + a `posts-archive` of seen ids, so a re-run is a no-op), their own LMDB sub-DB keyed `[createdAt, channelSlug, id]` (ISO-8601 sorts chronologically, fixing the intra-day ordering the `YYYYMMDD` video key has), and their own `/posts/<slug>/{manifest,page-NNNN}.json` page tree. Ingest is a pluggable `SocialFetcher` registry (`common/social/fetchers.ts`) modelled on the transcription-app registry: **`bluesky-atproto`** (pure `fetch` against the public AT Protocol — no auth, no binary, verified end-to-end against a live account), **`x-gallery-dl`** (the primary X path: a light headless subprocess, cookies via the existing `cookiePolicy.ts`), and **`x-playwright`** (the fallback, immune to the GraphQL query-id rotations that periodically break gallery-dl). One `fetch-posts` job kind carries it, drainable and bookmarkable, routed to `platform:x.com` / `platform:bsky.app` by the existing queue keys. A social channel gets a minimal two-stage rail (Fetch → Index) instead of the six video stages, and the channel form hides every video-only control (audio format, download format, keep-source-video, extraction mode, saved-video dir). Auto-sync works unchanged — but the scheduler's *dispatch* now routes social channels to a post fetch rather than a yt-dlp video sync. See `common/lib/posts.ts`, `common/social/*`, `common/controller/fetchPosts.ts`, `editor/e2e/social-channel.spec.ts`. diff --git a/editor/app/actionable/page.tsx b/editor/app/actionable/page.tsx @@ -287,6 +287,7 @@ function DuplicateClusterCard({ cluster }: { cluster: DuplicateCluster }) { const matchLabel: Record<DuplicateCluster["matchKind"], string> = { "transcript-exact": "exact transcript", "transcript-near": "near transcript", + "title-duration": "title + runtime", }; return ( <li @@ -295,6 +296,10 @@ function DuplicateClusterCard({ cluster }: { cluster: DuplicateCluster }) { > <div className="flex items-center gap-2 flex-wrap text-xs"> <Badge>{matchLabel[cluster.matchKind]}</Badge> + {/* Nothing compared these videos' content — one side has no transcript. + The cluster is a suspect for a human, stays out of the built site, + and shares no derived work until someone confirms it. */} + {cluster.needsReview && <Badge>needs review</Badge>} {cluster.score !== null && ( <Badge>score {cluster.score.toFixed(2)}</Badge> )} diff --git a/editor/e2e/duplicate-shorts.spec.ts b/editor/e2e/duplicate-shorts.spec.ts @@ -9,17 +9,26 @@ import { readJson, resetData, resolvePath, writeSite } from "./helpers"; // Drives Build index + Build stats, then runs detection from /actionable and // asserts on transcripts/duplicates.json. -type DuplicateRef = { slug: string; channelSlug: string; platform: string }; +type DuplicateRef = { + slug: string; + channelSlug: string; + platform: string; + aligned?: boolean; + offsetSeconds?: number | null; +}; type DuplicateCluster = { + clusterId: string; matchKind: string; score: number | null; contained: boolean; crossPlatform: boolean; crossChannel: boolean; + needsReview?: boolean; + canonicalSlug?: string; videoRefs: DuplicateRef[]; }; type DuplicateReport = { - runConfig: { thresholdSeconds: number | null }; + runConfig: { thresholdSeconds: number | null; blocking?: string }; totals: { clusters: number }; clusters: DuplicateCluster[]; }; @@ -97,6 +106,9 @@ type Seed = { duration: number; transcript: string | null; plain?: boolean; // emit plain (non-karaoke) VTT + // Defaults to a per-id unique title, so the duration-blocking seeds below + // never accidentally block by title too. The title-blocking seeds set it. + title?: string; }; const SEEDS: Seed[] = [ @@ -125,8 +137,8 @@ const PLATFORM_KEY: Record<Seed["platform"], string> = { rumble: "Rumble", }; -async function seed(): Promise<void> { - const channels = new Set(SEEDS.map((s) => s.channel)); +async function seed(seeds: Seed[] = SEEDS): Promise<void> { + const channels = new Set(seeds.map((s) => s.channel)); for (const channel of channels) { const dir = resolvePath(`test-transcripts/channels/${channel}`); await mkdir(dir, { recursive: true }); @@ -135,14 +147,14 @@ async function seed(): Promise<void> { JSON.stringify({ handling: "transcribe", name: channel }), ); } - for (const s of SEEDS) { + for (const s of seeds) { const dir = resolvePath(`test-transcripts/channels/${s.channel}/data/${s.id}`); await mkdir(dir, { recursive: true }); await writeFile( `${dir}/metadata.info.json`, JSON.stringify({ id: s.id, - title: `Title ${s.id}`, + title: s.title ?? `Title ${s.id}`, channel: s.channel, upload_date: "20240101", duration: s.duration, @@ -235,11 +247,22 @@ test("clusters cross-platform near-duplicate shorts (incl. plain VTT) and ignore expect(plain!.videoRefs.map((r) => r.slug)).not.toContain("yt-a/plainx"); expect(clusterWith(report, "yt-a/plainx")).toBeUndefined(); - // Every reported cluster is content-confirmed (no metadata-only matches). + // Every cluster that would SHIP is content-confirmed. (The original invariant + // was "every reported cluster", which was true when duration coincidence + // produced nothing at all. Title + near-identical runtime is a far stronger + // claim and now does produce clusters — but they are quarantined behind + // needsReview, never auto-share, and never reach a built site, so the + // guarantee is kept exactly where it matters. These seeds have distinct titles + // anyway, so nothing here is a suspect.) for (const c of report.clusters) { + if (c.needsReview) { + expect(c.matchKind).toBe("title-duration"); + continue; + } expect(["transcript-exact", "transcript-near"]).toContain(c.matchKind); expect(c.score).not.toBeNull(); } + expect(report.clusters.filter((c) => c.needsReview)).toHaveLength(0); // The detected cluster renders on the page after a refresh. await page.reload(); @@ -276,3 +299,95 @@ test("all-durations run pulls a long video into the short's cluster via containm expect(cluster!.contained).toBe(true); expect(cluster!.videoRefs.map((r) => r.slug)).toContain("yt-d/longvid"); }); + +// --------------------------------------------------------------------------- +// Title blocking: the pre-filter proposes, the transcript disposes +// --------------------------------------------------------------------------- + +// Three same-title, compatible-runtime pairs that differ only in what the +// TRANSCRIPTS say. The point of the whole design is that the pre-filter treats +// all three identically and the content cascade then splits them three ways. +const TITLE_SEEDS: Seed[] = [ + // (1) Same title, same content → CONFIRMED. A cross-platform mirror. + { channel: "t-a", id: "mirror1", platform: "youtube", duration: 300, transcript: BASE_WORDS, title: "The Weekly Roundup" }, + { channel: "t-b", id: "mirror2", platform: "rumble", duration: 303, transcript: variant(5, "foxtrotx"), title: "The Weekly Roundup" }, + // (2) Same title, same runtime, DIFFERENT content → REJECTED. Two episodes of + // a daily show. This is the case that makes nominating aggressively safe. + { channel: "t-a", id: "epis1", platform: "youtube", duration: 240, transcript: BASE_WORDS, title: "Daily Show Recap" }, + { channel: "t-b", id: "epis2", platform: "rumble", duration: 241, transcript: UNIQUE_WORDS, title: "Daily Show Recap" }, + // (3) Same title, same runtime, one side has NO transcript → untestable, so a + // needsReview SUSPECT rather than a match or a silent drop. + { channel: "t-a", id: "susp1", platform: "youtube", duration: 200, transcript: BASE_WORDS, title: "Archive Upload" }, + { channel: "t-b", id: "susp2", platform: "rumble", duration: 200, transcript: null, title: "Archive Upload" }, + // (4) Same title but a genuinely different cut (2x the runtime) → not even + // nominated, so the durations gate is doing its job. + { channel: "t-a", id: "cut1", platform: "youtube", duration: 300, transcript: BASE_WORDS, title: "Extended Interview" }, + { channel: "t-b", id: "cut2", platform: "rumble", duration: 700, transcript: BASE_WORDS, title: "Extended Interview" }, +]; + +test("title blocking confirms matching content, rejects differing content, and quarantines the untestable", async ({ + page, +}) => { + await resetData(null); + await seed(TITLE_SEEDS); + await writeSite("testsite", { + channels: [...new Set(TITLE_SEEDS.map((s) => s.channel))].map((slug) => ({ + slug, + groupId: "default", + })), + }); + await buildData(page); + + // Scope "all" runs corpus-wide, where the default blocking strategy is title. + await page.goto("/actionable"); + await page.getByLabel("duplicate detection scope").selectOption("all"); + await page.getByRole("button", { name: "detect duplicate shorts" }).click(); + await expect(page.getByLabel("detect duplicate shorts result")).toBeVisible({ + timeout: 30_000, + }); + + const report = await readJson<DuplicateReport>("test-transcripts/duplicates.json"); + expect(report.runConfig.blocking).toBe("title"); + + // (1) CONFIRMED — the content agreed, so this is a real cluster that may share. + const mirror = clusterWith(report, "t-a/mirror1"); + expect(mirror).toBeTruthy(); + expect(mirror!.needsReview).toBeFalsy(); + expect(["transcript-exact", "transcript-near"]).toContain(mirror!.matchKind); + expect(mirror!.videoRefs.map((r) => r.slug).sort()).toEqual([ + "t-a/mirror1", + "t-b/mirror2", + ]); + // Alignment is measured and persisted for confirmed clusters — without it the + // viewer's "jump to this moment" has nothing honest to key off. + expect(mirror!.canonicalSlug).toBeTruthy(); + for (const ref of mirror!.videoRefs) { + expect(typeof ref.aligned).toBe("boolean"); + } + + // (2) REJECTED — same title, same runtime, different words. No cluster at all. + expect(clusterWith(report, "t-a/epis1")).toBeUndefined(); + expect(clusterWith(report, "t-b/epis2")).toBeUndefined(); + + // (3) SUSPECT — nothing could compare the content, so it is a review item. + const suspect = clusterWith(report, "t-a/susp1"); + expect(suspect).toBeTruthy(); + expect(suspect!.needsReview).toBe(true); + expect(suspect!.matchKind).toBe("title-duration"); + expect(suspect!.score).toBeNull(); + expect(suspect!.videoRefs.map((r) => r.slug).sort()).toEqual([ + "t-a/susp1", + "t-b/susp2", + ]); + + // (4) NOT NOMINATED — a 300s and a 700s video are a different cut, not a + // mirror, however identical their titles. + expect(clusterWith(report, "t-a/cut1")).toBeUndefined(); + expect(clusterWith(report, "t-b/cut2")).toBeUndefined(); + + // The suspect is visibly flagged in the editor's review list. + await page.reload(); + const card = page.getByLabel(`duplicate cluster ${suspect!.clusterId}`); + await expect(card).toContainText("needs review"); + await expect(card).toContainText("title + runtime"); +}); diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md @@ -1,6 +1,7 @@ # Changelog ## [Unreleased] +- **Search results tell you when a video exists elsewhere in the archive, and take you there.** A result that belongs to a duplicate cluster now carries a **Dupe** badge, and a strip under the card header offers one button per other copy — the same recording mirrored to another platform, or re-uploaded on another channel. Clicking one opens that copy in the player. **The jump is honest about what it knows:** matching content does not imply matching timings (a mirror with a longer intro carries the same words at shifted times), so a button only carries your current timestamp when detection *measured* the two as aligned; otherwise it says so and opens the other copy from the start. A copy whose alignment was never measured is treated as not aligned. Only clusters whose transcripts were actually compared reach the site — a pair that merely shares a title and a runtime stays an internal review item and is never asserted to you. In hub mode the badge is limited to same-origin results, since the duplicate index is per-site. Sites with no duplicate report are entirely unaffected. - **One search now covers video transcripts *and* social posts.** Archived X/Twitter and Bluesky posts ship as a parallel corpus beside transcripts and live chat, and compose into the same boolean query tree — so `(transcripts:"foo" OR posts:"foo")` returns both kinds in one ranked, newest-first result set. `LayerScope` gains `"posts"` (whitelisted in `qt=` deserialization, so a shared link round-trips a posts leaf), the leaf scope selector gains **Posts**, and the filter row gains a **Posts** media kind beside Videos and Livestreams — a third kind, because a post is neither, and folding it into the video toggle would silently drop the whole corpus. Post and video slugs live in disjoint namespaces, partitioned per-leaf by the eval engine so a transcripts leaf never fetches a post and an AND across the two can't collapse to nothing. Post result cards drop what doesn't apply (no seek gutter, no livestream/age badges, no VOD expiry) and lead with the post body; opening one shows a new **PostModal** — a sibling of the transcript reader, not a generalization of it — with the post, its archived thread, its outbound links and its engagement counts. Date filters work unchanged: every post carries a derived `uploadDate`. Posts are cached and served under `/posts/`, offline-cached by the service worker, and CORS-readable so a federating hub merges them across origins. - **"Ask AI" is grounded in posts as well as transcripts.** Retrieval adds a posts leaf per keyword alongside the transcript and metadata leaves, sharing the same term key so ranking still counts a keyword once rather than three times. Post excerpts render without a `[clock]` line and are cited as a bare `[n]` (never `[n @ mm:ss]` — a post has no timeline), the source list shows a date instead of a meaningless 0:00 seek button, and "load more context" on a post returns its **thread** rather than a time window. - **MCP: posts are first-class.** `search_transcripts` gains `content_types` (defaulting to **both**, so existing agent flows pick posts up automatically) and renders post hits without moment links; new `get_post` and `get_thread` tools read one post or a whole conversation; `open_link` accepts a `posts` query scope; and the sweep prompt teaches the post citation form. All of it lands on the single `ShardSource` boundary, so local, remote and hub transports gain it at once. `corpus.json` bumps to spec 2 with a `postScheme` describing the new shards. diff --git a/export/app/duplicates/DuplicatesClient.tsx b/export/app/duplicates/DuplicatesClient.tsx @@ -28,11 +28,20 @@ const PLATFORM_LABEL: Record<string, string> = { const MATCH_LABEL: Record<DuplicateCluster["matchKind"], string> = { "transcript-exact": "exact transcript", "transcript-near": "near transcript", + "title-duration": "title + runtime", }; +// MATCH_KINDS is not just a label list — it seeds the default filter state, and +// the visible-cluster filter tests membership in it. A kind missing from here is +// invisibly dropped from the page with no checkbox to turn it back on, so every +// variant of DuplicateMatchKind must appear. (`title-duration` clusters are +// filtered out at compose time unless a human confirmed them, so in practice +// this checkbox only has anything to show once that happens — which is exactly +// the case that must not silently vanish.) const MATCH_KINDS: DuplicateCluster["matchKind"][] = [ "transcript-exact", "transcript-near", + "title-duration", ]; const RELATIONSHIPS = [ diff --git a/export/e2e/helpers.ts b/export/e2e/helpers.ts @@ -72,6 +72,17 @@ export async function installRoutes(page: Page) { await page.route("**/search-aliases.json", async (route) => { await fulfillJson(route, { aliases: [] }); }); + // Duplicate clusters — absent by default, which is also the real default: + // compose-site writes /duplicates.json only when a site has at least one + // shippable cluster. Tests that want the search-result dupe badge register + // their own route afterwards, which takes precedence. + await page.route("**/duplicates.json", async (route) => { + await route.fulfill({ + status: 404, + contentType: "application/json", + body: "{}", + }); + }); } // Charts page fetches: stats dataset + baked templates. Also installs the diff --git a/export/e2e/search-duplicates.spec.ts b/export/e2e/search-duplicates.spec.ts @@ -0,0 +1,199 @@ +import { expect, test, type Page } from "@playwright/test"; +import { installRoutes, urlParams } from "./helpers"; +import { + CHANNEL, + CHANNEL_SLUG, + VIDEO_CHAT_SMALL, + VIDEO_TRANSCRIPT_ONLY, +} from "./fixtures/data"; + +// The search-result duplicate affordance: a badge saying "this exists elsewhere +// in the archive", plus a strip of jump buttons to the other copies. +// +// The load-bearing assertion here is the HONESTY RULE. Matching content does not +// imply matching timings — a mirror with a longer intro carries the same words at +// shifted times — so a jump only carries the current timestamp when the detector +// MEASURED the two as aligned. An unaligned (or unmeasured) sibling opens at 0. +// Get this wrong and the viewer lands mid-sentence in the wrong place on a jump +// that looks like it worked, which is the exact failure measureAlignment exists +// to prevent. + +const ALIGNED_SIBLING = "mirror-chan/aligned-copy"; +const UNALIGNED_SIBLING = "mirror-chan/drifted-copy"; + +// "gamma" matches one transcript cue per fixture video, at start=100s. +const HIT_SECONDS = 100; + +function dupRef( + slug: string, + over: Record<string, unknown> = {}, +): Record<string, unknown> { + const [channelSlug, id] = slug.split("/"); + return { + slug, + channelSlug, + channel: channelSlug === CHANNEL_SLUG ? CHANNEL : "Mirror Chan", + platform: channelSlug === CHANNEL_SLUG ? "youtube" : "rumble", + id, + title: `Title ${id}`, + duration: 300, + uploadDate: "20260101", + hasTranscript: true, + ...over, + }; +} + +function duplicatesReport() { + return { + version: 1, + generatedAt: "2026-07-01T00:00:00.000Z", + runConfig: { + thresholdSeconds: null, + durationToleranceSeconds: 2, + nearThreshold: 0.6, + containmentThreshold: 0.8, + shingleSize: 5, + blocking: "title", + }, + totals: { videosScanned: 4, clusters: 2, videosInClusters: 4 }, + clusters: [ + { + // Measured as aligned: the jump may carry the moment. + clusterId: "cluster-aligned", + matchKind: "transcript-exact", + score: 1, + contained: false, + durationBucket: 300, + crossPlatform: true, + crossChannel: true, + canonicalSlug: `${CHANNEL_SLUG}/${VIDEO_TRANSCRIPT_ONLY}`, + videoRefs: [ + dupRef(`${CHANNEL_SLUG}/${VIDEO_TRANSCRIPT_ONLY}`, { + aligned: true, + offsetSeconds: 0, + }), + dupRef(ALIGNED_SIBLING, { aligned: true, offsetSeconds: 1.2 }), + ], + }, + { + // Same content, DIFFERENT timings — a 41s intro difference. The jump + // must fall back to the start of the sibling. + clusterId: "cluster-drifted", + matchKind: "transcript-near", + score: 0.91, + contained: false, + durationBucket: 300, + crossPlatform: true, + crossChannel: true, + canonicalSlug: `${CHANNEL_SLUG}/${VIDEO_CHAT_SMALL}`, + videoRefs: [ + dupRef(`${CHANNEL_SLUG}/${VIDEO_CHAT_SMALL}`, { + aligned: true, + offsetSeconds: 0, + }), + dupRef(UNALIGNED_SIBLING, { aligned: false, offsetSeconds: 41.5 }), + ], + }, + ], + }; +} + +async function installDuplicates(page: Page, report: unknown) { + await page.route("**/duplicates.json", async (route) => { + await route.fulfill({ + status: 200, + contentType: "application/json", + body: JSON.stringify(report), + }); + }); +} + +function qt(query: string): string { + return encodeURIComponent( + JSON.stringify({ + k: "g", + o: "AND", + c: [{ k: "l", q: query, s: "transcripts" }], + }), + ); +} + +function card(page: Page, slug: string) { + return page.locator(`[data-result-slug="${slug}"]`); +} + +test.describe("search result duplicates", () => { + test("badges a result that exists elsewhere and jumps to the moment when aligned", async ({ + page, + }) => { + await installRoutes(page); + await installDuplicates(page, duplicatesReport()); + await page.goto(`/?qt=${qt("gamma")}`); + + const alignedCard = card(page, `${CHANNEL_SLUG}/${VIDEO_TRANSCRIPT_ONLY}`); + await expect(alignedCard).toBeVisible(); + await expect(alignedCard.getByText("Dupe", { exact: true })).toBeVisible(); + + // The jump button states where it will land, and it lands there. + const jump = alignedCard.locator( + `[data-duplicate-slug="${ALIGNED_SIBLING}"]`, + ); + await expect(jump).toHaveAttribute("data-duplicate-aligned", "true"); + await expect(jump).toContainText("01:40"); + await jump.click(); + + const params = await urlParams(page); + expect(params.get("v")).toBe(ALIGNED_SIBLING); + expect(params.get("t")).toBe(String(HIT_SECONDS)); + }); + + test("an unaligned sibling opens at the start, not at the current moment", async ({ + page, + }) => { + await installRoutes(page); + await installDuplicates(page, duplicatesReport()); + await page.goto(`/?qt=${qt("gamma")}`); + + const driftedCard = card(page, `${CHANNEL_SLUG}/${VIDEO_CHAT_SMALL}`); + await expect(driftedCard).toBeVisible(); + await expect(driftedCard.getByText("Dupe", { exact: true })).toBeVisible(); + + const jump = driftedCard.locator( + `[data-duplicate-slug="${UNALIGNED_SIBLING}"]`, + ); + await expect(jump).toHaveAttribute("data-duplicate-aligned", "false"); + // No timestamp is offered, because none was measured. + await expect(jump).not.toContainText("@"); + await jump.click(); + + const params = await urlParams(page); + expect(params.get("v")).toBe(UNALIGNED_SIBLING); + expect(params.get("t")).toBe("0"); + }); + + test("a result in no cluster carries no badge and no jump strip", async ({ + page, + }) => { + await installRoutes(page); + await installDuplicates(page, duplicatesReport()); + await page.goto(`/?qt=${qt("gamma")}`); + + const lone = card(page, `${CHANNEL_SLUG}/vid-chat-large`); + await expect(lone).toBeVisible(); + await expect(lone.getByText("Dupe", { exact: true })).toHaveCount(0); + await expect(lone.locator("[data-duplicate-strip]")).toHaveCount(0); + }); + + test("no duplicates.json means no duplicate affordance at all", async ({ + page, + }) => { + // The overwhelmingly common case: the file is absent, and search must be + // completely unaffected rather than erroring on the 404. + await installRoutes(page); + await page.goto(`/?qt=${qt("gamma")}`); + + const anyCard = card(page, `${CHANNEL_SLUG}/${VIDEO_TRANSCRIPT_ONLY}`); + await expect(anyCard).toBeVisible(); + await expect(page.locator("[data-duplicate-strip]")).toHaveCount(0); + }); +}); diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -210,6 +210,250 @@ Transcript shards ship the **full `cues` array inline** — `buildIndex.ts:830-8 --- +## Digest bake-off (measured 2026-07-26) + +Harness: `common/bin/digest-bakeoff.ts`. Fixed stratified sample in +`plans/bakeoff/sample.json` (8 videos, 7 channels, 13 min - 8 h), never written +to the corpus. Full tables in `plans/bakeoff/round{1,2}.{json,md}`. + +Corpus totals captured by the same scan that picked the sample: +**76,354 videos, 73,367 with transcripts, 77,298 audio-hours**; 6,038 videos over +4 h holding 40,153 h (8.2% of videos, 52% of audio). Sweep days below = measured +seconds-per-audio-hour x 77,298 h, one lane, no parallelism. + +### Round 1 — screening (short+medium, absolute) + +| Candidate | Zero-yield | Chapters/h | Rejection | Max gap | Generic | Sweep days | +| --- | --- | --- | --- | --- | --- | --- | +| `gemma2:9b@8192` | 0/10 | 15.78 | 0% | 14:59 | 2.3% | 288.1 | +| `qwen3:8b@16384` | 1/7 | 15.42 | 23.6% | 39:57 | 11.9% | 204.9 | +| `mistral-nemo:12b@16384` | 1/7 | 10.28 | 28.2% | 31:20 | 7.1% | 447.3 | +| `qwen3:14b@8192` | 2/10 | 16.52 | 4.3% | 47:51 | 6.7% | 541.3 | +| `qwen2.5:7b@16384` (incumbent) | 2/7 | 6.61 | 55% | 41:37 | 11.1% | 59.0 | + +qwen3 candidates were run with `think: false`; with thinking on they spend most +of their output budget on a reasoning block the pinned schema then discards. +Two `qwen3:14b` chunks failed with `fetch failed` (ollama dropped the connection +under memory pressure — a 9.3 GB model on an 8 GB card) and are counted as +zero-yield, which inflates that row. It does not change its exclusion at 541 days. + +**The incumbent is the fastest by 3.5x and the worst on quality**: it discards 55% +of what it generates and lands at 6.6 chapters/h against a ~15 target. + +### Round 2 — the long tail (2 x ~3.5 h videos, both timestamp modes) + +| Candidate | Zero-yield | Chapters/h | Rejection | Max gap | Generic | Sweep days | +| --- | --- | --- | --- | --- | --- | --- | +| `gemma2:9b@8192` **chunk-local** | 0/17 (0%) | 12.93 | 14.3% | 24:04 | 10% | 132.4 | +| `qwen2.5:7b@8192` **chunk-local** | 2/17 (11.8%) | 13.37 | 19.1% | 24:13 | 31.2% | **24.2** | +| `qwen2.5:7b@16384` chunk-local | 1/9 (11.1%) | 9.34 | 41.4% | 1:27:48 | 20% | 25.1 | +| `gemma2:9b@8192` absolute | 3/17 (17.7%) | 11.07 | 28% | 1:04:07 | 5.2% | 127.9 | +| `qwen2.5:7b@8192` absolute | 5/17 (29.4%) | 10.35 | 47.5% | 1:05:16 | 18.1% | 27.7 | +| `qwen2.5:7b@16384` absolute (today's default) | 3/9 (33.3%) | 8.19 | 50.4% | 1:27:48 | 17.5% | 26.8 | + +**Three findings, all measured, none assumed:** + +1. **chunk-local beats absolute for every model tested**, on the metric this stage + exists to fix. Zero-yield chunks: 33.3% -> 11.1% (qwen2.5@16k), 29.4% -> 11.8% + (qwen2.5@8k), 17.7% -> 0% (gemma2). Rejection rate falls with it in every case. + The hypothesis — that the model cannot hold a large absolute offset across a + long chunk and reverts to counting from zero — is confirmed. + +2. **Smaller chunks are a second, independent win, and they are nearly free.** + Halving the context (16k -> 8k, 1200 -> 600 cues) took qwen2.5 from 9.34 to + 13.37 chapters/h and its max coverage gap from **1:27:48 to 24:13**. The + expected cost did not materialise: 24.2 vs 25.1 projected days. Twice the calls + at half the prompt each is the same seconds-per-audio-hour. The plan's working + assumption that a smaller context roughly doubles sweep cost is **wrong** — + it is flat. + +3. **Sweep days are far lower than the short-video estimate suggested.** Round 1 + measured 59 days on short+medium; Round 2 measures ~25 on the long tail, because + long videos amortise the fixed per-call overhead. The 77,298-hour corpus is + dominated by long videos, so ~25 days is the number to plan against. + +**Recommendation: `qwen2.5:7b` at `num_ctx` 8192 with `timestampMode: +"chunk-local"`.** It is a strict improvement over today's default on every axis +measured, including throughput (24.2 vs 26.8 days). Its one weakness is title +quality — 31.2% generic, the worst in the table; by hand its titles read +"Introduction and Context" / "Conclusion and Final Thoughts" where gemma2 writes +"Andrew Wilson and His Comparison to a Serial Killer". + +**`gemma2:9b@8192` + chunk-local is the quality option**: 0% zero-yield, 10% +generic, comparable density — at 5.5x the wall-clock (132 days). It is the right +engine for a targeted re-run of high-value channels, not for the first full sweep. +Note gemma2 is hard-capped at an 8192 context by the model itself, so it has no +16k row to compare. + +Not yet measured: the >4 h `verylong` stratum (Round 2 used the `long` bucket), +and the ~100-video validation run. + +--- + +## Pilot: community-notes (measured 2026-07-26) + +39 videos, local lane, `qwen2.5:7b` / `num_ctx` 8192 / `chunk-local`, launched from the +channel page's Digest stage through the real managed-job path. + +| | | +| --- | --- | +| Result | **39 generated, 0 failed**, 90 model calls, 140 warnings | +| Chunks | **90 / 90 usable (100%)** | +| Chapters | 435 across 39 videos | +| Warning codes | `out-of-range` 130, `non-monotonic` 9, `seam-duplicate` 1 | +| Generic titles | 103 / 435 (23.7%) | +| Re-run | `countMissingDigests` = 0; batch reports **0 generated, 39 fresh, 0 engine calls, 0.1 s** | + +**The defect video is fixed.** `community-notes/v2chrch` (3505 cues) is the 2.3 h video whose +third chunk previously produced NOTHING — nine of eleven starts rejected as out-of-range +because the model had reverted to counting from zero. At 8 k / chunk-local it is 7 chunks, +**7/7 usable, 42 chapters, one out-of-range warning**, covering 00:16:59 → 02:12:55 with no +gap larger than the leading one. + +The other two 6-chunk videos behave the same: `v2on7b3` 36 chapters 6/6, `v2apmfn` 20 +chapters 6/6. + +**Two honest caveats:** + +- 130 `out-of-range` rejections remain across the channel. The guard is catching them and + the chunks still yield, so this is degraded output rather than lost output — but it is not + zero, and it is the metric to watch in the validation run. +- **Coverage of the opening minutes is weak on long videos.** The first chapter lands at + 00:16:59 (`v2chrch`), 00:15:30 (`v2on7b3`) and **00:41:47** (`v2apmfn`) — the early chunks + yielded little or nothing, which is the mirror image of the original defect. Not + investigated in this stage. Worth checking whether these channels open with long + pre-shows, or whether the first chunk is systematically weaker. + +Generic-title rate on real data (23.7%) sits between the bake-off's long-tail sample figure +for this config (31.2%) and gemma2's (10%), which is consistent rather than surprising. + +--- + +## Duplicate detection at corpus scale (measured 2026-07-26) + +### Why the first attempt failed (the counts here are correct and are the reason blocking exists) + +`duplicate-shorts.ts --all-durations` originally died two different ways: + +| Heap | Outcome | Elapsed | +| --- | --- | --- | +| default (~4 GB) | `FATAL ERROR: Ineffective mark-compacts near heap limit` | 87 s | +| `--max-old-space-size=24576` | `RangeError: Invalid array length` at the containment `candidates.push` | 34 s | + +Counting the pairs the old generator would build, without building them: + +| | Count | +| --- | --- | +| Videos scanned | 76,354 | +| Participants (`duration > 0`) | 76,318 | +| Shorts (≤ 180 s) | 6,970 | +| Duration buckets (2 s) | 11,008 | +| Duration-bucket pairs (same + adjacent) | 9,131,916 | +| **Containment pairs** (short × every video ≥ 1.5× its length) | **497,860,972** | +| Total candidate pairs | 506,992,888 | + +Two independent defects, not one: + +1. **The containment pass was an unblocked cartesian product** — 98.2% of all + candidates. V8 throws `RangeError` because a fast-mode object array caps well + below 500 M elements. +2. **Fingerprints were held for the whole corpus at once.** `fingerprintFrom` + builds a `Set` of 5-word shingles over the entire transcript (~40 k strings + for a 3.5 h video), and the duration pass alone made `needed` ≈ the corpus. + That is what killed the 4 GB run before the array was even reached. + +### The conclusion that was wrong + +The earlier note here said corpus-wide detection "does not complete on this +corpus" and recorded that "as a cost, not worked around". **That was wrong.** +Neither failure is inherent to corpus-wide detection: blocking fixes (1) and +block-at-a-time streaming fixes (2). Both are now implemented, and corpus-wide +detection completes comfortably in minutes on a default heap. + +### Corpus-wide runs, measured (same corpus, 76,354 videos) + +`detectDuplicateShorts` now iterates blocks — fingerprint a block, evaluate its +pairs, keep the confirmed matches, release — so peak memory is O(largest block), +and no multi-million-element pair array is ever materialised. `--blocking` +selects the nomination strategy. + +| | `--blocking title` | `--blocking duration` | +| --- | --- | --- | +| Wall clock | **206 s** | **2,889 s** (48 min) | +| Peak RSS | **2,001 MB** | **2,450 MB** | +| Blocks | 7,551 title groups (of 67,064 distinct titles) | 11,008 buckets of 2 s | +| Oversized blocks skipped | **0** (cap 40) | **0** (cap 2,000) | +| Nominated pairs | 8,352 | 9,131,476 | +| Confirmed by content | 2,717 | **5,708** | +| Rejected by content | 5,250 | 9,125,768 | +| Untestable suspects | 385 | **0** (by design — see below) | +| Videos fingerprinted | 15,479 | 74,943 | +| Clusters written | 2,994 / 6,044 videos (335 `needsReview`) | 2,917 / 5,955 videos (0 `needsReview`) | +| Cluster kinds | 79 exact, 2,580 near, 335 suspect | 108 exact, 2,809 near | +| Alignment | 1,599 of 2,689 within 5 s | 1,818 of 3,028 within 5 s | + +The duration run's 9,131,476 nominated pairs land within 0.005% of the 9,131,916 +predicted above — the old counting was right, only the conclusion drawn from it +was wrong. + +**Duration blocking yields no suspects, and that is deliberate.** A shared runtime +alone was never evidence (it produced enormous false clusters of unrelated +same-length videos), so when a duration-nominated pair cannot be content-tested it +is dropped, not queued for review. Only title blocking produces suspects, because +only *title + near-identical runtime* is a claim worth a human's attention. + +Title-run cluster sizes are overwhelmingly pairs: 2,950×2, 39×3, 3×4, 1×6, 1×9. +2,728 cross-platform, 2,891 cross-channel. + +### Measured recall: what title blocking actually misses + +Comparing the two runs' clusters directly (title's suspects excluded, since they +are not confirmed): + +| | Count | +| --- | --- | +| Clusters with identical membership in both | 2,613 | +| Title-only clusters | 46 | +| Duration-only clusters | 304 | +| Videos in title clusters / in duration clusters | 5,351 / 5,955 | +| **Videos found only by duration** (re-titled mirrors) | **672** | +| Videos found only by title (runtime drift past the bucket) | 68 | + +So **title blocking has ~89% of duration blocking's video-level recall at 1/14th +the wall clock**, and the 672 it misses are exactly the re-titled mirrors it +structurally cannot see. `--blocking both` exists for when that 11% matters; the +title default is the right routine choice. + +The containment sweep is now budget-capped (`MAX_CONTAINMENT_PAIRS`, 5 M) and +reports itself skipped corpus-wide: `6,970 shorts × 76,318 videos ≈ 531,936,460 +pairs`. It still runs, unchanged, at shorts scale. Making it scale is the one +piece genuinely deferred — see STATE.md. + +### The finding that matters most: 63% of title matches are NOT duplicates + +Of 8,352 pairs nominated by an exact normalized title *and* a compatible runtime, +**5,250 were rejected once their transcripts were compared.** Only 2,717 held up. + +This directly corrects the planning assumption that ~8,204 title-level +redundancies were available to prune from the digest sweep. The title-level +*count* was right; the inference that they are duplicates was not. Banked saving +is at most 2,717 pairs, and only 1,599 of those are timing-aligned enough for a +shared digest to be placed correctly — roughly **a fifth of what was projected**. + +Caveat worth testing before treating that as final: the near threshold is a 5-gram +Jaccard at 0.6, and the two sides of a cross-platform mirror are often different +ASR engines (YouTube auto-captions vs. whisper). Some of the 5,250 may be real +mirrors whose transcriptions simply disagree at the 5-gram level. Sensitivity to +that threshold has **not** been measured. See STATE.md. + +### What this unblocks + +Corpus-wide detection is no longer off. `buildDigestClusterPlan` reads +`duplicates.json` and returns an empty plan when it is absent, so the digest batch +stays correct either way — it just re-generates mirrors instead of sharing them. + +--- + ## Editor surfaces | Surface | File | Note | @@ -271,12 +515,35 @@ client) and `mcp/src/search.ts` (server). MCP has 11 tools (`mcp/src/server.ts:1 ### Known-failing on base — not regressions -- The two `actionable.spec` "queues a job" tests -- The `settings.spec` social-links test -- `channels-actions.spec:42` -- The 4 raw-job-kind tests in bulk-actions / cleanup-actionable -- `widget.spec`'s "no buttons" assertion fails under `dev:test` because `next dev` injects a - Dev Tools button; it passes under `E2E_MODE=start` +**Re-verified 2026-07-26** by stashing all working changes and re-running the suspect specs +against `8c041fd`. Full suite in default dev mode: **21 failed / 363 passed of 384**, and +every one of the 21 is accounted for below. Do not chase these. + +| Spec | Count | Evidence | +| --- | --- | --- | +| `actionable.spec` — 2 "queues a job" + the "Needs attention" dashboard card | 3 | fails on base | +| `bulk-actions.spec` raw-job-kind tests | 3 | fails on base (known) | +| `channels-actions.spec:42` | 1 | fails on base (known) | +| `cleanup-actionable.spec` inline "Clean audio" | 1 | fails on base | +| `deploy-page.spec` (all three) | 3 | fails on base — `getByRole('button', {name:'Build & deploy'})` also matches "Build & deploy all sites"; `getByLabel('Deploy after build')` matches 2 checkboxes | +| `new-channel-onboarding.spec` (all five) | 5 | fails on base — `getByLabel('URL')` resolves to 4 elements (a `<label>` wrapping several controls) | +| `settings.spec` social-links | 1 | fails on base (known placeholder collision) | +| `site-scope.spec` dashboard stats | 1 | fails on base | +| `cut-release.spec` (both) | 2 | expected with a dirty working tree — the spec rewrites the real `editor/CHANGELOG.md` and makes a git commit | +| `sync-break-on-existing-tree.spec` | 1 | **passes in isolation** — order-dependent flake, not a base failure | + +The `deploy-page` (×3), `new-channel-onboarding` (×5) and `site-scope` failures — **9 specs** +— are all *strict-mode locator violations* ("resolved to N elements"): specs written against +an earlier form structure that has since gained sibling controls with overlapping accessible +names. They are **a real backlog item, not a regression**, and cheap to fix independently of +any feature work: each needs a tightened locator (`exact: true`, a `getByRole` scoped to the +form, or a `data-testid`), not a behaviour change. Fixing them is worth doing precisely +because 9 permanently-red specs train everyone to ignore the suite's exit code. + +`widget.spec`'s "no buttons" assertion fails under `dev:test` because `next dev` injects a +Dev Tools button; it passes under `E2E_MODE=start`. (It passed in this run.) + +The 9 `digest.spec.ts` tests all pass. Run `pnpm e2e` in **default dev mode** — `E2E_MODE=start` serves a stale build. Kill stale dev servers by port between runs. diff --git a/plans/STATE.md b/plans/STATE.md @@ -3,7 +3,9 @@ The working memory for the local-AI derived-corpus work. Rewritten at the end of every session, before context is cleared. See [`README.md`](README.md) for the protocol. -**Last updated:** 2026-07-26 — roadmap written, no code yet. +**Last updated:** 2026-07-26 — Stage B1 done (chunk-local + context-sized chunking measured +and adopted, digest e2e green, pilot run); corpus-wide duplicate detection rewritten to +stream, measured on both blocking strategies, and surfaced in search results. --- @@ -11,102 +13,232 @@ session, before context is cleared. See [`README.md`](README.md) for the protoco | Phase | Status | Notes | | --- | --- | --- | -| 0 · Benchmark transcription engines | not started | Measurement only. Parakeet is already selectable; what's missing is numbers. | -| 1 · Generation harness | not started | | -| 1.5 · Channel context | not started | **Blocks backfill** — `contextHash` invalidates digests made without it. | +| 0 · Benchmark transcription engines | not started | Still unmeasured. The DIGEST bake-off is done; this is the separate transcription one. | +| 1 · Generation harness | **done** | Spine in `8c041fd`; correctness + bake-off + e2e in Stage B1. | +| 1.5 · Channel context | not started | `contextHash` is plumbed and empty, so notes can be added without invalidating the corpus. | | 2 · Digest corpus in build | not started | | -| 2.5 · Observability | not started | **Blocks backfill** — an unobservable sweep can't be tuned. | +| 2.5 · Observability | not started | **Blocks backfill.** | | 3 · Viewer `?vm=summary` | not started | | | 4 · Search indexing | not started | | | 5 · Auto-queue | not started | | -| 6 · Ollama `/ask` provider | not started | **No dependencies.** Good first commit. | -| 7 · Tags + chat highlights | not started | | +| 6 · Ollama `/ask` provider | not started | **No dependencies.** | +| 7 · Tags + chat highlights | not started | Tag generation works; nothing consumes it. | | 8 · Visibility policy | not started | | -| 9 · Attribution + quote filtering | not started | Largest, least certain. Off by default. | +| 9 · Attribution + quote filtering | not started | | | 10 · Lead with the derived corpus | not started | | -| 11a · Review queue | not started | **Land before backfilling.** | -| 11b · Viewer feedback | not started | After the corpus is public. | +| 11a · Review queue | not started | **Land before backfilling.** `warnings[]` is the data it reads. | +| 11b · Viewer feedback | not started | | -**Recommended next:** Phase 6 (small, self-contained, zero dependencies) for an early win, -or Phase 0 → 1 to start the real spine. - -**Nothing has been implemented.** The only changes are `PLAN.md`, `plans/*`, and a -four-line `# Roadmap` block in `AGENTS.md`. +**Recommended next:** the Stage B2 surfaces that were deliberately deferred — settings form +fields (including `timestampMode` / `promptVariant`), the per-video review panel, the +`/actionable` cluster section, and the PipelineBand instrument — then the ~100-video +validation run before any sweep. --- ## Decisions log Recorded with reasons, because these are exactly what a cold agent would otherwise -relitigate. - -**Fabric is deferred as a digest engine.** Measured to ignore the `HH:MM:SS TOPIC` output -contract at 7B — twice, at two input sizes, with the pattern verified applied via -`--dry-run` and an input small enough to fit the window. It is a long discursive prompt -written for frontier models. Its 256-pattern library remains valuable as *prompt source -material*, and it can become a third registry entry later with no rework. - -**Ollama-direct with JSON-schema `format` is the structured default.** Measured clean and -parseable in 15.5 s on the same input fabric failed. Nothing structured should depend on a -7B voluntarily honoring a text format — that is now measured, not assumed. - -**Claude Code is a second engine on a separate queue key.** Both lanes are GPU/network -disjoint, so a backfill can drive them concurrently — Ollama saturating the RX 6600 while -Claude Code works the same backlog over the network. One shared key would serialize them and -waste the network lane; giving the local engine its own key would let Ollama and the -transcription engine thrash the same 8 GB of VRAM. Metered engines are opt-in and never the -auto-queue default. - -**Digests get their own page tree** rather than being inlined into `TranscriptDetail`. -Under `summary-only` visibility a transcript page ships no cues, so a digest must never live -inside a payload that policy may withhold. - -**Provenance is per-section, not per-file.** The digest controller does a read–modify–write -merge, so a video can legitimately carry Ollama chapters and Claude Code tags. A top-level -`engines` list is derived on read, never persisted as a second source of truth. - -**Transcripts are never rewritten.** Verbatim accuracy is what makes search, citations, and -MCP sweeps trustworthy; a paraphrase would still be derivative while adding hallucination -risk. Exposure is reduced by *withholding* cues, not altering them. +relitigate. Earlier entries (fabric deferred, ollama-direct as the structured default, +Claude Code on a second queue key, digests in their own page tree, per-section provenance, +transcripts never rewritten) still stand and are unchanged. + +**`chunk-local` timestamps are adopted, on measurement.** Each chunk's transcript is +re-based to `00:00:00` and the parser adds the offset back before any guard runs. Measured +on the long tail it cut zero-yield chunks from 33.3% to 11.1% (qwen2.5@16k), 29.4% to 11.8% +(qwen2.5@8k) and 17.7% to 0% (gemma2), with the rejection rate falling in every case. The +working hypothesis — the model cannot hold a large absolute offset across a long chunk and +reverts to counting from zero — is confirmed. It was shipped as a *scored variable*, not a +pre-applied fix, which is why there are numbers to quote. + +**Chunk size is sized to `num_ctx`, and it is a bigger lever than expected.** Halving the +context (16k → 8k, 1200 → 600 cues) took qwen2.5:7b from 9.34 to 13.37 chapters/hour and its +worst coverage gap from 1:27:48 to **24:13**. The cost the plan assumed did not appear: +24.2 vs 25.1 projected sweep days. Twice the calls at half the prompt each is the same +seconds-per-audio-hour. Production previously hardcoded 1200 cues regardless of `num_ctx`, +so lowering the context would have silently truncated every call — `maxCuesForContext()` +now derives it. + +**Anything that changes the output is in the freshness identity.** `promptVariant` folds in +`timestampMode` and a non-default chunk size, derived in ONE place +(`digestPromptVariant`). The default configuration maps to `undefined`, so every digest +written before the field existed still compares fresh. The alternative — an operator +remembering to bump a label — fails silently by skipping the whole corpus as "fresh". + +**The sweep is ~25 days, not ~64.** Round 1 measured 59 days on short+medium; Round 2 +measured ~25 on the long tail, because long videos amortise the fixed per-call overhead and +the corpus is dominated by them. Throughput is no longer the binding constraint it was +assumed to be. + +**Corpus-wide duplicate detection is ON — the earlier "it does not complete" conclusion was +wrong and has been reversed.** The 507 M-pair measurement was correct; the inference from it +was not. Two independent defects, neither inherent to corpus-wide detection: an unblocked +containment cartesian (498 M pairs, 98.2% of the work) and holding every video's shingle set +at once. Blocking fixes the first, block-at-a-time streaming fixes the second. **Both +strategies now complete corpus-wide on a default heap**: title in **206 s at 2.0 GB peak +RSS**, duration in **2,889 s at 2.45 GB** — the latter nominating 9,131,476 pairs, within +0.005% of the 9,131,916 the old counting predicted. The measurement was right; the inference +drawn from it was not. See +[`FACTS.md`](FACTS.md#duplicate-detection-at-corpus-scale-measured-2026-07-26) for the full +table and the measured recall comparison. + +**A blocking strategy NOMINATES; the transcript cascade DECIDES.** `--blocking +title|duration|both` chooses only how pairs are proposed, never what counts as a duplicate, +so choosing one trades recall against runtime rather than correctness. This is what makes +nominating aggressively safe — and it earns its keep: **5,250 of 8,352 title-nominated pairs +(63%) were rejected once their transcripts were compared.** + +**The projected digest saving was overstated by ~5x, and this is the headline finding.** The +~8,204 "title-level redundancies" were real as a *count* of same-title pairs, but only 2,717 +survive content comparison, and only 1,599 of those are timing-aligned enough for a shared +digest to land correctly. Plan the backfill on those numbers, not the old ones. + +**A previous decision is deliberately reversed: title+duration clusters DO form.** The old +rule was that only content-confirmed tiers cluster, because duration coincidence alone +produced enormous false clusters. Title **and** near-identical runtime is a far stronger +claim than duration alone, and such clusters are quarantined behind `needsReview`: they never +auto-share derived work (`clusterMaySharePartial`) and never reach a built site +(`clusterIsPublishable`) until a human records `confirmed`. The old invariant is kept exactly +where it mattered — every *shipped* cluster is still content-confirmed — and the e2e +assertion was narrowed to say so rather than deleted. + +**Alignment is measured at detection time and persisted.** `measureAlignment` already +existed but ran only in `digestSharing` and threw its result away. `DuplicateVideoRef` now +carries `offsetSeconds` + `aligned`, computed against the cluster's canonical member for +confirmed clusters only. Nothing that seeks INTO a sibling can be honest without it: a mirror +with a longer intro carries the same words at shifted times, so "jump to this moment" would +land wrong while looking right. Absent (older report, or no timed cues) reads as NOT aligned. + +**The containment sweep is budget-capped, not deleted.** It is the one pass blocking cannot +rescue — a clip and its parent share neither duration nor title. Under +`MAX_CONTAINMENT_PAIRS` (5 M) it runs unchanged; over it, it logs that it was skipped and +that clip-of-longer duplicates are not covered. Corpus-wide it reports ~531.9 M pairs and +skips. Silent truncation was never an option: it reads as "covered everything". + +**The ollama e2e stub is an HTTP server, not a fake binary.** `ollama-direct` POSTs to +`${ollamaUrl}/api/chat`, so the `e2e/fixtures/bin/` idiom does not apply; it is a third +playwright `webServer` with `OLLAMA_URL` pointed at it in **both** `dev:test` and +`start:test`. `fake-claude.mjs` IS a subprocess and follows the normal idiom. --- ## Open questions -- **Diarizer choice** — `sherpa-onnx` (ONNX, CPU-friendly, no HF token) vs `pyannote` - (higher quality, needs a HF token and GPU). Decide at Phase 9, on real audio. Quality here - cannot be assessed from code. -- **Phase 11b collector** — can the existing `r2-proxy/` package host the minimal feedback - POST endpoint? Worth evaluating before standing up new infrastructure. -- **Transcription engine default** — pending Phase 0 numbers. If parakeet wins, change - `DEFAULT_TRANSCRIPTION_APP_ID` and record it here. +- **Title quality vs throughput.** `qwen2.5:7b@8192` + chunk-local is the adopted default at + ~24 sweep days, but 31.2% of its titles are generic ("Introduction and Context", + "Conclusion and Final Thoughts"). `gemma2:9b@8192` + chunk-local writes markedly better + titles (10% generic, 0% zero-yield) at 5.5x the wall-clock (132 days). Worth revisiting + once Phase 11a can measure review effort: better titles may be cheaper than the review + they save. gemma2 is hard-capped at an 8192 context by the model. +- **The `verylong` (>6 h) stratum is unmeasured.** Round 2 used the `long` bucket (~3.5 h). + The sample already contains an 8 h video for it. +- **Is the 0.6 near threshold rejecting real mirrors?** 63% of title-nominated pairs failed + content comparison. Some of that is genuinely two episodes of a daily show — which is the + point — but the two sides of a cross-platform mirror are often *different ASR engines* + (YouTube auto-captions vs. whisper), and a 5-gram Jaccard is unforgiving of word-level + disagreement. Sensitivity to `--near` has **not** been measured. Cheap to answer: re-run + `--blocking title --near 0.35` and diff the confirmed set. Do this before treating the + 2,717 figure as the ceiling on the digest saving. +- **Recall of the title pre-filter — now measured, and the gap is 11%.** Running both + strategies corpus-wide and diffing the clusters: duration blocking finds **672 videos title + blocking misses** (the re-titled mirrors it structurally cannot see), title finds 68 that + duration misses (runtime drift past the bucket), and 2,613 clusters come out identical. + Title has **~89% of duration's video-level recall at 1/14th the wall clock** (206 s vs + 2,889 s). So the routine default is settled: title, with `--blocking both` when that 11% + matters. What remains open is whether to close the gap *cheaply* rather than by paying 14×: + (a) token-set / Jaccard over title *words* — cheap, but every loosening merges series + episodes and each merge is a false cluster a human must reject; (b) MinHash / LSH + signatures over transcript shingles — the general fix, and the same machinery that would + let the containment sweep scale. Neither is worth building until something needs better + than 89%. +- **`duplicates.json` has no offline story.** `export/service-worker/site-sw.js` keys its + runtime strategies off `manifest.json` / `page-NNNN.json` paths (`:55`, `:147`, `:205`), so + a root-level JSON is uncached. Affects the search-result duplicate badge in PWA mode only — + it silently does not appear offline, which is the correct failure but an unstated one. +- **The pilot's un-created job.** Clicking "Digest channel" a second time produced no job + record and an empty log, with no dedupe guard in `runManagedFunction` to explain it. + Verified via `runDigestBatch` instead. **Still unexplained** — reproduce before trusting + the button to drive a long sweep. +- **Diarizer choice** — unchanged, decide at Phase 9 on real audio. +- **Phase 11b collector** — **answered: no, not `r2-proxy/`.** `src/index.ts:35-41` rejects + every non-GET/HEAD method and `KEY_RE` at `:27` carries a comment stating the proxy must + never become a general oracle over the bucket. Making it a collector works directly against + that design intent. Its one reusable piece is the `RATE_LIMITER` binding. Use the + zero-infrastructure fallback the phase already names (copy-to-clipboard / download-JSON → + paste into an editor action), which needs no deployment and matches the existing idiom at + `PlayerProvider.tsx:131-136`. +- **Transcription engine default** — still pending Phase 0 numbers. --- ## Commands known to pass -Established against `main` @ `18c5a7a`: +Established against this branch: ```bash +pnpm --filter yt-dlp-transcript-common run test # 344 tests +pnpm --filter {yt-dlp-transcript-common,editor,export} exec tsc --noEmit +pnpm --filter editor build pnpm e2e # default dev mode; E2E_MODE=start serves a stale build -pnpm build:index -pnpm --filter export run build -pnpm --filter yt-dlp-transcript-mcp run test +pnpm --filter export exec playwright test # 141 tests, separate suite on :3020 ``` -Kill stale dev servers by port between e2e runs. The known-failing-on-base list lives in -[`FACTS.md`](FACTS.md#known-failing-on-base--not-regressions) — do not chase those as -regressions. +Duplicate detection (writes `transcripts/duplicates.json`; back it up first if the current +one matters): + +```bash +cd common +pnpm exec tsx bin/duplicate-shorts.ts --all-durations # title, ~3.5 min +pnpm exec tsx bin/duplicate-shorts.ts --all-durations --blocking duration # ~48 min +pnpm exec tsx bin/duplicate-shorts.ts --all-durations --blocking both +``` + +**Do not run a corpus-wide detection pass and `pnpm e2e` at the same time.** The detector +pins a core for its whole run and the contention alone failed +`partial-downloads-bucket`, `undownloaded` and `video-page` — all three pass in isolation. +That is three spurious entries to chase in a suite that already has 21 known-red ones. + +Bake-off (never writes to the corpus): + +```bash +cd common +pnpm exec tsx bin/digest-bakeoff.ts --pick # ONCE — fixes the sample +pnpm exec tsx bin/digest-bakeoff.ts --label rN --buckets long \ + --modes absolute,chunk-local --candidates 'qwen2.5:7b@8192' +``` + +**Kill stale dev servers AND orphaned fixture binaries between e2e runs.** Killing a +playwright run mid-flight leaves `e2e/fixtures/bin/fake-*.mjs` children orphaned; they spin +at ~24% CPU each and dozens of them will drive the load average past 35 and make every +subsequent HTTP request time out while the server still logs 200s. `pkill -9 -f +"fixtures/bin/fake-"` is part of cleanup, not just killing by port. + +For anything driving the editor against the REAL corpus (the pilot), use **production +mode** (`pnpm --filter editor build` then `pnpm start`): the dev server renders the +39-video channel page in 30 s - 9.9 min and holds ~21% of RAM, while the production build +serves the same page in 7.6 s. -Note there is currently **no root or `common` test script**; ~46 unit test files are -unrunnable as a suite. Phase 1 adds one to `common/package.json`. +The known-failing-on-base list lives in +[`FACTS.md`](FACTS.md#known-failing-on-base--not-regressions) — do not chase those. --- ## Surprises hit so far -The original draft plan named `windowCues()` as the transcript chunker. It is a center-based -search-hit helper and cannot do that — the chunker has to be written. Several other -confidently-stated claims in that draft were also wrong; all are corrected at the top of -[`FACTS.md`](FACTS.md#corrections-to-the-original-draft-plan). Treat plan prose written -without a `file:line` anchor as unverified. +`windowCues()` is not a chunker (corrected in FACTS.md, and `chunkCuesForContext` was +written for the job). Still true, plus: + +**Two bugs the unit tests could not have caught, both found by the new e2e.** +`taskHooks.ts` deliberately gives digest tasks no progress parser and then called +`downloadParser!.feed()` unconditionally — every digest job died on its first log line, and +a sweep would have reported "0 generated, N failed" looking like an engine fault. And +`editor/app/settings/actions.ts` rebuilds the digest settings block field by field, so the +first save after the settings form lands would have silently reset `timestampMode` and +invalidated every digest made under it. Both are fixed. The lesson is that the digest +layer's failure modes live at the *wiring*, not in the pure functions. + +**The bake-off's most useful finding was one nobody asked for.** The plan framed the choice +as model-vs-model with chunk-local as the variable. The data showed gemma2's clean sweep was +confounded with its forced 8k context, and testing chunk size directly turned out to matter +as much as the timestamp mode — and to be free. A comparison worth running is one that can +surprise you.