commit 6cb327e2121ce66203f5b156824a011aacb1aa9a
parent 4826443681a9fd288e55c2e459b3f089ecd0daab
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Wed, 29 Jul 2026 11:20:03 -0400
Lower the duplicate near-threshold to 0.35, on a bracketed measurement
0.6 was tuned for same-engine text and structurally failed the case detection
exists to catch: the two sides of a cross-platform mirror are transcribed by
DIFFERENT ASR engines, and a 5-gram Jaccard is unforgiving of word-level
disagreement, so two transcripts of the same audio land at ~0.35–0.60.
Bracketed corpus-wide at 0.6 / 0.45 / 0.35 / 0.25 on identical inputs. Every
step down is a strict superset — zero videos lost — and the returns fall off a
cliff: 0.6→0.45 adds 4,264 clusters, 0.45→0.35 adds 324, 0.35→0.25 adds 59.
The marginal bands were READ, not just counted. Of the 4,264 admitted at 0.45,
96.2% have byte-identical titles, 99.5% are cross-platform, and 3 (0.07%) are
same-channel; the riskiest are YouTube↔Rumble pairs agreeing on title, runtime
to the second, and upload date. The 324 at 0.35 are the same shape and all 11
same-channel-or-differing-title cases were inspected individually.
0.45 is where the recall knee is, and is the value to take if a more
conservative assertion is ever wanted. 0.35 is chosen because its band is still
clean and the report's job is to surface real mirrors to readers.
This is a PUBLISHING change in its own commit, not a digest optimisation: it
changes what every built site's Duplicates page and search badges assert. As a
cost lever it is negligible — 0.6→0.35 moves the sweep from 80.0 to 78.8 days,
because mirrors are 19% of the corpus by count but 4.3% of its audio-hours.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Diffstat:
1 file changed, 27 insertions(+), 1 deletion(-)
diff --git a/common/lib/duplicates.ts b/common/lib/duplicates.ts
@@ -16,7 +16,33 @@ export const DEFAULT_SHORT_THRESHOLD_SECONDS = 180;
// land in a shared comparison set (see controller).
export const DEFAULT_DURATION_TOLERANCE_SECONDS = 2;
// Phase-2 near-duplicate threshold (5-gram Jaccard).
-export const DEFAULT_NEAR_THRESHOLD = 0.6;
+//
+// 0.35, lowered from 0.6 on measurement (2026-07-29). 0.6 was tuned for
+// same-engine text and structurally failed the case detection exists to catch:
+// the two sides of a cross-platform mirror are transcribed by DIFFERENT ASR
+// engines, and a 5-gram Jaccard is unforgiving of word-level disagreement, so
+// two transcripts of the same audio land at ~0.35–0.60 rather than ≥ 0.6.
+//
+// Bracketed corpus-wide at 0.6 / 0.45 / 0.35 / 0.25 on identical inputs. Each
+// step down is a strict superset — zero videos are lost — and the returns fall
+// off a cliff: 0.6→0.45 adds 4,264 clusters, 0.45→0.35 adds 324, 0.35→0.25 adds
+// 59. The marginal band was read, not just counted: of the 4,264 admitted at
+// 0.45, 96.2% are byte-identical titles, 99.5% cross-platform, and 3 (0.07%)
+// are same-channel. The 324 admitted at 0.35 are the same shape (97.2%
+// byte-identical titles) and the riskiest 11 were inspected individually — all
+// same recording, same runtime to the second, mirrored platform.
+//
+// 0.45 is where the RECALL knee is, if a more conservative value is ever
+// wanted; 0.35 is chosen because its marginal band is still clean and the
+// report's job is to surface real mirrors to readers. 0.25 is the flat tail.
+//
+// This is NOT a meaningful digest-sweep cost lever, whatever the sharing code's
+// header says: cluster members are 19% of the corpus by video count but only
+// 4.3% of its AUDIO-HOURS (mirrors skew short, long-form VODs are unclustered),
+// so 0.6→0.35 moves the sweep from 80.0 to 78.8 days. It is a publishing change
+// — it is what the archive asserts to readers — and it is justified on that.
+// See bin/digest-plan.ts for the measurement.
+export const DEFAULT_NEAR_THRESHOLD = 0.35;
// Containment threshold for the "short is a clip of a longer video" case.
export const DEFAULT_CONTAINMENT_THRESHOLD = 0.8;
// Shingle (word n-gram) size for similarity. 5 tolerates word-level ASR