Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 6cb327e2121ce66203f5b156824a011aacb1aa9a
parent 4826443681a9fd288e55c2e459b3f089ecd0daab
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Wed, 29 Jul 2026 11:20:03 -0400

Lower the duplicate near-threshold to 0.35, on a bracketed measurement

0.6 was tuned for same-engine text and structurally failed the case detection
exists to catch: the two sides of a cross-platform mirror are transcribed by
DIFFERENT ASR engines, and a 5-gram Jaccard is unforgiving of word-level
disagreement, so two transcripts of the same audio land at ~0.35–0.60.

Bracketed corpus-wide at 0.6 / 0.45 / 0.35 / 0.25 on identical inputs. Every
step down is a strict superset — zero videos lost — and the returns fall off a
cliff: 0.6→0.45 adds 4,264 clusters, 0.45→0.35 adds 324, 0.35→0.25 adds 59.

The marginal bands were READ, not just counted. Of the 4,264 admitted at 0.45,
96.2% have byte-identical titles, 99.5% are cross-platform, and 3 (0.07%) are
same-channel; the riskiest are YouTube↔Rumble pairs agreeing on title, runtime
to the second, and upload date. The 324 at 0.35 are the same shape and all 11
same-channel-or-differing-title cases were inspected individually.

0.45 is where the recall knee is, and is the value to take if a more
conservative assertion is ever wanted. 0.35 is chosen because its band is still
clean and the report's job is to surface real mirrors to readers.

This is a PUBLISHING change in its own commit, not a digest optimisation: it
changes what every built site's Duplicates page and search badges assert. As a
cost lever it is negligible — 0.6→0.35 moves the sweep from 80.0 to 78.8 days,
because mirrors are 19% of the corpus by count but 4.3% of its audio-hours.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Diffstat:
Mcommon/lib/duplicates.ts | 28+++++++++++++++++++++++++++-
1 file changed, 27 insertions(+), 1 deletion(-)

diff --git a/common/lib/duplicates.ts b/common/lib/duplicates.ts @@ -16,7 +16,33 @@ export const DEFAULT_SHORT_THRESHOLD_SECONDS = 180; // land in a shared comparison set (see controller). export const DEFAULT_DURATION_TOLERANCE_SECONDS = 2; // Phase-2 near-duplicate threshold (5-gram Jaccard). -export const DEFAULT_NEAR_THRESHOLD = 0.6; +// +// 0.35, lowered from 0.6 on measurement (2026-07-29). 0.6 was tuned for +// same-engine text and structurally failed the case detection exists to catch: +// the two sides of a cross-platform mirror are transcribed by DIFFERENT ASR +// engines, and a 5-gram Jaccard is unforgiving of word-level disagreement, so +// two transcripts of the same audio land at ~0.35–0.60 rather than ≥ 0.6. +// +// Bracketed corpus-wide at 0.6 / 0.45 / 0.35 / 0.25 on identical inputs. Each +// step down is a strict superset — zero videos are lost — and the returns fall +// off a cliff: 0.6→0.45 adds 4,264 clusters, 0.45→0.35 adds 324, 0.35→0.25 adds +// 59. The marginal band was read, not just counted: of the 4,264 admitted at +// 0.45, 96.2% are byte-identical titles, 99.5% cross-platform, and 3 (0.07%) +// are same-channel. The 324 admitted at 0.35 are the same shape (97.2% +// byte-identical titles) and the riskiest 11 were inspected individually — all +// same recording, same runtime to the second, mirrored platform. +// +// 0.45 is where the RECALL knee is, if a more conservative value is ever +// wanted; 0.35 is chosen because its marginal band is still clean and the +// report's job is to surface real mirrors to readers. 0.25 is the flat tail. +// +// This is NOT a meaningful digest-sweep cost lever, whatever the sharing code's +// header says: cluster members are 19% of the corpus by video count but only +// 4.3% of its AUDIO-HOURS (mirrors skew short, long-form VODs are unclustered), +// so 0.6→0.35 moves the sweep from 80.0 to 78.8 days. It is a publishing change +// — it is what the archive asserts to readers — and it is justified on that. +// See bin/digest-plan.ts for the measurement. +export const DEFAULT_NEAR_THRESHOLD = 0.35; // Containment threshold for the "short is a clip of a longer video" case. export const DEFAULT_CONTAINMENT_THRESHOLD = 0.8; // Shingle (word n-gram) size for similarity. 5 tolerates word-level ASR