Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 6b097ff87b1cc826a1fb5abc758fbc4fb19115fe
parent 0164030fd1d5816083cc035851adf00d82fe16a2
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Mon, 21 Sep 2026 13:17:26 -0400

tags S1.5: write down what the next session must not re-derive

FACTS gains: the three unrelated meanings of "tags" in the naming-hazards
section (with the rule that the record field is curatedTags); the chat-cue
author contract (a `<author>: ` prefix, because Cue has no author field); the
meta-hash invalidation table and what curatedRulesHash deliberately excludes;
and the measured rule cost. SCHEMA_VERSION and CORPUS_SPEC_VERSION in the
constants table were stale at 12 and 2 — now 14 and 4, with their real anchors.

The measurement, taken offline against the production LMDB: on legal-mindset
(518 videos, 2.84 M cues) the three seed rules cost 1.28 s warm, of which 1.03 s
is cue decoding and 232 ms is regex. The cost is the decode, not the pattern —
which is why a scoped rule is worth scoping and a metadata-only vocabulary is
free.

AGENTS.md gets transcripts/tags.json in the corpus table and the rule that it
is written through applyTagAssignments and nothing else.

Co-Authored-By: Claude Opus <noreply@anthropic.com>

Diffstat:
MAGENTS.md | 13+++++++++++++
Mplans/FACTS.md | 72++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++--
2 files changed, 83 insertions(+), 2 deletions(-)

diff --git a/AGENTS.md b/AGENTS.md @@ -198,6 +198,19 @@ apps*, not [DEPLOY_DOCKER.md](DEPLOY_DOCKER.md), which is about *building sites* | `transcripts/index.mdb` | The LMDB transcript index. Key-only range scans over its `byChannel` sub-DB are cheap; see `common/controller/recencyIndex.ts`. | | `transcripts/saved-videos/` | Persisted source-video store. | | `transcripts/search-aliases.json`, `duplicates*.json` | Corpus-wide curated data. | +| `transcripts/tags.json` | **Curated per-video tags** — the cross-channel vocabulary AND every assignment. Curated data: never hand-edited, never a scratch file. | + +**Curated tags are written through ONE path, and it is not your text editor.** +`transcripts/tags.json` holds the operator's vocabulary (`eva-collab`, rules that +re-evaluate at index build) and every pin/suppression with its provenance; a site's +`sites/<id>/tags.json` is presentation only. Every writer — editor UI, `pnpm ops`, umtool — +goes through `applyTagAssignments` in `common/lib/curatedTagsStore.ts`, which validates, +records who claimed what, and writes the file once, atomically. Hand-editing it loses +provenance and races whatever is running. Tests use temp dirs. + +Three unrelated things in this repo are called "tags": the yt-dlp keywords on a record +(`TranscriptSummary.tags`), AI digest topic tags, and these. The record field for these is +**`curatedTags`**, never `tags` — see `plans/FACTS.md`, "Naming hazards". **The roots a channel's media may be moved to are named entities**, `settings.storage.locations` (id, label, root, `autoRepoint`, and the volume UUID learned at the last probe) — managed on diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -139,6 +139,17 @@ track. The exclusion list in `readSubTracks` (`:238-256`) is hardcoded. Use **`report` already means three things**: `reportDebouncePreset` in settings, the `refresh-report` job kind, and `/ask`'s `ReportPanel`. Use `feedback` / `review` instead. +**`tags` means three unrelated things, and the record field is `curatedTags`.** + +| Which | Where | What it is | +| --- | --- | --- | +| yt-dlp keywords | `common/lib/transcripts.ts:21` — `TranscriptSummary.tags` | The platform's own keywords from `metadata.info.json`, searchable via the `"tags"` scope. The UI label for this scope is being changed to **"Keywords"** (`QueryLeafView.tsx:46,156-159`); the code token, the URL and the MCP `scopes` enum stay `"tags"`. | +| AI digest topic tags | `common/lib/digests.ts` — `VideoDigest.tags` | Model-generated topics on a digested video. | +| **Curated tags** | `common/lib/curatedTags.ts`; record field **`curatedTags`** | The operator's cross-channel vocabulary ("every stream where X collabs"), published at `/tags.json`. | + +Never name the curated field `tags`. Never assume a `tags.json` is the keyword list: +`transcripts/tags.json` and `sites/<id>/tags.json` are curated tags. + --- ## Reusable helpers (do not rewrite these) @@ -160,6 +171,9 @@ track. The exclusion list in `readSubTracks` (`:238-256`) is hardcoded. Use | ENOENT → friendly message | `common/social/xGalleryDlFetcher.ts:207-213` | `/ENOENT/.test(message) ? "X not found (set X_BIN)" : message`. | | ETA estimator | `editor/app/jobs/active/buildActiveJobs.ts:38` | `computeEtaSeconds`. | | Binary search over cue starts | `common/components/TranscriptModal.tsx:473-491` | `findActiveIndex`. | +| Curated-tag model (pure) | `common/lib/curatedTags.ts` | `sanitizeTagsConfig` / `mergeTagDefs` / `effectiveTagsFor` / `compileTagRules` + `evaluateCompiledRules`. Compile ONCE per build, evaluate per video. | +| Curated-tag persistence | `common/lib/curatedTagsStore.ts:176` | `applyTagAssignments` — the ONE write path for every tag writer (editor, ops, umtool); one atomic write per call. | +| A chat cue's author | `common/lib/curatedTags.ts:488` — `chatAuthorOf` | The `<author>: ` prefix; there is no `author` field on `Cue`. | --- @@ -187,9 +201,9 @@ Exposed to the settings form via `listTranscriptionApps()` (`:272`), consumed at | Constant | Where | Value | | --- | --- | --- | -| `SCHEMA_VERSION` | `buildIndex.ts:116` | 12 | +| `SCHEMA_VERSION` | `buildIndex.ts:157` | **14** (13 → 14 for curated tags, 2026-09-21) | | `CUES_FILE_VERSION` | `normalizeTranscript.ts:33` | 2 | -| `CORPUS_SPEC_VERSION` | `common/lib/corpus.ts:16` | 2 | +| `CORPUS_SPEC_VERSION` | `common/lib/archive/contract.ts:35` (re-exported `corpus.ts:59`) | **4** (3 → 4 for `/tags.json`, 2026-09-21) | | `MANIFEST_VERSION` (site summaries) | `common/lib/manifest.ts:36` | 3 | | `TRANSCRIPTS_MANIFEST_VERSION` | `manifest.ts:43` | 1 | | `SUBS_MANIFEST_VERSION` | `manifest.ts:59` | 4 | @@ -210,6 +224,60 @@ three times (`:667` transcripts, `:675` subs, `:686` posts); it passes Transcript shards ship the **full `cues` array inline** — `buildIndex.ts:830-832`. +### Curated tags invalidate through `meta`, not through mtimes (2026-09-21) + +`curatedTags` is DERIVED at build time — `(rule hits ∪ pins) − suppressions` — and **rule +hits are never persisted**. So nothing in the per-video mtime diff can see that a rule was +edited or a video pinned, and the mechanism is two keys in the existing `meta` sub-DB +(`common/controller/curatedTagsIndex.ts:44-46`), checked **unconditionally on every build**: + +| Key | Covers | Deliberately excludes | +| --- | --- | --- | +| `curatedRulesHash` | every def's id + `order`, and every **enabled** rule's kind, pattern, channel scope and date range | label / colour / group / hidden — relabelling a tag must not re-page 30,000 videos | +| `curatedAssignHash` | every pin and suppression | — | +| `curatedAssignSigs` | per-key `manual|suppressed` signature | not a hash: it is what makes an assignment-only edit re-derive just the videos that moved | + +`reapplyCuratedTags` (`:269`) runs after the mtime diff: rules changed → every video is a +candidate (cue reads gated per video by channel/date scope); assignments changed → only the +symmetric difference of the signatures, located by a key-only `byChannel` range scan per +affected channel. It re-derives out of LMDB — **no video directory is read twice** — and +whatever it changes flips `sharedNeedsBuild`, because the page writer's sha1 skip is what +then leaves the untouched pages untouched. The per-site fingerprint (`buildIndex.ts:1741`) +includes both hashes, or a tag-only edit would leave a site reporting "up to date" with +yesterday's counts. **No new sub-DB**, so the `clearAsync()` enumeration is unchanged. + +Only the CORPUS vocabulary is evaluated: a site-only rule contributes its definition to that +site's `/tags.json` but does **not** tag records, because records are shared by every site +carrying the channel. Promote a rule to `transcripts/tags.json` to make it bind. + +**A chat cue's author is a string prefix, not a field.** `common/lib/liveChat.ts:100-105` +emits every live-chat cue as ``text: author ? `${author}: ${text}` : text`` — `Cue` is +`{start, end, text}` (`vtt.ts:1`) and has no `author`. So a `chat-author` rule matches +`chatAuthorOf(text)` (`curatedTags.ts:488`): everything before the FIRST `": "`, and `null` +when there is none (a cue with no author can never match). Anything that wants the author +of a chat line must use this, not a second copy of the split. + +### Measured: what a curated rule costs (2026-09-21, real corpus) + +`legal-mindset`, 518 videos / 1.37 M caption cues / 1.47 M live-chat cues, the three seed +rules (metadata + chat-author + caption), run offline against the production LMDB (warm): + +| Rules | Wall | LMDB decode | Regex | +| --- | --- | --- | --- | +| metadata only | 18 ms | 0 | 4 ms | +| chat-author only | 848 ms | 727 ms | 110 ms | +| caption only | 421 ms | 302 ms | 109 ms | +| all three | 1.28 s | 1.03 s | 232 ms | + +**The cost is the cue decode, not the regex** (4:1). A metadata-only vocabulary is free; a +cue-backed rule costs about 2.5 ms per video with cues, so a whole-corpus re-derive at +Jeralyzer's scale (30,886 videos) projects to ~80 s — a rule edit, not a rebuild. Scoping a +rule with `channels`/`dateFrom` skips the decode entirely for everything out of scope. + +Hits on that channel: 7 `eva-collab` (metadata), 96 `eva-in-chat` (chat author), 1 +`eva-topic` (caption `\belf ?pire\b|elfpyre|legal loli`) — the caption number is a pattern +problem, not a plumbing one: a spoken name rarely appears spelled that way in ASR output. + --- ## Digest bake-off (measured 2026-07-26)