commit 6b097ff87b1cc826a1fb5abc758fbc4fb19115fe
parent 0164030fd1d5816083cc035851adf00d82fe16a2
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Mon, 21 Sep 2026 13:17:26 -0400
tags S1.5: write down what the next session must not re-derive
FACTS gains: the three unrelated meanings of "tags" in the naming-hazards
section (with the rule that the record field is curatedTags); the chat-cue
author contract (a `<author>: ` prefix, because Cue has no author field); the
meta-hash invalidation table and what curatedRulesHash deliberately excludes;
and the measured rule cost. SCHEMA_VERSION and CORPUS_SPEC_VERSION in the
constants table were stale at 12 and 2 — now 14 and 4, with their real anchors.
The measurement, taken offline against the production LMDB: on legal-mindset
(518 videos, 2.84 M cues) the three seed rules cost 1.28 s warm, of which 1.03 s
is cue decoding and 232 ms is regex. The cost is the decode, not the pattern —
which is why a scoped rule is worth scoping and a metadata-only vocabulary is
free.
AGENTS.md gets transcripts/tags.json in the corpus table and the rule that it
is written through applyTagAssignments and nothing else.
Co-Authored-By: Claude Opus <noreply@anthropic.com>
Diffstat:
| M | AGENTS.md | | | 13 | +++++++++++++ |
| M | plans/FACTS.md | | | 72 | ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-- |
2 files changed, 83 insertions(+), 2 deletions(-)
diff --git a/AGENTS.md b/AGENTS.md
@@ -198,6 +198,19 @@ apps*, not [DEPLOY_DOCKER.md](DEPLOY_DOCKER.md), which is about *building sites*
| `transcripts/index.mdb` | The LMDB transcript index. Key-only range scans over its `byChannel` sub-DB are cheap; see `common/controller/recencyIndex.ts`. |
| `transcripts/saved-videos/` | Persisted source-video store. |
| `transcripts/search-aliases.json`, `duplicates*.json` | Corpus-wide curated data. |
+| `transcripts/tags.json` | **Curated per-video tags** — the cross-channel vocabulary AND every assignment. Curated data: never hand-edited, never a scratch file. |
+
+**Curated tags are written through ONE path, and it is not your text editor.**
+`transcripts/tags.json` holds the operator's vocabulary (`eva-collab`, rules that
+re-evaluate at index build) and every pin/suppression with its provenance; a site's
+`sites/<id>/tags.json` is presentation only. Every writer — editor UI, `pnpm ops`, umtool —
+goes through `applyTagAssignments` in `common/lib/curatedTagsStore.ts`, which validates,
+records who claimed what, and writes the file once, atomically. Hand-editing it loses
+provenance and races whatever is running. Tests use temp dirs.
+
+Three unrelated things in this repo are called "tags": the yt-dlp keywords on a record
+(`TranscriptSummary.tags`), AI digest topic tags, and these. The record field for these is
+**`curatedTags`**, never `tags` — see `plans/FACTS.md`, "Naming hazards".
**The roots a channel's media may be moved to are named entities**, `settings.storage.locations`
(id, label, root, `autoRepoint`, and the volume UUID learned at the last probe) — managed on
diff --git a/plans/FACTS.md b/plans/FACTS.md
@@ -139,6 +139,17 @@ track. The exclusion list in `readSubTracks` (`:238-256`) is hardcoded. Use
**`report` already means three things**: `reportDebouncePreset` in settings, the
`refresh-report` job kind, and `/ask`'s `ReportPanel`. Use `feedback` / `review` instead.
+**`tags` means three unrelated things, and the record field is `curatedTags`.**
+
+| Which | Where | What it is |
+| --- | --- | --- |
+| yt-dlp keywords | `common/lib/transcripts.ts:21` — `TranscriptSummary.tags` | The platform's own keywords from `metadata.info.json`, searchable via the `"tags"` scope. The UI label for this scope is being changed to **"Keywords"** (`QueryLeafView.tsx:46,156-159`); the code token, the URL and the MCP `scopes` enum stay `"tags"`. |
+| AI digest topic tags | `common/lib/digests.ts` — `VideoDigest.tags` | Model-generated topics on a digested video. |
+| **Curated tags** | `common/lib/curatedTags.ts`; record field **`curatedTags`** | The operator's cross-channel vocabulary ("every stream where X collabs"), published at `/tags.json`. |
+
+Never name the curated field `tags`. Never assume a `tags.json` is the keyword list:
+`transcripts/tags.json` and `sites/<id>/tags.json` are curated tags.
+
---
## Reusable helpers (do not rewrite these)
@@ -160,6 +171,9 @@ track. The exclusion list in `readSubTracks` (`:238-256`) is hardcoded. Use
| ENOENT → friendly message | `common/social/xGalleryDlFetcher.ts:207-213` | `/ENOENT/.test(message) ? "X not found (set X_BIN)" : message`. |
| ETA estimator | `editor/app/jobs/active/buildActiveJobs.ts:38` | `computeEtaSeconds`. |
| Binary search over cue starts | `common/components/TranscriptModal.tsx:473-491` | `findActiveIndex`. |
+| Curated-tag model (pure) | `common/lib/curatedTags.ts` | `sanitizeTagsConfig` / `mergeTagDefs` / `effectiveTagsFor` / `compileTagRules` + `evaluateCompiledRules`. Compile ONCE per build, evaluate per video. |
+| Curated-tag persistence | `common/lib/curatedTagsStore.ts:176` | `applyTagAssignments` — the ONE write path for every tag writer (editor, ops, umtool); one atomic write per call. |
+| A chat cue's author | `common/lib/curatedTags.ts:488` — `chatAuthorOf` | The `<author>: ` prefix; there is no `author` field on `Cue`. |
---
@@ -187,9 +201,9 @@ Exposed to the settings form via `listTranscriptionApps()` (`:272`), consumed at
| Constant | Where | Value |
| --- | --- | --- |
-| `SCHEMA_VERSION` | `buildIndex.ts:116` | 12 |
+| `SCHEMA_VERSION` | `buildIndex.ts:157` | **14** (13 → 14 for curated tags, 2026-09-21) |
| `CUES_FILE_VERSION` | `normalizeTranscript.ts:33` | 2 |
-| `CORPUS_SPEC_VERSION` | `common/lib/corpus.ts:16` | 2 |
+| `CORPUS_SPEC_VERSION` | `common/lib/archive/contract.ts:35` (re-exported `corpus.ts:59`) | **4** (3 → 4 for `/tags.json`, 2026-09-21) |
| `MANIFEST_VERSION` (site summaries) | `common/lib/manifest.ts:36` | 3 |
| `TRANSCRIPTS_MANIFEST_VERSION` | `manifest.ts:43` | 1 |
| `SUBS_MANIFEST_VERSION` | `manifest.ts:59` | 4 |
@@ -210,6 +224,60 @@ three times (`:667` transcripts, `:675` subs, `:686` posts); it passes
Transcript shards ship the **full `cues` array inline** — `buildIndex.ts:830-832`.
+### Curated tags invalidate through `meta`, not through mtimes (2026-09-21)
+
+`curatedTags` is DERIVED at build time — `(rule hits ∪ pins) − suppressions` — and **rule
+hits are never persisted**. So nothing in the per-video mtime diff can see that a rule was
+edited or a video pinned, and the mechanism is two keys in the existing `meta` sub-DB
+(`common/controller/curatedTagsIndex.ts:44-46`), checked **unconditionally on every build**:
+
+| Key | Covers | Deliberately excludes |
+| --- | --- | --- |
+| `curatedRulesHash` | every def's id + `order`, and every **enabled** rule's kind, pattern, channel scope and date range | label / colour / group / hidden — relabelling a tag must not re-page 30,000 videos |
+| `curatedAssignHash` | every pin and suppression | — |
+| `curatedAssignSigs` | per-key `manual|suppressed` signature | not a hash: it is what makes an assignment-only edit re-derive just the videos that moved |
+
+`reapplyCuratedTags` (`:269`) runs after the mtime diff: rules changed → every video is a
+candidate (cue reads gated per video by channel/date scope); assignments changed → only the
+symmetric difference of the signatures, located by a key-only `byChannel` range scan per
+affected channel. It re-derives out of LMDB — **no video directory is read twice** — and
+whatever it changes flips `sharedNeedsBuild`, because the page writer's sha1 skip is what
+then leaves the untouched pages untouched. The per-site fingerprint (`buildIndex.ts:1741`)
+includes both hashes, or a tag-only edit would leave a site reporting "up to date" with
+yesterday's counts. **No new sub-DB**, so the `clearAsync()` enumeration is unchanged.
+
+Only the CORPUS vocabulary is evaluated: a site-only rule contributes its definition to that
+site's `/tags.json` but does **not** tag records, because records are shared by every site
+carrying the channel. Promote a rule to `transcripts/tags.json` to make it bind.
+
+**A chat cue's author is a string prefix, not a field.** `common/lib/liveChat.ts:100-105`
+emits every live-chat cue as ``text: author ? `${author}: ${text}` : text`` — `Cue` is
+`{start, end, text}` (`vtt.ts:1`) and has no `author`. So a `chat-author` rule matches
+`chatAuthorOf(text)` (`curatedTags.ts:488`): everything before the FIRST `": "`, and `null`
+when there is none (a cue with no author can never match). Anything that wants the author
+of a chat line must use this, not a second copy of the split.
+
+### Measured: what a curated rule costs (2026-09-21, real corpus)
+
+`legal-mindset`, 518 videos / 1.37 M caption cues / 1.47 M live-chat cues, the three seed
+rules (metadata + chat-author + caption), run offline against the production LMDB (warm):
+
+| Rules | Wall | LMDB decode | Regex |
+| --- | --- | --- | --- |
+| metadata only | 18 ms | 0 | 4 ms |
+| chat-author only | 848 ms | 727 ms | 110 ms |
+| caption only | 421 ms | 302 ms | 109 ms |
+| all three | 1.28 s | 1.03 s | 232 ms |
+
+**The cost is the cue decode, not the regex** (4:1). A metadata-only vocabulary is free; a
+cue-backed rule costs about 2.5 ms per video with cues, so a whole-corpus re-derive at
+Jeralyzer's scale (30,886 videos) projects to ~80 s — a rule edit, not a rebuild. Scoping a
+rule with `channels`/`dateFrom` skips the decode entirely for everything out of scope.
+
+Hits on that channel: 7 `eva-collab` (metadata), 96 `eva-in-chat` (chat author), 1
+`eva-topic` (caption `\belf ?pire\b|elfpyre|legal loli`) — the caption number is a pattern
+problem, not a plumbing one: a spoken name rarely appears spelled that way in ASR output.
+
---
## Digest bake-off (measured 2026-07-26)