commit 0bd2cfd44cad3fc0a0c948548cf4baab263c6118
parent 58dbf5ad2bdc141cb9a7b96ce812c2497f545e98
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Mon, 21 Sep 2026 12:48:44 -0400
plans: curated per-video tags (approved 2026-09-21)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Diffstat:
| A | plans/curated-tags.md | | | 209 | +++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ |
1 file changed, 209 insertions(+), 0 deletions(-)
diff --git a/plans/curated-tags.md b/plans/curated-tags.md
@@ -0,0 +1,209 @@
+# Plan — universal per-video tags ("all streams where Eva collabs", on any site, any channel)
+
+Approved by the operator 2026-09-21. Implemented as three slices (S1 core, S2 site, S3 editor) after
+a frozen S1.1 lands on main.
+
+## Context
+
+The operator wants to open Anilyzer and narrow any video list to "streams where Eva collabs", including
+streams on channels that are not hers (Legal Mindset, Destiny, Nux Taku…), with a mechanism that is
+generic: any person/topic, any site, applied by the operator, by an agent, or by umtool. Today the only
+cross-channel curation on a site is channel **groups** (`sites/<id>/site.json` `groups`; Anilyzer's
+"Elfpire Eva" group holds only her own two channels) — whole channels, nothing per video. The corpus has
+three unrelated things called "tags" already (yt-dlp keyword `TranscriptSummary.tags`, digest topic tags,
+the `"tags"` search scope), which is the naming hazard this plan works around. Research on 2026-09-20
+produced the first concrete tag set: 7 metadata-flagged LM collabs, ~19 LM streams discussing her, ~64
+with her in chat.
+
+## Operator decisions (2026-09-21)
+
+1. **Assignment — all three**: by hand (video toggle + bulk "Tag selected"), by **rules** on a tag
+ (re-evaluated at index build; operator pins/suppresses exceptions), by **import from a umtool project**.
+2. **One tag per kind**: `eva-collab` (on mic), `eva-in-chat`, `eva-topic`; filter is OR across chips.
+3. **Vocabulary both corpus-wide and per site, like aliases** — corpus file authoritative for
+ definitions AND assignments; the site file is presentation (relabel/hide/order/colour) + site-only rules.
+4. **Surfaces — all four**: export lists (channel + all-videos), export search, MCP filter, editor lists.
+5. **Universal + flexible, three writers**: free-form tag ids for any subject, optional `group` so chips
+ read "Eva: Collab · In chat · Discussed"; operator UI, `pnpm ops` (agents; MCP stays read-only per
+ AGENTS.md), umtool — all through ONE server action, every assignment carrying provenance.
+6. **The export lists selectable tags dynamically** (operator, 2026-09-21): a site's chip row is built
+ from its published `/tags.json`, and compose writes only tags with a non-zero count ON THAT SITE — a
+ global tag with no matches on a site does not appear there, nor does a hidden one. No static list anywhere.
+
+## Naming
+
+| Surface | Name |
+|---|---|
+| UI word, files, CLI, wire | **tags**: `transcripts/tags.json`, `sites/<id>/tags.json`, published `/tags.json`, `pnpm ops tags`, `pnpm ops tag-videos` |
+| Record field on shards | **`curatedTags?: string[]`** (never `tags` — `common/lib/transcripts.ts:21` is yt-dlp keywords) |
+| Modules / types | `common/lib/curatedTags.ts`, `common/lib/curatedTagsStore.ts`, `common/controller/curatedTagsIndex.ts`; `CuratedTagDef`, `CuratedTagRule`, `CuratedTagAssignment`, `CuratedTagsConfig`, `PublishedTags` |
+| Existing keyword scope | UI label "Tags" → **"Keywords"** at `common/components/QueryLeafView.tsx:46` (`SCOPE_LABELS`) and `:156-159` (placeholder); code token `"tags"`, URL and MCP `scopes` enum unchanged |
+
+Tag ids: `^[a-z0-9][a-z0-9._-]*$`. Nothing in code names Eva; she is seed data (§ Seed).
+
+## Data model
+
+**`transcripts/tags.json`** (authoritative; `writeJsonAtomic`; `sanitizeTagsConfig` on read):
+```json
+{ "version": 1,
+ "tags": [ { "id": "eva-collab", "label": "Collab", "group": "eva", "groupLabel": "Eva",
+ "description": "Eva is on mic", "color": "#b48ead", "order": 1, "hidden": false,
+ "rules": [ { "id": "meta", "kind": "metadata", "pattern": "elfpire|elf ?pire ?eva",
+ "channels": [], "dateFrom": null, "dateTo": null, "enabled": true } ] } ],
+ "assignments": {
+ "legal-mindset/XZqL6k9IHGA": {
+ "manual": ["eva-collab"], "suppressed": ["eva-topic"],
+ "sources": { "eva-collab": { "source": "umtool:elfpire-eva", "setAt": "2026-09-21T18:04:11Z" },
+ "eva-topic": { "source": "operator", "setAt": "2026-09-21T18:05:02Z" } } } } }
+```
+- `source` ∈ `operator | agent:<label> | umtool:<project> | rule:<ruleId>`; recorded for pins AND suppressions.
+- `rules` optional (pure-manual tags are legal); `kind` ∈ `metadata | chat-author | caption`; a rule whose
+ regex fails to compile is disabled with a reason, never thrown. If a tag is in both `manual` and
+ `suppressed` for one key, `manual` wins.
+- **Rule hits are never persisted** — only pins/suppressions. Edit a rule → rebuild index → done.
+
+**`sites/<id>/tags.json`**: same shape. `mergeTagDefs(global, site)`: for an existing id the site entry is a
+field-wise overlay of `label/groupLabel/color/order/hidden` and may append rules; it can never delete a
+global rule or assignment (an assignment is a fact about the video). A new id is a full site-only tag.
+(Deliberately NOT `mergeAliases`' wholesale replacement.)
+
+**Pure functions** (`common/lib/curatedTags.ts`): `effectiveTagsFor(ruleHits, assignment) = (hits ∪ manual)
+− suppressed`, sorted by def `order` then id; `evaluateTagRules(input, rules)` with
+`input = {channelSlug, id, title, description, tags, uploadDate, captionCues?, chatCues?}`:
+`metadata` → one regex over `title\ndescription\ntags.join(" ")`; `chat-author` → the author is the
+`<author>: ` prefix of each chat cue (`common/lib/liveChat.ts:11,59-60,100` — no `author` field on `Cue`;
+the prefix is the contract); `caption` → regex over caption cue text, first hit short-circuits. Channel
+scope and date range are applied before compiling; each kind compiles once per build.
+
+## Derivation pipeline
+
+- `curatedTags?: string[]` on `TranscriptSummary` (`common/lib/transcripts.ts:10-29`) and `DisplaySummary`
+ (`:31-55`, optional like `state?` at `:52`) and `toDisplaySummary` (`common/lib/transcripts-server.ts:195`).
+ **Omitted when empty** so untagged corpora's pages stay byte-identical (`createPageWriter` sha1 skip,
+ `buildIndex.ts:869`).
+- Computed in `common/controller/buildIndex.ts` in the per-video worker just before `sums.put(indexKey,
+ summary)` (`:697`) — caption cues and the `live_chat` track are already in hand there (`:705`, `:720`).
+- **Invalidation** (`common/controller/curatedTagsIndex.ts`): store `curatedRulesHash` (sha1 of enabled
+ rules) and `curatedAssignHash` (sha1 of assignments) in the existing `meta` sub-DB (no new sub-DB, so no
+ new `clearAsync()` at `:528-548`). `reapplyCuratedTags` runs after the mtime diff (`:561-587`): assignments
+ changed → re-derive only the symmetric difference of keys (no cue reads); rules changed → re-derive every
+ video from the `cues`/`subs` sub-DBs (no disk re-read). Returns the changed `indexKey`s, which are unioned
+ into the dirty-channel set so transcript/subs/summaries pages rewrite — `common/lib/channelSignature.ts`
+ hashes only mtimes (FACTS ~:203-207), so this explicit dirtying is what keeps compose from skipping them.
+ `SCHEMA_VERSION` 13 → 14 (`:143`). Build log prints `curated tags: rules <hash8> (changed), re-derived N`.
+- **Counts**: while streaming `sums` into per-site summaries pages (`:1712`) accumulate per-tag totals and
+ per-channel breakdown → `tag-counts.json` in `siteIndexDir(paths, siteId)` (`common/lib/site.ts:138`).
+- **Compose** (`common/bin/compose-site.ts:744-753`, beside aliases): `effectiveSiteTags(paths, siteId)` +
+ counts → `/tags.json` `{version, tags:[{id,label,group,groupLabel,color,order,count,channels:{slug:n}}]}`;
+ hidden and zero-count tags dropped; nothing left → file not written (404 is legitimate, like `duplicates.json`).
+- **Contract**: `CONTRACT.corpusSpec` 3 → **4** and `"tags.json"` in `ROOT_FILES` (`common/lib/archive/
+ contract.ts:31-51, :121`) — per the rule at `common/lib/corpus.ts:25-34` a bump announces a new fetchable
+ document; the per-record `curatedTags` key is additive and bumps no manifest version. `buildSiteCorpus`
+ (`corpus.ts:204`) gains `tags: {url, videoField: "curatedTags"}`; `llms.txt` (`:320`) gains one line.
+
+## Slices (three Opus implementers; S1.1 lands first and alone, then S1 ∥ S2 ∥ S3)
+
+### S1 — common store, rules, index, export data (branch `tags/core`)
+1. **S1.1 FREEZE**: `common/lib/curatedTags.ts` (types, `sanitizeTagsConfig`, `mergeTagDefs`,
+ `effectiveTagsFor`, `evaluateTagRules`, `PublishedTags`, `TAGS_FILENAME`) + tests; the two record fields;
+ `export/e2e/fixtures/data.ts` emits `curatedTags` on two summaries and a `/tags.json` fixture. Merge to
+ main immediately — S2 and S3 branch from it.
+2. `common/lib/curatedTagsStore.ts` (mirror `common/lib/aliasesStore.ts`): `readGlobalTags/writeGlobalTags/
+ readSiteTags/writeSiteTags/effectiveSiteTags` + **`applyTagAssignments(paths, {op: add|remove|replace, tag,
+ videos, source, at})`**; `paths.globalTagsFile` (`common/lib/paths.ts:105,219`), `siteTagsFile`
+ (`common/lib/site.ts:132`); tests: sanitize, layering, both-lists conflict, provenance.
+3. `curatedTagsIndex.ts` + `buildIndex.ts` hook at `:697`, meta hashes, `reapplyCuratedTags`,
+ `SCHEMA_VERSION` 14, per-site counts at `:1712`; `toDisplaySummary`. Measure rule cost on legal-mindset
+ (478 chat streams) and record it in FACTS.
+4. Contract + compose: `contract.ts` spec 4 + `ROOT_FILES`; `corpus.ts:204,:320`; `compose-site.ts:744-753`
+ writes `/tags.json`; `corpus.test.ts`.
+5. `plans/FACTS.md` (the three meanings of "tags" in the naming-hazards table; chat-author prefix contract;
+ the two meta hashes; SCHEMA_VERSION 14) + `AGENTS.md` (tags.json is curated data; never hand-edit).
+
+### S2 — export UI + MCP (branch `tags/site`, needs only S1.1)
+6. `common/lib/search/evalTree.ts`: `curatedTags?` on `SearchFilters` (`:311`) and `FilterableRecord`
+ (`:328`), OR test in `passesFilters` (`:335`), `filterIsSelective` (`:365`) true for a non-empty set; tests.
+7. `common/components/tagsCache.ts` (mirror `aliasesCache.ts`) + `SearchDataContext.tsx`; `urlState.ts`
+ param `tg` (the `tk` convention: `p.getAll("tk")` `:94`, delete/append `:185-186`).
+8. `SearchSessionContext.tsx`: `draftTags/committedTags` (`:389` pattern), `passesFilter` (`:703`), filter
+ key, snapshot/commit (`:986, :1013`), `curatedTags` on `ResultGroup` (`:141-158`).
+9. UI: grouped chip row in `common/components/FiltersPanel.tsx` beside the channel-group chips (`:225-250`),
+ rendered only when `/tags.json` has a tag with `count > 0`, chips show counts, OR semantics; tag chips on
+ the video card (`SearchResults.tsx:607`); "Keywords" relabel (`QueryLeafView.tsx:46,156-159`).
+10. MCP: `tags: string[]` in `SEARCH_FILTER_ARGS` (`mcp/src/server.ts:196`), parsed in `parseSearchArgs`
+ (`:1382`, filters literal ~`:1435`) for `search_transcripts` (`:309`) and `enumerate_matches` (`:413`);
+ `curatedTags` on `IndexedVideo` (`common/lib/archive/reader.ts:100`) + `buildVideoIndex` (`:242`) so the
+ scan-plan pruner (`mcp/src/search.ts:383`) prunes by tag; `loadTags()` on `ArchiveReader` (sibling of
+ `readAliasConfig` `:761`); new `list_tags` tool after `list_channels` (`:285`); a defs line in
+ `resolve_source` (`:682`); footer warning when a source ships no `/tags.json` but a tag filter was given
+ (pre-spec-4 sites: `loadTags()` → `[]`, pruner treats "index never heard of it" as "read the page",
+ the invariant at `search.ts:319`).
+11. Tests: `export/e2e/tag-chips.spec.ts` (model: `export/e2e/channel-group-chips.spec.ts`), URL round-trip,
+ `mcp/src/search.test.ts` filter case.
+
+### S3 — editor UI + ops + umtool (branch `tags/editor`, needs S1.1; stubs `applyTagAssignments` until S1.2)
+12. `editor/app/tags/page.tsx` + `components/EditorTagsClient.tsx` + `actions.ts` (`saveTagDefsAction`,
+ `previewTagRuleAction` = dry-run `evaluateTagRules` over LMDB `sums/cues/subs`, listing title/date/channel
+ with Pin all / Pin / Unpin, `applyTagAssignmentsAction`) — pattern `editor/app/sites/[siteId]/aliases/`.
+ Nav entry (twelve → thirteen; `editor/scripts/measure-nav.mjs`).
+13. Per-site overlay tab `editor/app/sites/[siteId]/tags/page.tsx` (presentation fields + site-only rules).
+14. Video page `videos/[id]/components/TagsPanel.tsx` beside `OperationPanel` (`page.tsx:251`);
+ `toggleVideoTagAction` in `videoActions.ts` (shows each tag's provenance).
+15. `VideoListPane.tsx`: `tag`/`untag` in `BULK_ACTION_OPTIONS` (`:38-49`) with a tag picker via
+ `doSummaryBulk` (`:254`); `bulkApplyTagAction` in `bulkVideoActions.ts` = ONE `applyTagAssignments` call
+ (single file write, not a per-id loop); a Tags chip filter on the list.
+16. Ops: `editor/app/api/ops/tag-videos/route.ts` (`POST {tag, op, videos:[{slug,id}], source?}`) and
+ `editor/app/api/ops/tags/route.ts` (`GET` defs+counts, `?tag=` → assignments with provenance; `POST`
+ define/remove defs and rules) — adapters per `_lib.ts:7-27`; register `tag-videos`, `tags` in
+ `scripts/archilyzer-ops.mjs:51-65`; add `--file <path>` for large bodies (agents write big id lists via
+ script, not model output) and a GET entry beside `channel` (`:45-49`). Default `source` for the CLI:
+ `agent:<$ARCHILYZER_AGENT or "cli">`.
+17. umtool: `umtool/app/api/report/tag/route.ts` + project-page action "Tag cited videos as …" (whole
+ project or per section) → `${ARCHILYZER_EDITOR_URL}/api/ops/tag-videos` with `WORKER_TOKEN`
+ (`umtool/report-to-video/fetch-via-editor.mjs:24,70-80`), `source: "umtool:<projectId>"`, honouring
+ per-clip `siteChannel`/`siteVideo` (Rumble's two ids, `cues.mjs:206-207`).
+18. `editor/e2e/tags.spec.ts` (defs CRUD, rule preview, pin/unpin, bulk tag, list filter) + cases in
+ `editor/e2e/ops-api.spec.ts`; umtool spec for the project action against a fixture editor stub
+ (`EDITOR_STUB_PORT` pattern from the fetch-window work).
+
+**Three writers, one action**: editor UI, ops route and umtool all reach `applyTagAssignmentsAction` →
+`applyTagAssignments` in `curatedTagsStore.ts`. The MCP gains no write tool.
+
+**Graph**: S1.1 → { S1.2 → S1.3 → S1.4 → S1.5 } ‖ { S2.1 → S2.2 → S2.3 → S2.4, S2.5, S2.6 } ‖
+{ S3.1 → S3.2, S3.3, S3.4 → S3.5 → S3.6 → S3.7 }. Review: Opus for S1 (index + contract) and S3.5/S3.6
+(writers), Sonnet for S2. Merge order S1 → S2 → S3 (S3 re-merges main once S1.2 lands).
+
+## Risks
+- Invalidation: a skipped `reapplyCuratedTags` makes a rule edit silently inert until an unrelated mtime
+ moves — the two meta hashes are checked unconditionally, and tag-only changes explicitly dirty channels.
+- Name collision (three "tags"): `curatedTags` is the only field name; the UI says "Keywords" for yt-dlp tags.
+- Shard growth: omitted-when-empty keeps Jeralyzer (30k videos) byte-identical until tagged; fully tagged
+ ≈ 25 B/record ≈ 750 KB across its pages.
+- Chat regex cost: author-prefix only, compiled once, channel/date scoped first, short-circuit; measured in S1.3.
+- Remote sources built before spec 4: `/tags.json` 404 → empty with an explicit warning, never a silent miss.
+- Site-layer overreach: the field-wise overlay cannot delete global rules or assignments (tested in S1.2).
+
+## Seed (applied through the new surfaces — never by hand-writing `transcripts/`)
+```sh
+pnpm ops tags --json '{"op":"define","tag":{"id":"eva-collab","label":"Collab","group":"eva","groupLabel":"Eva","rules":[{"id":"meta","kind":"metadata","pattern":"elfpire|elf ?pire ?eva","enabled":true}]}}'
+pnpm ops tags --json '{"op":"define","tag":{"id":"eva-in-chat","label":"In chat","group":"eva","groupLabel":"Eva","rules":[{"id":"chat","kind":"chat-author","pattern":"elfpire","enabled":true}]}}'
+pnpm ops tags --json '{"op":"define","tag":{"id":"eva-topic","label":"Discussed","group":"eva","groupLabel":"Eva","rules":[{"id":"cap","kind":"caption","pattern":"\\belf ?pire\\b|elfpyre|legal loli","enabled":true}]}}'
+```
+Expected from the 2026-09-20 research: 7 metadata hits on `eva-collab`, ~64 streams on `eva-in-chat`, ~19 on
+`eva-topic` (Legal Mindset alone). Review in `/tags` → Preview, pin survivors, suppress false positives;
+then the umtool project page "Tag cited videos as eva-collab"; then `pnpm ops build-index --wait` and
+`pnpm ops build-deploy --json '{"siteId":"anilyzer"}' --wait`.
+
+## Verification
+- Unit: `pnpm --filter yt-dlp-transcript-common run test` (curatedTags, store, evalTree, corpus); tsc for
+ common/editor/umtool/mcp; `pnpm run test:scripts`; `mcp` tests.
+- Index/export on the fixture and then the real corpus: `pnpm ops build-index --wait`;
+ `pnpm ops build-site --json '{"siteId":"anilyzer"}' --wait`; `jq .tags export/.export-index/builds/anilyzer/
+ tags.json`; find a legal-mindset id's page via `slugToPage` and assert `.curatedTags` on the record;
+ `curl -H "authorization: Bearer $WORKER_TOKEN" localhost:3001/api/ops/tags | jq '.tags[].count'`.
+- Browser: `?tg=eva-collab` narrows the all-videos list and a search; chip count == `/tags.json` count;
+ clearing the chip drops the param. e2e: `export/e2e/tag-chips.spec.ts`, `editor/e2e/tags.spec.ts`,
+ ops-api cases, umtool tag spec; full editor suite on the final main.
+- MCP: `list_tags {source:"remote:https://anilyzer.pages.dev"}` after deploy; `search_transcripts {…,
+ tags:["eva-collab"]}` returns only tagged videos; the same against a pre-spec-4 site (Jeralyzer until
+ rebuilt) returns the empty-with-warning path.