Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit f387da569a0809466bd6fc86c61af54b8f4beb7b
parent 29bfb7c42d2c4b1f110c0848fb0fa73c61189a24
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Mon, 21 Sep 2026 16:01:31 -0400

plans: curated tags — rollout log and the two things the seed got wrong

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Diffstat:
Mplans/curated-tags.md | 21++++++++++++++++++---
1 file changed, 18 insertions(+), 3 deletions(-)

diff --git a/plans/curated-tags.md b/plans/curated-tags.md @@ -196,14 +196,15 @@ pnpm ops tags --json '{"op":"define","tag":{"id":"eva-topic","label":"Discussed" Expected from the 2026-09-20 research: 7 metadata hits on `eva-collab`, ~64 streams on `eva-in-chat`, ~19 on `eva-topic` (Legal Mindset alone). Review in `/tags` → Preview, pin survivors, suppress false positives; then the umtool project page "Tag cited videos as eva-collab"; then `pnpm ops build-index --wait` and -`pnpm ops build-deploy --json '{"siteId":"anilyzer"}' --wait`. +`pnpm ops build-deploy --json '{"siteIds":["anilyzer"]}' --wait` (the routes take `siteIds`, a list). ## Verification - Unit: `pnpm --filter yt-dlp-transcript-common run test` (curatedTags, store, evalTree, corpus); tsc for common/editor/umtool/mcp; `pnpm run test:scripts`; `mcp` tests. - Index/export on the fixture and then the real corpus: `pnpm ops build-index --wait`; - `pnpm ops build-site --json '{"siteId":"anilyzer"}' --wait`; `jq .tags export/.export-index/builds/anilyzer/ - tags.json`; find a legal-mindset id's page via `slugToPage` and assert `.curatedTags` on the record; + `pnpm ops build-site --json '{"siteIds":["anilyzer"]}' --wait` (note: `--wait` streams the job log and can drop + with `fetch failed` while the in-process index build is busy — the job keeps running; poll `/api/jobs/active`); + the site composes into `export/public` and exports to `export/out`, so `jq .tags export/out/tags.json`; find a legal-mindset id's page via `slugToPage` and assert `.curatedTags` on the record; `curl -H "authorization: Bearer $WORKER_TOKEN" localhost:3001/api/ops/tags | jq '.tags[].count'`. - Browser: `?tg=eva-collab` narrows the all-videos list and a search; chip count == `/tags.json` count; clearing the chip drops the param. e2e: `export/e2e/tag-chips.spec.ts`, `editor/e2e/tags.spec.ts`, @@ -211,3 +212,17 @@ then the umtool project page "Tag cited videos as eva-collab"; then `pnpm ops bu - MCP: `list_tags {source:"remote:https://anilyzer.pages.dev"}` after deploy; `search_transcripts {…, tags:["eva-collab"]}` returns only tagged videos; the same against a pre-spec-4 site (Jeralyzer until rebuilt) returns the empty-with-warning path. + +## Rollout log (2026-09-21, local only, no deploy) + +Seeded through `pnpm ops tags --file`, index rebuilt (first derivation: examined 77,224, re-derived 831, +623 s; 103 transcript pages rewritten, 906 unchanged), Anilyzer composed locally. The seed patterns were +then corrected from what the corpus actually contains: + +- `eva-collab` metadata: `@elfpire|#elfpire|elf ?pire ?eva|^[^\n]*elf ?pire` — Legal Mindset's 7 hits are + all `@ElfpireEva` guest credits; MommaOcco's 200 were a friends-list boilerplate (`[EvaElfpire]` + URL), + which the bare `elfpire` matched. `^[^\n]*` = "in the title" (the haystack's first line; no `m` flag). +- `eva-in-chat` chat-author: `^@?elfpire ?eva$` — the bare `elfpire` also matched `@PapaElfpire`. +- `eva-topic` caption: `\b(alpire|elpire|elfire|alpier|elpyre|elfpire|elfpyre)\b|legal[ -]?lol+i|\blawli\b` — + ASR never spells "Elfpire"; without `\b` it matched "Jason Angelfire", "channelfireball" and + "RekietaLawLive" on four other sites.