Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 8346ea674cdb8e76aa4e9fa4310b4e7d6ea85b94
parent fb9050a94e0ff14257df997094ae589d6bf8a3cf
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Mon,  5 Oct 2026 05:25:45 -0400

records: the feed backfill's [Unreleased] bullet at the end of the section

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
Meditor/CHANGELOG.md | 2+-
1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,7 +1,6 @@ # Changelog ## [Unreleased] -- **A podcast channel's episodes imported by their enclosure URL get their titles and dates back from the feed.** Such a record carried only its file name as its title and no upload date, so the index skipped it and every surface showed the file id. `archilyzer feeds backfill-metadata <slug> [--feed <url>] [--dry-run]`, `POST /api/ops/feed-metadata {slug, dryRun?}` and `pnpm ops feed-metadata` (the new `feed-metadata` job, on the channel's download queue) fetch the channel's RSS feed once — no media — and match each record lacking a date or a real title to an item by guid, enclosure URL or enclosure file name; a record more than one item fits is reported, not guessed. A match fills in what the record lacks: the title, the description as plain text, the duration (a measured one is kept), the upload date and timestamp in UTC, and `webpage_url` only when the new URL keeps the record's id. Every change is a `feed-backfill` entry in the video's metadata history; a dry run counts matched, unmatched and already complete and writes nothing. Rebuild the index afterwards for the titles and dates to reach search and the sites. - **A site's build composes its reports, and a site that publishes only its reports ships nothing else.** Every site's compose now writes the reports its `site.json` publishes: each report's page and its citations as `citations.json` and `citations.csv` under `/reports/<id>/`, its cited stills, a page per cited moment with the record, the transcript lines around the span and every report that cites it, and the clips and post captures `archilyzer reports prepare` made for it, only the cited ones. Each quote is checked against the record as it is composed (a span's against its cues within 5 s either side, read from `en-orig` when the `en` track has no cues; a post's against its text) and the score, time and method are written into the citation, replacing any typed by hand. The build stops with the list of every problem before anything is written: an invalid report, a citation of a channel outside the site or of a post the site may not carry, a missing record, still or post, a quote that matches less than 60 % of what the record says, and a citation without prepared media or with media cut for another span (`--allow-missing-media` on `archilyzer compose site` and `build site` lets those two through, without a clip). A site with `publish: "cited"` removes everything corpus-shaped from `export/public` before it writes its reports, and its built `out/` is checked against what a cited site may hold: anything else, a file over 25 MiB or more than 20,000 files fails the build, and every deploy path (the Publish tab, `deploy site`, Build & deploy, Build & deploy all, the container build) refuses it, as it refuses a site set to cited whose last build was a full one. The hub's compose removes a report site's files too. - **A site has a Reports tab.** `/sites/<site>/reports` lists every report under the site's `reports/` directory — the published ones in their order, then the drafts — with its kind, dates, sections, claims, citations by kind and, for a fact-check, how many claims carry each verdict. Each report's problems, from the same checker the prepare step and the build use, open under it. A draft with no problems can be published, and a published report moved up or down or unpublished; each writes only the site's `reports` list, applied to the list as it is on disk at that moment, so it never overwrites another change to the site. "Prepare evidence media" queues the `reports-prepare` job, and beside it the tab shows the last prepared media (moments by kind, total size, problems by kind) and links the last prepare job. What the site publishes (full or cited) is shown with a link to Settings, where it is changed. - **A site can say what it publishes, and which reports.** `site.json` takes `publish` — `"full"`, the searchable corpus every site has been (the default, never written), or `"cited"`, only the site's reports and the moments they cite — and `reports`, the ordered ids of its published reports (each a slug; invalid and repeated ids are dropped). The site form has a Publish control and lists the site's reports read-only; saving the form keeps the stored list. A cited site still builds as a full one until the reports pipeline applies the scope. SITE.md documents both keys. @@ -65,6 +64,7 @@ - **`archilyzer storage migrate-tier` brings a channel moved the old way onto the media tier: its text comes home to this disk, its big files stay on the drive.** Run it with the editor **stopped** — a real run refuses while the editor answers on its port (`ARCHILYZER_EDITOR_URL`, default `http://localhost:3001`) or while a job on disk belongs to a live process; a dry run only notes it. `archilyzer storage migrate-tier <channel>` does one channel, `--all` every channel still shown as "Media layout retired". Start with `--dry-run`: it writes nothing and prints, per channel, how much text, clip windows, scratch and source video would be copied to this disk, how many audio files and live-chat replays stay on the drive, and whether it fits. A real run copies everything but the audio, the raw live-chat replays and any dead `*.temp.*` left by a download's postprocessor into the channel's `data/` on this disk, checks the copy file by file and kind by kind, links each audio file and replay back under its old name with its own time, renames the drive's `<root>/<channel>/data` to `<root>/<channel>/media`, and records `mediaDir` in the channel's `config.json` in place of `dataDir`; then it prints the bytes copied by kind and the links made. Nothing needs re-indexing: every file keeps its time. It needs the copy plus the resume margin free above the disk gate's floor on this disk, and refuses before writing anything when that is not there. `--all` takes the smallest channels first and **stops before omnibased, rekietalaw and the-quartering-rumble**, printing the free space and what each of them would copy; go on with the channel's name, or `--all --include-large`. A run that is stopped or fails partway carries on where it left off when run again, and a channel already done is left alone. The drive keeps its copy of the text, and those dead temps, until you run it again with `--reclaim`, which deletes them and keeps the media. `--reclaim` takes only channels this command migrated, and deletes a copy only where the same file, at the same size, is in the channel's `data/` on this disk; anything else stays on the drive and is listed. Leave it until the editor has run on the migrated channels for a while: until then the drive's copy is a second one. **Superseded auto-subtitles are purged after the migration, not before:** a channel on the retired layout refuses `purge-superseded-auto-subs` like every text job, so a channel's superseded `en-orig` subtitles (7.6 GB on omnibased) are copied to this disk first and purged from there. - **A video whose YouTube subtitles answer "Too Many Requests" (HTTP 429) is downloaded anyway, and YouTube is not put in a cooldown for it.** YouTube refuses a subtitle file per video while the video itself downloads fine; every one of the day's 429s on 2026-10-01 was a subtitle fetch, and each failed its download, put all of YouTube in a cooldown that reached 30 minutes, and deferred the video for 6 hours to fail the same way again. Now that refusal is noted and the download goes on to the audio, as for a video with no captions, so the transcription lane transcribes it; nothing platform-wide is backed off. The video's subtitles are deferred: **Download missing subs** skips them for 6 hours, and from the third time they are refused, for 7 days; a download from the video's own page still fetches them. The download lane's page lists them under **Deferred subtitles**, and the video page says how many times and when. **Download missing subs** also goes on to the next video when one video's subtitles are refused, instead of stopping. Any other subtitle failure (a 403, a missing file, a chat replay that fails) still fails the download, as before. Needs a rebuild and restart of the editor. - **The download pace adapts to rate limits, a rate limit that outlasts the cooldown holds the platform, and the auto-download lane waits between downloads.** Every yt-dlp run against a platform now waits its platform's current pace between requests: 1 second for YouTube and Rumble, doubled by each real rate limit (up to 16 seconds) and eased back one step after every 5 clean downloads, and one step for every hour with no rate limit; a subtitle-only refusal never raises it. When a platform has failed three times in a row at the 30-minute cooldown, it is **held**: auto-download tries it once an hour instead of every 30 minutes, and until that try is due a manual **Sync**, download or metadata scan on it is refused with a sentence giving its time. A clean try lifts the hold — the lane's, or a manual Sync, download or scan once the try is due, which is how a hold ends while auto-download is off or has nothing to fetch on that platform — and so does a try whose video came down although its subtitles were refused. The lane page keeps a held platform listed until then (saying when the lane is off), with a **Clear hold** button that drops the hold, the cooldown and the raised pace at once and says so in a job log. The auto-download lane now waits **Sleep between downloads** between two downloads on one platform, as a channel's batch downloads always did, plus whatever the pace was raised by; batch downloads add that too. The lane page's **Rate-limit cooldown** box shows held platforms, the raised paces and the deferred subtitles, the lane says when it is idle because a platform is held or it is pausing between downloads, and `archilyzer doctor` warns about a platform in a cooldown or held, and about a raised pace. **Download missing subs** also waits that gap between videos. The four numbers are the new `pacing` block in `settings.json` (SETTINGS.md). **After the restart, YouTube may be held at its first real failure:** its cooldown count from before the update (the subtitle refusals) still stands, so one failure puts it straight past the cap — **Clear hold** on the download lane's page resets it, and a clean download does too. Needs a rebuild and restart of the editor. +- **A podcast channel's episodes imported by their enclosure URL get their titles and dates back from the feed.** Such a record carried only its file name as its title and no upload date, so the index skipped it and every surface showed the file id. `archilyzer feeds backfill-metadata <slug> [--feed <url>] [--dry-run]`, `POST /api/ops/feed-metadata {slug, dryRun?}` and `pnpm ops feed-metadata` (the new `feed-metadata` job, on the channel's download queue) fetch the channel's RSS feed once — no media — and match each record lacking a date or a real title to an item by guid, enclosure URL or enclosure file name; a record more than one item fits is reported, not guessed. A match fills in what the record lacks: the title, the description as plain text, the duration (a measured one is kept), the upload date and timestamp in UTC, and `webpage_url` only when the new URL keeps the record's id. Every change is a `feed-backfill` entry in the video's metadata history; a dry run counts matched, unmatched and already complete and writes nothing. Rebuild the index afterwards for the titles and dates to reach search and the sites. ## [0.11.0] - 2026-09-30 - **Transcripts that arrived after a video was first seen are counted.** The stats behind the homepage, the hub and every site's charts were cached per video and refreshed only when the video's metadata changed, so a transcript that came later — a Whisper run days after the download, or a video downloaded after the last index build — never reached them, and a video with YouTube captions alone had no transcription date. Counts and charts were low; the homepage could show a site with 0 transcripts, 0 channels and 0 hours while it served its videos. A stat is now also redone whenever the index re-reads the video, every transcript has a date, and a captioned video is dated by when its captions arrived rather than by a later Normalize run, so its place on "Transcribed over time" can move. **After updating, rebuild and restart the editor before anything else:** until then, **Build stats dataset** runs the old code and would undo the new stats, while a site, hub or homepage build already runs the new code — and the first stats build of any kind re-reads every video once (about 10–30 minutes on a large archive; it can be stopped and picks up where it stopped). Then build the index, the stats, the homepage, the hub, and the sites.