Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit d75c63c989dcd750945c3101c086cf2d1e074e7f
parent bb156c29beaa3bd7e0c048dc82447bd6d047e788
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Thu,  1 Oct 2026 21:10:17 -0400

plans, editor: FACTS and the changelog bullet say what the index does with a tiered live chat; the bullet needs a rebuild and restart (review L11)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Diffstat:
Meditor/CHANGELOG.md | 2+-
Mplans/FACTS.md | 10++++++++--
2 files changed, 9 insertions(+), 3 deletions(-)

diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -16,7 +16,7 @@ - **A form whose save is refused keeps what you typed.** Every editor form put its plain fields back to the stored values when its save was refused — a site's ID rejected, a page size out of range, a slug already taken — so everything typed had to be typed again. A refused save now leaves every field as you left it, beside the reason: **Settings**; a site's form (new and existing); the hub's config on `/sites`; **Cut release**; a channel's form (new and **Configure**), **Rename** and **Delete**; a video's **Delete directory**; **Drive health timing** on `/storage`; the backup config on `/saved-videos`; the sync operation's controls; the **Digest**, **Diarization**, **Speaker attribution** and **Speaker work lane** settings; and the worker list on `/workers`. A save that succeeds behaves as before, with one difference you may notice: a drop-down, and a checkbox or choice that the page tracks as you change it (a cadence, a worker's **Enabled**, a social link's **Keep in header**, a site membership, a site's accent), now shows what was saved. A form's own drop-downs used to go back to what the page had loaded with until a reload, and a second save from the same page sent that old choice again; the others went back until the page next refreshed itself (every 5 seconds by default). - **A media move no longer starts over a job that is writing into the channel, holds the channel's writers while it runs, and makes its copy match the source before it verifies — so a transcription or a download during a move cannot fail it.** A move that has waited its turn behind other moves now checks again when it starts: if a job is running on the channel, or an auto-queue lane is working on one of its videos, it stops at once and says which ("a transcription of abc123 is running (Transcribe all, job …) — wait for it or cancel it"), with nothing copied — and a job you have just cancelled counts until it has actually stopped ("is stopping … — wait for it to stop"); **Preview** says the same, and the Storage panel's blocked message now names the job too. While a move's marker stands, the channel is held: every lane skips it, and every job that reads or writes its media (single-video transcriptions, downloads and transcodes and the availability checks now included) refuses to start, including one that was already queued when the move began. The rack shows a **media held** chip in the channel's Tier cell and the Storage panel says "Held: its media is moving"; both go when the move finishes or its marker is cleared. The copy is now followed by a pass that makes the destination copy match the source — files the source no longer has are removed from the copy, never from the source — so a file written or deleted during the copy (a transcriber's scratch folder, say) no longer fails the check, and **Resume move** finishes a move whose copy holds such leftovers. Every file removed from a copy is listed in the move's log, and **Preview** says so when a copy from an earlier attempt is already there. If the source keeps changing, the move stops and lists what differs: extra on the destination, missing there, or changed. A new **Reconcile and resume** button beside **Resume move** lists those differences, makes the copy match and finishes the move, so no file has to be deleted by hand. The saved-video store's move does the same matching and the same check before it starts. Needs a rebuild and restart of the editor. - **Connecting an X account opens your own browser, and the X fetchers can use your everyday browser's X login instead.** **Settings → X account session → Connect X account** used to open Playwright's bundled Chromium with its automation signals on (the "controlled by automated test software" bar, `navigator.webdriver`): Google's sign-in refused it and X's own login form stalled in it. It now opens your Chromium or Chrome when one is installed (`ARCHILYZER_X_BROWSER` names another; Playwright's bundled Chromium otherwise), without those signals. Google's sign-in may still refuse an embedded browser; X's password login is the reliable path. A new **Login source** choice (`social.x.cookieSource` in `settings.json`) says where the X fetchers' login comes from: **Browser login** hands gallery-dl `--cookies-from-browser` with your `cookiesFromBrowser` on every fetch, so the login lasts as long as you stay logged in to x.com in that browser and no window is needed; **Connected profile** is the session broker, as before. Left on **Automatic**, it is the browser login when `cookiesFromBrowser` is set and no profile is connected, and the profile otherwise. **Check** says which source is in use, whether an X login is visible in it and when it was last used (the browser's cookies are read from a private copy, never written; this reads Firefox's, and gallery-dl reads Chromium's itself). Needs a rebuild and restart of the editor. -- **A channel's text stays on the fast disk when its media moves, so a slow or unplugged media drive no longer holds its transcripts.** A channel's big files — the audio and the raw live-chat replay — can now live in the channel's own `media` folder, on this disk or another, while its transcripts, cues, metadata and every other small file stay in `data/` where they always were; each big file that moves leaves a small link behind, so everything that opens it by name still finds it. New downloads, transcodes and live-chat normalizes put their big files there as they finish, and every cleanup that deletes audio removes the file the link points to, not just the link. What that changes when a media drive is stalled, unplugged or mid-move: **the index and stats builds read the text only and are never held by it**, the channel's report still refreshes (its media size reads as unknown until the drive answers), **digests keep running — even during a move of that channel's media** — and so do normalize, the availability checks, the metadata scan and clip eviction. Transcription, downloads, the backfill lane and anything else that opens the audio are held as before. **A channel moved the old way — its whole `data/` on the other drive — is now shown as "Media layout retired" and held by everything, the builds and digests included, until `archilyzer storage migrate-tier <channel>` brings its text home;** every refusal says so. +- **A channel's text stays on the fast disk when its media moves, so a slow or unplugged media drive no longer holds its transcripts.** A channel's big files — the audio and the raw live-chat replay — can now live in the channel's own `media` folder, on this disk or another, while its transcripts, cues, metadata and every other small file stay in `data/` where they always were; each big file that moves leaves a small link behind, so everything that opens it by name still finds it. New downloads, transcodes and live-chat normalizes put their big files there as they finish, and every cleanup that deletes audio removes the file the link points to, not just the link. What that changes when a media drive is stalled, unplugged or mid-move: **the index and stats builds never wait on it or are held by it** (a live chat whose transcript cues are out of date keeps the cues the last build read until the drive answers), the channel's report still refreshes (its media size reads as unknown until the drive answers), **digests keep running — even during a move of that channel's media** — and so do normalize, the availability checks, the metadata scan and clip eviction. Transcription, downloads, the backfill lane and anything else that opens the audio are held as before. **A channel moved the old way — its whole `data/` on the other drive — is now shown as "Media layout retired" and held by everything, the builds and digests included, until `archilyzer storage migrate-tier <channel>` brings its text home;** every refusal says so. Deleting a video from its page is refused while its channel's media drive is not reachable, so its audio is never left behind on the drive. Needs a rebuild and restart of the editor. ## [0.11.0] - 2026-09-30 - **Transcripts that arrived after a video was first seen are counted.** The stats behind the homepage, the hub and every site's charts were cached per video and refreshed only when the video's metadata changed, so a transcript that came later — a Whisper run days after the download, or a video downloaded after the last index build — never reached them, and a video with YouTube captions alone had no transcription date. Counts and charts were low; the homepage could show a site with 0 transcripts, 0 channels and 0 hours while it served its videos. A stat is now also redone whenever the index re-reads the video, every transcript has a date, and a captioned video is dated by when its captions arrived rather than by a later Normalize run, so its place on "Transcribed over time" can move. **After updating, rebuild and restart the editor before anything else:** until then, **Build stats dataset** runs the old code and would undo the new stats, while a site, hub or homepage build already runs the new code — and the first stats build of any kind re-reads every video once (about 10–30 minutes on a large archive; it can be stopped and picks up where it stopped). Then build the index, the stats, the homepage, the hub, and the sites. diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -3883,8 +3883,14 @@ as they were measured. recorded as `config.mediaDir` (relocated), or absent (classic). `data/` is always a real dir on the corpus disk. The hook (`tierMediaFile` & co., `lib/mediaTier-server.ts`) is called after every media finalisation (transcode, both download-outcome writes, a batch `runYtdlp` mode, a live-chat - normalize that wrote) and never throws; the link carries the file's mtime (`lutimes`), and the - live-chat freshness check `lstat`s the raw, so the index never stats a media file. Every deleter of + normalize that wrote — not the export build's archive pass) and never throws; the bytes are placed by + a hard link (same disk) or an atomic copy before one rename swaps the name, so a crash never leaves the + name missing; the link carries the file's mtime (`lutimes`), and the + live-chat freshness check and the index's sub-track key (`subsMs`) `lstat` a tierable name, so the + index's scan never reaches the media drive; when a tiered live chat's cues are stale the index reads + the raw through `onDrive(mediaDir)` only while the channel's media is `ok`/`in-place`, and otherwise + keeps the cues the last build held (`buildIndex.test.ts` (k): a stalled drive, nothing asked of it). + An EXDEV tier gives the copy the file's times, so `stat` and `lstat` agree. Every deleter of a video-dir entry goes through `removeMediaFile` / `removeVideoDirMedia`. - **`inspectChannelMedia`** now describes the `media` link + `mediaDir`, returns `mediaLink` and `text: { dir, readable }`, and has a seventh status, **`legacy`**: a `data` link or a recorded