Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 2d9b260d4939bb95a959646b7d16febb14217e87
parent b385a8fb849dc85c401a56583b4e69458a9a0e86
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Mon, 29 Jun 2026 22:07:45 -0400

Bulk-fix incomplete transcripts: e2e coverage + changelog

Extend incomplete-transcript.spec.ts with bulk clear (asserts audio +
transcript removed, runners enabled, video re-bucketed as undownloaded),
bulk re-download queuing, and the actionable per-channel + global buttons
(including global clear-all emptying the section). Add the editor changelog
entry.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

Diffstat:
Meditor/CHANGELOG.md | 1+
Meditor/e2e/incomplete-transcript.spec.ts | 148++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-
2 files changed, 148 insertions(+), 1 deletion(-)

diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -3,6 +3,7 @@ ## [Unreleased] - **The Deploy page is reworked around a clearer build/deploy lifecycle, with one-click build-then-deploy and batch multi-site builds.** The page now reads top-to-bottom as you'd actually ship: **Release notes** (the `## [Unreleased]` changelog preview + Cut release) → **Build & deploy** → optional **Individual steps** → **Build multiple sites**. A new **Build & deploy** button runs the build and, only if it succeeds (and wasn't cancelled), deploys it — as a single managed job with one combined streamed log and one Cancel (`buildAndDeployAction`, a composite `runManagedFunction`; cancelling mid-build skips the deploy). The new **Build multiple sites** panel kicks off a build (optionally build+deploy) for several sites at once, each rendered as its own live status-chipped log lane (`BuildSitesPanel` + `JobLane`); in Basic mode the jobs serialize on the shared build/deploy queue (the `export/` output tree is shared), with a note that true parallelism arrives with Docker mode. A **Build mode** toggle (Basic | Docker) on the page persists the choice as the default (`settings.buildPipeline`, also editable on Settings); Docker mode is a follow-up and currently falls back to a basic build with an inline notice. The build/deploy commands now share a child-streaming helper (`common/jobs/runChild.ts`) and mode-routing core (`editor/app/deploy/buildDeployCore.ts`). See `editor/app/deploy/{page.tsx,buildAction usage,components/*}`, `editor/app/build/buildAction.ts`, and `common/lib/settings.ts`. - **Truncated transcripts are now detected and flagged for re-download.** When an audio download silently stops early (yt-dlp exits `ok`, `download-outcome.json` records success), whisper transcribes only the few minutes that landed — so a 2h22m video ends up with a ~7-minute transcript and nothing warns you. A new coverage check (last cue end ÷ video duration) flags any non-livestream video ≥10min whose transcript covers <50% of its runtime. The single source of truth is `common/lib/transcriptCoverage.ts` (`transcriptCoverage` + `isIncompleteTranscript`, with named thresholds), read from each video's `transcript.cues.json` so the existing corpus is flagged with no migration. Surfaced everywhere: a new **`incompleteTranscript`** channel-snapshot bucket → an **"Incomplete transcript"** filter chip and an **amber transcribed-dot** in the per-channel video list; a warning banner on the video page ("Transcript covers 6:52 of 2:22:21 (4.8%)…") with a one-click **Re-download & re-transcribe** button; and an **"Channels with incomplete (truncated) transcripts"** section on `/actionable`. The fix action (`redownloadIncompleteTranscriptAction`) deletes the truncated audio first, then re-downloads and re-transcribes — re-running whisper alone would just reproduce the short transcript. See `common/controller/channelSnapshot.ts`, `editor/app/channels/[slug]/{lib/videoRows.ts,lib/videoRowsServer.ts,lib/stageStatus.ts,components/VideoListPane.tsx,videos/[id]/{components/VideoPanel.tsx,videoActions.ts,page.tsx},page.tsx}`, and `editor/app/actionable/{lib/loadActionable.ts,page.tsx}`. +- **Fix truncated transcripts in bulk — two buttons, in three places.** The per-video fix now has channel-wide and cross-channel counterparts, each offered as a **batch re-fix** (queues one job that removes the truncated audio → re-downloads → re-transcribes every flagged video in place; the transcript is never gapped) **and** a **clear & re-queue** (deletes the truncated audio + transcript so the videos drop back into the normal *undownloaded → needs-transcript* pipeline, then enables + starts the auto-download/auto-transcribe runners so they reprocess automatically). Both appear on the **`/actionable`** "incomplete transcripts" section — per-channel **Re-download & re-transcribe** / **Clear & re-queue** buttons (replacing the old "Review"-only link) plus a section-header **Re-fix all** / **Clear & re-queue all** that acts across every affected channel — and on the **channel page bulk bar** as two new Action options with a new **Select incomplete** quick-select. The clear path needs no archive pruning: `undownloadedIds` is derived purely from on-disk artifacts, and a single-video re-download isn't archive-gated. Destructive clears are confirm-gated everywhere; enabling the runners is disclosed in the confirm (note: the auto-queue policy must cover the channel for auto-reprocessing — cleared videos also surface in the existing "Download missing" / "Transcribe pending" sections as a fallback). New shared helper `editor/app/channels/[slug]/lib/fixIncompleteTranscript.ts` is the single source of truth for the per-video fix/clear, reused by the per-video action, the new `redownload-incomplete-bucket` batch job (bookmarkable; re-derives the live `incompleteTranscript` bucket), the bulk-bar wrappers, and the global actions. See `editor/app/channels/[slug]/{incompleteTranscriptActions.ts,bulkVideoActions.ts,components/VideoListPane.tsx}`, `editor/app/actionable/{actions.ts,page.tsx,components/{InlineActionButton.tsx,FixAllIncompleteButton.tsx}}`, `common/jobs/{jobKinds.ts,jobSpec.ts}`, `editor/app/jobs/jobReplayRegistry.ts`, and `editor/e2e/incomplete-transcript.spec.ts`. - **Auto-queue rules with no bucket now draw from *all* of a runner's buckets, and auto-download can resume partial downloads.** A policy-tree rule left at the **"all buckets (default)"** setting (previously just labeled *default*) now draws from the **union** of every bucket that runner kind tracks — deduped, in priority order — instead of only the single primary bucket. This fixes channels (e.g. an Odysee channel mid-download) that quietly stopped being auto-downloaded once their remaining work drifted entirely into **partially-downloaded** videos: those have a `.part` file but no completed audio, so they live in the `partialDownloads` bucket and were **absent from `undownloadedIds`** — the only bucket auto-download used to load. The download runner now loads `partialDownloads` alongside `undownloadedIds` (partials first, so in-progress downloads resume via `downloadOneManaged` before fresh ones start), and exposes `partialDownloads` as a selectable bucket in the policy editor so you can dedicate a high-priority rule to resuming partials. The per-kind bucket lists are consolidated behind a single `bucketsForKind` source of truth shared by the runner, the per-rule pending-count helper, and the editor's bucket picker (so they can't drift). Note: a *bucketless* auto-transcribe rule now also drains `failedListed` after `downloadedNoTranscript` (it already loaded both); platform rate-limit backoff is unchanged and remains an independent reason a throttled platform may pause. See `common/jobs/autoQueuePolicy.ts` (`buildPendingByLeaf` + `bucketsForKind` + unit tests), `common/controller/autoRunner.ts`, and `editor/app/auto-queue/{page.tsx,components/PolicyTreeEditor.tsx}`. - **Hub homepage redesigned into a cross-site landing; the homepage page-creator is removed.** The hub's home page is now a single mobile-first cross-site landing (headline KPIs and one stacked activity chart with Metric [Transcribed/Downloaded] · Breakdown [By site/By channel] · Bucket [Week/Month/Cumulative] · Range [90d/12mo/All] · Display [Share/Counts] controls, plus a metric-aware site-links grid with sparklines and a "#1 this month" badge), built from a small `homepage-summary.json` pre-computed by `compose-homepage`. The separate `/stats` dashboard route folds into it. Consequently the hub's **Markdown-pages subsystem is dropped**: **Manage → Homepage** now edits only branding (the Pages list, New-page, and the page editor are gone), and the homepage config no longer carries a `nav`. The page server actions (`saveHomepagePageAction`/`deleteHomepagePageAction`), `editor/app/homepage/pages/*`, `PageEditor.tsx`, `common/lib/{homepagePages,homepageConstants}.ts`, and `paths.homepagePagesDir` are removed. See `editor/app/homepage/{page.tsx,actions.ts}`, `common/bin/compose-homepage.ts`, `common/lib/{homepageSummary,homepageChart}.ts`, and the `homepage/` package. (Re-addable later if needed.) - **One-click Retry for failed jobs (plus "Retry all failed").** A failed job that carries a replay descriptor (any bookmarkable kind — sync, download-missing, transcribe-all, retry-bucket, …) now shows a **Retry** button on the Jobs history table, and the page header gains a **Retry all failed** button whenever at least one such job is listed. Retry re-runs the job from its stored spec exactly like a bookmark re-run (so bucket jobs re-derive from the channel's *current* state), and the re-run **jumps ahead of other queued work** (it's promoted to the front of its queue, reusing the new reorder machinery) so a fix-and-retry runs next rather than at the back of the line. The spec is resolved from the live registry or, for an evicted/archived job, from its on-disk `<id>.meta.json` sidecar — so even a failure the 100-job cap has dropped is still retryable. Kinds with no replay descriptor (e.g. `import-one`) intentionally offer no Retry. See `editor/app/jobs/actions.ts` (`retryJobAction` / `retryAllFailedAction`), the new `RetryJobButton` / `RetryAllFailedButton`, and `editor/e2e/jobs-retry.spec.ts`. diff --git a/editor/e2e/incomplete-transcript.spec.ts b/editor/e2e/incomplete-transcript.spec.ts @@ -9,7 +9,7 @@ import { mkdir, writeFile } from "node:fs/promises"; import { test, expect } from "@playwright/test"; -import { resetData, resolvePath } from "./helpers"; +import { pathExists, readJson, resetData, resolvePath } from "./helpers"; const CHANNEL = "test-transcribe"; const DATA = `test-transcripts/channels/${CHANNEL}/data`; @@ -126,3 +126,149 @@ test("incomplete-transcript filter, glyph, panel banner, and actionable", async section.getByLabel(`incomplete-transcripts row ${CHANNEL}`), ).toBeVisible(); }); + +test("channel bulk bar: clear incomplete resets the video and enables auto-runners", async ({ + page, +}) => { + await seed(); + await page.goto(`/channels/${CHANNEL}`); + + // Select the flagged video via the new quick-select, then clear it. + await page + .getByRole("button", { name: "Select incomplete", exact: true }) + .click(); + const bar = page.getByLabel("bulk action bar"); + await expect(bar).toBeVisible(); + await page.getByLabel("bulk action", { exact: true }).selectOption("clear_incomplete"); + page.once("dialog", (d) => d.accept()); + await page.getByLabel("apply bulk action").click(); + // Selection clears on success → the bar hides. + await expect(bar).toBeHidden(); + + // The truncated audio + transcript are gone on disk. + await expect + .poll(() => pathExists(`${DATA}/vidTrunc/audio.m4a`)) + .toBe(false); + await expect + .poll(() => pathExists(`${DATA}/vidTrunc/transcript.json`)) + .toBe(false); + await expect + .poll(() => pathExists(`${DATA}/vidTrunc/transcript.cues.json`)) + .toBe(false); + + // The auto-download + auto-transcribe runners are now enabled. + const settings = await readJson<{ + autoQueue: { + transcription: { enabled: boolean }; + download: { enabled: boolean }; + }; + }>("test-settings.json"); + expect(settings.autoQueue.transcription.enabled).toBe(true); + expect(settings.autoQueue.download.enabled).toBe(true); + + // The video is no longer flagged and now reads as undownloaded ("No audio"). + // The channel snapshot regenerates on a ~1s debounce after the clear, and the + // channel page serves the persisted snapshot, so re-navigate until it's fresh. + const list = page.getByLabel("videos", { exact: true }); + await expect + .poll( + async () => { + await page.goto(`/channels/${CHANNEL}?filter=incomplete_transcript`); + return list.getByLabel("open vidTrunc").count(); + }, + { timeout: 15000 }, + ) + .toBe(0); + await page.goto(`/channels/${CHANNEL}?filter=no_audio`); + await expect(list.getByLabel("open vidTrunc")).toBeVisible(); +}); + +test("channel bulk bar: re-download & re-transcribe queues a batch fix", async ({ + page, +}) => { + await seed(); + await page.goto(`/channels/${CHANNEL}`); + + await page + .getByRole("button", { name: "Select incomplete", exact: true }) + .click(); + const bar = page.getByLabel("bulk action bar"); + await expect(bar).toBeVisible(); + await page.getByLabel("bulk action", { exact: true }).selectOption("redownload_incomplete"); + await page.getByLabel("apply bulk action").click(); + // A streaming bulk action clears the selection on success (the batch job runs + // in the background) → the bar hides with no error surfaced. + await expect(bar).toBeHidden(); + await expect(page.getByLabel("bulk action error")).toBeHidden(); +}); + +test("actionable: section exposes per-channel + global fix buttons; per-channel re-download queues a job", async ({ + page, +}) => { + await seed(); + // Visiting the channel materializes its snapshot so /actionable lists it. + await page.goto(`/channels/${CHANNEL}`); + await page.goto(`/actionable`); + const section = page.getByRole("region", { + name: "incomplete-transcripts", + exact: true, + }); + await expect(section).toBeVisible(); + + // Per-channel and global buttons are all present. + await expect( + section.getByLabel(`re-download & re-transcribe ${CHANNEL}`), + ).toBeVisible(); + await expect( + section.getByLabel(`clear & re-queue ${CHANNEL}`), + ).toBeVisible(); + await expect( + section.getByLabel("re-download all incomplete transcripts"), + ).toBeVisible(); + await expect( + section.getByLabel("clear all incomplete transcripts"), + ).toBeVisible(); + + // Per-channel re-download queues a job (the job id is returned synchronously). + await section.getByLabel(`re-download & re-transcribe ${CHANNEL}`).click(); + await expect( + section.getByLabel(`re-download & re-transcribe ${CHANNEL} job`), + ).toBeVisible(); +}); + +test("actionable: global clear-all clears every flagged video and empties the section", async ({ + page, +}) => { + await seed(); + // Visiting the channel materializes its snapshot so /actionable lists it. + await page.goto(`/channels/${CHANNEL}`); + await page.goto(`/actionable`); + const section = page.getByRole("region", { + name: "incomplete-transcripts", + exact: true, + }); + await expect(section).toBeVisible(); + + // Global clear-all clears every flagged video across all channels (synchronous + // fs op — no background job to race the assertions below). + page.once("dialog", (d) => d.accept()); + await section.getByLabel("clear all incomplete transcripts").click(); + await expect( + section.getByLabel("fix all incomplete transcripts result"), + ).toContainText(/Cleared 1/); + await expect + .poll(() => pathExists(`${DATA}/vidTrunc/transcript.cues.json`)) + .toBe(false); + + // The snapshot regenerates on a ~1s debounce after the clear; /actionable + // reads the persisted snapshot, so re-navigate until the section is empty. + await expect + .poll( + async () => { + await page.goto(`/actionable`); + return page.getByLabel("incomplete-transcripts empty").count(); + }, + { timeout: 15000 }, + ) + .toBe(1); +});