commit 2d9b260d4939bb95a959646b7d16febb14217e87
parent b385a8fb849dc85c401a56583b4e69458a9a0e86
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Mon, 29 Jun 2026 22:07:45 -0400
Bulk-fix incomplete transcripts: e2e coverage + changelog
Extend incomplete-transcript.spec.ts with bulk clear (asserts audio +
transcript removed, runners enabled, video re-bucketed as undownloaded),
bulk re-download queuing, and the actionable per-channel + global buttons
(including global clear-all emptying the section). Add the editor changelog
entry.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Diffstat:
2 files changed, 148 insertions(+), 1 deletion(-)
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -3,6 +3,7 @@
## [Unreleased]
- **The Deploy page is reworked around a clearer build/deploy lifecycle, with one-click build-then-deploy and batch multi-site builds.** The page now reads top-to-bottom as you'd actually ship: **Release notes** (the `## [Unreleased]` changelog preview + Cut release) → **Build & deploy** → optional **Individual steps** → **Build multiple sites**. A new **Build & deploy** button runs the build and, only if it succeeds (and wasn't cancelled), deploys it — as a single managed job with one combined streamed log and one Cancel (`buildAndDeployAction`, a composite `runManagedFunction`; cancelling mid-build skips the deploy). The new **Build multiple sites** panel kicks off a build (optionally build+deploy) for several sites at once, each rendered as its own live status-chipped log lane (`BuildSitesPanel` + `JobLane`); in Basic mode the jobs serialize on the shared build/deploy queue (the `export/` output tree is shared), with a note that true parallelism arrives with Docker mode. A **Build mode** toggle (Basic | Docker) on the page persists the choice as the default (`settings.buildPipeline`, also editable on Settings); Docker mode is a follow-up and currently falls back to a basic build with an inline notice. The build/deploy commands now share a child-streaming helper (`common/jobs/runChild.ts`) and mode-routing core (`editor/app/deploy/buildDeployCore.ts`). See `editor/app/deploy/{page.tsx,buildAction usage,components/*}`, `editor/app/build/buildAction.ts`, and `common/lib/settings.ts`.
- **Truncated transcripts are now detected and flagged for re-download.** When an audio download silently stops early (yt-dlp exits `ok`, `download-outcome.json` records success), whisper transcribes only the few minutes that landed — so a 2h22m video ends up with a ~7-minute transcript and nothing warns you. A new coverage check (last cue end ÷ video duration) flags any non-livestream video ≥10min whose transcript covers <50% of its runtime. The single source of truth is `common/lib/transcriptCoverage.ts` (`transcriptCoverage` + `isIncompleteTranscript`, with named thresholds), read from each video's `transcript.cues.json` so the existing corpus is flagged with no migration. Surfaced everywhere: a new **`incompleteTranscript`** channel-snapshot bucket → an **"Incomplete transcript"** filter chip and an **amber transcribed-dot** in the per-channel video list; a warning banner on the video page ("Transcript covers 6:52 of 2:22:21 (4.8%)…") with a one-click **Re-download & re-transcribe** button; and an **"Channels with incomplete (truncated) transcripts"** section on `/actionable`. The fix action (`redownloadIncompleteTranscriptAction`) deletes the truncated audio first, then re-downloads and re-transcribes — re-running whisper alone would just reproduce the short transcript. See `common/controller/channelSnapshot.ts`, `editor/app/channels/[slug]/{lib/videoRows.ts,lib/videoRowsServer.ts,lib/stageStatus.ts,components/VideoListPane.tsx,videos/[id]/{components/VideoPanel.tsx,videoActions.ts,page.tsx},page.tsx}`, and `editor/app/actionable/{lib/loadActionable.ts,page.tsx}`.
+- **Fix truncated transcripts in bulk — two buttons, in three places.** The per-video fix now has channel-wide and cross-channel counterparts, each offered as a **batch re-fix** (queues one job that removes the truncated audio → re-downloads → re-transcribes every flagged video in place; the transcript is never gapped) **and** a **clear & re-queue** (deletes the truncated audio + transcript so the videos drop back into the normal *undownloaded → needs-transcript* pipeline, then enables + starts the auto-download/auto-transcribe runners so they reprocess automatically). Both appear on the **`/actionable`** "incomplete transcripts" section — per-channel **Re-download & re-transcribe** / **Clear & re-queue** buttons (replacing the old "Review"-only link) plus a section-header **Re-fix all** / **Clear & re-queue all** that acts across every affected channel — and on the **channel page bulk bar** as two new Action options with a new **Select incomplete** quick-select. The clear path needs no archive pruning: `undownloadedIds` is derived purely from on-disk artifacts, and a single-video re-download isn't archive-gated. Destructive clears are confirm-gated everywhere; enabling the runners is disclosed in the confirm (note: the auto-queue policy must cover the channel for auto-reprocessing — cleared videos also surface in the existing "Download missing" / "Transcribe pending" sections as a fallback). New shared helper `editor/app/channels/[slug]/lib/fixIncompleteTranscript.ts` is the single source of truth for the per-video fix/clear, reused by the per-video action, the new `redownload-incomplete-bucket` batch job (bookmarkable; re-derives the live `incompleteTranscript` bucket), the bulk-bar wrappers, and the global actions. See `editor/app/channels/[slug]/{incompleteTranscriptActions.ts,bulkVideoActions.ts,components/VideoListPane.tsx}`, `editor/app/actionable/{actions.ts,page.tsx,components/{InlineActionButton.tsx,FixAllIncompleteButton.tsx}}`, `common/jobs/{jobKinds.ts,jobSpec.ts}`, `editor/app/jobs/jobReplayRegistry.ts`, and `editor/e2e/incomplete-transcript.spec.ts`.
- **Auto-queue rules with no bucket now draw from *all* of a runner's buckets, and auto-download can resume partial downloads.** A policy-tree rule left at the **"all buckets (default)"** setting (previously just labeled *default*) now draws from the **union** of every bucket that runner kind tracks — deduped, in priority order — instead of only the single primary bucket. This fixes channels (e.g. an Odysee channel mid-download) that quietly stopped being auto-downloaded once their remaining work drifted entirely into **partially-downloaded** videos: those have a `.part` file but no completed audio, so they live in the `partialDownloads` bucket and were **absent from `undownloadedIds`** — the only bucket auto-download used to load. The download runner now loads `partialDownloads` alongside `undownloadedIds` (partials first, so in-progress downloads resume via `downloadOneManaged` before fresh ones start), and exposes `partialDownloads` as a selectable bucket in the policy editor so you can dedicate a high-priority rule to resuming partials. The per-kind bucket lists are consolidated behind a single `bucketsForKind` source of truth shared by the runner, the per-rule pending-count helper, and the editor's bucket picker (so they can't drift). Note: a *bucketless* auto-transcribe rule now also drains `failedListed` after `downloadedNoTranscript` (it already loaded both); platform rate-limit backoff is unchanged and remains an independent reason a throttled platform may pause. See `common/jobs/autoQueuePolicy.ts` (`buildPendingByLeaf` + `bucketsForKind` + unit tests), `common/controller/autoRunner.ts`, and `editor/app/auto-queue/{page.tsx,components/PolicyTreeEditor.tsx}`.
- **Hub homepage redesigned into a cross-site landing; the homepage page-creator is removed.** The hub's home page is now a single mobile-first cross-site landing (headline KPIs and one stacked activity chart with Metric [Transcribed/Downloaded] · Breakdown [By site/By channel] · Bucket [Week/Month/Cumulative] · Range [90d/12mo/All] · Display [Share/Counts] controls, plus a metric-aware site-links grid with sparklines and a "#1 this month" badge), built from a small `homepage-summary.json` pre-computed by `compose-homepage`. The separate `/stats` dashboard route folds into it. Consequently the hub's **Markdown-pages subsystem is dropped**: **Manage → Homepage** now edits only branding (the Pages list, New-page, and the page editor are gone), and the homepage config no longer carries a `nav`. The page server actions (`saveHomepagePageAction`/`deleteHomepagePageAction`), `editor/app/homepage/pages/*`, `PageEditor.tsx`, `common/lib/{homepagePages,homepageConstants}.ts`, and `paths.homepagePagesDir` are removed. See `editor/app/homepage/{page.tsx,actions.ts}`, `common/bin/compose-homepage.ts`, `common/lib/{homepageSummary,homepageChart}.ts`, and the `homepage/` package. (Re-addable later if needed.)
- **One-click Retry for failed jobs (plus "Retry all failed").** A failed job that carries a replay descriptor (any bookmarkable kind — sync, download-missing, transcribe-all, retry-bucket, …) now shows a **Retry** button on the Jobs history table, and the page header gains a **Retry all failed** button whenever at least one such job is listed. Retry re-runs the job from its stored spec exactly like a bookmark re-run (so bucket jobs re-derive from the channel's *current* state), and the re-run **jumps ahead of other queued work** (it's promoted to the front of its queue, reusing the new reorder machinery) so a fix-and-retry runs next rather than at the back of the line. The spec is resolved from the live registry or, for an evicted/archived job, from its on-disk `<id>.meta.json` sidecar — so even a failure the 100-job cap has dropped is still retryable. Kinds with no replay descriptor (e.g. `import-one`) intentionally offer no Retry. See `editor/app/jobs/actions.ts` (`retryJobAction` / `retryAllFailedAction`), the new `RetryJobButton` / `RetryAllFailedButton`, and `editor/e2e/jobs-retry.spec.ts`.
diff --git a/editor/e2e/incomplete-transcript.spec.ts b/editor/e2e/incomplete-transcript.spec.ts
@@ -9,7 +9,7 @@
import { mkdir, writeFile } from "node:fs/promises";
import { test, expect } from "@playwright/test";
-import { resetData, resolvePath } from "./helpers";
+import { pathExists, readJson, resetData, resolvePath } from "./helpers";
const CHANNEL = "test-transcribe";
const DATA = `test-transcripts/channels/${CHANNEL}/data`;
@@ -126,3 +126,149 @@ test("incomplete-transcript filter, glyph, panel banner, and actionable", async
section.getByLabel(`incomplete-transcripts row ${CHANNEL}`),
).toBeVisible();
});
+
+test("channel bulk bar: clear incomplete resets the video and enables auto-runners", async ({
+ page,
+}) => {
+ await seed();
+ await page.goto(`/channels/${CHANNEL}`);
+
+ // Select the flagged video via the new quick-select, then clear it.
+ await page
+ .getByRole("button", { name: "Select incomplete", exact: true })
+ .click();
+ const bar = page.getByLabel("bulk action bar");
+ await expect(bar).toBeVisible();
+ await page.getByLabel("bulk action", { exact: true }).selectOption("clear_incomplete");
+ page.once("dialog", (d) => d.accept());
+ await page.getByLabel("apply bulk action").click();
+ // Selection clears on success → the bar hides.
+ await expect(bar).toBeHidden();
+
+ // The truncated audio + transcript are gone on disk.
+ await expect
+ .poll(() => pathExists(`${DATA}/vidTrunc/audio.m4a`))
+ .toBe(false);
+ await expect
+ .poll(() => pathExists(`${DATA}/vidTrunc/transcript.json`))
+ .toBe(false);
+ await expect
+ .poll(() => pathExists(`${DATA}/vidTrunc/transcript.cues.json`))
+ .toBe(false);
+
+ // The auto-download + auto-transcribe runners are now enabled.
+ const settings = await readJson<{
+ autoQueue: {
+ transcription: { enabled: boolean };
+ download: { enabled: boolean };
+ };
+ }>("test-settings.json");
+ expect(settings.autoQueue.transcription.enabled).toBe(true);
+ expect(settings.autoQueue.download.enabled).toBe(true);
+
+ // The video is no longer flagged and now reads as undownloaded ("No audio").
+ // The channel snapshot regenerates on a ~1s debounce after the clear, and the
+ // channel page serves the persisted snapshot, so re-navigate until it's fresh.
+ const list = page.getByLabel("videos", { exact: true });
+ await expect
+ .poll(
+ async () => {
+ await page.goto(`/channels/${CHANNEL}?filter=incomplete_transcript`);
+ return list.getByLabel("open vidTrunc").count();
+ },
+ { timeout: 15000 },
+ )
+ .toBe(0);
+ await page.goto(`/channels/${CHANNEL}?filter=no_audio`);
+ await expect(list.getByLabel("open vidTrunc")).toBeVisible();
+});
+
+test("channel bulk bar: re-download & re-transcribe queues a batch fix", async ({
+ page,
+}) => {
+ await seed();
+ await page.goto(`/channels/${CHANNEL}`);
+
+ await page
+ .getByRole("button", { name: "Select incomplete", exact: true })
+ .click();
+ const bar = page.getByLabel("bulk action bar");
+ await expect(bar).toBeVisible();
+ await page.getByLabel("bulk action", { exact: true }).selectOption("redownload_incomplete");
+ await page.getByLabel("apply bulk action").click();
+ // A streaming bulk action clears the selection on success (the batch job runs
+ // in the background) → the bar hides with no error surfaced.
+ await expect(bar).toBeHidden();
+ await expect(page.getByLabel("bulk action error")).toBeHidden();
+});
+
+test("actionable: section exposes per-channel + global fix buttons; per-channel re-download queues a job", async ({
+ page,
+}) => {
+ await seed();
+ // Visiting the channel materializes its snapshot so /actionable lists it.
+ await page.goto(`/channels/${CHANNEL}`);
+ await page.goto(`/actionable`);
+ const section = page.getByRole("region", {
+ name: "incomplete-transcripts",
+ exact: true,
+ });
+ await expect(section).toBeVisible();
+
+ // Per-channel and global buttons are all present.
+ await expect(
+ section.getByLabel(`re-download & re-transcribe ${CHANNEL}`),
+ ).toBeVisible();
+ await expect(
+ section.getByLabel(`clear & re-queue ${CHANNEL}`),
+ ).toBeVisible();
+ await expect(
+ section.getByLabel("re-download all incomplete transcripts"),
+ ).toBeVisible();
+ await expect(
+ section.getByLabel("clear all incomplete transcripts"),
+ ).toBeVisible();
+
+ // Per-channel re-download queues a job (the job id is returned synchronously).
+ await section.getByLabel(`re-download & re-transcribe ${CHANNEL}`).click();
+ await expect(
+ section.getByLabel(`re-download & re-transcribe ${CHANNEL} job`),
+ ).toBeVisible();
+});
+
+test("actionable: global clear-all clears every flagged video and empties the section", async ({
+ page,
+}) => {
+ await seed();
+ // Visiting the channel materializes its snapshot so /actionable lists it.
+ await page.goto(`/channels/${CHANNEL}`);
+ await page.goto(`/actionable`);
+ const section = page.getByRole("region", {
+ name: "incomplete-transcripts",
+ exact: true,
+ });
+ await expect(section).toBeVisible();
+
+ // Global clear-all clears every flagged video across all channels (synchronous
+ // fs op — no background job to race the assertions below).
+ page.once("dialog", (d) => d.accept());
+ await section.getByLabel("clear all incomplete transcripts").click();
+ await expect(
+ section.getByLabel("fix all incomplete transcripts result"),
+ ).toContainText(/Cleared 1/);
+ await expect
+ .poll(() => pathExists(`${DATA}/vidTrunc/transcript.cues.json`))
+ .toBe(false);
+
+ // The snapshot regenerates on a ~1s debounce after the clear; /actionable
+ // reads the persisted snapshot, so re-navigate until the section is empty.
+ await expect
+ .poll(
+ async () => {
+ await page.goto(`/actionable`);
+ return page.getByLabel("incomplete-transcripts empty").count();
+ },
+ { timeout: 15000 },
+ )
+ .toBe(1);
+});