Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit ba12c240b94c2118fc69d5975a4743e04437352a
parent 29907945395ab14b6aed83ec615c7a764dcf482c
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Tue,  8 Sep 2026 21:29:29 -0400

editor: three e2e cases that read a run's log after the run's own result lands

`a6e6d46` fixed the two speaker bodies and named the rule: a component
carrying a StreamActionLog must not be UNMOUNTED when the record its own
run produced lands, because StreamActionLog keeps its log in useState,
fires router.refresh() itself at the end of a run, and renders no
<pre role="log"> at all once its state is empty. `plans/FACTS.md` lists
three more instances of it in VideoPanel.tsx, none of them under a spec.
These are the specs. Each starts the operation from the surface that owns
the log, waits for the STATE THE RUN PRODUCES to appear (the proof the
refresh landed, as attribution.spec.ts does with the freshness pill), and
only then reads the log's closing line:

  - Persist source video: the saved-video pointer lands and the panel
    draws the persisted view (saved-videos.spec.ts).
  - Re-download & re-transcribe: the transcript is no longer truncated,
    so "Transcript looks truncated" is gone (incomplete-transcript.spec.ts).
  - Re-download as Original: a successful download outcome, so "Download
    was truncated at the source" is gone (download-format-guard.spec.ts).

All three are red at this commit and green at the next.

The persist case needed the fake yt-dlp to grow an app-extraction branch.
That mode is how the app persists a source video — no -x, so yt-dlp leaves
the full container and the app extracts audio from it with ffmpeg — and the
fake had no branch for it, so the invocation fell through to "unknown
invocation" and exit 2, no container, no pointer, nothing for the refresh
to land. It is the LAST branch checked, matched on the media output
template (`source-media`, which only that mode uses), so no invocation the
fake already recognised changes shape.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Diffstat:
Meditor/e2e/download-format-guard.spec.ts | 59+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/e2e/fixtures/bin/fake-ytdlp.mjs | 47+++++++++++++++++++++++++++++++++++++++++++++++
Meditor/e2e/incomplete-transcript.spec.ts | 92+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/e2e/saved-videos.spec.ts | 41+++++++++++++++++++++++++++++++++++++++++
4 files changed, 239 insertions(+), 0 deletions(-)

diff --git a/editor/e2e/download-format-guard.spec.ts b/editor/e2e/download-format-guard.spec.ts @@ -160,6 +160,65 @@ test("duration guard flags a source-truncated download and keeps the stub", asyn await expect(section.getByLabel(`short-audio row ${SLUG}`)).toBeVisible(); }); +// The re-download offered by the short-audio banner is what CLEARS the flag: a +// full download rewrites download-outcome.json, and the banner's condition is +// gone by the time the run ends. StreamActionLog fires that refresh itself, so a +// parent that stops rendering the banner on it destroys the run's log, the +// <pre role="log"> included. The warning text disappearing is the proof the +// refresh landed; the log is read only after that. See plans/FACTS.md, "A run +// log lives in the panel's React state". +test("the short-audio re-download's log survives the refresh that clears the banner", async ({ + page, +}) => { + const SLUG = "shortaudio-log"; + // No `truncaud` sentinel: the re-download lands full-length audio, so the + // duration guard passes and the outcome is rewritten as a success. + const ID = "shortaudiofix1"; + await resetData(); + await makeChannel(SLUG, [ID]); + // Seed the flagged state directly (the kept short stub plus the guard's + // sidecar) rather than re-running a truncated download for it. + const dir = resolvePath(`test-transcripts/channels/${SLUG}/data/${ID}`); + await mkdir(dir, { recursive: true }); + await writeFile(`${dir}/audio.mp3`, "fake truncated audio __DUR=120__\n"); + await writeFile( + `${dir}/download-outcome.json`, + JSON.stringify({ + videoId: ID, + status: "failed-short-audio", + startedAt: "2026-06-01T00:00:00.000Z", + finishedAt: "2026-06-01T00:00:10.000Z", + attempts: [], + shortAudio: { + audioDurationSec: 120, + expectedDurationSec: 6000, + coverage: 0.02, + }, + }), + ); + await invalidate(); + + await page.goto(`/channels/${SLUG}/videos/${ID}`); + const banner = page.getByLabel("short audio"); + await expect(banner).toContainText("Download was truncated at the source"); + const run = banner.getByRole("button", { name: /Re-download as Original/ }); + await expect(run).toBeEnabled(); + await run.click(); + + const log = page.getByLabel(`Re-download ${ID} as Original output`); + await expect(log).toContainText(`[download] Fetching ${ID}`, { + timeout: 15_000, + }); + + // The state the run produces: a successful download outcome, so the truncated + // warning is gone. + await expect( + page.getByText("Download was truncated at the source"), + ).toHaveCount(0, { timeout: 15_000 }); + // And the log is STILL on screen, carrying the run's closing line. + await expect(log).toContainText("download complete"); +}); + test("download format: Odysee auto picks original; YouTube uses bestaudio; channel override forces original", async ({ page, }) => { diff --git a/editor/e2e/fixtures/bin/fake-ytdlp.mjs b/editor/e2e/fixtures/bin/fake-ytdlp.mjs @@ -677,6 +677,53 @@ async function main() { return; } + // App-extraction mode (transcribe handling, persisting the source video): + // no -x, no -c, no --skip-download — yt-dlp is asked for the FULL container + // and the app extracts audio from it with ffmpeg afterwards + // (downloadOneManaged's finalizeAppExtraction). The only reliable marker is + // the media output template, which names source-media instead of audio + // (runYtdlp.ts outputArgsForUrl). Last branch before the error, so no + // previously-recognised invocation changes shape. + if ((arg("-o") ?? "").includes("source-media")) { + // As in the other single-URL modes, the real download reuses the prefetched + // metadata via --load-info-json and passes NO positional URL; derive the id + // from the info.json path (our layout is `data/<id>/metadata.info.json`). + const infoJsonPath = arg("--load-info-json"); + let url; + if (infoJsonPath) { + url = `https://www.youtube.com/watch?v=${path.basename(path.dirname(infoJsonPath))}`; + await appendFile( + "fake-ytdlp.invocations", + `load-info-json:${infoJsonPath}\n`, + ); + } else { + url = lastNonFlag(); + } + if (!url) { + process.stderr.write(`[fake-ytdlp] app-extraction mode missing URL\n`); + process.exit(2); + } + const id = urlIdYouTube(url) ?? "unknown"; + const videoDir = path.join("data", id); + await ensureDir(videoDir); + await appendFile( + "fake-ytdlp.invocations", + `download-source-media:${url} format:${arg("-f") ?? ""}\n`, + ); + if (cookieGateBlocked(url)) failCookieGate(url); + if (!existsSync(path.join(videoDir, "metadata.info.json"))) { + await writeMetadata(videoDir, id, urlSentinels(url)); + } + process.stdout.write(`[download] Fetching ${id}\n`); + await writeFile( + path.join(videoDir, "source-media.mp4"), + `fake-ytdlp synthesised source container for ${id}\n`, + ); + process.stdout.write(`[download] ${id} done\n`); + process.stdout.write(`[fake-ytdlp] download complete\n`); + return; + } + process.stderr.write( `[fake-ytdlp] unknown invocation: ${argv.join(" ")}\n`, ); diff --git a/editor/e2e/incomplete-transcript.spec.ts b/editor/e2e/incomplete-transcript.spec.ts @@ -10,6 +10,7 @@ import { mkdir, writeFile } from "node:fs/promises"; import { test, expect } from "@playwright/test"; +import { baseUrl } from "./baseUrl"; import { channelVideos, generateReport, @@ -136,6 +137,97 @@ test("incomplete-transcript filter, glyph, panel banner, and the transcription p ).toBeVisible(); }); +// The fix run re-downloads the full audio and re-transcribes, so the transcript +// is no longer truncated and the banner's own condition is gone by the time the +// run ends. StreamActionLog fires that refresh itself, so a parent that stops +// rendering the banner on it destroys the run's log — the <pre role="log"> +// included. The warning text disappearing is the proof the refresh landed; the +// log is read only after that. See plans/FACTS.md, "A run log lives in the +// panel's React state". +// +// Two things about the setup, both forced: +// +// - Its own channel, not this file's shared `test-transcribe`. This is the +// only case here that runs a JOB, and a finished job arms the debounced +// snapshot regen (snapshotScheduler). Landing late, that regen writes +// channels/<slug>/snapshot.json after the NEXT test's resetData has copied +// its fixture in — and generateReport RETURNS EARLY when a snapshot file +// exists (helpers.ts:146), so that test would then read a snapshot built +// before its own seed and find no flagged video. Measured: with this case +// sharing the slug, "clear incomplete" below failed 1 run in 10 on +// "Select incomplete" never appearing; the baseline without this case is +// 50/50. A separate slug puts the late write somewhere nobody reads. +// - A truncated VTT rather than the whisper transcript the other cases seed: +// transcribeOneVideo returns "already-exists" the moment a transcript.json +// is on disk (common/controller/transcribeOne.ts:107), so with one there the +// fix run re-downloads the audio, writes no new transcript, and nothing +// regenerates transcript.cues.json — the flag would never clear. Not this +// test's subject. +test("the fix run's log survives the refresh that clears the banner", async ({ + page, +}) => { + await resetData(); + const slug = "incomplete-log"; + const id = "vidVttTrunc"; + const root = resolvePath(`test-transcripts/channels/${slug}`); + const dir = `${root}/data/${id}`; + await mkdir(dir, { recursive: true }); + await writeFile( + `${root}/config.json`, + JSON.stringify({ + handling: "transcribe", + name: slug, + platform: "youtube", + url: `https://www.youtube.com/@${slug}`, + audioFormat: "m4a", + }), + ); + // fixIncompleteTranscriptOne resolves the source URL from metadata.info.json + // or the playlist, and there is no metadata here yet. + await writeFile( + `${root}/playlist`, + `https://www.youtube.com/watch?v=${id}\n`, + ); + await writeFile( + `${dir}/transcript.en.vtt`, + "WEBVTT\n\n00:00:00.000 --> 00:06:52.000\nhello\n", + ); + await writeFile(`${dir}/audio.m4a`, "fake-audio"); + await writeFile( + `${dir}/transcript.cues.json`, + JSON.stringify({ + version: 1, + source: "vtt", + duration: 8541, + isLivestream: false, + cues: [{ start: 0, end: 412, text: "hello" }], + }), + ); + await fetch(`${baseUrl}/api/test/invalidate-cache`).catch(() => {}); + + await page.goto(`/channels/${slug}/videos/${id}`); + const banner = page.getByLabel("incomplete transcript"); + await expect(banner).toContainText("Transcript looks truncated"); + const run = banner.getByRole("button", { + name: /Re-download & re-transcribe/, + }); + await expect(run).toBeEnabled(); + await run.click(); + + const log = page.getByLabel(`Re-download & re-transcribe ${id} output`); + await expect(log).toContainText(`Re-downloading audio for ${id}`, { + timeout: 15_000, + }); + + // The state the run produces: a full transcript over the re-downloaded audio, + // so the truncation warning is gone. + await expect(page.getByText("Transcript looks truncated")).toHaveCount(0, { + timeout: 15_000, + }); + // And the log is STILL on screen, carrying the run's closing line. + await expect(log).toContainText(`Normalized ${slug}/${id}`); +}); + test("channel bulk bar: clear incomplete resets the video and enables auto-runners", async ({ page, }) => { diff --git a/editor/e2e/saved-videos.spec.ts b/editor/e2e/saved-videos.spec.ts @@ -114,3 +114,44 @@ test("video page shows persisted source status and an unpersist control", async page.getByLabel("unpersist source video vidA"), ).toBeVisible(); }); + +// The persist run writes saved-video.json and StreamActionLog then calls +// router.refresh() ITSELF, which re-renders this section with the pointer now on +// disk. The section used to early-return the persisted view at that point, +// unmounting the log — and an unmounted StreamActionLog renders no +// <pre role="log"> at all, so the operator's log does not go stale, it +// disappears. The assertion order is the point: read the log only AFTER the +// persisted view is on screen, so the refresh has provably landed. See +// plans/FACTS.md, "A run log lives in the panel's React state". +test("the persist run's log survives the refresh that lands the pointer", async ({ + page, +}) => { + await resetData("one-transcribe-channel-with-audio"); + // redownloadToArchiveAction resolves the source URL from metadata.info.json + // or the playlist, and this fixture ships neither. + await writeFile( + resolvePath(`test-transcripts/channels/${SLUG}/playlist`), + "https://www.youtube.com/watch?v=vidA\n", + ); + await fetch(`${baseUrl}/api/test/invalidate-cache`).catch(() => {}); + + await page.goto(`/channels/${SLUG}/videos/vidA`); + // Nothing is persisted yet, so the card starts closed. + await page.getByLabel("Source video stage summary").click(); + const run = page.getByRole("button", { name: "Persist source video" }); + await expect(run).toBeEnabled(); + await run.click(); + + const log = page.getByLabel("Persist source video for vidA output"); + await expect(log).toContainText("Re-downloading vidA to archive", { + timeout: 15_000, + }); + + // The state the run produces: the pointer is on disk, so the panel draws the + // persisted view. That is the proof the refresh landed. + await expect(page.getByLabel("unpersist source video vidA")).toBeVisible({ + timeout: 15_000, + }); + // And the log is STILL on screen, carrying the run's closing line. + await expect(log).toContainText("Persisted source video to"); +});