Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit d8fc333ddeefd7db919e7eed2c34c233e202ffd6
parent bda3169b920fdaa09f707a0c4d9fcd4716a84f8c
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Thu,  1 Oct 2026 14:41:19 -0400

report-to-video: the crossfade concat pins every clip's sound to its picture's length

xfade places a segment by the picture's length and acrossfade by the sound's;
an encoded segment's audio is routinely a few to ~20 ms off its video, so the
sound drifted ahead clip by clip (0.30 s by the end of a 17-clip cut, 2.0 s on
one with cards). Each input's sound is now apad/atrim'd to exactly its length
in the cut before the crossfade. Every crossfaded graph changes; a real-ffmpeg
test (av-sync.test.mjs) puts the third clip's tone within 10 ms of its picture,
and shows the old graph ~60 ms early on three clips.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
Meditor/CHANGELOG.md | 1+
Aumtool/report-to-video/av-sync.test.mjs | 87+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mumtool/report-to-video/build-video.mjs | 16++++++++++++++--
Mumtool/report-to-video/cut-edits.test.mjs | 6++++--
Mumtool/report-to-video/deck-overlay.test.mjs | 4+++-
Mumtool/report-to-video/deck-room.test.mjs | 26++++++++++++++++++--------
6 files changed, 127 insertions(+), 13 deletions(-)

diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,6 +1,7 @@ # Changelog ## [Unreleased] +- **umtool's report videos keep every clip's sound on its picture.** In a crossfaded cut each clip's audio was placed by the audio's own length and its picture by the picture's, and an encoded clip's audio is routinely a few to twenty milliseconds shorter or longer than its video, so the sound drifted further ahead clip by clip: by the end of a seventeen-clip cut it was a third of a second early, and two seconds on one with title and sources cards. Each clip's sound is now padded or trimmed to exactly its picture's length before the crossfade. Every crossfaded report video changes when it is rebuilt, and is in sync; a hard-cut video was not affected. - **umtool's report videos can wear an on-screen deck: one panel under the footage for the whole cut, with a pip timeline, a title per clip, its source and date, and its QR.** A report manifest whose `render` says `"chrome": { "engine": "hyperframes", "layout": "deck" }` scales the footage into a box above a 190 px panel (both sizes are settings) and draws, over the whole cut, one unlabelled pip per clip on a track that fills as the cut plays, the clip's own title from `onscreen.title`, a subtitle naming the recording and its date (the channel too when the cut spans more than one; `onscreen.subtitle` replaces it), and the clip's QR. At each clip change the marker travels to the next pip and the title, subtitle and QR hand over; over a card the panel slides away and comes back after. The citation header, the corner QR and the section footer are not drawn on such a cut, and chapters take the clip's on-screen title. Every setting (sizes, spacing, date format, what the subtitle names, whether cards keep the panel, the motion's timings) is in `render.chrome.deck` and checked when it is saved; an unknown or out-of-range one is refused with a sentence saying why. The panel is rendered once per cut by HyperFrames (pinned to 0.8.24; `HYPERFRAMES_PKG` or `HYPERFRAMES_BIN` override it) and reused until its text or settings change. `build-video.mjs --chrome-only` redraws it over the built segments without rebuilding or fetching anything, `--no-chrome` builds the framed cut without it, and `--chrome-preview <at> <dur>` renders a short window. In umtool, the report page has an **On-screen** section — a switch, the settings, a table of every entry's title and subtitle with the automatic subtitle as its placeholder and a character counter, a live preview with a scrubber, a true still, **Re-render on-screen** and the built video — and the clip bench has on-screen title and subtitle fields with the panel previewed over the clip. A manifest without `render.chrome` builds exactly as before, byte for byte. - **A report cut that wears the on-screen deck can show posts — Bluesky or X statements — as cards over the footage.** A report manifest's `posts` list (each with its platform, handle, date, words and link) is drawn near the end of the clip each post belongs with: the clip whose recording most closely precedes it by date, unless the post names one with `attachTo`; `hide` leaves one out. A clip's posts appear four seconds apart and stack down a column at the frame's top right; as the first appears, the footage eases aside (to 86 % of its box, at the far side) to make room, and the clip's last frame is held, in silence, for 2.5 seconds so the last post can be read; then they all leave together in the change to the next clip, which comes in at the normal size. When the column is full the oldest slide up and out. Each card slides in from the edge of the frame and flares in the deck's accent as it lands; it has an accent rail down its edge and shows the post's date, a platform label ("Bluesky" or "X") beside `@handle`, its words in paragraphs up to seven lines with an ellipsis, and a QR of the post's link, in the deck's colours and faces. The hold and the move are made where the cut is joined, not in a clip, so `--chrome-only` changes them without rebuilding one; chapters and the deck's timing count the hold. The timing, the hold (`hold`, 0 turns it off), the move (`shift`: its scale and seconds, or `false`), the column's side, width and inset, the QR size and the line limit are settings under `render.chrome.deck.posts`, and a bad post or setting is refused with a sentence before a build fetches anything. Only the seconds the cards are up are rendered, one short sequence per clip, cached like the deck; `--chrome-only`, `--chrome-preview` and a hard-cut cut lay them as they lay the deck, and `--no-chrome` draws neither — though it still holds and moves the footage, which are part of the cut rather than the chrome. A first post that appears inside the hold still moves the footage, and a hold is a whole number of frames. A cut without `posts` builds exactly as before, and without the deck `posts` is not drawn at all. - **A report clip can go silent partway through, a report cut can fade out at its end, and the deck's QR names its site in larger type.** A clip's `muteFrom` (in the recording's own seconds, inside the clip) silences it from that second to its end while the picture plays on, after a 40 ms fade that ends there, so nothing clicks and no next word leaks in; a hold on that clip stays silent. `render.endFade` (seconds; 0, the default, is off) fades the cut's last segment to the background colour and to silence over its final seconds, the hold included, and the deck stays drawn over it; once a closing card follows the last clip, the ordinary crossfade into it does that job instead. Both are applied where the cut is joined, so `--chrome-only` changes them without rebuilding a clip, and a value out of range is refused with a sentence before a build fetches anything. Each clip build now writes `<id>.cut.json` beside its segment, saying where in the recording the segment really starts after its cut was snapped to a silence; `muteFrom` is measured from it, and a segment built before this measures from the clip's unsnapped start and says so. The site's name beside the deck's QR is now exactly as long as the code is tall, for any site. A cut without `muteFrom` or `endFade` builds exactly as before. diff --git a/umtool/report-to-video/av-sync.test.mjs b/umtool/report-to-video/av-sync.test.mjs @@ -0,0 +1,87 @@ +// The crossfade concat keeps every clip's sound on its picture. +// +// An encoded segment's audio is routinely a few to ~20 ms off its video, and +// `acrossfade` joins by the SOUND's length while `xfade` offsets come from the +// PICTURE's. Unpinned, the difference accumulates clip by clip. These segments +// make it large on purpose -- each one's audio is 30 ms short -- and put a tone +// exactly 1.0 s into each: in the joined cut, the third tone must still start +// 1.0 s after the third segment does. +// +// Run with: pnpm test:scripts +import assert from "node:assert/strict"; +import { spawnSync } from "node:child_process"; +import { mkdtempSync, rmSync } from "node:fs"; +import { tmpdir } from "node:os"; +import path from "node:path"; +import test from "node:test"; + +import { xfadeGraph } from "./build-video.mjs"; + +const have = spawnSync("ffmpeg", ["-version"]).status === 0; +const RENDER = { width: 320, height: 180, fps: 30, audioRate: 48000, audioChannels: 2 }; + +function segment(dir, name, seconds) { + const out = path.join(dir, `${name}.mp4`); + const a = (seconds - 0.03).toFixed(3); + const r = spawnSync("ffmpeg", [ + "-v", "error", "-y", + "-f", "lavfi", "-i", `color=c=gray:s=320x180:r=30:d=${seconds}`, + // Silence, then a tone from exactly 1.0 s; the stream is 30 ms short. + "-f", "lavfi", "-i", `sine=f=1000:r=48000:d=${a}`, + "-filter_complex", `[1:a]volume=enable='lt(t,1)':volume=0,aformat=channel_layouts=stereo[a]`, + "-map", "0:v", "-map", "[a]", "-c:v", "libx264", "-pix_fmt", "yuv420p", "-c:a", "aac", "-shortest", + out, + ]); + assert.equal(r.status, 0, String(r.stderr)); + return out; +} + +/** The time (s) the tone's third onset starts in a file: the first loud window after `after`. */ +function onsetAfter(file, after) { + const r = spawnSync("ffmpeg", ["-v", "error", "-i", file, "-ac", "1", "-ar", "8000", "-f", "s16le", "-"], { maxBuffer: 1 << 26 }); + const pcm = new Int16Array(r.stdout.buffer, r.stdout.byteOffset, r.stdout.length / 2); + const w = 40; // 5 ms windows + for (let i = Math.floor(after * 8000); i + w < pcm.length; i += w) { + let sum = 0; + for (let k = i; k < i + w; k += 1) sum += pcm[k] * pcm[k]; + if (Math.sqrt(sum / w) > 1000) return i / 8000; + } + return null; +} + +test("the crossfade concat keeps every clip's sound on its picture", { skip: !have && "no ffmpeg" }, () => { + const dir = mkdtempSync(path.join(tmpdir(), "av-sync-")); + try { + const durs = [4, 4, 4]; + const segs = durs.map((d, i) => segment(dir, `s${i}`, d)); + const D = 0.5; + const run = (graph, out) => { + const r = spawnSync("ffmpeg", [ + "-v", "error", "-y", ...segs.flatMap((s) => ["-i", s]), + "-filter_complex", graph.parts.join(";"), + "-map", graph.vlab, "-map", graph.alab, "-c:v", "libx264", "-pix_fmt", "yuv420p", "-c:a", "aac", out, + ]); + assert.equal(r.status, 0, String(r.stderr)); + return out; + }; + // The third segment starts at 2 × (4 − 0.5) = 7.0 s; its tone at 8.0 s. + const pinned = onsetAfter(run(xfadeGraph(durs, D, null, RENDER), path.join(dir, "pinned.mp4")), 7.6); + assert.ok(Math.abs(pinned - 8.0) <= 0.01, `pinned onset ${pinned}, expected 8.0`); + + // The unpinned graph (what the concat wrote before) lands each later clip + // 30 ms earlier per clip before it: the drift this guards against. + const unpinned = { + parts: [ + `[0:v][1:v]xfade=transition=fade:duration=${D}:offset=3.500[v1]`, + `[0:a][1:a]acrossfade=d=${D}:c1=tri:c2=tri[a1]`, + `[v1][2:v]xfade=transition=fade:duration=${D}:offset=7.000[v2]`, + `[a1][2:a]acrossfade=d=${D}:c1=tri:c2=tri[a2]`, + ], + vlab: "[v2]", alab: "[a2]", + }; + const drift = onsetAfter(run(unpinned, path.join(dir, "unpinned.mp4")), 7.6); + assert.ok(8.0 - drift >= 0.045, `unpinned onset ${drift} should be ~60 ms early`); + } finally { + rmSync(dir, { recursive: true, force: true }); + } +}); diff --git a/umtool/report-to-video/build-video.mjs b/umtool/report-to-video/build-video.mjs @@ -2367,10 +2367,22 @@ export async function cutOffsets(segments, D, fps, joins = null) { */ export function xfadeGraph(durs, D, joins = null, render = null) { const parts = []; - const ins = durs.map((_, i) => { + const ins = durs.map((d, i) => { const c = joinInputChain(i, joins?.[i] ?? null, render); parts.push(...c.parts); - return c; + // The sound is pinned to the picture's length before it is crossfaded. + // `xfade` places segment i+1 by the PICTURE's length, `acrossfade` by the + // SOUND's, and an encoded segment's audio is routinely a few to ~20 ms off + // its video (AAC frames do not end on video frames). Unpinned, the error + // accumulates segment by segment: measured 0.30 s early by the last clip of + // a 17-clip cut, 2.0 s on one with title and sources cards. Padded with + // silence and trimmed to exactly `d`, every clip's sound starts with its + // picture. `d` is the segment's length in the cut (frames ÷ fps, hold + // included), the same number the xfade offsets are summed from. + const len = d.toFixed(6); + const a = `[p${i}a]`; + parts.push(`${c.a}apad=whole_dur=${len},atrim=end=${len},asetpts=PTS-STARTPTS${a}`); + return { v: c.v, a }; }); let vlab = ins[0].v; let alab = ins[0].a; diff --git a/umtool/report-to-video/cut-edits.test.mjs b/umtool/report-to-video/cut-edits.test.mjs @@ -147,10 +147,12 @@ test("withCutEdits: nothing to add leaves the joins as they were (null stays nul }); test("the graphs: unchanged without edits; the hard cut's record names a mute and a fade, so a changed one is never reused", () => { - // No edits: the crossfade graph the build always wrote. + // No edits: the plain crossfade graph (each sound pinned to its picture). assert.deepEqual(xfadeGraph([10, 12], 0.5, withCutEdits(null, 2), RENDER).parts, [ + "[0:a]apad=whole_dur=10.000000,atrim=end=10.000000,asetpts=PTS-STARTPTS[p0a]", + "[1:a]apad=whole_dur=12.000000,atrim=end=12.000000,asetpts=PTS-STARTPTS[p1a]", "[0:v][1:v]xfade=transition=fade:duration=0.5:offset=9.500[v1]", - "[0:a][1:a]acrossfade=d=0.5:c1=tri:c2=tri[a1]", + "[p0a][p1a]acrossfade=d=0.5:c1=tri:c2=tri[a1]", ]); const segs = ["/s/a.mp4", "/s/b.mp4"]; const holdOnly = [null, { hold: 2.5, move: null }]; diff --git a/umtool/report-to-video/deck-overlay.test.mjs b/umtool/report-to-video/deck-overlay.test.mjs @@ -86,8 +86,10 @@ test("previewFromSegmentsArgs: the window's segments crossfaded as the full conc assert.equal( args[args.indexOf("-filter_complex") + 1], [ + "[0:a]apad=whole_dur=10.000000,atrim=end=10.000000,asetpts=PTS-STARTPTS[p0a]", + "[1:a]apad=whole_dur=10.000000,atrim=end=10.000000,asetpts=PTS-STARTPTS[p1a]", "[0:v][1:v]xfade=transition=fade:duration=0.5:offset=9.500[v1]", - "[0:a][1:a]acrossfade=d=0.5:c1=tri:c2=tri[a1]", + "[p0a][p1a]acrossfade=d=0.5:c1=tri:c2=tri[a1]", "[v1]trim=start=2.500:duration=10.000,setpts=PTS-STARTPTS[vw]", "[a1]atrim=start=2.500:duration=10.000,asetpts=PTS-STARTPTS[aw]", "[2:v]format=rgba[hfa0];[vw][hfa0]overlay=x=0:y=890:format=yuv444:shortest=1[hf0];[hf0]format=yuv420p[vout]", diff --git a/umtool/report-to-video/deck-room.test.mjs b/umtool/report-to-video/deck-room.test.mjs @@ -183,10 +183,14 @@ function near(a, b, msg) { assert.ok(Math.abs(a - b) < 0.01, `${msg}: ${a} != ${ test("xfadeGraph: without joins, the graph the crossfade concat always wrote", () => { const g = xfadeGraph([10, 12, 4], 0.5, null, RENDER); assert.deepEqual(g.parts, [ + // Each input's sound pinned to its picture's length before the crossfade. + "[0:a]apad=whole_dur=10.000000,atrim=end=10.000000,asetpts=PTS-STARTPTS[p0a]", + "[1:a]apad=whole_dur=12.000000,atrim=end=12.000000,asetpts=PTS-STARTPTS[p1a]", + "[2:a]apad=whole_dur=4.000000,atrim=end=4.000000,asetpts=PTS-STARTPTS[p2a]", "[0:v][1:v]xfade=transition=fade:duration=0.5:offset=9.500[v1]", - "[0:a][1:a]acrossfade=d=0.5:c1=tri:c2=tri[a1]", + "[p0a][p1a]acrossfade=d=0.5:c1=tri:c2=tri[a1]", "[v1][2:v]xfade=transition=fade:duration=0.5:offset=21.000[v2]", - "[a1][2:a]acrossfade=d=0.5:c1=tri:c2=tri[a2]", + "[a1][p2a]acrossfade=d=0.5:c1=tri:c2=tri[a2]", ]); assert.equal(g.vlab, "[v2]"); assert.equal(g.alab, "[a2]"); @@ -197,13 +201,18 @@ test("xfadeGraph: without joins, the graph the crossfade concat always wrote", ( test("xfadeGraph: a held input joins through its chain, and the offsets after it move by the hold", () => { const joins = [null, { hold: 2.5, move: MOVE(6) }, null]; const g = xfadeGraph([10, 14.5, 4], 0.5, joins, RENDER); - assert.deepEqual(g.parts.slice(2), [ + const chain = joinInputChain(1, joins[1], RENDER).parts; + assert.deepEqual(g.parts, [ + "[0:a]apad=whole_dur=10.000000,atrim=end=10.000000,asetpts=PTS-STARTPTS[p0a]", + ...chain, + // The held input's sound is pinned to its length WITH the hold. + "[j1a]apad=whole_dur=14.500000,atrim=end=14.500000,asetpts=PTS-STARTPTS[p1a]", + "[2:a]apad=whole_dur=4.000000,atrim=end=4.000000,asetpts=PTS-STARTPTS[p2a]", "[0:v][j1v]xfade=transition=fade:duration=0.5:offset=9.500[v1]", - "[0:a][j1a]acrossfade=d=0.5:c1=tri:c2=tri[a1]", + "[p0a][p1a]acrossfade=d=0.5:c1=tri:c2=tri[a1]", "[v1][2:v]xfade=transition=fade:duration=0.5:offset=23.500[v2]", - "[a1][2:a]acrossfade=d=0.5:c1=tri:c2=tri[a2]", + "[a1][p2a]acrossfade=d=0.5:c1=tri:c2=tri[a2]", ]); - assert.deepEqual(g.parts.slice(0, 2), joinInputChain(1, joins[1], RENDER).parts); }); test("hardCutFilterArgs: the concat filter over each input's chain, one encode", () => { @@ -264,12 +273,13 @@ test("a preview window over a held clip: picked by the cut's lengths, its inputs render: RENDER, chromePlan: plan, outPath: "/p.mp4", joins, }); assert.match(one[one.indexOf("-filter_complex") + 1], /\[j0v\]trim=start=12\.000:duration=1\.500,setpts=PTS-STARTPTS\[vw\];\[j0a\]atrim=/); - // Without joins, the window's graph is unchanged. const plain = previewFromSegmentsArgs({ segments: ["/s/a.mp4", "/s/b.mp4"], durs: [10, 10], starts: [0, 9.5], D: 0.5, at: 8, dur: 4, render: RENDER, chromePlan: plan, outPath: "/p.mp4", }); - assert.match(plain[plain.indexOf("-filter_complex") + 1], /^\[0:v\]\[1:v\]xfade=transition=fade:duration=0\.5:offset=9\.500\[v1\];\[0:a\]\[1:a\]acrossfade/); + // Without joins, the window's graph is the crossfade concat's: each sound + // pinned to its picture, then the fades. + assert.match(plain[plain.indexOf("-filter_complex") + 1], /^\[0:a\]apad=whole_dur=10\.000000,atrim=end=10\.000000,asetpts=PTS-STARTPTS\[p0a\];\[1:a\]apad=whole_dur=10\.000000[^;]*\[p1a\];\[0:v\]\[1:v\]xfade=transition=fade:duration=0\.5:offset=9\.500\[v1\];\[p0a\]\[p1a\]acrossfade/); }); // ---- ffmpeg, for real -------------------------------------------------------