Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 9fcb5b8dd6e43c0832e42cb64f6ee7cf362ac5db
parent 08651344cc988856cc130bfa0b8f3f87ec2b70cf
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Tue, 18 Aug 2026 10:57:01 -0400

facecrop.py --detect/--crop, and make-thumb honours the judgements

Both python modes are ADDITIVE: the positional call is documented in the
thumbnails spec sheet, so it is parsed exactly as it was once the flags are
pulled out, and its printed line keeps its first five fields byte for byte.
Verified on the Super Mario RPG corner -- `830 68 449 0.933 113.0` before and
after.

  --detect          detection only, as JSON: the face box, the auto crop and
                    the frame size. Writes nothing. ~0.94s.
  --crop x,y,w,h    no detection at all, cut exactly that box at exactly that
                    frame. ~0.27s. Reports a confidence of -1.000, because the
                    crop was CHOSEN and printing 1.000 would claim a certainty
                    no detector produced.

The crop maths moves into auto_crop() so --detect reports the same box the
normal path cuts rather than a second copy of the arithmetic.

make-thumb.mjs then:

  * skips rejected videos, one clause beside the existing usedEver/EXCLUDE test
    -- which is already exactly this shape
  * passes an approved crop straight through to --crop, so a framing chosen to
    exclude a strip of YouTube chrome is not silently re-derived by the very
    clamp that put the chrome there
  * RECORDS the box in corners[].crop

That last one closes the hole for every future generation. The geometry was
computed and thrown away -- the script kept only the score and the time -- so a
corner could be cut and never cut the same way again, which is the state two
accepted corners are in today: no face is found at all at their recorded
frameAt. `crop` is optional and its absence means "never recorded", the same
way AsrWord.conf's absence means unknown. REUSE=1 now honours a recorded box
too, so a rebuild reproduces rather than re-detects.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Diffstat:
Mumtool/song/facecrop.py | 174++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-----------
Mumtool/song/make-thumb.mjs | 82+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++----------
2 files changed, 223 insertions(+), 33 deletions(-)

diff --git a/umtool/song/facecrop.py b/umtool/song/facecrop.py @@ -7,9 +7,26 @@ returns a screenshot of a news article, which is what the first thumbnail did. So the face has to be located, not assumed. facecrop.py <video> <time> <out.jpg> <size> [window] [step] [aspect] + facecrop.py --detect <video> <time> [window] [step] [aspect] + facecrop.py --crop x,y,w,h <video> <time> <out.jpg> <size> [aspect] Samples a few frames around <time>, runs YuNet on each, and keeps the largest, -most confident face. Prints "x y w h score t" or "NONE". +most confident face. Prints "x0 y0 side score t cwid ch" or "NONE" -- the first +five fields are the original contract and must not move; cwid/ch were appended +so a caller can RECORD the box that was cut instead of inferring it. + +--detect runs the detection and prints the face box, the auto crop and the frame +size as JSON. It WRITES NOTHING. This is what the face judger asks, so a person +can see where the detector looked before deciding whether the crop it implies is +the picture they want. + +--crop x,y,w,h skips detection entirely and cuts exactly that box at exactly +<time>. This is how a human's framing survives into a rebuild: the clamp below +preserves the crop's SIZE at the cost of moving it OFF the face, which is how a +corner ends up showing YouTube chrome, and re-deriving a framing that was chosen +to avoid that would silently undo the choice. A hand-framed crop reports a +confidence of -1.000 -- it was chosen, not detected, and reporting 1.000 would +be claiming a certainty no detector produced. <aspect> is width/height, default 1.0 (square). Pass 1.7778 for 16:9. The output is <size> wide and <size>/<aspect> tall, so `480 ... 1.7778` fills a 480x270 @@ -17,7 +34,7 @@ secondary box in the shorts layout exactly. The crop is taken about the same fac centre either way -- only the box around it changes shape -- so the two options are a fair comparison rather than two different framings. """ -import sys, os, subprocess, tempfile +import sys, os, json, subprocess, tempfile import cv2 import numpy as np @@ -40,12 +57,34 @@ def grab(video, t, path): check=False, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL) return os.path.exists(path) and os.path.getsize(path) > 0 -def main(): - video, t0, out, size = sys.argv[1], float(sys.argv[2]), sys.argv[3], int(sys.argv[4]) - window = float(sys.argv[5]) if len(sys.argv) > 5 else 0.8 - step = float(sys.argv[6]) if len(sys.argv) > 6 else 0.4 - aspect = float(sys.argv[7]) if len(sys.argv) > 7 else 1.0 +def auto_crop(x, y, fw, fh, w, h, aspect): + """The crop facecrop.py takes around a face box, and its geometry alone. + Lifted out of main() unchanged so --detect can report the SAME box the + normal path cuts without running the cut. lib/face-types.ts re-implements + exactly this, and the detect route checks the two agree on every call. + """ + cx, cy = x + fw / 2, y + fh / 2 + side = max(fw, fh) * 2.2 + side = min(side, min(w, h)) + # The face box sets the HEIGHT; the width follows the aspect. Driving it off the + # height keeps the head the same size in both options, so a 16:9 crop is a square + # one with more room either side rather than the same picture zoomed out. + ch = side + cwid = side * aspect + if cwid > w: # too wide for the frame: give up width, not head + cwid = w + ch = min(ch, cwid / aspect) + # THE CLAMP, and the reason this tool exists. It keeps the crop inside the + # frame by SLIDING it, so a webcam inset near an edge gets a crop that stays + # the right size and stops being centred on the face. + x0 = int(max(0, min(w - cwid, cx - cwid / 2))) + y0 = int(max(0, min(h - ch, cy - ch / 2))) + return x0, y0, int(cwid), int(ch), side + + +def detect_best(video, t0, window, step): + """The largest, most confident face near t0, with the frame it was found on.""" det = cv2.FaceDetectorYN.create(MODEL, "", (320, 320), 0.7, 0.3, 5000) best = None with tempfile.TemporaryDirectory() as td: @@ -72,29 +111,118 @@ def main(): # is gone by the time the crop is taken best = (q, float(x), float(y), float(fw), float(fh), float(score), t, img.copy(), w, h) + return best + + +def frame_size(video, t): + """The frame size, decoded rather than probed -- one code path for both.""" + with tempfile.TemporaryDirectory() as td: + p = os.path.join(td, "f.jpg") + if not grab(video, t, p): + return None, None + img = cv2.imread(p) + if img is None: + return None, None + return img, img.shape[:2] + + +def mode_detect(video, t0, window, step, aspect): + """Report what the detector found. Writes nothing.""" + best = detect_best(video, t0, window, step) + if best is None: + _img, size = frame_size(video, t0) + h, w = size if size else (0, 0) + print(json.dumps({"face": None, "auto": None, "frame": {"w": int(w), "h": int(h)}})) + return 1 + _, x, y, fw, fh, score, t, img, w, h = best + x0, y0, cwid, ch, _side = auto_crop(x, y, fw, fh, w, h, aspect) + print(json.dumps({ + "face": {"x": x, "y": y, "w": fw, "h": fh, "score": round(float(score), 4), "t": t}, + "auto": {"x": x0, "y": y0, "w": cwid, "h": ch}, + "frame": {"w": int(w), "h": int(h)}, + })) + return 0 + + +def mode_crop(spec, video, t0, out, size, aspect): + """Cut exactly the box a person chose, at exactly the frame they chose it on.""" + try: + x0, y0, cwid, ch = [int(round(float(v))) for v in spec.split(",")] + except ValueError: + print("NONE") + return 1 + img, dims = frame_size(video, t0) + if img is None: + print("NONE") + return 1 + h, w = dims + # Clamped, not refused: the caller validated this box against the frame, and + # a numpy slice that runs off the edge returns a smaller picture in silence + # rather than an error. Better to cut what was asked for as closely as the + # frame allows and print what was actually cut. + cwid = max(1, min(w, cwid)) + ch = max(1, min(h, ch)) + x0 = max(0, min(w - cwid, x0)) + y0 = max(0, min(h - ch, y0)) + crop = img[y0:y0 + ch, x0:x0 + cwid] + crop = cv2.resize(crop, (size, max(1, int(round(size / aspect)))), interpolation=cv2.INTER_AREA) + cv2.imwrite(out, crop, [cv2.IMWRITE_JPEG_QUALITY, 95]) + # -1.000 where a detector confidence would go: this crop was CHOSEN. + print(f"{x0} {y0} {ch} -1.000 {t0} {cwid} {ch}") + return 0 + + +def main(): + # Both flags are ADDITIVE. The positional call is documented in the + # thumbnails spec sheet and used by make-thumb.mjs, so it is parsed exactly + # as it always was once the flags are pulled out. + argv, detect_only, crop_spec = [], False, None + i = 1 + while i < len(sys.argv): + a = sys.argv[i] + if a == "--detect": + detect_only = True + i += 1 + elif a == "--crop": + crop_spec = sys.argv[i + 1] + i += 2 + else: + argv.append(a) + i += 1 + + if detect_only: + video, t0 = argv[0], float(argv[1]) + window = float(argv[2]) if len(argv) > 2 else 0.8 + step = float(argv[3]) if len(argv) > 3 else 0.4 + aspect = float(argv[4]) if len(argv) > 4 else 1.0 + return mode_detect(video, t0, window, step, aspect) + + video, t0, out, size = argv[0], float(argv[1]), argv[2], int(argv[3]) + + if crop_spec is not None: + aspect = float(argv[4]) if len(argv) > 4 else 1.0 + return mode_crop(crop_spec, video, t0, out, size, aspect) + + window = float(argv[4]) if len(argv) > 4 else 0.8 + step = float(argv[5]) if len(argv) > 5 else 0.4 + aspect = float(argv[6]) if len(argv) > 6 else 1.0 + + best = detect_best(video, t0, window, step) if best is None: print("NONE") return 1 _, x, y, fw, fh, score, t, img, w, h = best # crop centred on the face, padded out so it is a portrait not a close-up of a - # nose; clamped to stay inside the frame - cx, cy = x + fw / 2, y + fh / 2 - side = max(fw, fh) * 2.2 - side = min(side, min(w, h)) - # The face box sets the HEIGHT; the width follows the aspect. Driving it off the - # height keeps the head the same size in both options, so a 16:9 crop is a square - # one with more room either side rather than the same picture zoomed out. - ch = side - cwid = side * aspect - if cwid > w: # too wide for the frame: give up width, not head - cwid = w - ch = min(ch, cwid / aspect) - x0 = int(max(0, min(w - cwid, cx - cwid / 2))) - y0 = int(max(0, min(h - ch, cy - ch / 2))) - crop = img[y0:y0 + int(ch), x0:x0 + int(cwid)] + # nose; clamped to stay inside the frame -- and the clamp is what the face + # judger exists to make visible, so the maths is auto_crop()'s alone now. + x0, y0, cwid, ch, side = auto_crop(x, y, fw, fh, w, h, aspect) + crop = img[y0:y0 + ch, x0:x0 + cwid] crop = cv2.resize(crop, (size, max(1, int(round(size / aspect)))), interpolation=cv2.INTER_AREA) cv2.imwrite(out, crop, [cv2.IMWRITE_JPEG_QUALITY, 95]) - print(f"{x0} {y0} {int(side)} {score:.3f} {t}") + # The first five fields are the original contract. cwid/ch are appended so + # make-thumb.mjs can RECORD the box rather than infer it from `side`, which + # is only the same number while the aspect is square. + print(f"{x0} {y0} {int(side)} {score:.3f} {t} {cwid} {ch}") return 0 if __name__ == "__main__": diff --git a/umtool/song/make-thumb.mjs b/umtool/song/make-thumb.mjs @@ -22,6 +22,7 @@ const CW = Number(process.env.CORNER_W ?? 300), CH = Number(process.env.CORNER_H const M = Number(process.env.MARGIN ?? 26); const MF = path.join(DIR, "thumb-manifest.json"); // every generation, including tests const AF = path.join(DIR, "thumb-accepted.json"); // ONLY what was accepted -- the authority +const VF = path.join(DIR, "face-verdicts.json"); // what a person said about the FACES let manifest = { version: 1, thumbs: {}, used: [] }; try { manifest = JSON.parse(readFileSync(MF, "utf8")); } catch {} @@ -32,6 +33,27 @@ const usedEver = new Set(accepted.used ?? []); if (usedEver.size) console.log(` ${usedEver.size} source videos spoken for by accepted thumbnails`); const EXCLUDE = new Set((process.env.EXCLUDE ?? "").split(",").map((x) => x.trim()).filter(Boolean)); +// ---- what the face judger decided ------------------------------------------ +// /browse/faces writes this file and NOTHING ELSE writes it -- the same split +// thumb-manifest.json has, with the sides swapped. There the CLI is the sole +// writer and the app reads; here the app is the sole writer and the CLI reads. +// Either way one program owns the file and the other honours it, so there is +// never a question of whose copy is right. +// +// Two things come out of it. A REJECTED video is skipped entirely -- somebody +// looked at that face and said it is not usable, and re-picking it every build +// would make the judging pointless. An approved CROP is passed to facecrop.py +// verbatim, because a framing chosen to exclude a strip of YouTube chrome must +// not be silently re-derived by the very clamp that put the chrome there. +let verdicts = { version: 1, faces: {} }; +try { verdicts = JSON.parse(readFileSync(VF, "utf8")); } catch {} +const faces = verdicts.faces ?? {}; +const keyOf = (video, srcStart) => `${video}@${Number(srcStart).toFixed(2)}`; +const REJECTED = new Set( + Object.values(faces).filter((j) => j.verdict === "reject").map((j) => j.video), +); +if (REJECTED.size) console.log(` ${REJECTED.size} source videos rejected by the face judger`); + // ---- which clips play in this song ----------------------------------------- const rows = readFileSync(CSV, "utf8").trim().split("\n").slice(1) .map((l) => l.split(",")) @@ -43,7 +65,7 @@ if (!rows.length) throw new Error("no clips in " + CSV); // mid-syllable is less likely to catch a blink or a mouth mid-consonant. const byVideo = new Map(); for (const r of rows) { - if (usedEver.has(r.video) || EXCLUDE.has(r.video)) continue; + if (usedEver.has(r.video) || EXCLUDE.has(r.video) || REJECTED.has(r.video)) continue; const dur = r.srcEnd - r.srcStart; const cur = byVideo.get(r.video); if (!cur || dur > cur.dur) byVideo.set(r.video, { ...r, dur }); @@ -110,16 +132,43 @@ const actionScore = (file, sr) => { // middle of the frame is a browser window. A centre crop returns a screenshot // of a news article -- which is exactly what the first attempt produced. const PY = path.join(SONG_DATA, "..", "facedet", "bin", "python"); -const grabFace = (c, i) => { + +// A human framing, if one was recorded for this exact corner. `fallback` is the +// box the manifest recorded for a REUSE rebuild -- also a crop worth honouring, +// because re-detecting at the same frameAt is exactly what turned out not to be +// reproducible: two accepted corners find no face at all at their recorded time. +const cropFor = (c, fallback) => { + const j = faces[keyOf(c.video, c.srcStart)]; + if (j && j.verdict !== "reject" && j.crop) return { box: j.crop, at: j.frameAt ?? c.srcStart, human: true }; + if (fallback) return { box: fallback, at: c.frameAt ?? c.srcStart, human: false }; + return null; +}; + +const grabFace = (c, i, fallback = null) => { const src = path.join(SONG_DATA, "media", `${c.video}.mp4`); if (!existsSync(src)) return null; const f = path.join(tmp, `c${i}.jpg`); + const chosen = cropFor(c, fallback); + const args = chosen + ? ["--crop", `${chosen.box.x},${chosen.box.y},${chosen.box.w},${chosen.box.h}`, + src, String(chosen.at), f, String(CW)] + : [src, String(c.srcStart), f, String(CW)]; try { - const out = execFileSync(PY, [path.join(DIR, "facecrop.py"), src, - String(c.srcStart), f, String(CW)], { encoding: "utf8", stdio: ["ignore", "pipe", "ignore"] }).trim(); + const out = execFileSync(PY, [path.join(DIR, "facecrop.py"), ...args], + { encoding: "utf8", stdio: ["ignore", "pipe", "ignore"] }).trim(); if (out === "NONE" || !existsSync(f)) return null; const p = out.split(/\s+/); - return { f, s: Number(p[3]), t: Number(p[4]) }; + // cwid/ch when facecrop.py reports them, `side` when it does not -- an older + // copy of the script still round-trips, it just records a square box. + const w = p[5] !== undefined ? Number(p[5]) : Number(p[2]); + const h = p[6] !== undefined ? Number(p[6]) : Number(p[2]); + return { + f, + s: Number(p[3]), + t: Number(p[4]), + crop: { x: Number(p[0]), y: Number(p[1]), w, h }, + human: Boolean(chosen && chosen.human), + }; } catch { return null; } }; @@ -131,8 +180,8 @@ const picked = []; if (process.env.REUSE === "1" && manifest.thumbs[NAME]) { const prev = manifest.thumbs[NAME].corners ?? []; prev.forEach((p, i) => { - const g = grabFace({ video: p.video, srcStart: p.srcStart }, i); - if (g) picked.push({ ...p, file: g.f, frameAt: g.t, sharp: g.s }); + const g = grabFace({ video: p.video, srcStart: p.srcStart, frameAt: p.frameAt }, i, p.crop ?? null); + if (g) picked.push({ ...p, file: g.f, frameAt: g.t, sharp: g.s, crop: g.crop, human: g.human }); }); console.log(` reusing ${picked.length} corners from the manifest`); } @@ -140,8 +189,12 @@ for (const c of picked.length >= 4 ? [] : cands) { if (picked.length >= 4) break; const g = grabFace(c, picked.length); if (!g) continue; - picked.push({ ...c, file: g.f, frameAt: g.t, sharp: g.s }); - console.log(` corner ${picked.length}: ${c.video}@${c.srcStart.toFixed(2)} (face at ${g.t.toFixed(2)}s, confidence ${g.s.toFixed(3)})`); + picked.push({ ...c, file: g.f, frameAt: g.t, sharp: g.s, crop: g.crop, human: g.human }); + // A hand-framed crop has no detector confidence -- facecrop.py reports -1 for + // it precisely so this line can say what actually happened rather than print + // a number nobody measured. + const how = g.human ? "framed by hand" : `confidence ${g.s.toFixed(3)}`; + console.log(` corner ${picked.length}: ${c.video}@${c.srcStart.toFixed(2)} (face at ${g.t.toFixed(2)}s, ${how})`); } if (picked.length < 4) throw new Error(`only found ${picked.length} usable faces`); @@ -192,7 +245,16 @@ rmSync(tmp, { recursive: true, force: true }); manifest.thumbs[NAME] = { out: OUT, bgAt: +bgAt.toFixed(2), - corners: picked.map((p) => ({ video: p.video, srcStart: +p.srcStart.toFixed(2), frameAt: +p.frameAt.toFixed(2) })), + // THE BOX IS RECORDED NOW. It used to be computed and thrown away -- the + // script kept only p[3] and p[4] -- so a corner could be cut and never cut + // again the same way, which is the state two accepted corners are in today. + // `crop` is additive: entries without it stay valid and mean "never recorded". + corners: picked.map((p) => ({ + video: p.video, + srcStart: +p.srcStart.toFixed(2), + frameAt: +p.frameAt.toFixed(2), + ...(p.crop ? { crop: p.crop } : {}), + })), }; manifest.used = [...new Set([...(manifest.used ?? []), ...picked.map((p) => p.video)])]; writeFileSync(MF, JSON.stringify(manifest, null, 1));