commit 936956ee6bd7d643085d02bd58fc3b76910631e4
parent d29a47e556c2f76f349ffe336551af8f860b00bb
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Sun, 4 Oct 2026 16:34:05 -0400
editor: POST /api/ops/persist-videos and pnpm ops persist-videos
The adapter over persistVideosAction: { items: [{slug, id}], format?,
replace?, gapMs?, minFreeMemMb?, dryRun?, queueKey? }. A dry run answers with
the plan; a real run returns one job per channel (jobs, jobIds, and jobId when
there is one) plus the channels skipped. reqVideoItems checks every slug as a
channel slug and every id as one path segment; optNonNegativeInt reads a value
where 0 means off. A long list goes through the CLI's --file. The e2e case
pins the dry-run buckets, the body refusals and the disk-floor refusal, none
of which start a job.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Diffstat:
8 files changed, 279 insertions(+), 1 deletion(-)
diff --git a/RUNNING_IN_DOCKER.md b/RUNNING_IN_DOCKER.md
@@ -256,6 +256,7 @@ pnpm ops refresh-report --json '{"all":true}'
pnpm ops keep-videos --json '{"slug":"paramount-tactical","match":"TheQuartering","dryRun":true}'
pnpm ops fetch-posts --json '{"slug":"example-x","older":true}' --wait
pnpm ops capture-posts --json '{"slug":"example-x","ids":["1234567890"]}' --wait
+pnpm ops persist-videos --json '{"items":[{"slug":"example-channel","id":"abc123"}],"dryRun":true}'
pnpm ops get channel the-quartering
pnpm ops list # every action name
```
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -1,6 +1,7 @@
# Changelog
## [Unreleased]
+- **Persist a list of videos, across channels, to the saved-video store.** `pnpm ops persist-videos --json '{"items":[{"slug":"<channel>","id":"<video id>"}, …]}'` (or `--file list.json` for a long list) re-fetches the source container of each video that is not saved yet, the way **Persist source video** does on a video's page. `"format": "original"` or `"video_720"` picks the quality (default: each channel's own), and `"replace": "above-height"` also re-fetches a saved video whose recorded height is unknown or above that quality — the old file is removed only after the new one is saved, and the saved video keeps its retention class. Downloads run one at a time, one job per channel on that channel's download queue, with the usual gap between them (`"gapMs"` overrides it); `"minFreeMemMb"` holds each download until that much memory is free. A low disk or a rate limit stops the job, and running the same list again picks up where it left off: saved videos are skipped. `"dryRun": true` answers with what a run would do — saved, saved above the height, to fetch, no source URL, unknown — and starts nothing. Jobs show as **Persist videos**, can be drained, and can be retried from /jobs.
- **Capture specific X posts: a screenshot of each, and its attached media.** `pnpm ops capture-posts --json '{"slug":"<channel>","ids":["<post id>", …]}'` shoots each post as X shows it, through the connected X profile, and downloads its pictures and videos with gallery-dl, into the channel's `posts-media/<post id>/` beside a `capture.json` that records when, from which URLs, and each file's size and SHA-256. Every id must already be in the channel's posts archive; one that is not is refused by name and nothing runs. `"shots": false` or `"media": false` skips that half, and posts already captured are skipped unless `"force": true`. The job runs on the X queue with a post fetch, so the two never run at once, and waits a random 4–10 seconds before each request to X, as fetches do. A deleted post, or one behind its account's wall (protected, suspended, gone), is recorded as such in the channel's deleted-post record; a post behind a sensitive-media warning is opened and shot. If X asks to log in, or answers "Something went wrong", the job stops at that post and leaves the rest for a later run. Captures are never published: the export does not read them.
- **A source video can be saved at 720p for clip and editing work.** **Persist source video** on a video's page has a **Quality** select: **Original** (the best video and audio, as every persist has been) or **Video 720p (H.264, for clips/editing)**, which saves an H.264 mp4 at most 720p tall — smaller, and quick to cut. When a source has nothing at or under 720p in H.264 it takes 480p, and when it has neither it takes whatever is best and says so in the job's log, with the height it got. The default is the new **Source video quality** under **Settings**, which a channel can override in its Advanced settings; the whole-recording fetch (`full: true`) and **Persist kept now** follow the channel's, else the global, choice. A persisted video's **Source video** card now shows the format that was saved (height and codec) and the quality asked for; videos persisted before this show nothing new. "Video 720p" is also offered as a download format.
- **X posts are fetched more slowly, with random gaps.** Every read of X now waits a random 4 to 10 seconds before each request to X, where it used to page as fast as X answered, and always waits out a rate limit rather than pushing through. When fetching older posts, the pause between one three-month window and the next is a random 45 to 120 seconds instead of a fixed 15. A deep walk of an account's history takes longer; a routine fetch of new posts takes a few seconds more.
diff --git a/editor/app/api/ops/_lib.test.ts b/editor/app/api/ops/_lib.test.ts
@@ -1,6 +1,12 @@
import test from "node:test";
import assert from "node:assert/strict";
-import { OpsInputError, optPositiveInt, optSubset } from "./_lib";
+import {
+ OpsInputError,
+ optNonNegativeInt,
+ optPositiveInt,
+ optSubset,
+ reqVideoItems,
+} from "./_lib";
// Run with:
// pnpm -C editor exec tsx --test "app/**/*.test.ts"
@@ -52,3 +58,57 @@ test("optPositiveInt: zero, negatives, fractions and strings are refused", () =>
);
}
});
+
+test("optNonNegativeInt: absent is undefined; zero and whole numbers come back", () => {
+ assert.equal(optNonNegativeInt({}, "gapMs"), undefined);
+ assert.equal(optNonNegativeInt({ gapMs: 0 }, "gapMs"), 0);
+ assert.equal(optNonNegativeInt({ gapMs: 30000 }, "gapMs"), 30000);
+});
+
+test("optNonNegativeInt: negatives, fractions and strings are refused", () => {
+ for (const gapMs of [-1, 1.5, "10", null, Number.NaN]) {
+ assert.throws(
+ () => optNonNegativeInt({ gapMs }, "gapMs"),
+ (e: unknown) =>
+ e instanceof OpsInputError && /"gapMs" must be a whole number, zero or above/.test(e.message),
+ JSON.stringify(gapMs),
+ );
+ }
+});
+
+test("reqVideoItems: a list of { slug, id } comes back trimmed, in order", () => {
+ assert.deepEqual(
+ reqVideoItems(
+ { items: [{ slug: "demo-channel", id: " abc123 " }, { slug: "other", id: "x" }] },
+ "items",
+ ),
+ [
+ { slug: "demo-channel", id: "abc123" },
+ { slug: "other", id: "x" },
+ ],
+ );
+});
+
+test("reqVideoItems: an empty list, a non-object, an extra key or a missing id is refused", () => {
+ for (const items of [
+ undefined,
+ [],
+ "demo-channel/abc123",
+ ["abc123"],
+ [{ slug: "demo-channel" }],
+ [{ slug: "demo-channel", id: "abc123", height: 720 }],
+ ]) {
+ assert.throws(() => reqVideoItems({ items }, "items"), OpsInputError, JSON.stringify(items));
+ }
+});
+
+test("reqVideoItems: a traversing slug or id is refused at the door", () => {
+ for (const item of [
+ { slug: "../../escape", id: "abc123" },
+ { slug: "demo-channel", id: "../abc123" },
+ { slug: "demo-channel", id: ".." },
+ { slug: "demo-channel", id: "a\\b" },
+ ]) {
+ assert.throws(() => reqVideoItems({ items: [item] }, "items"), OpsInputError, JSON.stringify(item));
+ }
+});
diff --git a/editor/app/api/ops/_lib.ts b/editor/app/api/ops/_lib.ts
@@ -157,6 +157,49 @@ export function optPositiveInt(body: OpsBody, key: string): number | undefined {
return v;
}
+// A whole number, zero or above, or undefined when absent — for a duration or
+// a floor where 0 means "off".
+export function optNonNegativeInt(body: OpsBody, key: string): number | undefined {
+ const v = body[key];
+ if (v === undefined) return undefined;
+ if (typeof v !== "number" || !Number.isInteger(v) || v < 0) {
+ throw new OpsInputError(`"${key}" must be a whole number, zero or above`);
+ }
+ return v;
+}
+
+// A list of videos across channels: `[{ slug, id }, …]`. Every slug is a
+// CHANNEL SLUG (see reqSlug) and every id one path segment, checked here so a
+// traversing entry is refused at the door rather than read as an unknown video.
+export function reqVideoItems(
+ body: OpsBody,
+ key: string,
+): { slug: string; id: string }[] {
+ const v = body[key];
+ if (!Array.isArray(v) || v.length === 0) {
+ throw new OpsInputError(
+ `"${key}" is required and must be a non-empty array of { "slug", "id" }`,
+ );
+ }
+ return v.map((entry, i) => {
+ if (typeof entry !== "object" || entry === null || Array.isArray(entry)) {
+ throw new OpsInputError(`"${key}[${i}]" must be an object { "slug", "id" }`);
+ }
+ const extra = Object.keys(entry).filter((k) => k !== "slug" && k !== "id");
+ if (extra.length) {
+ throw new OpsInputError(
+ `"${key}[${i}]" has unknown key(s): ${extra.join(", ")} — an item is { "slug", "id" }`,
+ );
+ }
+ const slug = reqSlug(entry as OpsBody, "slug");
+ const id = reqString(entry as OpsBody, "id");
+ if (id === "." || id === ".." || /[/\\\0]/.test(id)) {
+ throw new OpsInputError(`"${id}" is not a video id (one path segment)`);
+ }
+ return { slug, id };
+ });
+}
+
export function reqStringArray(body: OpsBody, key: string): string[] {
const v = body[key];
if (
diff --git a/editor/app/api/ops/persist-videos/route.ts b/editor/app/api/ops/persist-videos/route.ts
@@ -0,0 +1,65 @@
+import { NextResponse } from "next/server";
+import { SOURCE_VIDEO_QUALITIES } from "yt-dlp-transcript-common/ytdlp/downloadFormat";
+import { PERSIST_REPLACE_POLICIES } from "yt-dlp-transcript-common/controller/persistVideos";
+import { persistVideosAction } from "../../../channels/[slug]/persistActions";
+import {
+ oneOf,
+ ops,
+ opsFail,
+ optBool,
+ optNonNegativeInt,
+ optString,
+ reqVideoItems,
+} from "../_lib";
+
+export const dynamic = "force-dynamic";
+
+// POST { items: [{ slug, id }, …], format?, replace?, gapMs?, minFreeMemMb?,
+// dryRun?, queueKey? }
+// dryRun -> { ok: true, dryRun: true, plan }
+// else -> { ok: true, dryRun: false, plan, jobs: [{ slug, jobId }], jobIds,
+// jobId (one job only), skipped: [{ slug, reason }] }
+//
+// Persist specific videos, across channels, to the saved-video store: the
+// video page's "Persist source video" over a list. `plan` buckets every item —
+// saved, wrongHeight, toFetch, noUrl, unknown — with counts and ids, and
+// `willFetch` is how many downloads a run makes. A real run starts one job per
+// channel with something to fetch, on that channel's download queue.
+//
+// `format` ("original" | "video_720") defaults to each channel's own quality.
+// `replace: "above-height"` also re-fetches a saved container whose height is
+// unknown or above that quality's ceiling (default "never"). `gapMs` is the
+// pause between downloads (default the batch downloads' gap), `minFreeMemMb`
+// holds each download until that much memory is available (default 0, off).
+// Re-running the same body is the resume: saved videos are skipped.
+//
+// A big list goes in a file: `pnpm ops persist-videos --file list.json`.
+export async function POST(request: Request) {
+ return ops(
+ request,
+ ["items", "format", "replace", "gapMs", "minFreeMemMb", "dryRun", "queueKey"],
+ async (body) => {
+ const result = await persistVideosAction({
+ items: reqVideoItems(body, "items"),
+ format:
+ body.format === undefined
+ ? undefined
+ : oneOf(body, "format", SOURCE_VIDEO_QUALITIES),
+ replace:
+ body.replace === undefined
+ ? undefined
+ : oneOf(body, "replace", PERSIST_REPLACE_POLICIES),
+ gapMs: optNonNegativeInt(body, "gapMs"),
+ minFreeMemMb: optNonNegativeInt(body, "minFreeMemMb"),
+ dryRun: optBool(body, "dryRun"),
+ queueKey: optString(body, "queueKey"),
+ });
+ if (!result.ok) return opsFail(result.error);
+ if (result.dryRun) return NextResponse.json(result);
+ return NextResponse.json({
+ ...result,
+ ...(result.jobIds.length === 1 ? { jobId: result.jobIds[0] } : {}),
+ });
+ },
+ );
+}
diff --git a/editor/e2e/ops-api.spec.ts b/editor/e2e/ops-api.spec.ts
@@ -194,6 +194,7 @@ test("a traversing slug is refused at the door, on every route that takes one",
["relocate-back", { slugs: ["../../escape"] }],
["fetch-posts", { slug: "../../escape", older: true }],
["capture-posts", { slug: "../../escape", ids: ["1"] }],
+ ["persist-videos", { items: [{ slug: "../../escape", id: "abc123" }] }],
];
for (const [action, data] of cases) {
const { status, body } = await ops(request, action, data);
@@ -1042,6 +1043,79 @@ test("keep-videos marks the matching video only, after a dry run that writes not
expect(badRegex.body.error).toContain("invalid pattern");
});
+test("persist-videos buckets a list in a dry run, and refuses a bad body and a low disk — before any job", async ({
+ request,
+}) => {
+ await resetData("one-youtube-channel-with-data");
+ await settings();
+ const slug = "test-youtube";
+ const video = "20240101_test1234567";
+ const before = await listJobIds();
+
+ type Bucket = { count: number; items: { slug: string; id: string }[] };
+ type PersistBody = OpsResponse & {
+ dryRun?: boolean;
+ plan?: Record<"saved" | "wrongHeight" | "toFetch" | "noUrl" | "unknown", Bucket> & {
+ willFetch: number;
+ };
+ };
+ const items = [
+ { slug, id: video },
+ { slug, id: "missing12345" },
+ { slug: "no-such-channel", id: "abc123" },
+ ];
+ const dry = await ops(request, "persist-videos", {
+ items,
+ format: "video_720",
+ dryRun: true,
+ });
+ expect(dry.status, JSON.stringify(dry.body)).toBe(200);
+ const plan = (dry.body as PersistBody).plan!;
+ expect((dry.body as PersistBody).dryRun).toBe(true);
+ expect(plan.toFetch.items).toEqual([{ slug, id: video }]);
+ expect(plan.unknown.count).toBe(2);
+ expect(plan.saved.count).toBe(0);
+ expect(plan.willFetch).toBe(1);
+
+ // Nothing to fetch is an answer, not a job.
+ const none = await ops(request, "persist-videos", {
+ items: [{ slug: "no-such-channel", id: "abc123" }],
+ });
+ expect(none.status, JSON.stringify(none.body)).toBe(200);
+ expect(none.body.jobIds).toEqual([]);
+
+ // The body's shape.
+ const cases: [Record<string, unknown>, RegExp][] = [
+ [{}, /"items" is required/],
+ [{ items: [] }, /"items" is required/],
+ [{ items: [{ slug, id: video, height: 720 }] }, /unknown key\(s\): height/],
+ [{ items: [{ slug, id: "../escape" }] }, /is not a video id/],
+ [{ items, format: "1080p" }, /"format" must be one of original, video_720/],
+ [{ items, replace: "always" }, /"replace" must be one of never, above-height/],
+ [{ items, gapMs: -1 }, /"gapMs" must be a whole number, zero or above/],
+ [{ items, minFreeMemMb: "4096" }, /"minFreeMemMb" must be a whole number/],
+ [{ items, dryRun: "yes" }, /"dryRun" must be a boolean/],
+ [{ items, ids: ["x"] }, /unknown key\(s\): ids/],
+ ];
+ for (const [data, error] of cases) {
+ const res = await ops(request, "persist-videos", data);
+ expect(res.status, JSON.stringify(data)).toBe(400);
+ expect(res.body.error, JSON.stringify(data)).toMatch(error);
+ }
+
+ // A real run asks the disk floor before it queues anything.
+ await writeSettings({ minFreeDiskGB: 1_000_000 });
+ try {
+ const low = await ops(request, "persist-videos", { items });
+ expect(low.status, JSON.stringify(low.body)).toBe(400);
+ expect(low.body.error).toMatch(/^Low disk space/);
+ } finally {
+ await settings();
+ }
+
+ expect(await listJobIds()).toEqual(before);
+});
+
// --- the two build routes speak the same body ---------------------------------
//
// They did not. build-site took `siteIds` (a list) and build-deploy took
diff --git a/scripts/archilyzer-ops.mjs b/scripts/archilyzer-ops.mjs
@@ -42,6 +42,7 @@
// pnpm ops get channel the-quartering
// pnpm ops tags --json '{"op":"define","tag":{"id":"eva-collab","label":"Collab"}}'
// pnpm ops tag-videos --file ids.json
+// pnpm ops persist-videos --file list.json --wait
// pnpm ops get tags eva-collab
// pnpm ops cut-release --json '{"workspace":"all","version":"next","commit":true}'
//
@@ -132,6 +133,9 @@ const ACTIONS = [
// Set the per-video do-not-clean marker on every video of a channel whose
// title/description matches a download-filter pattern ({slug, match}).
"keep-videos",
+ // Persist specific videos, across channels, to the saved-video store
+ // ({items: [{slug, id}]}), paced and gated; a re-run resumes.
+ "persist-videos",
// Cut a changelog's [Unreleased] into a dated release heading (release 10
// slice P). Synchronous. The same writer as `archilyzer release cut`, which
// needs no editor at all — this route exists only on an editor built from
@@ -335,6 +339,17 @@ export function usage() {
' "shots": false or "media": false skips that half; posts already captured',
' are skipped unless "force": true. Paced like a post fetch, on its queue.',
"",
+ 'persist-videos saves specific videos, across channels, to the saved-video',
+ ' store: {"items": [{"slug", "id"}, ...]}. "format": "original" |',
+ ' "video_720" (default: each channel\'s own). "replace": "above-height" also',
+ ' re-fetches a saved one whose height is unknown or above that quality',
+ ' (default "never"). "gapMs" pauses between downloads (default the batch',
+ ' gap), "minFreeMemMb" waits for that much free memory before each.',
+ ' "dryRun": true answers with the buckets (saved, wrongHeight, toFetch,',
+ ' noUrl, unknown) and starts nothing. One job per channel, on its download',
+ " queue; a low disk or a rate limit stops it, and running the same body",
+ " again resumes — saved videos are skipped.",
+ "",
'"preview": "<branch>" on deploy-site or build-deploy makes it a Cloudflare',
" Pages PREVIEW instead of production: the same bundle goes to a branch",
" alias, https://<branch>.<project>.pages.dev, and the live site is left",
diff --git a/scripts/archilyzer-ops.test.mjs b/scripts/archilyzer-ops.test.mjs
@@ -405,3 +405,22 @@ test("capture-posts is a POST to its route, named in the usage", () => {
assert.match(usage(), /Actions:.*fetch-posts, capture-posts/);
assert.match(usage(), /Every id must be in the channel's posts archive/);
});
+
+test("persist-videos is a POST to its route, named in the usage", () => {
+ const p = parseArgs([
+ "persist-videos",
+ "--json",
+ '{"items":[{"slug":"demo-channel","id":"abc123"}],"format":"video_720","dryRun":true}',
+ ]);
+ assert.equal(p.method, "POST");
+ assert.equal(p.path, "/api/ops/persist-videos");
+ assert.deepEqual(p.body, {
+ items: [{ slug: "demo-channel", id: "abc123" }],
+ format: "video_720",
+ dryRun: true,
+ });
+ assert.equal(p.defaultSource, undefined);
+ // A long list arrives from a file.
+ assert.equal(parseArgs(["persist-videos", "--file", "list.json"]).bodyFile, "list.json");
+ assert.match(usage(), /persist-videos/);
+});