Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit aa295e2cf4372e30b05ad00779c5cd81b6b92ccf
parent 980624e9c2e7496bb19956e891ce6d092ae46ae4
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Fri, 21 Aug 2026 23:22:50 -0400

digest/backfill: order a channel's candidates by upload date

Newest-first existed only on the auto-queue runners. The digest batch
ordered by duration and the backfill batch by readdir order, so on a
channel with 11,329 pending digests a video uploaded this morning sat
behind every one of them with no setting that changed it.

Adds digest.recencyOrder and backfill.order (AutoQueueOrder, sanitized
to "listed"), resolved INSIDE each batch so the armed sweep, a
hand-clicked run and an ids-scoped run all inherit it. Digest's
existing DigestOrder stays a separate field: shortest-first and
newest-first are different axes and one enum would make them look
mutually exclusive.

The two COMPOSE rather than choose. Sort by duration, then stable-sort
by the recency comparator: makeRecencyComparator returns 0 for equal
YYYYMMDD keys and Array#sort is stable, so this is newest day first,
shortest video within a day. With "listed" the comparator is null and
the second sort never runs, reproducing today byte-for-byte.

Sorting backfill's `wanted` is not a break of "re-derive eligibility
from disk on every pull" — that invariant is about kind.state(), which
next() still reads per item.

Measured on the live corpus: 11,329 the-quartering candidates ordered
in 118 ms; the-quartering-rumble's 7,858 in 707 ms, correctly, which
is the case the previous commit's key-space fix exists for.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Diffstat:
Mcommon/controller/backfillBatch.ts | 29++++++++++++++++++++++++++++-
Acommon/controller/batchRecency.test.ts | 129+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/controller/batchRecency.ts | 51+++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/controller/digestBatch.ts | 25++++++++++++++++++++++++-
Mcommon/jobs/autoQueuePolicy.ts | 12+++++++++++-
Mcommon/lib/settings.ts | 22++++++++++++++++++++++
Meditor/CHANGELOG.md | 1+
7 files changed, 266 insertions(+), 3 deletions(-)

diff --git a/common/controller/backfillBatch.ts b/common/controller/backfillBatch.ts @@ -50,6 +50,8 @@ import { } from "../lib/backfillKinds"; import { transcriptionActivity } from "./digestYield"; import { reacquireMediaFor, type ReacquireOutcome } from "./backfillReacquire"; +import type { AutoQueueOrder } from "../jobs/autoQueuePolicy"; +import { batchRecencyComparator } from "./batchRecency"; // How many slots the backfill lane may use right now. // @@ -85,6 +87,10 @@ export type BackfillBatchOptions = { kindIds?: string[]; // When set, only consider these video ids (intersected with what's on disk). ids?: string[]; + // Upload-date ordering over this channel's candidates. Absent = whatever + // settings.backfill.order says, resolved inside the batch so every caller + // inherits it. "listed" is today's behaviour and the default. + order?: AutoQueueOrder; // Redo videos that are already present at the current identity. force?: boolean; // Stop after this many successful videos. @@ -250,7 +256,27 @@ export async function runBackfillBatch( const dataDir = path.join(opts.paths.channelsDir, opts.channelSlug, "data"); const allDirs = await readdir(dataDir).catch(() => [] as string[]); const onDisk = new Set(allDirs); - const wanted = opts.ids ? opts.ids.filter((id) => onDisk.has(id)) : allDirs; + const wantedListed = opts.ids + ? opts.ids.filter((id) => onDisk.has(id)) + : allDirs; + + // Upload-date ordering. Applied ONCE, here, to the frozen list the cursor + // walks — not inside next(), which must stay O(1) per pull. Sorting `wanted` + // is not a break of the "re-derive eligibility from disk on every pull" + // invariant: that invariant is about kind.state(), which next() still reads + // per item; this only decides what order the cursor reaches them in. + // + // Resolved inside the batch so the armed sweep, a hand-clicked channel run + // and an ids-scoped run all inherit settings.backfill.order. "listed" gives a + // null comparator and no sort at all, reproducing today byte-for-byte. + const order = opts.order ?? settings.backfill.order; + const byRecency = await batchRecencyComparator( + opts.paths, + opts.channelSlug, + wantedListed, + order, + ); + const wanted = byRecency ? [...wantedListed].sort(byRecency) : wantedListed; // Resolved once per run, not per video: the identity is a settings read plus // some string work, and deriving it per item is how a counter and a runner end @@ -270,6 +296,7 @@ export async function runBackfillBatch( log( `Backfill ${opts.channelSlug}: ${kinds.map((k) => k.id).join(", ")} over ` + `${wanted.length} video dir(s)` + + (byRecency ? `, ${order} first` : "") + (allowRedownload ? ", re-acquiring media where it is gone." : ", retained media only (re-download is off)."), diff --git a/common/controller/batchRecency.test.ts b/common/controller/batchRecency.test.ts @@ -0,0 +1,129 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { mkdir, mkdtemp, rm, writeFile } from "node:fs/promises"; +import { tmpdir } from "node:os"; +import path from "node:path"; +import type { Paths } from "../lib/paths"; +import { batchRecencyComparator } from "./batchRecency"; +import { clearRecencyCache } from "./recencyIndex"; + +// Run with: node_modules/.bin/tsx --test common/controller/batchRecency.test.ts + +async function writeMeta( + channelsDir: string, + slug: string, + id: string, + uploadDate: string, +): Promise<void> { + const dir = path.join(channelsDir, slug, "data", id); + await mkdir(dir, { recursive: true }); + await writeFile( + path.join(dir, "metadata.info.json"), + JSON.stringify({ id, upload_date: uploadDate }), + ); +} + +// One channel, three videos, dates deliberately out of id order so a sort that +// silently fell back to "listed" would be visible. +async function fixture(): Promise<{ paths: Paths; cleanup: () => Promise<void> }> { + const dir = await mkdtemp(path.join(tmpdir(), "batch-recency-")); + clearRecencyCache(); + const channelsDir = path.join(dir, "channels"); + await writeMeta(channelsDir, "ch", "a-old", "20240101"); + await writeMeta(channelsDir, "ch", "b-new", "20260812"); + await writeMeta(channelsDir, "ch", "c-mid", "20250601"); + return { + paths: { lmdbPath: path.join(dir, "none.mdb"), channelsDir } as Paths, + cleanup: () => rm(dir, { recursive: true, force: true }), + }; +} + +const IDS = ["a-old", "b-new", "c-mid"]; + +test('order "listed" reproduces today\'s order exactly', async () => { + // The compatibility claim for this stage in one assertion. "listed" must not + // return a comparator that happens to be order-preserving — it must return + // NOTHING, so the call sites skip the sort entirely and a list is byte-for- + // byte the list they already had. + const { paths, cleanup } = await fixture(); + try { + const cmp = await batchRecencyComparator(paths, "ch", IDS, "listed"); + assert.equal(cmp, null); + } finally { + await cleanup(); + } +}); + +test("newest/oldest sort a channel's candidates by upload date", async () => { + const { paths, cleanup } = await fixture(); + try { + const newest = await batchRecencyComparator(paths, "ch", IDS, "newest"); + assert.deepEqual([...IDS].sort(newest!), ["b-new", "c-mid", "a-old"]); + clearRecencyCache(); + const oldest = await batchRecencyComparator(paths, "ch", IDS, "oldest"); + assert.deepEqual([...IDS].sort(oldest!), ["a-old", "c-mid", "b-new"]); + } finally { + await cleanup(); + } +}); + +test("recency COMPOSES with a prior sort rather than replacing it", async () => { + // This is how digestBatch keeps both of its rules: sort by duration, then + // stable-sort by date. makeRecencyComparator returns 0 for two videos sharing + // a YYYYMMDD key and Array#sort is stable, so the duration order survives + // inside a day. "newest day first, shortest video within a day." + const dir = await mkdtemp(path.join(tmpdir(), "batch-recency-compose-")); + try { + clearRecencyCache(); + const channelsDir = path.join(dir, "channels"); + // Two videos share the newer day; the older day has one. + await writeMeta(channelsDir, "ch", "new-long", "20260812"); + await writeMeta(channelsDir, "ch", "new-short", "20260812"); + await writeMeta(channelsDir, "ch", "old-short", "20240101"); + const paths = { + lmdbPath: path.join(dir, "none.mdb"), + channelsDir, + } as Paths; + const durations = new Map([ + ["new-long", 9000], + ["new-short", 60], + ["old-short", 30], + ]); + const rows = [...durations.keys()]; + rows.sort((a, b) => durations.get(a)! - durations.get(b)!); + assert.deepEqual(rows, ["old-short", "new-short", "new-long"]); + const cmp = await batchRecencyComparator(paths, "ch", rows, "newest"); + rows.sort(cmp!); + // Newest day first; within that day the shortest-first order is intact. + assert.deepEqual(rows, ["new-short", "new-long", "old-short"]); + } finally { + await rm(dir, { recursive: true, force: true }); + } +}); + +test("an unreadable corpus degrades to the listed order, never throws", async () => { + // A recency lookup failing is not a reason to refuse to digest a channel. + const cmp = await batchRecencyComparator( + { + lmdbPath: "/nonexistent/none.mdb", + channelsDir: "/nonexistent/channels", + } as Paths, + "ch", + IDS, + "newest", + ); + // Every id falls to layer 4's "" key, so every pair compares equal and the + // sort is a no-op — the listed order, reached honestly rather than by + // dropping undatable videos out of the queue. + assert.deepEqual([...IDS].sort(cmp!), IDS); +}); + +test("an empty candidate list short-circuits", async () => { + const cmp = await batchRecencyComparator( + { lmdbPath: "x", channelsDir: "y" } as Paths, + "ch", + [], + "newest", + ); + assert.equal(cmp, null); +}); diff --git a/common/controller/batchRecency.ts b/common/controller/batchRecency.ts @@ -0,0 +1,51 @@ +import type { Paths } from "../lib/paths"; +import type { AutoQueueOrder } from "../jobs/autoQueuePolicy"; +import { buildRecencyKeys, makeRecencyComparator } from "./recencyIndex"; + +// Recency ordering for the per-channel BATCH controllers (digestBatch, +// backfillBatch), as distinct from the auto-queue runners. +// +// The runners already sort by upload date through recencyIndex; the batches did +// not, so "newest first" meant one thing on two of the four pipelines and +// nothing on the other two. This is the shared half — one call site's worth of +// glue, so digest and backfill cannot drift into two different notions of what +// "newest" means. +// +// Three things are settled here rather than at each call site: +// +// * `order: "listed"` returns null, and a null comparator means DO NOT SORT. +// That is what reproduces today byte-for-byte: the historical order is +// reproduced by not sorting at all, never by sorting with an identity +// comparator, which would still perturb ties on some engines. +// * `interpolate: false`. A batch's candidates come from a readdir of the +// channel's data/ dir, so every one of them is downloaded and has a +// metadata.info.json for the tail read. Playlist interpolation exists for +// videos with nothing on disk, which by construction cannot be here. +// * every id is owned by this channel. That is not a convenience — it is what +// keeps layer 1 out of the wrong key space (see recencyIndex's layer-1 +// comment: candidate ids here are DIRECTORY names). +// +// Never throws. A recency lookup failing is not a reason to refuse to digest a +// channel, so a broken or mid-rebuild index degrades to the listed order. +export async function batchRecencyComparator( + paths: Paths, + channelSlug: string, + ids: ReadonlyArray<string>, + order: AutoQueueOrder, +): Promise<((a: string, b: string) => number) | null> { + if (order === "listed" || ids.length === 0) return null; + try { + const owner = new Map<string, string>(); + for (const id of ids) owner.set(id, channelSlug); + const keys = await buildRecencyKeys({ + paths, + meta: [{ slug: channelSlug }], + candidateIds: new Set(ids), + owner, + interpolate: false, + }); + return makeRecencyComparator(keys, order); + } catch { + return null; + } +} diff --git a/common/controller/digestBatch.ts b/common/controller/digestBatch.ts @@ -37,6 +37,8 @@ import { import { CUES_JSON_FILENAME } from "../lib/videoStatus"; import { digestVideo } from "./digestVideo"; import { transcriptionActivity } from "./digestYield"; +import type { AutoQueueOrder } from "../jobs/autoQueuePolicy"; +import { batchRecencyComparator } from "./batchRecency"; import { buildDigestClusterPlan, planSlugForDir, @@ -61,6 +63,11 @@ export type DigestBatchOptions = { // default: it converts the backlog into visible coverage fastest, and the // long-tail 8% is where a prompt bug is most expensive to discover late. order?: DigestOrder; + // Upload-date ordering, applied AFTER `order` above. Absent = whatever + // settings.digest.recencyOrder says, resolved inside the batch so the armed + // sweep, a hand-clicked channel run and an ids-scoped run all inherit the + // same setting instead of three call sites each remembering to pass it. + recencyOrder?: AutoQueueOrder; // Duration window, for splitting the corpus between lanes (e.g. local takes // ≤ longTailSeconds, the metered lane takes the tail). minDurationSeconds?: number; @@ -224,10 +231,26 @@ export async function runDigestBatch( : a.duration - b.duration || a.id.localeCompare(b.id), ); + // Then by upload date, if asked. COMPOSED with the duration sort above rather + // than replacing it: Array#sort is stable and makeRecencyComparator returns 0 + // for two videos sharing a YYYYMMDD key, so this reads as "newest day first, + // shortest video within a day" — both rules intact. With recencyOrder + // "listed" the comparator is null and this second sort never runs, which is + // what reproduces the historical order byte-for-byte. + const recencyOrder = opts.recencyOrder ?? digestSettings.recencyOrder; + const byRecency = await batchRecencyComparator( + opts.paths, + opts.channelSlug, + candidates.map((c) => c.id), + recencyOrder, + ); + if (byRecency) candidates.sort((a, b) => byRecency(a.id, b.id)); + log( `Digest ${opts.channelSlug} (${lane} lane, ${app.id}/${modelRequested}, sections: ${sections.join(", ")}): ` + `${candidates.length} candidate(s) of ${wanted.length} on disk ` + - `(${noTranscript} without a transcript, ${outOfWindow} outside the duration window), ${order}.` + + `(${noTranscript} without a transcript, ${outOfWindow} outside the duration window), ${order}` + + (byRecency ? `, ${recencyOrder} first.` : ".") + (clusterPlan ? ` Duplicate plan: ${clusterPlan.clusters} cluster(s), ${clusterPlan.bySlug.size} member(s) mapped.` : " Duplicate sharing off."), diff --git a/common/jobs/autoQueuePolicy.ts b/common/jobs/autoQueuePolicy.ts @@ -44,6 +44,16 @@ export const AUTO_QUEUE_ORDERS: ReadonlyArray<AutoQueueOrder> = [ "oldest", ]; +// Coerce a stored/raw value to a legal order. Anything unrecognised — including +// a missing field on a settings file written before the field existed — means +// "listed", i.e. today's behaviour. One sanitizer, because the same enum is now +// stored in three places (the two auto-queue policies, digest.recencyOrder and +// backfill.order) and three copies of this line would eventually disagree about +// what an absent field means. +export function sanitizeAutoQueueOrder(value: unknown): AutoQueueOrder { + return value === "newest" || value === "oldest" ? value : "listed"; +} + export type AutoQueueMatch = { type: AutoQueueMatchType; // Channel slug (type=channel) or platform name (type=platform). Ignored for @@ -472,7 +482,7 @@ function sanitizePolicy(value: unknown): AutoQueuePolicy { // a pre-existing settings.json) leaves the lane off. replaceAutoSubs: r.replaceAutoSubs === true, // Anything unrecognised (including a missing field) means today's behaviour. - order: r.order === "newest" || r.order === "oldest" ? r.order : "listed", + order: sanitizeAutoQueueOrder(r.order), snoozeUntil: sanitizeSnooze(r.snoozeUntil), root: sanitizeRoot(r.root, seen), }; diff --git a/common/lib/settings.ts b/common/lib/settings.ts @@ -20,9 +20,11 @@ import { validateWorkers, } from "./workers"; import { + type AutoQueueOrder, type AutoQueueSettings, defaultAutoQueue, sanitizeAutoQueue, + sanitizeAutoQueueOrder, } from "../jobs/autoQueuePolicy"; import { DEFAULT_DIARIZATION_ENGINE, @@ -310,6 +312,11 @@ export type BackfillSettings = { // Empty = every registered lane-tier kind / every channel. sweepKinds: string[]; sweepChannels: string[]; + // The order the backfill batch hands out a channel's videos. Same three + // values and the same "listed means don't sort" contract as the auto-queue + // policies' `order`, so an operator learns the control once. Default + // "listed", i.e. today's behaviour. + order: AutoQueueOrder; // Re-acquire media for videos whose input is GONE (audio deleted after // transcription). OFF by default and deliberately so: measured on this corpus, // 836 videos still have media and ~76,270 would need a re-download — 91x the @@ -504,6 +511,15 @@ export type DigestSettings = { // Was a scored variable in the bake-off rather than a pre-applied fix; the // measurement is in and "chunk-local" is now the shipped default. timestampMode: DigestTimestampMode; + // The order the digest batch hands out a channel's videos: newest upload + // first, oldest first, or the listed order (today's behaviour, and the + // default). Deliberately a SEPARATE field from DigestBatchOptions.order — + // shortest-first/longest-first is a duration axis and this is a date axis, + // and they compose rather than exclude: the batch sorts by duration and then + // stable-sorts by date, so "newest first" reads as newest day first, shortest + // video within a day. One enum carrying both would make them look mutually + // exclusive, which they are not. + recencyOrder: AutoQueueOrder; // A free-text label for a non-default prompt shape, folded into the recorded // provenance by digestPromptVariant(). Setting it invalidates every digest // generated under a different label, which is exactly what makes a bake-off @@ -867,6 +883,8 @@ export function defaultDigest(): DigestSettings { spendCapUsd: 0, sections: ["chapters"], timestampMode: DEFAULT_DIGEST_TIMESTAMP_MODE, + // "listed" — today's behaviour exactly. Ordering is opt-in. + recencyOrder: "listed", promptVariant: "", }; } @@ -958,6 +976,7 @@ export function sanitizeDigest(value: unknown): DigestSettings { timestampMode: isDigestTimestampMode(r.timestampMode) ? r.timestampMode : d.timestampMode, + recencyOrder: sanitizeAutoQueueOrder(r.recencyOrder), // Trimmed and length-capped: it goes into provenance on every record, and a // runaway value would bloat 119k sidecars. promptVariant: @@ -1069,6 +1088,8 @@ export function defaultBackfill(): BackfillSettings { sweepEnabled: false, sweepKinds: [], sweepChannels: [], + // "listed" — today's behaviour exactly. Ordering is opt-in. + order: "listed", // See BackfillSettings.allowRedownload — this one holds disk. allowRedownload: false, }; @@ -1094,6 +1115,7 @@ export function sanitizeBackfill(value: unknown): BackfillSettings { sweepEnabled: r.sweepEnabled === true, sweepKinds: slugs(r.sweepKinds), sweepChannels: slugs(r.sweepChannels), + order: sanitizeAutoQueueOrder(r.order), allowRedownload: r.allowRedownload === true, }; } diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -5,6 +5,7 @@ - **Nothing the container runs is reachable from outside the machine until you say so.** The editor has no authentication of any kind, shells out to yt-dlp, and deletes media — so **no application container publishes a port at all**. Caddy is the single front door, and every one of the four ports it publishes binds `127.0.0.1` by default, the public sites included; opening one is a deliberate edit of a single line in `.env`. Because a bind address is exactly the sort of thing that gets changed in a hurry, there is also a rail: if a private app is bound off-loopback with nothing checking credentials, **the containers refuse to start** — both the app and Caddy, which is the process that actually opens the ports — and print the four ways to fix it. `basic_auth` is built into Caddy so a password needs nothing installed; Tinyauth and Authelia attach through `forward_auth` as documented drop-in overlays; `ARCHILYZER_AUTH_MODE=none` is the one explicit escape hatch for people who already have their own front door. - **The GPU is usable from the container, including on AMD.** Alongside the default CPU whisper.cpp image there is a **Vulkan** target running parakeet.cpp — one build that covers AMD (RADV), Intel and NVIDIA, needing nothing on the host but a render node at `/dev/dri` and no vendor container toolkit — and a CUDA target for NVIDIA whisper. Measured on an RX 6600 XT, through the app's own overlapping-window wrapper: **3.4 s against 36.3 s** for the same 33-second clip pinned to the CPU. That gap is also the thing to watch for, because the failure here is silent — a Vulkan container with no `/dev/dri` does not error, it transcribes correctly on the CPU about ten times slower. The entrypoint prints which one it got on every boot, and the image ships `vulkaninfo` so you can ask directly. - **The editor can be started without resuming whatever it was in the middle of.** Booting arms the sync heartbeat, both auto-queue runners and the digest and backfill sweeps — right for a host install, where a restart interrupts work you own, and wrong the first time a container is pointed at a corpus somebody else configured: its stored policies may say *sweep*, and a corpus-wide digest sweep is GPU-**weeks** that would start seconds after `docker compose up`. `ARCHILYZER_IDLE_BOOT=1` starts the server with all of that stopped, and you can start any of it from the UI afterwards. The shutdown reaper and the persisted-pause restore stay armed either way, because both only ever *stop* work. +- **The digest and backfill lanes can do the newest uploads first too.** *Newest first* only ever existed on two of the four pipelines. The digest batch ordered a channel by **duration** — shortest first, which is the right default and turns a backlog into visible coverage fastest — and the backfill batch used whatever order a `readdir` happened to return, which is neither of the two orders anyone would choose. So on a channel with eleven thousand undigested videos, a video uploaded this morning sat behind every one of them, and there was no setting that changed that. Both lanes now take the same three-value **Order** the auto-queue runners already use: *Listed order* (exactly today's behaviour, and still the default), *Newest first*, *Oldest first*. It is opt-in, and with it left alone every list is byte-for-byte the list it was before — that is asserted, not assumed. Digest **composes** the two rules rather than choosing between them: videos are ordered by date first and by duration inside a date, so *newest first* reads as "newest day first, shortest video within that day" and the duration rule you already had is still doing its job. Measured on the live corpus, ordering a channel's 11,329 pending digests costs about a tenth of a second. - **Newest-first could quietly hand a video another channel's upload date.** The date lookup's first and cheapest layer scans the transcript index, which is keyed by the id the *platform* reports; every video the auto-queue actually asks about is named by its **folder**, and on this archive those two disagree for about one video in seven — a Rumble folder is named for the URL, while its recorded id is the embed id. Matched across the whole corpus, one channel's folder name could collide with a different channel's recorded id and inherit its date, which is a *wrong* answer rather than a missing one, and wrong dates are exactly what an ordering setting cannot survive. The lookup is now scoped to the channel the video belongs to, so a folder name that is not an id in its own channel falls through to reading the date out of the folder it actually names. Separately, the first pass over a very large backlog can only date so many videos at once; it now says so once in the log, because the remainder sorting to the back and being picked up on the next pass is the design, not a fault. - **The auto-queue can be told to do the newest uploads first, and it now genuinely does.** Both runners always took the first video off a rule's pile, and that pile's order came straight from the channel snapshot, which sorts most buckets **alphabetically by video id** — arbitrary for YouTube ids, and oldest-first for the date-prefixed folder names some sites use. So when a channel uploaded today, nothing made that video jump the nine-thousand-video backlog in front of it; the only reason auto-download roughly worked was that its one bucket happens to be left in playlist order. There is now an **Order** setting per runner — *Listed order* (what you have today, and still the default), *Newest first*, *Oldest first*. It sorts the videos **inside** each rule, across every channel and bucket that rule claims; the rule list still decides which rule goes first, because that is what the rule list is for. For a straight newest-first archive, use one catch-all rule. Worth knowing before you switch it on: under *Newest first* the retry and partial-download buckets lose their head start, so a half-finished download can end up waiting behind fresh work. The page says so next to the setting. - **Working out how recent 79,000 videos are turned out to be nearly free, once we stopped guessing where the dates were.** The obvious source — reading each video's metadata file — is about six and a half minutes and 41 GB of reading, on every scheduling decision, which is a non-starter. The transcript index already holds a date per video in a form that can be scanned without decoding anything: **78,583 videos in well under a second**. That covers the corpus, but it turned out **not** to cover the videos auto-transcribe actually queues, because the index only holds videos that already *have* a transcript and auto-transcribe's whole job is the ones that don't — of the 870 videos genuinely pending here, it knew the date of **122**. The gap is closed by reading the last 8 KB of each remaining video's metadata file, where the upload date happens to sit: **868 of the 870, at a fifth of a millisecond each**, and remembered afterwards so it is paid once rather than every few seconds. Videos not downloaded yet have no date anywhere on disk at all, so auto-download estimates one from the video's position in the channel's newest-first listing; those show with a `≈`, and a brand-new upload with nothing dated above it goes to the front, which is the entire point. Anything still undatable sorts to the back rather than disappearing, and a missing or busy index degrades to the old ordering instead of stopping the runner.