# settings.json keys Global operational settings shared by every site this editor powers, persisted to `settings.json` at the repo root (or `$SETTINGS_FILE`). Per-site presentation lives in `sites//site.json`. Every key is optional: a missing key reads as its default, an ill-typed one is coerced to its default or clamped, and an unknown one is dropped on the next save. Regenerate this file and `settings.json.example` with `pnpm --filter yt-dlp-transcript-common exec tsx bin/settings-example.ts`. `settings.json.example` is the default object with one key left out, `workers`: a file that does not name it gets a worker list synthesized from `transcriptionApp` on read. A file that spells `workers: []` READS as no transcription at all — until the next save, when the writer synthesizes a worker the same way. A copied example PINS every default it spells — including each lane's `autoQueue..held` — so a default changed in a later release will not reach that file. Delete any key you would rather have track the defaults. | Key | Default | |---|---| | [`adminTitle`](#admintitle) | `"Transcript Browser Admin"` | | [`maxTranscriptPageBytes`](#maxtranscriptpagebytes) | `8388608` | | [`transcriptionApp`](#transcriptionapp) | `"whisper-cpp"` | | [`transcriptionApps`](#transcriptionapps) | `{}` | | [`workers`](#workers) | `[]` | | [`cookiesFromBrowser`](#cookiesfrombrowser) | `""` | | [`cookieMode`](#cookiemode) | `"when-required"` | | [`social`](#social) | object — see below | | [`sleepBetweenDownloadsSeconds`](#sleepbetweendownloadsseconds) | `10` | | [`pacing`](#pacing) | object — see below | | [`downloadFormat`](#downloadformat) | `"auto"` | | [`sourceVideoQuality`](#sourcevideoquality) | `"original"` | | [`minFreeDiskGB`](#minfreediskgb) | `5` | | [`resumeMarginGB`](#resumemargingb) | `2` | | [`parallelTranscriptions`](#paralleltranscriptions) | `2` | | [`inlineTranscribeOnFallback`](#inlinetranscribeonfallback) | `false` | | [`skipLiveDownloads`](#skiplivedownloads) | `true` | | [`verifyAvailabilityBeforeClean`](#verifyavailabilitybeforeclean) | `true` | | [`buildArchives`](#buildarchives) | `true` | | [`archiveStorage`](#archivestorage) | object — see below | | [`reportDebouncePreset`](#reportdebouncepreset) | `"fast"` | | [`autoRefreshIntervalSeconds`](#autorefreshintervalseconds) | `5` | | [`syncScheduler`](#syncscheduler) | object — see below | | [`autoQueue`](#autoqueue) | object — see below | | [`channelPriority`](#channelpriority) | object — see below | | [`socialLinks`](#sociallinks) | `[]` | | [`homepageUrl`](#homepageurl) | `""` | | [`savedVideoBackup`](#savedvideobackup) | object — see below | | [`storage`](#storage) | object — see below | | [`buildPipeline`](#buildpipeline) | object — see below | | [`digest`](#digest) | object — see below | | [`diarization`](#diarization) | object — see below | | [`backfill`](#backfill) | object — see below | | [`attribution`](#attribution) | object — see below | | [`archiveOrg`](#archiveorg) | object — see below | | [`publish`](#publish) | object — see below | | [`seeder`](#seeder) | object — see below | ## `adminTitle` Title for the EDITOR admin shell only (the editor manages all sites and so is not tied to any one site's branding). Public sites get their own titles from site.json. Default: `"Transcript Browser Admin"` ## `maxTranscriptPageBytes` Target size (bytes) of one exported transcript page shard — the unit the export site fetches. Clamped into [TRANSCRIPT_PAGE_MIN_BYTES, TRANSCRIPT_PAGE_HARD_CAP_BYTES] (256 KiB – 20 MiB); default 8 MiB. Default: `8388608` ## `transcriptionApp` Active transcription app id (key into TRANSCRIPTION_APPS, e.g. "whisper-cpp" or "chough"). Selected globally; see common/lib/transcriptionApps.ts. Default: `"whisper-cpp"` ## `transcriptionApps` Per-app configuration, keyed by app id. Each app reads only its own block; a missing block means "use the app's defaults". DEPRECATED in favor of `workers` (each local worker carries its own config); kept one release to drive migration and allow rollback. See common/lib/workers.ts. #### `transcriptionApps.` Per entry — each entry spells its own values. | Key | Description | |---|---| | `bin` | Binary path/name override. Empty/undefined falls back to app.defaultBin(). | | `model` | whisper.cpp model path (substituted for {model}); for chough this is the optional CHOUGH_MODEL env (chough auto-downloads a model when unset). | | `remoteUrl` | chough remote server URL (CHOUGH_URL). Empty/undefined = local transcription. | | `chunkSize` | chough chunk size in seconds (-c). Undefined = chough's own default. | | `customArgs` | whisper.cpp custom argv template using the {audioFile}/{outputBase}/{model} placeholders. Undefined = DEFAULT_TRANSCRIBE_ARGS. | | `device` | parakeet compute device passed to parakeet-cli (--device / PARAKEET_DEVICE), e.g. "cuda:0", "cpu". Undefined = parakeet-cli's default device. | Default: ```json {} ``` ## `workers` Configured transcription workers (named processing slots). The scheduler distributes each video to the highest-priority free worker. A settings.json predating this field is migrated to a single enabled worker from the active app (see defaultWorkersFromApps). See common/lib/workers.ts. #### `workers[]` Per entry — each entry spells its own values. | Key | Description | |---|---| | `id` | Stable slug; used in settings, task ids, and logs. | | `name` | Human label shown in the UI. | | `kind` | "local" runs an app from TRANSCRIPTION_APPS on this machine (`appId` + `config`); "remote" delegates to another instance on the LAN (`remote`); "llm" is a bare ollama endpoint serving digest/attribution calls only (`llm`). | | `enabled` | Whether the scheduler may hand this slot work. Each worker is one slot, so parallelism is toggled per slot on the Workers page. Anything but an explicit `false` reads as enabled. | | `priority` | Lower = preferred. Ties broken by array order in the scheduler. | | `tags` | Capability routing. A tag is an OPERATION id from the backfill catalog ("diarization", "attribution-text", …) or a contended RESOURCE (WORKER_RESOURCE_TAGS). The scheduler consults them through workerMatches below: an untagged worker takes anything, a tagged worker takes only work whose requirement intersects its tags. Unknown tags are tolerated (they match nothing and warn in the settings UI), never fatal. | | `appId` | LOCAL: an instance of a TRANSCRIPTION_APPS entry + its per-worker config. | | `config` | LOCAL: the per-worker engine config (binary, model, device, …) — an AppInstanceConfig, see `transcriptionApps.`. | | `remote` | REMOTE: how to reach the delegate instance. | | `llm` | LLM: how to reach the bare model endpoint. | #### `workers[].config` Per entry — each entry spells its own values. | Key | Description | |---|---| | `bin` | Binary path/name override. Empty/undefined falls back to app.defaultBin(). | | `model` | whisper.cpp model path (substituted for {model}); for chough this is the optional CHOUGH_MODEL env (chough auto-downloads a model when unset). | | `remoteUrl` | chough remote server URL (CHOUGH_URL). Empty/undefined = local transcription. | | `chunkSize` | chough chunk size in seconds (-c). Undefined = chough's own default. | | `customArgs` | whisper.cpp custom argv template using the {audioFile}/{outputBase}/{model} placeholders. Undefined = DEFAULT_TRANSCRIBE_ARGS. | | `device` | parakeet compute device passed to parakeet-cli (--device / PARAKEET_DEVICE), e.g. "cuda:0", "cpu". Undefined = parakeet-cli's default device. | #### `workers[].remote` Per entry — each entry spells its own values. | Key | Description | |---|---| | `baseUrl` | e.g. http://gpu-box.lan:3011 | | `token` | Outbound bearer token sent with every /api/worker request to this remote. The accepting side validates against its own WORKER_TOKEN env, never this. | | `sharedFs` | When true the remote shares the transcripts mount, so we send {channelSlug, videoId} instead of uploading the audio bytes. | | `slots` | How many units this remote takes in parallel. The pool expands one remote config into this many independently-schedulable slot entries at reconfigure time (the defaultWorkersFromApps trick, applied live). Absent = probed from the remote's /api/worker/health (its enabled worker count) — see controller/remoteCapacity.ts; 1 until the probe answers. | #### `workers[].llm` Per entry — each entry spells its own values. | Key | Description | |---|---| | `baseUrl` | e.g. http://macbook.lan:11434 | | `slots` | Concurrent generations to allow this endpoint. Defaults to 1 — one model instance, one generation — unless the operator knows better. | Default: ```json [] ``` ## `cookiesFromBrowser` Browser spec (e.g. "firefox", "chrome:Default") passed to `yt-dlp --cookies-from-browser`. WHEN it is passed is governed by `cookieMode` below. Empty string = no cookies configured. Per-channel override available (ChannelConfig.cookiesFromBrowser). Default: `""` ## `cookieMode` How yt-dlp invocations use the configured cookies (see common/lib/cookiePolicy.ts): "always" passes them on every invocation, "when-required" (default; the historical behavior) only to retry an auth/age failure, "defer" never in normal runs — auth-gated videos are excluded from batches and collected into the per-channel "Needs cookies" bucket for a manual cookie run. Per-channel override available (ChannelConfig.cookieMode). Default: `"when-required"` ## `social` Per-platform settings of the social posts. Today two keys, both X's, both chosen in the X account session section of /settings: where the X fetchers' login comes from (`social.x.cookieSource`) and where X posts may appear (`social.x.visibility`). See common/social/xCookieSource.ts. #### `social` | Key | Default | Description | |---|---|---| | `x` | `{}` | X / Twitter. See `social.x` below. | #### `social.x` | Key | Default | Description | |---|---|---| | `cookieSource` | absent | Where the X fetchers' login comes from. `"browser"`: the operator's everyday browser, named by `cookiesFromBrowser` — gallery-dl is handed `--cookies-from-browser ` and reads it on every run, and the Playwright fallback reads the same store (Firefox only; common/social/xBrowserLogin.ts), so the login lasts as long as the browser's. `"profile"`: the session broker's persistent profile ("Connect X account" on /settings) and the cookie jar it exports. ABSENT (the default) is resolved at read time, never stored: `"browser"` when `cookiesFromBrowser` is set and no profile is connected (no exported jar carrying an auth_token), else `"profile"`. The source, not `cookieMode`, governs the X fetchers. | | `visibility` | absent | Where X posts may appear. `"public"` (the default; absent): an X channel's posts are built into every site that has the channel. `"private"`: every X channel's posts (a channel with `sourceKind: "social"` and `platform: "twitter"`) are left out of every PUBLIC site build — the channel with them, since posts are all an X channel holds — and built only into PRIVATE sites (`site.json` `audience`). Nothing on disk changes and fetching does not; a site already published changes on its next build and deploy, and flipping back is a rebuild. Chosen in the X account session section of /settings; the rule is common/lib/postsVisibility.ts. | Default: ```json { "x": {} } ``` ## `sleepBetweenDownloadsSeconds` Pause (seconds) inserted between per-video yt-dlp invocations in managed batch downloads, and between two auto-download units on one platform (release 17; the lane ignored it before). yt-dlp's own `-t sleep` only paces requests within one invocation, so without this the managed loop hammers the source IP back-to-back. The adaptive pace above its base (see `pacing`) is added to it. 0 disables. Per-channel override available for batch downloads. Default: `10` ## `pacing` How the download pace adapts to rate limits, per platform (release 17). Every yt-dlp spawn against a platform paces its requests (`--sleep-requests`) at the platform's adaptive pace, and the download lane waits sleepBetweenDownloadsSeconds plus the pace above its fixed value between units. A rate limit doubles the pace; clean units ease it back; a rate limit that outlasts the cooldown cap holds the platform to one probe at a time. The live pace, cooldowns and holds are in `.auto-queue/state.json`, shown on /operations/download. See common/jobs/platformBackoff.ts. #### `pacing` | Key | Default | Description | |---|---|---| | `sleepRequestsCapSeconds` | `16` | Ceiling (seconds) on a platform's adaptive `--sleep-requests`. The pace starts at the platform's fixed value (1 s for YouTube and Rumble) and doubles on every platform-level rate limit until it reaches this. A subtitle-only 429 never moves it. Clamped 1–120; default 16. | | `decayAfterCleanUnits` | `5` | How many clean auto-download units on a platform halve its pace one step back toward the fixed value. Clamped 1–1000; default 5. | | `holdAfterFailsAtCap` | `3` | How many consecutive failures AT the 30-minute cooldown cap put a platform in a hold: the lane then runs one probe unit per holdProbeMinutes instead of one per cooldown, and a manual Sync or download on it is refused with the next probe's time. A clean probe clears the hold and the backoff. Clamped 1–100; default 3. | | `holdProbeMinutes` | `60` | Minutes between probes while a platform is held. Clamped 1–1440; default 60. | Default: ```json { "sleepRequestsCapSeconds": 16, "decayAfterCleanUnits": 5, "holdAfterFailsAtCap": 3, "holdProbeMinutes": 60 } ``` ## `downloadFormat` Default yt-dlp `-f` download format for every channel that doesn't set its own (ChannelConfig.downloadFormat). "auto" picks per-source: `original` for Odysee (whose HLS rungs are CDN-truncated), `bestaudio/worst` elsewhere. See common/ytdlp/downloadFormat.ts. Default: `"auto"` ## `sourceVideoQuality` Quality of the source container a full persist keeps — "Persist source video", the whole-recording fetch (`full: true`) and "Persist kept now" — for every channel that doesn't set its own (ChannelConfig.sourceVideoQuality). "original" (default) = `bestvideo*+bestaudio/best`; "video_720" = ≤720p H.264/AAC mp4 for clip and editing work, falling back to 480p and then to anything (logged). A single persist can override it. See common/ytdlp/downloadFormat.ts. Default: `"original"` ## `minFreeDiskGB` Minimum free disk space (GB) required on the transcripts data directory for downloads to run. When free space is below this floor, a download job is prevented from starting and a running batch stops launching new videos (the in-flight one finishes). 0 disables the gate. See common/lib/diskSpace.ts. Default: `5` ## `resumeMarginGB` Extra headroom (GB) above minFreeDiskGB that a stopped pipeline must see before it resumes. Resuming at the same number we stopped at flaps — the first restarted download pushes free space back under the floor. This is the hysteresis margin, so "resumed" means the operator actually freed something rather than a scratch file being cleaned up. 0 disables the hysteresis (resume at the floor). See diskGate() in common/lib/diskSpace.ts. Default: `2` ## `parallelTranscriptions` Default number of videos transcribed in parallel when a "Transcribe missing" / bucket run doesn't specify its own concurrency. The per-run Concurrency input in the channel UI overrides this for a single run. Default: `2` ## `inlineTranscribeOnFallback` When true, the no-subs fallback in the managed downloader runs whisper inline immediately after the audio download succeeds. When false (default), audio is left for the next "Transcribe missing" pass so a batch download finishes faster and whisper can parallelize. Default: `false` ## `skipLiveDownloads` When true (default), managed downloads skip videos that are currently live or scheduled/upcoming, decided from a metadata-only prefetch pass. Finished livestream VODs (was_live) are NOT skipped and download normally. A skip is recorded but not archived, so the next sync/download-missing retries the video once the stream ends. Per-channel override available (ChannelConfig.skipLiveDownloads). Default: `true` ## `verifyAvailabilityBeforeClean` Whether the transcribed-audio cleanup sweep checks each candidate is still available upstream before deleting its audio, pinning (do-not-clean) any video found permanently gone. The delete is irreversible and a gone video's audio is irreplaceable, so this defaults to true. Turn it off for an offline or URL-less setup, where the check can never resolve and cleanup would otherwise never delete anything. See verifyBeforeClean.ts. Default: `true` ## `buildArchives` Whether site builds generate downloadable transcript/live-chat archive zips (into public/archives, linked on the Downloads page). Global default; a site can opt out via site.json `archives: false`, and a single build can skip via the "Skip archive zips" build control. Opt-out: default true. Default: `true` ## `archiveStorage` Overflow object storage (Cloudflare R2) for archive zips that exceed the Pages per-file size cap (see Site.archiveMaxBytes). When both fields are set, an oversize archive is uploaded here on deploy — via `wrangler r2 object put`, keyed `/archives/` — instead of being dropped, and the Downloads page links to `/`. Blank/absent → no overflow, so oversize archives stay unavailable ("Too large to host"). #### `archiveStorage` | Key | Default | Description | |---|---|---| | `bucket` | `""` | Cloudflare R2 bucket an oversize archive zip is uploaded to on deploy (`wrangler r2 object put`, keyed `/archives/`). Blank = no overflow. | | `publicBaseUrl` | `""` | Public base URL of that bucket; the Downloads page links `/`. Both fields must be set for overflow to happen. | Default: ```json { "bucket": "", "publicBaseUrl": "" } ``` ## `reportDebouncePreset` Debounce preset for the global snapshot scheduler: how long it waits after the last report-changing action before regenerating affected channel reports. See REPORT_DEBOUNCE_PRESETS. Default "fast" (~1s, no cap). Default: `"fast"` ## `autoRefreshIntervalSeconds` How often (seconds) the editor UI passively re-fetches the current page's server-rendered data via router.refresh(), so sidebar badges and reports stay live without a manual reload. Mounted globally; pauses while the tab is hidden. 0 disables passive refresh entirely. See AUTO_REFRESH_INTERVAL_*. Default: `5` ## `syncScheduler` Global configuration for the scheduled (cron-driven) channel sync system. The per-channel cadence lives on ChannelConfig.syncIntervalMinutes; this block holds the defaults and guard rails the scheduler applies across all channels. See common/jobs/syncScheduler.ts. #### `syncScheduler` | Key | Default | Description | |---|---|---| | `enabled` | `false` | Master switch. When false, a tick selects nothing (manual sync still works). | | `defaultIntervalMinutes` | `1440` | Fallback cadence (minutes) for channels with no per-channel override. | | `maxConcurrentSyncs` | `2` | Cap on sync jobs running/queued at once. A tick queues at most (cap - currently-active) channels; the rest roll to the next tick. This is also the stagger mechanism that keeps a big due-batch from hitting the source all at once. | | `quietHoursStart` | `null` | Optional local-clock quiet window during which auto-sync is suppressed. Both null = always allowed. The window may wrap past midnight (e.g. start=22, end=6). Hours are [0,23]; the window is [start, end). | | `quietHoursEnd` | `null` | End hour of the quiet window, [0,23], exclusive. See `quietHoursStart`: both must be valid hours or the window is cleared (null = always allowed). | | `backoffBaseMinutes` | `30` | Failure backoff bounds. After N consecutive failed scheduled syncs a channel waits min(base * 2^(N-1), max) minutes before it's eligible again. | | `backoffMaxMinutes` | `1440` | Ceiling on the failure backoff (see `backoffBaseMinutes`): a channel waits min(base * 2^(N-1), max) minutes after N consecutive failures. Never below the base. | | `heartbeatSeconds` | `0` | Cadence (seconds) for the editor's in-process heartbeat — the internal timer armed by the instrumentation hook (editor/instrumentation.ts) that calls the scheduler tick directly, so no external cron is needed. 0 = off: rely on the external `pnpm sync:tick` heartbeat instead. Any positive value is clamped to [SYNC_HEARTBEAT_MIN_SECONDS, SYNC_HEARTBEAT_MAX_SECONDS]. The env var SYNC_HEARTBEAT_SECONDS overrides this at runtime. See SCHEDULED_SYNC.md. | | `keepLatestCheckIntervalMinutes` | `1440` | Cadence (minutes) for the scheduled keep-latest deletion check. For each channel with ChannelConfig.keepLatest > 0, the tick re-probes the kept window for source deletion (checkKeptDeletedAction) at most this often and pins any gone videos as do-not-clean. Clamped into the sync-interval window; default daily. The check shares the same concurrency cap and quiet-hours window as scheduled syncs. See editor/app/scheduler/runTick.ts. | | `fullSweepIntervalMinutes` | `1440` | Default cadence (minutes) for the sync FULL SWEEP — the deep pass that re-enumerates a channel's whole listing in one yt-dlp spawn, refreshes the stored `playlist` file, and flags videos that have left the listing into maybe-missing.json. Ordinary syncs stay on the cheap newest-first paged walk; a sync only upgrades itself to a sweep when this interval has elapsed since the channel's lastFullSweepAt. Per-channel override: ChannelConfig.fullSweepIntervalMinutes. 0 = never sweep. Default daily. See common/jobs/deepSync.ts. | | `fullSweepConfirmMaxSuspects` | `25` | Upper bound on how many maybe-missing suspects a full sweep will resolve in-line with the per-video availability probe (deleted vs private vs unlisted). At or under the cap the sweep runs the targeted check itself, so "Sync all" surfaces upstream deletions with no extra clicks; over it, the suspects are flagged and left for a manual check rather than firing hundreds of probes inside a sync. 0 = never auto-confirm. | | `fullSweepShrinkGuardPercent` | `10` | Shrink guard: how far a fresh listing may fall below the stored one before it is treated as suspect rather than acted on. Expressed as a percentage of the previous count, floored at SHRINK_ABS_FLOOR entries so ordinary churn on a small channel doesn't trip it. A suspect listing does not rewrite `playlist` or maybe-missing.json and does not count as a sweep — but a SECOND enumeration reporting a similar count confirms it and is accepted, so a genuine mass deletion costs at most one cadence period. 0 = off (the empty-listing rejection still applies). See controller/acceptListing.ts. | Default: ```json { "enabled": false, "defaultIntervalMinutes": 1440, "maxConcurrentSyncs": 2, "quietHoursStart": null, "quietHoursEnd": null, "backoffBaseMinutes": 30, "backoffMaxMinutes": 1440, "heartbeatSeconds": 0, "keepLatestCheckIntervalMinutes": 1440, "fullSweepIntervalMinutes": 1440, "fullSweepConfirmMaxSuspects": 25, "fullSweepShrinkGuardPercent": 10 } ``` ## `autoQueue` Configuration for the automatic priority-queue runners (auto-transcribe / auto-download). Each holds a tree policy that decides which channel's video to process next, cross-channel, by priority/round-robin/weighted-fair rules. Independent of syncScheduler (which decides staleness, not work order). See common/jobs/autoQueuePolicy.ts. #### `autoQueue.` | Key | Default | Description | |---|---|---| | `enabled` | `false` | Master switch for this runner (transcription / download independently). | | `maxWorkers` | `null` | Overall ceiling on concurrent in-flight workers for this runner. null = no runner-level cap (the worker pool / platform queues are the real throttle). | | `replaceAutoSubs` | `false` | Opt in to the lowest-priority "replace YouTube auto-captions" lane: append this kind's opt-in buckets (autoSubsOnly / downloadedAutoSubsOnly) to the tail of the default union, so videos whose only transcript is YouTube ASR get re-done with our own engine whenever nothing more important is pending. Default false — the corpus-wide cost is large (an audio download plus a transcription per video). A leaf can also target the bucket by name for per-channel opt-in without flipping this switch. Optional: settings written before this field existed lack it; the sanitizer defaults it to false. | | `order` | transcription `"listed"`
download `"listed"`
digest `"cheapest"`
backfill `"listed"` | Ordering within each rule (see AutoQueueOrder). Optional exactly like replaceAutoSubs: settings files written before this field existed lack it, and the sanitizer defaults them to "listed" (today's behaviour). | | `snoozeUntil` | `null` | Epoch ms until which this runner idles WITHOUT stopping: next() returns null so the loop stays up, re-reads settings each iteration, and resumes by itself when the moment passes. null/absent/past = not snoozed. Survives a restart because it lives in settings.json, not in runner memory. | | `held` | transcription `false`
download `false`
digest `false`
backfill `true` | THE LANE'S PAUSE GATE. Shut means the lane holds: every dispatch path asks lib/pauseGates.ts, whose limit()/guard returns 0 so runPool idle-waits. A hold, never a stop — see that file's header.

OPTIONAL IN THE TYPE, FILLED BY THE SANITIZER. Until slice 1.4 four separate settings fields carried this — `transcriptionsPaused`, `downloadsPaused`, `digest.digestsPaused` and (inverted) `backfill.enabled` — so `undefined` meant "ask the legacy field" and `sanitizePolicy` deliberately refused to default it: a default would have read a paused corpus as running. S0-pause deleted those four, on the precondition that the live settings.json already carried every `held` key, and the default came in with them (`defaultHeldFor` — free everywhere except backfill, whose field was inverted and shipped held).

It stays optional because a reader may be handed a PARTIAL settings object (laneGuards.test.ts casts one), and `isGateHeld` answers `false` for a lane that carries no key at all rather than throwing. | | `root` | object — see below | The lane's rule tree: a group whose children are groups and leaves (see the node table). A missing root is the lane's default — empty for the runner lanes, one catch-all leaf for digest and backfill. While a channel-priority document exists, the four roots are compiled from it and not hand-edited. | #### `autoQueue..root (tree nodes)` Per entry — each entry spells its own values. | Key | Description | |---|---| | `id` | Stable node id, unique within the lane's tree. Preserved on save when valid and not taken, so persisted fairness state survives an unrelated edit; a missing or duplicate id is replaced with a generated one. | | `match` | LEAF ONLY: which videos this leaf owns (see the match table). | | `weight` | Relative share under a weighted-fair parent. Default 1. Ignored otherwise. | | `maxWorkers` | Optional ceiling on concurrent in-flight workers drawn from this node (and, for a group, its whole subtree). A capped node reads as "no work" and the parent falls through to the next sibling, like an HTB class ceiling. null = no cap. | | `mode` | GROUP ONLY: how the children compete — "strict" (first child with work wins), "round-robin", or "weighted-fair" (by each child's `weight`). Unknown values read as "strict". | | `children` | GROUP ONLY: the child nodes, in priority order for a strict group. A node with a `children` array is a group; any other node is a leaf. | #### `autoQueue..root … .match` Per entry — each entry spells its own values. | Key | Description | |---|---| | `type` | What the leaf matches: "channel" (one channel slug in `value`), "platform" (a platform name in `value`), or "all". | | `value` | Channel slug (type=channel) or platform name (type=platform). Ignored for type=all. A type=channel leaf with no value matches nothing. | | `bucket` | Optional snapshot bucket this leaf draws from, narrowing the default for the runner kind (transcription → downloadedNoTranscript, download → undownloadedIds). E.g. bucket="failedListed" prioritizes retries. | | `operation` | Optional OPERATION this leaf draws from — a registered backfill kind id, or "digest". Same meaning as `bucket` one level up: it narrows what the leaf claims, and it draws from ChannelWork.operations rather than ChannelWork.buckets.

It lives on the MATCH, beside `bucket`, and not on the node. A field on the node would need group inheritance — "this group is the digest subtree" — and inheritance is resolution logic buildPendingByLeaf does not have. Here it needs exactly one sanitizer and exactly one claiming path.

It is a SEPARATE id space from `bucket`, and the sanitizer enforces that a leaf names at most one of the two (operation wins): `defaultBuckets` is a priority-ordered union, so a name that meant a bucket to one leaf and an operation to another would silently mix two id spaces, and selectableBucketsForKind feeds the editor's bucket dropdown, where an operation must not appear as a bucket. | Default: ```json { "transcription": { "enabled": false, "maxWorkers": null, "replaceAutoSubs": false, "order": "listed", "snoozeUntil": null, "held": false, "root": { "id": "root", "mode": "strict", "weight": 1, "maxWorkers": null, "children": [] } }, "download": { "enabled": false, "maxWorkers": null, "replaceAutoSubs": false, "order": "listed", "snoozeUntil": null, "held": false, "root": { "id": "root", "mode": "strict", "weight": 1, "maxWorkers": null, "children": [] } }, "digest": { "enabled": false, "maxWorkers": null, "replaceAutoSubs": false, "order": "cheapest", "snoozeUntil": null, "held": false, "root": { "id": "root", "mode": "strict", "weight": 1, "maxWorkers": null, "children": [ { "id": "all", "match": { "type": "all" }, "weight": 1, "maxWorkers": null } ] } }, "backfill": { "enabled": false, "maxWorkers": null, "replaceAutoSubs": false, "order": "listed", "snoozeUntil": null, "held": true, "root": { "id": "root", "mode": "strict", "weight": 1, "maxWorkers": null, "children": [ { "id": "all", "match": { "type": "all" }, "weight": 1, "maxWorkers": null } ] } } } ``` ## `channelPriority` THE OPERATOR-FACING PRIORITY MODEL: one tier per channel plus one corpus-wide focus selector. It is the SOURCE the four `autoQueue[lane].root` trees are compiled from (common/lib/channelPriority.ts), not a second mechanism beside them — and its `paused` tier is the one part that is not a tree shape, filtering the runner's channel list instead. An empty document (the default) is today's behaviour exactly: no focus, every channel normal, the stored trees stand. #### `channelPriority` | Key | Default | Description | |---|---|---| | `focus` | object — see below | The corpus-wide focus selector: none, one site's channels, or a list of channels. A focus is compiled into a leading `prio-focus` group in every lane's tree. | | `channels` | `{}` | ONLY channels that differ from the default appear. An absent slug is `normal`, unranked — so the default document is empty and "absent document = today's behaviour" holds byte for byte. | #### `channelPriority.focus` Per entry — each entry spells its own values. | Key | Description | |---|---| | `kind` | "none" (no focus), "site" (the channels of one site, resolved at compile time so it tracks membership) or "channels" (an explicit list, from "Focus these"). | | `siteId` | kind "site" only: the site whose channels are focused. A blank id reads as no focus; an unknown one survives and focuses nothing. | | `slugs` | kind "channels" only: the focused channel slugs, trimmed and de-duplicated. An empty list reads as no focus. | #### `channelPriority.channels.` Per entry — each entry spells its own values. | Key | Description | |---|---| | `tier` | THE BASE TIER: what every operation gets unless an override says otherwise. | | `rank` | Order WITHIN the tier, ascending. Absent = unranked, which sorts after every ranked sibling and then by slug. ONE rank per channel, not one per lane — the two hand-made lane orders collapse into this on migration. | | `overrides` | PER-OPERATION OVERRIDES of the base tier. Only operations that DIFFER from the base appear: the sanitizer normalises an override equal to `tier` away, so the on-disk document stays a list of exceptions to a list of exceptions.

`{tier:"normal", overrides:{sync:"paused"}}` is "everything but sync" — the lossless reading of the retired `excludeFromSync`. Its inverse, `{tier:"paused", overrides:{sync:"normal"}}`, is "sync only": keep the playlist and metadata current, dispatch nothing. | | `autoPaused` | PAUSED BY THE MACHINE, NOT BY THE OPERATOR, and what to put back.

Set when the drive a channel's media is on stops being there: the watch pass records the tier the channel HAD and forces `paused`, so nothing in any lane dispatches against a `data/` nobody can read. Cleared — and the tier restored — when the drive comes back.

WHY IT IS A FIELD AND NOT A DERIVED STATE. The lanes read `tier`; making them all ask a second question would be four more places to forget. And the tier the channel is to be RESTORED to is not derivable from anything once it has been overwritten — that is the whole content of this field.

OPTIONAL, and an older binary that drops it leaves the channel Paused with nothing lost but the automatic restore. The operator's own word always wins: a MANUAL tier change clears it (see clearAutoPause), so a drive coming back can never un-pause a channel somebody paused on purpose. | #### `channelPriority.channels..autoPaused` Per entry — each entry spells its own values. | Key | Description | |---|---| | `reason` | One reason today. A union so a second one has somewhere to go, and so a surface can say WHICH machine decided rather than "automatic". | | `since` | ISO, for "auto-paused — media unreachable since ". | | `previousTier` | The base tier the channel had before the machine paused it; what a restore puts back. Never `paused` (that would restore to paused — a no-op dressed as a restore). | | `cause` | `not-there` (the drive is unmounted or unplugged) or `not-answering` (it is there and does not answer: a stalled disk). Only what the words on /review, the rack and the channel page say. Optional: a record written before it existed reads as `not-there`. | Default: ```json { "focus": { "kind": "none" }, "channels": {} } ``` ## `socialLinks` Default social links applied to every site that doesn't define its own. A site inherits these unless its site.json carries an explicit `socialLinks` array — see Site.socialLinks / resolveSocialLinks in common/lib/site.ts. The one presentation field that lives globally so a shared footer doesn't have to be repeated per site. #### `socialLinks[]` Per entry — each entry spells its own values. | Key | Description | |---|---| | `label` | The link's name: the icon's accessible name and its tooltip, never text beside it. Shown as text only in place of an icon that fails the check at render. | | `url` | Link target: http(s), mailto: or a site-relative path. | | `svg` | Inline SVG markup: ONE well-formed `` element, checked when it is saved new or edited and again every time it is rendered (a link whose icon fails at render shows its label instead; `archilyzer doctor` names it). It may contain shapes, groups, defs, gradients, patterns, clip paths, masks, filters, text and animate/animateTransform/set — no script, style block, foreignObject, a, image, title, desc or any HTML element (a title or desc holding text only is removed); SVG presentation attributes plus aria-*, data-* and xmlns:* — no event handler (on…); a `style` attribute of presentation properties only; an href or url(…) only to an id inside the icon, written plainly; no CSS escape, comment or function that loads anything (image-set, image, cross-fade, element, src, paint, @import); ids plain names. A leading XML declaration, a DOCTYPE without an internal subset and comments are removed. Normalized on save: width/height stripped, aria-hidden added, a single-colour icon's fills made fill="currentColor" (an icon of two or more colours keeps them). A root with no viewBox but a numeric width W and height H (unitless or px) is given `viewBox="0 0 W H"`, so a file pasted as downloaded is accepted. A refused save names why; export from a drawing program with presentation attributes rather than a style block (in Inkscape, save as Plain SVG). | | `featured` | Keep this link in the header on small screens (the editor's "Keep in header on small screens"). A narrow header shows only the featured links (up to 4, the last 4 if more are marked; none marked → none, so the name has the room); a wide header shows every link, up to 4, the featured ones kept first, then the last of the rest. The footer shows every link. Written only when true. | Default: ```json [] ``` ## `homepageUrl` Absolute public URL of the family hub (e.g. "https://archilyzer-hub.pages.dev"). The default for every site's `hubUrl` (a site's own wins): published as `hubUrl` in the site's public `/site.json` and `/corpus.json`, so the hub can tell its member sites from arbitrary added origins. No page links to it (the header's Hub link was removed in release 14). Empty = none published. Normalized to a trailing-slash-free http(s) URL. Default: `""` ## `savedVideoBackup` Backup configuration for the saved-video store (Phase 4 of the video-persistence feature). When enabled with a destination, the store is mirrored there (additively, no deletes) with a per-backup manifest, and the sync scheduler runs the backup on the configured cadence. See common/controller/backupSavedVideos.ts. #### `savedVideoBackup` | Key | Default | Description | |---|---|---| | `enabled` | `false` | Master switch for the scheduled backup. A backup can still be run manually when this is false, as long as a destination is set. | | `dest` | `""` | Destination root the store is mirrored into (a local path or any rsync target). Empty disables both scheduled and manual backups. | | `intervalMinutes` | `1440` | Cadence (minutes) for the scheduled backup when enabled. Clamped into the sync-interval window; default daily. | Default: ```json { "enabled": false, "dest": "", "intervalMinutes": 1440 } ``` ## `storage` Where a channel's downloaded media goes when it is relocated off the corpus disk. A DEFAULT ONLY: the relocate controller never reads it and always takes an explicit root, so this is the value the per-channel Storage panel prefills and the /channels bulk move falls back to. Blank = no default. See StorageSettings. #### `storage` | Key | Default | Description | |---|---|---| | `locations` | `[]` | The named storage locations a channel's media may be relocated to — one entry per root, each with an id, label, root, `autoRepoint` and the learned volume identity. Order is display order. Managed on /storage. | | `defaultLocationId` | `""` | The location prefilled as the destination of a move. "" = no default. | | `savedVideosLocationId` | absent | WHERE THE SAVED-VIDEO STORE IS, by location id. "" = in place, under the corpus at `paths.savedVideosDir`.

A RECORD OF WHAT IS ON DISK, never an intention — the same contract as a channel's `config.mediaDir`. It is written by the move, on success, after the copy has verified and the symlink is in place; nothing else writes it, and a reader that disagrees with the disk trusts the disk. Optional so an older settings.json parses (and an older binary that drops it leaves a store that still works, because the symlink is what every reader follows). | | `health` | absent | THE DRIVE-HEALTH TIMINGS: how long a read may take before a drive counts as not answering, how often the health pass looks, how long its look may take, how many clean looks clear a stall, and how many reads may be on one drive at once. Edited on /storage (Drive health timing). Absent = every default, and only a value that differs from its default is written, so an untuned install follows a default changed later. See `storage.health` below. | #### `storage.locations[]` Per entry — each entry spells its own values. | Key | Description | |---|---| | `id` | /^[a-z0-9][a-z0-9-]{0,63}$/, unique within the list. Stable: it is what `defaultLocationId` and every form and action refer to. | | `label` | Human name. Blank sanitizes to the id. | | `root` | Absolute directory, trailing "/" stripped. NEVER existence-checked on read — the whole point of a cold location is a drive that may not be mounted when settings are parsed. | | `autoRepoint` | Opt-in: when the volume is found mounted somewhere else, re-point without asking (if the preflight passes). Off by default — re-point rewrites every channel symlink on the location, and that is not something to do silently unless the operator asked for it. | | `volume` | Identity learned at the last successful probe. Optional because a location may never have been probed, and because in a container block devices are invisible and identity is permanently unknown. | #### `storage.locations[].volume` Per entry — each entry spells its own values. | Key | Description | |---|---| | `uuid` | Filesystem UUID, the one stable name a disk has across mountpoints. This is what makes "the platter came up somewhere else" a recoverable situation. | | `fstype` | Filesystem type reported by the probe (e.g. "ext4"). Informational; omitted when unknown. | | `label` | Filesystem label reported by the probe. Informational; omitted when unknown. | | `mountpoint` | Where the volume was mounted at the last successful probe, and the path of the location's root RELATIVE to that mountpoint. Invariant: `root === join(mountpoint, relPath)`. Keeping the two halves is what lets a probe compute a candidate root when the volume reappears elsewhere. | | `relPath` | The location root's path RELATIVE to `mountpoint` (see there). Invariant: `root === join(mountpoint, relPath)`. | #### `storage.health` | Key | Default | Description | |---|---|---| | `budgetMs` | `3000` | How long one read may take before the drive counts as not answering, in ms (default 3000, 500–60000). The watchdog's budget (`onDrive`, lib/storageHealth.ts) for one unit of work — a video directory's reads, a page's reads of one video: a unit that has not answered by then is refused, and marks its location not answering unless the disk's request counters show it still completing others (slow, not stalled). A read waiting for a slot is refused when nothing on the drive has returned for this long plus a quarter of it (at most 250 ms). Takes effect on the next read after a save. | | `passIntervalMs` | `15000` | How often the health pass reads each location's disk counters, in ms (default 15000, 5000–300000). A stall that starts between two passes is seen by the next, or at once by a page's read. A save on /storage re-arms the pass's timer at once; a hand edit, at the next pass. Two counter samples are compared only when at least min(10 s, this − 5 s) apart, a spacing never less than half of this. | | `probeTimeoutMs` | `3000` | How long the health pass waits, in ms (default 3000, 500–30000), for the child `stat` of a root where no disk can be named (a timeout counts as not answering) and for the `findmnt` that names a root's disk (a timeout names none that pass). Takes effect on the next pass. | | `clearAfterCleanPasses` | `2` | How many clean answers in a row clear a location marked not answering (default 2, 1–10). Each health pass is one answer, and so is a Refresh on /storage; a miss in between starts the count again. Takes effect on the next answer. | | `inFlightPerLocation` | `4` | How many reads through the watchdog may be on one location's drive at once (default 4, 1–8); the rest wait in the editor's own queue, so a stall mid-walk holds this many of Node's threads, not all of them. At most 8, half of `UV_THREADPOOL_SIZE` (16 in the editor's start script and the container), so one drive that stops answering cannot hold every thread. Takes effect on the next read. | Default: ```json { "locations": [], "defaultLocationId": "" } ``` ## `buildPipeline` The docker build runner's settings. A site's build is a publish stage (release 18): by default (`publish.runner: "local"`) every site builds in turn as a child of the editor, one stage at a time on the `publish` queue. With `publish.runner: "docker"` (or `archilyzer publish build all --runner docker`) every stale site builds in its own container — Dockerfile.build, the image and maxParallelBuilds below — on a Linux host whose container engine answers; it is refused inside a container. Deploys are their own stages. See PUBLISH.md. There is no mode switch; a `mode` key left in an older file is dropped on the next save. #### `buildPipeline` | Key | Default | Description | |---|---|---| | `maxParallelBuilds` | `2` | Cap on concurrent per-site container builds when the docker build runner builds every site (`publish.runner: "docker"`, or `archilyzer publish build all --runner docker`). Clamped to [1, BUILD_MAX_PARALLEL_MAX]. | | `dockerImage` | `"yt-dlp-transcript-browser-build"` | Tag of the reusable build image (built once, reused for every site). | | `dockerfile` | `"Dockerfile.build"` | Dockerfile path relative to the monorepo root, used to (re)build the image. | Default: ```json { "maxParallelBuilds": 2, "dockerImage": "yt-dlp-transcript-browser-build", "dockerfile": "Dockerfile.build" } ``` ## `digest` AI digest generation (chapters + topic tags over the existing transcripts). Local-first: the metered lane is off by default. See DigestSettings. #### `digest` | Key | Default | Description | |---|---|---| | `remoteEnabled` | `false` | Master switch for the metered (remote-api) lane. OFF by default — an opt-in overflow for the long tail or a channel where local quality is poor, never the default path. | | `longTailSeconds` | `14400` | Videos longer than this are "long tail": 8.2% of the corpus by count, 46% of all transcript tokens. The batch's duration-aware ordering and the optional remote overflow both key off it. | | `localAppId` | `"ollama-direct"` | The engine each lane uses (ids from common/lib/digestApps.ts). | | `remoteAppId` | `"claude-code"` | The engine the metered (remote-api) lane uses — an id from common/lib/digestApps.ts. Unknown ids degrade to the default app rather than failing. | | `apps` | `{}` | Per-app config, keyed by app id — the same id-keyed sub-record shape as transcriptionApps. | | `yieldToTranscription` | `true` | Yield the GPU to the transcription lane: while transcription is working, the digest batch's limit() returns 0 and the pool idle-waits. ON by default, because `digest:local` is deliberately on a different queue from TRANSCRIPTION_QUEUE and so would otherwise run ollama and the transcription engine on the same 8 GB card. See controller/digestYield.ts. | | `yieldToCpuWorkers` | `false` | Whether a busy worker pinned to `device: "cpu"` counts as GPU contention.

OFF by default, which is the FIX for a real bug: the yield originally tested only `kind === "local"`, so on a box with one GPU worker and two CPU-pinned ones (this box, at parallelTranscriptions 2) the digest lane stopped dead for transcription that competes for zero GPU shaders.

Only an EXPLICIT "cpu" is treated as non-contending. A worker with no device set is using the engine binary's own default, which may be the GPU, so it still triggers the yield — the unknown case fails safe.

Composes with `yieldToTranscription`: that is the master switch, this only narrows which workers it reacts to. | | `spendCapUsd` | `0` | Hard ceiling on cumulative metered spend per job, USD. 0 = no cap. Only ever consulted for a metered app. | | `sections` | list — see below | Which sections a sweep generates.

Tags DOUBLE THE CALL COUNT but cost only 5–15% more TIME, measured, and that is not a contradiction: a tag call sends the same transcript as the chapter call before it, so it hits the engine's cached prefix and pays essentially no prefill (+0.4 s across 4 extra calls, against 22.4 s for the first 4). All it pays is decode, and a tag list is ~30 output tokens where a chapter list is ~200–290.

The corollary matters more than the number: run them in the SAME pass. Tags generated later, on their own, pay full prefill again — measured at 44% of a whole chapters pass, i.e. 3–9× the marginal cost of just including them now. | | `timestampMode` | `"chunk-local"` | How each chunk's transcript markers are numbered — see DigestTimestampMode. Was a scored variable in the bake-off rather than a pre-applied fix; the measurement is in and "chunk-local" is now the shipped default. | | `promptVariant` | `""` | A free-text label for a non-default prompt shape, folded into the recorded provenance by digestPromptVariant(). Setting it invalidates every digest generated under a different label, which is exactly what makes a bake-off round re-run its sample instead of skipping it as fresh. Empty = default. | #### `digest.apps.` Per entry — each entry spells its own values. | Key | Description | |---|---| | `bin` | Binary path/name override (process-based apps only). | | `baseUrl` | Base URL override (HTTP apps only). | | `model` | Model id, e.g. "qwen2.5:7b" or "haiku". | | `numCtx` | Context window in tokens. MUST reach the engine explicitly for ollama: its 4096 default silently truncates the input and the model then summarizes whatever fragment survived — measured, and the single easiest way to get quietly-wrong output at scale. | | `temperature` | Sampling temperature. 0 for a structured extraction task. | | `think` | Reasoning-model toggle (ollama's top-level `think`). Only sent when set, so a model that does not support thinking is never handed a field it rejects.

It matters for throughput, not correctness: measured on this box, qwen3:8b with thinking on spends most of its output budget on a `thinking` block before the JSON body the schema constrains. For an extraction task with a pinned schema that reasoning buys little and costs a multiple of the tokens, and tokens are what a multi-week sweep is priced in. | | `timeoutMs` | Per-request wall-clock ceiling (ms). A wedged engine must not stall a sweep. | Default: ```json { "remoteEnabled": false, "longTailSeconds": 14400, "localAppId": "ollama-direct", "remoteAppId": "claude-code", "apps": {}, "yieldToTranscription": true, "yieldToCpuWorkers": false, "spendCapUsd": 0, "sections": [ "chapters" ], "timestampMode": "chunk-local", "promptVariant": "" } ``` ## `diarization` Speaker diarization captured right after transcription, while the audio is still on disk. OFF by default. See DiarizationSettings. #### `diarization` | Key | Default | Description | |---|---|---| | `enabled` | `false` | Master switch. OFF by default so a transcription batch can start before this lands, with diarization backfilled over the retained audio afterwards.

Turning it ON also arms the cleanup guard: the Clean-audio sweep stops deleting audio for a transcribed video that has no diarization.json yet. That is the point — it is what keeps the perishable input alive long enough to be captured — but it means enabling this holds disk. | | `inlineAfterTranscribe` | `false` | Run diarization inline in the post-transcribe hook.

OFF by default, and that default is a MEASURED decision, not caution. Measured on this box: GPU transcription runs at 221 s/audio-hour (16.3x realtime, over 3,602 real videos), CPU diarization at ~500-680 s/audio-hour. Diarization is therefore ~2-3x SLOWER than the transcription it follows, so running it inline drops whole-pipeline throughput by roughly 3-4x and leaves the GPU idle while the CPU catches up.

The intended sequence for a large batch is the opposite: leave this off, let the batch transcribe at full GPU speed with `enabled` holding the audio, and diarize afterwards with the backfill pass. Turn it on for steady state, once the arrival rate is a few videos a day rather than a corpus. | | `threshold` | `0.9` | Clustering threshold — the single most consequential knob, since it decides how many speakers come out. Larger merges more aggressively.

The default is 0.9, NOT sherpa-onnx's own 0.5, and that is measured on this corpus. On a 6-minute excerpt of a two-person interview (known ground truth: 2 speakers), sherpa's default produced 22 clusters; 0.9 produced 6, with the top two at 40%/40% of talk time — recognizably the two hosts. Sweep on the same clip: 0.4→23, 0.5→22, 0.6→17, 0.7→12, 0.8→10, 0.9→6.

It still over-splits, which is why this is a capture lane and not an answer: the turns are recorded with the threshold that produced them, so a later attribution pass can re-cluster or re-run without needing the audio back. | | `threads` | `4` | Engine threads per diarize run. | | `engine` | `"sherpa-onnx"` | Which engine runs. "sherpa-onnx" is the shipped default and what every sidecar on disk was produced by; "sortformer" is the ggml engine built by scripts/build-sortformer.sh.

CHANGING THIS RESTATES THE FRESHNESS IDENTITY (see diarizationTarget), so every sidecar written by the other engine becomes stale and the backfill lane offers to redo it. That is intended — the two disagree about how many speakers exist, and a corpus half-diarized by each is not one corpus — but on the retained audio it is weeks of work, not a toggle.

Why anyone would: on the same file, sherpa at its tuned threshold returns 13 speakers and sortformer returns 4, agreeing on the dominant speaker's share to within half a point (73.1% vs 73.5%). On the corpus's worst case sherpa returns 35 and sortformer 4. Over-splitting is the failure mode this lane has always had, and sortformer is end-to-end rather than clustered, so it does not have it. The cost is a hard ceiling of 4 speakers and ~1.8x the wall clock. | | `backend` | `"vulkan"` | Compute device for the sortformer engine; ignored by sherpa-onnx, which has no Vulkan compute path on Linux.

"vulkan" is 1.5x faster than a thread-tuned CPU run (894 vs 1305 s/audio-hour, measured on this box) and holds 558 MB resident instead of 4.84 GB by keeping weights and activations in VRAM. It also takes ~4.4 GB of an 8 GB card, which is why the lane YIELDS to transcription rather than sharing — see controller/digestYield.ts. | | `python` | `"python3"` | Python interpreter for the default sherpa-onnx engine. sherpa-onnx ships wheels only up to cp313, and this box's system python is 3.14 — so this usually points at a dedicated venv rather than `python3`. | | `segModel` | `""` | ONNX model paths for the default engine. Empty = the lane cannot run, which is reported as a skip rather than a failure. | | `embModel` | `""` | ONNX speaker-embedding model path for the sherpa-onnx engine. Empty = the lane cannot run, reported as a `not-configured` skip rather than a failure (same as `segModel`). | | `sortformerBin` | `""` | Binary and model for the sortformer engine, both produced by scripts/build-sortformer.sh. Empty = that engine cannot run, reported as the same "not-configured" skip as an unset segModel/embModel. | | `sortformerModel` | `""` | Model for the sortformer engine, produced by scripts/build-sortformer.sh. Empty = that engine cannot run, reported as the same `not-configured` skip as an unset `sortformerBin`. | | `concurrency` | `1` | How many diarize runs may execute at once in the backfill pass. Kept low by default: diarization is CPU-bound and competes with GPU feeding and the digest sweep for the same 8 threads. | | `maxAudioHours` | `0` | Videos longer than this are DEFERRED rather than diarized: reported as a third number that is never summed into reachable work, so a capped corpus can never read as finished.

THIS IS A STOPGAP AND IT IS NOT THE FIX. sherpa-onnx's clustering holds a pairwise distance matrix over speech-segment embeddings — O(n^2) in SEGMENT count — and speaker-turn density varies 40x across this corpus (33-1364 turns/hour), so duration does not actually predict the blowup: a sparse 7h42m video completed while a dense 6h12m one was OOM-killed. Duration is merely the only predictor available for free, from metadata already on disk, BEFORE spending 45 minutes to find out. n^2 at 30k segments is 6.7 GiB and at 40k is 11.9 GiB, which brackets the 10.6 GB and 9.6 GB peaks measured on this 16 GB box.

0 disables the cap. That is where this goes once windowed diarization lands: windowing divides per-window n by the window count, so the matrix falls by its square, and the cap stops being needed rather than being tuned. | Default: ```json { "enabled": false, "inlineAfterTranscribe": false, "threshold": 0.9, "threads": 4, "engine": "sherpa-onnx", "backend": "vulkan", "python": "python3", "segModel": "", "embModel": "", "sortformerBin": "", "sortformerModel": "", "concurrency": 1, "maxAudioHours": 0 } ``` ## `backfill` The generic catch-up lane for derived data the existing corpus predates. OFF by default, and idle-only when on. See BackfillSettings. #### `backfill` | Key | Default | Description | |---|---|---| | `concurrency` | `1` | Slots the lane may use when it is not standing aside. Kept at 1 by default for the same reason diarization.concurrency is: this is CPU-bound work competing with GPU feeding and the digest sweep for the same 8 threads. | | `allowRedownload` | `false` | Re-acquire media for videos whose input is GONE (audio deleted after transcription). OFF by default and deliberately so: measured on this corpus, 836 videos still have media and ~76,270 would need a re-download — 91x the reachable work, against 45 GB free at 97% full. When on, each re-fetched file is removed in a `finally` as soon as the backfill has used it, unless the video is marked do-not-clean, or unless the auto-transcribe policy would replace its auto-captions (`replaceAutoSubs`, or a leaf on `downloadedAutoSubsOnly`), in which case the audio is kept for that runner.

WHAT IT DOWNLOADS IS AUDIO, on every channel. On a `handling: "youtube"` channel — which normally only fetches subtitles — the re-acquire applies a PER-VIDEO transcribe override so yt-dlp lands audio a diarizer can read; the channel's stored config is not changed. Without that override the fetch re-downloads the captions the video already has and lands nothing, which is what happened to ~16,000 videos on eight channels in 2026-08. | Default: ```json { "concurrency": 1, "allowRedownload": false } ``` ## `attribution` Naming the speakers diarization found (or reconstructing them from the transcript when it found none). OFF by default. See AttributionSettings. #### `attribution` | Key | Default | Description | |---|---|---| | `enabled` | `false` | Master switch. Off means the backfill registry reports no attribution work at all — the feature gate every Operation has. | | `appId` | `"ollama-direct"` | Which digest app runs the naming. Attribution IS a digest-app workload — constrained JSON decoding over transcript text — so it reuses that registry and that per-app config (settings.digest.apps[appId]) rather than growing a second copy of the ollama URL, context size and timeout. | | `model` | `""` | Model override. Empty = the app's configured model, then its default. It is separate from the digest's because the two workloads may want different sizes, and because it is part of the freshness identity: sharing the digest's model field would make a digest bake-off invalidate every attribution record on disk as a side effect. | | `diarizedEnabled` | `false` | The lanes, separately. Both default OFF even when `enabled` is on, so turning the feature on to look at it cannot start a corpus sweep.

They are not a fallback pair. `diarized` is one call per video and grounded in acoustic clustering; `textOnly` is ~30 calls and guesses at identity across chunk seams. An operator may reasonably want the first forever and the second never. | | `textOnlyEnabled` | `false` | The text-only attribution lane: names speakers from the transcript alone (~30 model calls per video, guessing identity across chunk seams). Default OFF even when `enabled` is on. See `diarizedEnabled` — the two are separate lanes, not a fallback pair. | | `promptVersion` | `1` | The prompt generation a record must match to count as fresh.

Defaults to (and is floored at) ATTRIBUTION_PROMPT_VERSION, the shipped constant. Raising it forces a corpus-wide regeneration without a code change, which is the honest way to redo everything after a prompt tweak. It cannot be set BELOW the shipped constant, and that floor is the lesson from digestPrompt.ts's version 1 -> 2 note: pinning freshness to an older generation freezes output from a superseded prompt into the corpus, looking identical to output from the current one. | Default: ```json { "enabled": false, "appId": "ollama-direct", "model": "", "diarizedEnabled": false, "textOnlyEnabled": false, "promptVersion": 1 } ``` ## `archiveOrg` How archive.org files are fetched (controller/archiveOrgDownload.ts). Over BitTorrent with aria2c when the item's torrent carries the file — archive.org is the torrent's web seed, so the swarm takes load off archive.org — then seeded for a while; otherwise, or when the torrent stalls, a direct download from archive.org. Either way the file is verified against archive.org's sha1/md5. No yt-dlp. #### `archiveOrg` | Key | Default | Description | |---|---|---| | `torrent` | `true` | Fetch an archive.org file over BitTorrent when the item's `_archive.torrent` carries it and aria2c is installed (ARIA2C_BIN). archive.org is the torrent's web seed, so what other peers give never touches archive.org. False = always the direct download. | | `seedMinutes` | `10` | Minutes to seed the file after it is complete, as a good swarm citizen. The import holds archive.org's queue while it seeds. 0 = do not seed. Clamped to [0, 1440]; default 10. | | `seedRatio` | `1` | Stop seeding sooner once this much has been uploaded relative to the file's size (aria2c --seed-ratio), whichever of the two comes first. 0 = no ratio limit. Clamped to [0, 100]; default 1. | | `stallMinutes` | `5` | No download progress for this long stops aria2c and the file is downloaded directly instead. Clamped to [1, 120]; default 5. | | `maxPeers` | `30` | Most peers per torrent (aria2c --bt-max-peers). Clamped to [1, 500]; default 30. | | `maxDownloadKiBps` | `0` | Download rate cap for a torrent fetch, KiB/s (aria2c --max-overall-download-limit). 0 = unlimited. | | `maxUploadKiBps` | `0` | Upload rate cap while downloading and seeding, KiB/s (aria2c --max-overall-upload-limit). 0 = unlimited. | Default: ```json { "torrent": true, "seedMinutes": 10, "seedRatio": 1, "stallMinutes": 5, "maxPeers": 30, "maxDownloadKiBps": 0, "maxUploadKiBps": 0 } ``` ## `publish` The publish LANE (release 18): a runner that, when the index is stale, updates it and then builds — and, where a site's own `publish.auto` (site.json) says so, deploys — what changed, one stage at a time on the `publish` queue. OFF by default; the manual stages (Publish now, Build, Deploy) work either way. See PublishSettings and common/publish/publishRunner.ts. #### `publish` | Key | Default | Description | |---|---|---| | `enabled` | `false` | Whether the publish lane's runner runs (`auto-publish` on /jobs). Off by default. Turning it on starts nothing by itself until the index is stale (or there is no index stamp yet) — see `refreshEveryMinutes`. | | `held` | `false` | The lane's pause gate (lib/pauseGates.ts, lane `publish`). A hold stops the runner DISPATCHING: the stage in flight finishes, no next one starts. It never kills a stage. | | `checkEveryMinutes` | `10` | How often (minutes) the runner wakes to ask whether a pass is due. Clamped to [1, 1440]; default 10. | | `refreshEveryMinutes` | `360` | The least time (minutes) between two index updates the lane starts: a pass runs when the index is stale and the last update is at least this old, or when there is no index stamp. Clamped to [0, 43200]; default 360. 0 = whenever the index is stale. | | `quietHours` | `null` | A local-clock window in which the lane starts no pass, `{ "start": 22, "end": 6 }` (hours [0,23], `[start, end)`, may wrap midnight), or null (default). A pass already running finishes its stage and then waits. | | `runner` | `"local"` | Which build runner the lane's builds ask for: "local" (default; each site in turn, as a child of the editor) or "docker" (every stale site in containers — a host with a container engine only; in a container it is refused). See PUBLISH.md. | | `previewBranch` | `"preview"` | The Pages preview branch a `preview` policy deploys to (`wrangler pages deploy --branch `). Lowercase letters, digits and dashes, never `main`; an invalid name reads as the default, "preview". | | `hub` | `"off"` | The hub's policy: "off" (default), "build", "preview" or "production". The hub is built when the index or the listed sites changed, and deployed as the policy says — to production only with a Pages project in homepage.json. | | `homepage` | `"off"` | The homepage's policy, as `hub` (default "off"). Building the homepage also publishes the source mirror (the operator's scrub and denylist files must exist). | Default: ```json { "enabled": false, "held": false, "checkEveryMinutes": 10, "refreshEveryMinutes": 360, "quietHours": null, "runner": "local", "previewBranch": "preview", "hub": "off", "homepage": "off" } ``` ## `seeder` The home seeder of last resort (release 21): `archilyzer seed` seeds the sites' playable torrents (`archilyzer media playable`), each one only while no other seeder has it, and only behind a VPN (the docker profile `seeder`, docker-compose.seeder.yml). Not configured until `sites` names one. See SeederSettings and common/controller/seeder.ts. #### `seeder` | Key | Default | Description | |---|---|---| | `sites` | `[]` | The site ids whose playable torrents the seeder seeds: every channel of each site that has a playable manifest (`archilyzer media playable`). Empty = the seeder is not configured, and `archilyzer seed` refuses to start. | | `trackers` | `[]` | Announce URLs the seeder announces to and scrapes (wss:// or ws:// for browsers, http(s):// or udp:// for desktop clients). Empty = each torrent's own announce list. | | `maxUploadKiBps` | `0` | Upload cap across every torrent, KiB/s. 0 = unlimited. | | `maxConnections` | `50` | Most peer connections per torrent. Clamped to [1, 500]; default 50. | | `pollSeconds` | `60` | How often the seeder scrapes the trackers and decides, per torrent, whether to seed or stand by. A viewer whose last other seeder leaves mid-stream waits up to this long. Clamped to [10, 3600]; default 60. | | `standbyAfterSeconds` | `600` | Other seeders seen on every poll for this long → the seeder stands by on that torrent (stops announcing, closes its peers, keeps the data). The same window, with no other seeder seen, brings it back when nobody is waiting; a waiting leecher brings it back at once. Clamped to [60, 86400]; default 600. | | `bindInterface` | `""` | The network interface the seeder must have — the VPN tunnel's (e.g. `wg0`, `tun0`). While it is missing the seeder does not start, and stops every torrent if it disappears. Empty = not checked (the docker profile `seeder` confines the process to the tunnel either way; see docker-compose.seeder.yml). | Default: ```json { "sites": [], "trackers": [], "maxUploadKiBps": 0, "maxConnections": 50, "pollSeconds": 60, "standbyAfterSeconds": 600, "bindInterface": "" } ```