Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit f123ee6bf23ca62b92460129e766e6f2095aff7d
parent bda861cacd0e4c14dea093356315bc0ddcd9b814
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Fri,  2 Oct 2026 10:21:23 -0400

records: post links after review -- the README says what a stalled archive costs (one timeout per URL per build; the preview's 10 s bound), that hidden posts are not looked up and that null is unset; plans/deck-posts.md gains the failure memo, the shared cache and the two known limits (siteOrigin only; corpus-order tie-break)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
Mplans/deck-posts.md | 44++++++++++++++++++++++++++++++++++++--------
Mumtool/report-to-video/README.md | 18+++++++++++++-----
2 files changed, 49 insertions(+), 13 deletions(-)

diff --git a/plans/deck-posts.md b/plans/deck-posts.md @@ -695,12 +695,21 @@ the original ("Open original" in the post view; unchanged). archive (`piratesoftware-bsky`), which a manifest rarely names, so `resolvePostLinks` finds it for every post that pins neither `siteChannel` nor `siteUrl`: `/corpus.json`'s channels with `manifests.posts`, handle-matching - slugs/names first, the first whose `slugToPage` has the id. One note per + slugs/names first, the first whose `slugToPage` has the id. One note per shown post. A miss or an archive that does not answer links the original; it never - fails a build. Reads go through `createJsonCache`, factored out of - `cues.mjs` (the cue walk's memory + disk cache): a miss in a cached copy is - re-read fresh once per process (per `refreshAfterMs` for a server); a failed - read is remembered for the resolver's life. 15 s a request. + fails a build. Hidden posts and a cut without the deck are not looked up. + Reads go through `createJsonCache`, factored out of `cues.mjs` (the cue walk's + memory + disk cache, which now also shares an in-flight fetch between callers): + a miss in a cached copy is re-read fresh once per process (per + `refreshAfterMs` for a server). 15 s a request. +- **The failure memo** (after review). One memo, `url -> { err, at, refresh }`, + read by both passes and expiring after `refreshAfterMs` (Infinity in a build, + ten minutes in the server); a success clears it. A failed cached read went to + the network, so it bars both passes; a failed fresh read bars only fresh + reads, so the copy in hand still answers the cached pass (and a cached hit + does not clear it). An archive that stalls therefore costs one timeout per + URL per build, not one per post. `resolvePostLinks` takes `deadlineMs` for a + bound on the whole call; the preview passes 10 s. - **The build.** `linkPosts()` runs once, just before the first `writeChromeSchedule` (full build, `--chrome-only`, `--chrome-preview`): after every refusal, so a refused run asks nothing of the network. It sets @@ -709,9 +718,12 @@ the original ("Open original" in the post view; unchanged). so a changed link is a new chrome key. `opts.postResolver` is the injection point. - **umtool's preview: resolved, not pinned.** `scheduleForPreview` links the - variant's posts through the same resolver (one per server process: memory - across requests, the build's disk cache, 8 s a request, a miss re-read at most - every ten minutes) before `previewSchedule`, built schedule or estimate. Chosen + variant's posts, with the request's posts draft applied (so an un-hidden post + is linked), through the same resolver (one per server process: memory across + requests, the build's disk cache, 8 s a request, 10 s for the whole request, + failures and misses retried at most every ten minutes) before + `previewSchedule`, built schedule or estimate. The found `siteChannel` goes on + the request's copy of the posts, not the draft's. Chosen over a writer action that pins `siteChannel`: a pin needs a route, a control and the same network lookup, and until someone pressed it the preview and the build would disagree. Pinning by hand still works and skips the lookup. @@ -722,6 +734,22 @@ the original ("Open original" in the post view; unchanged). - **MCP.** `get_post` adds `- archive: <url>` (common `viewerPostUrl`) when the source has a viewer origin: the channel's member site on a hub, the site's own otherwise. +- **`null` is unset** for `siteChannel`, `siteUrl` and `postId`, as for + `attachTo`, so the README's example validates. +- **The shared cache.** A post miss refreshes the cue walk's on-disk + `corpus.json` (and the posts manifests); `cues.mjs`'s header says so. Manifest + URLs carry no version, so no cue window moves; a later cue lookup can find a + channel a stale copy lacked. +- **Known, left as is.** + - The archive is only ever `provenance.siteOrigin`: a hub origin (whose + `corpus.json` lists no channels), a manifest naming its archive only as + `corpus: "remote:…"`, or a build run with `--site-origin` all link the + originals. Clip QRs behave the same way. The no-origin note says to set + `provenance.siteOrigin`. + - With no handle match, corpus order decides between channels that hold the + same id, and a handle whose first label is generic (`bsky.app` → `bsky`) + matches every `*-bsky` channel. X ids are global and bsky rkeys are TIDs, so + a collision is unlikely. First use: the ferret-rescue manifest's seven Bluesky posts all resolve to `piratesoftware-bsky` on its archive (each id on page 0 of that channel's posts diff --git a/umtool/report-to-video/README.md b/umtool/report-to-video/README.md @@ -353,7 +353,8 @@ an archive channel and the post's id; else the post's own `url`, as before. rkey (`/post/<rkey>`) or an X status id (`/status/<id>`), or `postId` — channels whose slug or name carries the post's handle first. Through the cue walk's cache (memory and disk, `REPORT_CACHE_DIR`); a post missing from a - cached copy is looked for once more in a fresh one. + cached copy is looked for once more in a fresh one. Hidden posts, and a cut + without the deck, are not looked up. - **In memory only.** The found channel is set on the run's copy of the post and written to `schedule.json` as the post's `qrUrl` (beside its own `url`), so a `--chrome-only` re-render draws the link the build found; the manifest @@ -361,11 +362,18 @@ an archive channel and the post's id; else the post's own `url`, as before. up — pin one to make a cut that never asks the network. - **Never a failure.** Each post gets a note: `<id>: archive link via <slug>`, `<id>: not in the archive at <origin> -- QR links the original`, or the - archive did not answer (15 s a request) and the QR links the original. + archive did not answer and the QR links the original. Each request may take + 15 s, and a URL that failed is not asked again in that build — whichever + pass or post asks next — so an archive that does not answer costs at most one + timeout per URL it has (its `corpus.json` and each posts manifest), not one + per post. - **umtool's preview** links its posts through the same resolver and the same - disk cache (8 s a request; a miss re-read at most every ten minutes), so the - On-screen preview draws the build's QRs. The deck form's `qr links` field - is `posts.links`. + disk cache, so the On-screen preview draws the build's QRs. A request waits + at most 10 s for all its posts together (8 s a request); a post not found by + then links the original for that request, while its lookup finishes in the + background for the next. A URL that failed is not asked again for ten + minutes, and a miss is re-read fresh at most that often; a post a draft + un-hides is linked too. The deck form's `qr links` field is `posts.links`. - **The chrome cache** follows: a QR is an asset of the page, so a post whose link changes is a new key and a re-render.