Skip to content
This repository was archived by the owner on Sep 17, 2026. It is now read-only.

Record what the layer cache measured: 23s warm, against a 7m34s baseline - #26

Merged
Peeja merged 1 commit into
mainfrom
claude/cache-measured
Sep 17, 2026
Merged

Peeja merged 1 commit into
mainfrom
claude/cache-measured

Conversation

@Peeja

@Peeja Peeja commented Sep 17, 2026

Copy link
Copy Markdown

⏭️ Safe to merge without waiting for CI — all of it

Refreshed on every push. Checks are not required yet, so this is advice, not a gate.

One file changed: MONOREPO_TODO.md. Nothing in the repository reads it.

$ git diff --name-only origin/main HEAD
MONOREPO_TODO.md
$ grep -rn "MONOREPO_TODO" --include='*.yml' --include='*.sh' --include='*.go' --include=Makefile .
.github/scripts/check-base-images.sh:12:# not installed yet (MONOREPO_TODO.md), so that half is inert. This check is

One hit, prose inside a comment. No workflow, guard, Go package or Makefile target consumes the file.

What this costs if it's wrong: the checks still run; we just don't wait. If an ignored check goes red after merge, main is red.


The entry landed by #23 said to re-measure the layer cache after a warm run. That number arrived sooner than expected, so this records it.

Why it arrived early: #25's e2e job hit the pre-existing TestUploadAndRetrieve/filesystem flake, I re-ran it once — and a pull request writes its own cache scope, so the re-run read back what run 1 wrote. The flake handed us the warm measurement.

The numbers

all 8 images
baseline, plain docker build 7m34s (itest ingot), 6m08s (itest hilt)
cold + cache write 11m21s — ~3 min worse
warm 23s
build piri 4s · hilt 5s · ingot 3s · sprue 2s · delegator 2s
      piri-signing-service 2s · swarf 3s · indexing-service 2s

The whole e2e job went 13m10s → 5m29s.

Three caveats, recorded with the number rather than under it

  1. The warm figure is a re-run of the same commit, so every layer hit. That is the ceiling, not the average. A real change invalidates the go build layer for the services it touches; the base image, apt and go mod download layers — the bulk — still hit, so expect minutes rather than seconds. The entry says to re-measure on a real source change.
  2. The cold path is ~3 minutes worse than before, because mode=max exports every layer. The first run on main after Layer-cache the image builds in itest and e2e, keeping images.yml cold #25 merges will be slower, by construction.
  3. Nobody has checked the cache against GitHub's 10 GB per-repository limit. Eight images at mode=max is not small, eviction is LRU, and an overflowing cache degrades quietly back to cold builds — which is exactly the kind of silent regression this repo keeps deleting.

One option dies as a consequence

Build-once-and-load was the other way at the same ~6 minutes: have images.yml hand its images over as an artifact. The entry already said "try the cache first". 23 seconds is below anything a 1–2 GB artifact round-trip could achieve, so it is now marked effectively dead — it only comes back if the cache turns out to evict often.

Worth your eye

The 10 GB limit is the one I'd actually check before treating 23 seconds as durable. I have not, and the entry says so rather than implying otherwise.

🤖 Generated with Claude Code

https://claude.ai/code/session_01CGAGAib517Ae1kg8SCdcEt


Generated by Claude Code

The entry said to re-measure after a warm run. #25's e2e re-run was warm, because
a pull request writes its own cache scope and a re-run reads it back, so the
number arrived earlier than expected: all 8 images in 23 seconds, and the whole
e2e job from 13m10s to 5m29s.

Recorded with the caveats rather than the headline alone. The warm figure is a
re-run of the same commit, so it is the ceiling and not the average -- a real
change invalidates the go build layer for the services it touches, though the
base, apt and go mod download layers still hit. The cold path is about 3 minutes
worse than before, because mode=max exports every layer. And the cache has not
been checked against GitHub's 10 GB per-repository limit, where eviction is LRU
and an overflow degrades quietly back to cold builds.

Build-once-and-load is marked effectively dead as a consequence: 23 seconds is
below anything an artifact round-trip of 1-2 GB could achieve, so that option
only comes back if the cache evicts often.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGAGAib517Ae1kg8SCdcEt
Peeja pushed a commit that referenced this pull request Sep 17, 2026
Petra merged #20, #17, #22, #19, #23 and #24; main is 144b316 and the open
list is #25 and #26. Dropped the merged rows from both tables rather than
letting them accumulate.

#25's e2e red turned out to be the pre-existing filesystem flake, and the one
re-run that cleared it was warm -- so it produced the measurement as well as the
green: 8 images in 23s against a 7m34s baseline, and the e2e job from 13m10s to
5m29s. #26 puts that in MONOREPO_TODO.md with the caveats attached rather than
the headline alone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGAGAib517Ae1kg8SCdcEt
@Peeja
Peeja merged commit 7688820 into main Sep 17, 2026
21 of 22 checks passed
@Peeja
Peeja deleted the claude/cache-measured branch September 17, 2026 17:25
Peeja pushed a commit that referenced this pull request Sep 17, 2026
Chasing #25's e2e red to the bottom found a real bug. pg_isready without -h
probes the Unix socket; the official postgres image runs initdb against a
temporary server started with listen_addresses='', so the probe passes while TCP
is still refused. plc-postgres's log against plc's crash shows the healthcheck
going green 2.25 seconds before the port existed, and plc booting into the gap.

That explains the shape that made it look like a generic race: plc already
declared condition: service_healthy, so the ordering was never wrong -- the
readiness signal was, and whichever dependent happened to boot inside the window
was the one that failed. Hence a different container each run.

#27 fixes six sites and adds a list-free guard that also covers Go, because the
generator builds one of these strings, and that fails if it ever finds nothing
to check. Postgres was the only class member; every other healthcheck is HTTP
over localhost or redis-cli ping, TCP by construction.

Recorded with what is still unproven: the mechanism, yes; the frequency, no. A
5% flake cannot be shown fixed by one green run.

Also: main is 95e8366, #25 and #26 merged, the layer cache is live.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGAGAib517Ae1kg8SCdcEt
Peeja pushed a commit that referenced this pull request Sep 17, 2026
…heck

The previous commit updated Current State but its companion edit failed before
writing, so this page still claimed five open PRs and #25 unmerged. Now: #27
only, with #25 and #26 merged and the cache numbers carried across.

Names the one thing on that list the agent genuinely cannot do rather than has
not done: checking the cache against GitHub's 10 GB per-repository limit needs
gh cache list or the Actions cache API, and neither is reachable from here. It
matters because eviction is LRU and an overflowing cache degrades quietly back
to cold builds, which is the silent regression shape this repo keeps deleting.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CGAGAib517Ae1kg8SCdcEt
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants