Repository navigation
Bound operator memory: GC terminal sessions and drop the in-memory audit mirror - #21
Merged
Merged
Conversation
…mirror
The operator OOMKilled on long-lived clusters from two unbounded,
retain-forever-by-design sources, neither of which serves its running work.
1. Terminal AgentSessions and the ToolCalls they own were never deleted. The
pod reap and PVC reclaim free pods and scratch volumes but keep the CR, so
every session and every ToolCall ever created stayed resident in etcd and in
the operator's informer cache (observed: 10k+ ToolCalls, oldest months old).
Add a session-lifetime GC (--session-gc-after, default 30d): a terminal
session past the window is deleted, cascade-collecting its ToolCalls, pods
and Secrets. finalize keeps the append-only audit log (it lives in the
durable backend; DeleteScope refuses append-only kinds), so the record
survives — only the CR and its owned ephemera go. This is the retention the
reap already anticipated ("until retention GCs the AgentSession").
2. Under MEMORY_BACKEND=postgres the ShadowBackend dual-wrote every append-only
entry into an in-memory backend as well, while serving reads only from
postgres. That inmem mirror was never read and scope-delete can never free it
(append-only scopes are undeletable), so it grew with the whole cluster's
audit history for the operator's lifetime. Use the postgres backend directly;
reads are byte-identical to what the shadow already served from its secondary.
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
samkim
approved these changes
Oct 6, 2026
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
On long-lived clusters the operator steadily climbed to its memory limit, was OOMKilled, and crash-looped. The memory wasn't going to its actual work — it came from two retain-forever-by-design sources, so the limit was reached within days and reset on each restart.
Two causes, two fixes
1. Terminal sessions (and their ToolCalls) were never garbage-collected.
The pod reap and PVC reclaim free pods and scratch storage but intentionally keep the AgentSession record — and with it every ToolCall it owns. So every session and every tool call ever run stayed resident in etcd and in the operator's informer cache, growing without bound.
This adds a session-lifetime GC (
--session-gc-after, default 30d): a terminal session past the window is deleted, which cascade-collects its owned ToolCalls, pods and Secrets. The append-only audit log is untouched — it lives in the durable memory backend and session teardown already refuses to delete it — so the record survives; only the CR and its ephemera go. This is the retention the existing reap logic already anticipated.2. A redundant in-memory copy of the entire audit log.
With a Postgres memory backend, every append-only write was also mirrored into an in-memory backend, while reads were served only from Postgres. That mirror was never read and could never be freed, so it grew with the whole cluster's audit history for the operator's lifetime.
This uses the Postgres backend directly. Reads are byte-identical to what the mirror already served, at none of the RAM.
Testing
Operational notes
--session-gc-afterdefaults to 30 days and is tunable (0 disables). On first rollout to an existing cluster it will collect sessions already older than the window.status.auditChainHeads); the audit records and their signing-key witness remain durable, so signature verification is unaffected.