Repository navigation
feat(sandbox): support stop requests conditional on sandbox execution identity #4009
Description
Activity
- addedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Sep 30, 2026 One additional invariant that may help tighten this design is:
"target identity != target generation"
A stable sandbox identity/name can tell us which logical sandbox is being addressed, but it does not prove that the execution being acted on is the same execution that produced the triggering observation.
For an event-driven controller, I would treat the stop authorization as a compare-and-act operation over an immutable execution generation:
observation
- sandbox identity
- execution generation/incarnation
- observation provenance
-> conditional lifecycle effect
The important property is that the execution-generation value must come from an authority that can actually establish the lifecycle transition, not merely from a client-supplied correlation field.
That also suggests keeping two replay dimensions separate:
- "request_id" — has this operation request already been admitted?
- "expected_execution" — is the execution that justified the operation still current?
Those answer different questions.
A particularly useful adversarial test would be:
- execution A emits observation O(A);
- A stops;
- execution B starts under the same sandbox identity/name;
- delayed O(A) produces conditional stop S(A);
- the request is retried several times with the same "request_id".
Every retry should deterministically report a stale-target result and B must remain running.
I would also test delete/recreate separately from stop/start, because a persistent sandbox identifier may fence one transition but not necessarily both.
A compact evidence chain would be:
"observed execution -> attested execution generation -> serialized lifecycle precondition -> effect"
rather than allowing:
"observed sandbox name -> current sandbox by name -> effect"
The latter preserves addressing, but not causal binding between the observation and the execution being stopped.
I added a live SDK repro of the scenario behind this feature request: https://github.com/danehans/openshell-4009-repro
Use case: an event-driven controller that stops a sandbox in response to a delayed observation. The observation is bound to one execution, but the stop request is addressed only by workspace and name.
What it does (public Go SDK only, unpatched gateway
374c035, SQLite): observe execution A, replace it, then issueSandboxes().Stop(workspace, name)for A.- stop/start: the sandbox ID is unchanged, only
Status.MainProcessStartedAtMsdiffers. The stale stop returned no error and stopped the replacement. - delete/recreate: the ID and start time both differ. The stale stop again returned no error and stopped the replacement.
This isn't a bug in unconditional
Stop. It shows the API has no way to express "stop only if this is still the execution I observed", and that the sandbox ID alone fences delete/recreate but not stop/start. I can extend the repo to exercise a conditional-stop API once the design settles.- stop/start: the sandbox ID is unchanged, only
Thanks for publishing the live SDK repro. It closes the main uncertainty cleanly: this is not an unconditional "Stop" correctness issue; it is a missing compare-and-act primitive at the execution boundary.
The two scenarios also show that the public precondition should not be defined as “sandbox ID equality”:
- stop/start: sandbox ID is stable while the execution changes;
- delete/recreate: sandbox ID changes, but the current "Stop" API cannot carry an expected target at all.
I would therefore make the public contract an opaque execution-incarnation token (name is unimportant) issued by the Gateway/supervisor authority and changed on every transition that creates a new execution. "MainProcessStartedAtMs" is useful evidence for the repro, but I would avoid making a timestamp the durable concurrency token unless its uniqueness/lifetime semantics are explicitly guaranteed.
The critical implementation property is:
conditional_stop(expected_execution = E)
== compare current_execution with E
and perform Stop
within the same serialized lifecycle operationA client-side "Get -> compare -> Stop" sequence must not be considered equivalent.
I would also make the idempotency semantics explicit because "request_id" and the execution precondition protect different races:
- "request_id" identifies one admitted operation;
- "expected_execution" identifies the execution that operation is allowed to affect.
A useful invariant is:
request_id -> {expected_execution, terminal_result}
Once a request ID is admitted, retries should replay that same terminal result rather than re-evaluate the request against a newer execution. Reusing the same "request_id" with a different "expected_execution" should fail as a request conflict, not silently retarget.
That prevents an important case:
- conditional stop for A succeeds;
- B starts;
- the client retries the same request because the original response was lost;
- the retry must return the original success result and must not evaluate against or stop B.
For the stale path, I would prefer a typed result such as "STALE_TARGET" / "FAILED_PRECONDITION", carrying the current execution identity only to the extent the API already authorizes its disclosure. Missing or unverifiable preconditions in the conditional API should fail closed and never degrade to unconditional "Stop".
The repro repo is already a strong basis for the conformance suite. I would extend it with at least these cases:
- matching token stops A;
- stale token after stop/start leaves B running;
- stale token after delete/recreate leaves B running;
- replacement racing with conditional stop is serialized to exactly one valid outcome;
- retry after successful conditional stop cannot affect a replacement;
- retry after stale-target returns the same stale result;
- same "request_id" + different execution token is rejected;
- missing/invalid token never falls back to unconditional stop.
If those properties hold, the API establishes the stronger guarantee the controller actually needs:
an observation-bound lifecycle action cannot cross an execution-generation boundary.
Happy to review the eventual token/result semantics or help turn the current repro into that conformance matrix once the API shape is proposed.
User Story
As an operator building an event-driven controller, I want to stop only the sandbox execution that produced an observation, so delayed events cannot stop a newer execution.
Problem Statement
The public
StopSandboxRequestacceptsworkspace_scope,name, andrequest_id. It cannot express an expected sandbox incarnation or execution. The request ID provides durable at-most-once admission; it is not a target precondition.The live public Go SDK reproduction demonstrated this scenario against an unpatched Gateway at
374c035with SQLite, for both stop/start and same-name delete/recreate:Deleting and recreating a sandbox with the same name presents a similar stale-target problem.
Impact / Why This Matters
Controllers must either disable automatic stop or risk stopping a replacement execution. Reading current state before stopping leaves a race between the read and the mutation. Direct Kubernetes Pod UID checks/deletes are backend-specific and bypass the Gateway lifecycle contract.
This blocks portable automation that must bind an action to the execution observed. The finding is a missing API capability; this investigation has not established an authorization bypass or sandbox-isolation vulnerability.
Proposed Design
Retain the public workspace/name addressing convention and support a conditional stop workflow:
The exact field names, token representation, and internal implementation are open for maintainers to choose.
Acceptance Criteria
Alternatives Considered
Agent Investigation
Original source investigation reviewed
mainat912a077bd641272016fb8b2fd58209f6c7c6f194:Related: #3556 concerns driver-facing sandbox references; #3153 concerns durable activation holds. Neither defines this conditional execution-targeting workflow.
Checklist
Upstream Refresh (2026-10-08)
Reviewed upstream
fe3942f2d. The public stop API still has no execution precondition. PR #4235 revalidates delayed internal driver exit observations; it does not fence controller-issued stop requests. PR #4094 introduces persistent SSH identity, while #3661 and #4321 introduce supervisor placement and handoff. PR #4217 must preserve those behaviors alongside execution targeting. The independent policy-version change remains in #4219.