fix: route /codex:rescue through the Agent tool to stop Skill recursion (#234) - #235
Conversation
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: eb24951c89
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
Address openai#235 review comment (Codex bot, P2): the new character-for-character model rule in `agents/codex-rescue.md` conflicts with the `spark` alias line immediately above it. For a user request like "use spark", line 31 says to emit `--model gpt-5.3-codex-spark`, while line 33 says never to emit a model name the user did not literally type. That is nondeterministic and undermines the no-duplicate-call goal. Rephrase the literal-copy rule to explicitly exempt the documented `spark` alias, and mirror the fix in `skills/codex-cli-runtime/SKILL.md`. Pin the exemption with a regression assertion so a future edit can't silently reintroduce the contradiction.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c582eda025
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
…on (openai#234) `/codex:rescue` previously combined two things that together caused a hang: - `context: fork` in the frontmatter, which spawns a `general-purpose` subagent for the command body. - Body prose "Route this request to the `codex:codex-rescue` subagent." without naming the transport. When the main agent called `Skill(codex:rescue)` programmatically, the fork resolved the ambiguous prose by trying `Skill(codex:codex-rescue)` (unknown skill) and then falling back to `Skill(codex:rescue)`, which re-entered this command and hung the session until the user cancelled. No Codex job was ever created. Naming the transport as `Agent(codex:codex-rescue)` alone is not enough: forked general-purpose subagents do not expose the `Agent` tool, so the forked runner cannot reach the subagent that way either. The minimal fix is therefore two coordinated changes: - Drop `context: fork` so the command body runs inline in the calling agent's context, where `Agent` is in scope. - Say explicitly "use the `Agent` tool with `subagent_type: "codex:codex-rescue"`", and call out that `Skill(codex:codex-rescue)` and `Skill(codex:rescue)` are not valid routing paths. Add `Agent` to `allowed-tools` so the call does not prompt for permission. Everything else in rescue.md (resume-candidate check, flag handling, background/foreground semantics, operating rules) is unchanged. The `codex:codex-rescue` subagent itself is unchanged. Tests pin the new allow-list, the explicit `subagent_type`, the ban on `Skill(codex:codex-rescue)`, and the absence of `context: fork`. The existing "run the `codex:codex-rescue` subagent in the background" assertion continues to hold since that sentence still reads correctly with the Agent-tool transport. Fixes openai#234
c582eda to
cd4cb60
Compare
|
A note on the Original topology (with The bug: forked After the fix: the command body runs inline in the caller, which does have Consistency check: every other command in Trade-off acknowledged: the calling agent now sees the rescue.md body and the Alternative considered: keeping the fork and having it call |
Replace context:fork with explicit Agent tool routing to prevent Skill-recursion hangs when Claude Code misroutes to Skill(codex:rescue). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* Fix working-tree review crash on untracked directories (openai#166) * fix: skip untracked directories in review context * fix: skip broken untracked symlinks in reviews * fix: respect SHELL on Windows for Git Bash (openai#178) * Use app-server auth status for Codex readiness (openai#177) * Use app-server auth status for Codex readiness * fix: reuse existing app server for auth checks * fix: inherit process.env in app-server spawn when no explicit env is provided (openai#159) `SpawnedCodexAppServerClient.initialize()` passes `this.options.env` to `spawn()`, but no caller ever sets `env` in options. In Node.js, passing `undefined` for `env` gives the child process **no** environment variables, breaking any model provider that relies on env vars (e.g. DATABRICKS_TOKEN). Fall back to `process.env` when `this.options.env` is not set, matching the existing pattern in `broker-lifecycle.mjs` and `codex-companion.mjs`. Co-authored-by: Isaac Co-authored-by: Bhuvanesh Sridharan <bhuvanesh.sridharan@databricks.com> * fix: gracefully handle unsupported thread/name/set on older Codex CLI (openai#126) * fix: gracefully handle unsupported thread/name/set on older Codex CLI Codex CLI v0.118.0 does not recognize the thread/name/set JSON-RPC method, causing startThread() to throw. Thread naming is cosmetic (for job log labels) and should not block thread creation. Wraps the call in try/catch so it fails silently on older CLI versions. Fixes openai#119 * refactor: only suppress unsupported-method errors for thread/name/set Address Codex review feedback: the bare catch swallowed all errors including auth, network, and server failures. Now only suppresses errors containing 'unknown variant' or 'unknown method' (the specific error older CLI versions return) and rethrows everything else. * codex: scope implicit resume-last selection to the current Claude session (openai#83) Co-authored-by: VOIDXAI <VOIDXAI@users.noreply.github.com> * codex: scope default cancel selection to the current Claude session (openai#84) Co-authored-by: VOIDXAI <VOIDXAI@users.noreply.github.com> * fix: avoid embedding large adversarial review diffs (openai#179) * fix: avoid embedding large adversarial review diffs * fix: preserve untracked content in lightweight review * address comments * fix: handle ENOBUFS type check in git diff sizing * bump: update plugin version to 1.0.3 (openai#180) * codex: honor --cwd when reporting session runtime (openai#35) * codex: honor --cwd when reporting session runtime * codex: keep session runtime lookup scoped to callers --------- Co-authored-by: VOIDXAI <VOIDXAI@users.noreply.github.com> * fix: declare model in codex-rescue agent frontmatter (openai#169) * fix: declare model in codex-rescue agent frontmatter The codex-rescue agent had no model field, leaving Claude Code to assign whatever default it chooses. As a thin forwarding wrapper that issues a single Bash call, this agent is well-suited to the haiku tier; declaring it explicitly ensures a predictable cost profile and tier guarantee. Co-Authored-By: Claude Code <noreply@anthropic.com> * Update codex-rescue.md --------- Co-authored-by: claude[bot] <claude-bot@anthropic.com> Co-authored-by: Claude Code <noreply@anthropic.com> Co-authored-by: Dominik Kundel <dkundel@openai.com> * fix: correct invalid 'xhigh' reasoning effort in README (openai#99) Closes openai#77 Updated README to replace unsupported 'xhigh' value with 'high' to prevent configuration errors. * fix: quote \$ARGUMENTS in cancel, result, and status commands (openai#168) Unquoted \$ARGUMENTS in the ! shell commands allowed shell metacharacters in user-supplied job IDs to be expanded before Node received them (e.g., `task-123; malicious-cmd` would execute the trailing command). This is inconsistent with review.md and adversarial-review.md, which both wrap "$ARGUMENTS" in double quotes. Co-authored-by: claude[bot] <claude-bot@anthropic.com> Co-authored-by: Claude Code <noreply@anthropic.com> * fix: route /codex:rescue through the Agent tool to stop Skill recursion (openai#234) (openai#235) * fix: route /codex:rescue through the Agent tool to stop Skill recursion (openai#234) `/codex:rescue` previously combined two things that together caused a hang: - `context: fork` in the frontmatter, which spawns a `general-purpose` subagent for the command body. - Body prose "Route this request to the `codex:codex-rescue` subagent." without naming the transport. When the main agent called `Skill(codex:rescue)` programmatically, the fork resolved the ambiguous prose by trying `Skill(codex:codex-rescue)` (unknown skill) and then falling back to `Skill(codex:rescue)`, which re-entered this command and hung the session until the user cancelled. No Codex job was ever created. Naming the transport as `Agent(codex:codex-rescue)` alone is not enough: forked general-purpose subagents do not expose the `Agent` tool, so the forked runner cannot reach the subagent that way either. The minimal fix is therefore two coordinated changes: - Drop `context: fork` so the command body runs inline in the calling agent's context, where `Agent` is in scope. - Say explicitly "use the `Agent` tool with `subagent_type: "codex:codex-rescue"`", and call out that `Skill(codex:codex-rescue)` and `Skill(codex:rescue)` are not valid routing paths. Add `Agent` to `allowed-tools` so the call does not prompt for permission. Everything else in rescue.md (resume-candidate check, flag handling, background/foreground semantics, operating rules) is unchanged. The `codex:codex-rescue` subagent itself is unchanged. Tests pin the new allow-list, the explicit `subagent_type`, the ban on `Skill(codex:codex-rescue)`, and the absence of `context: fork`. The existing "run the `codex:codex-rescue` subagent in the background" assertion continues to hold since that sentence still reads correctly with the Agent-tool transport. Fixes openai#234 * test: match quoted result and cancel command arguments --------- Co-authored-by: Dominik Kundel <dkundel@openai.com> * fix: bump plugin version to 1.0.4 (openai#244) * feat: add Claude session transfer command (openai#374) * fix: bump plugin version to 1.0.5 (openai#398) * Remove shell expansion for git commands (openai#447) * Remove shell expansion for git commands * Version bump --------- Co-authored-by: Dominik Kundel <dkundel@openai.com> Co-authored-by: Bhuvanesh Sridharan <bhuvan.sridharan@gmail.com> Co-authored-by: Bhuvanesh Sridharan <bhuvanesh.sridharan@databricks.com> Co-authored-by: Trevin Chow <trevin@trevinchow.com> Co-authored-by: VOIDXAI <oliviawilliams199206@gmail.com> Co-authored-by: VOIDXAI <VOIDXAI@users.noreply.github.com> Co-authored-by: xiaolai <lixiaolai@gmail.com> Co-authored-by: claude[bot] <claude-bot@anthropic.com> Co-authored-by: Claude Code <noreply@anthropic.com> Co-authored-by: Manav Agarwal <23bai70003@cuchd.in> Co-authored-by: Friende <35026241+pengyou200902@users.noreply.github.com> Co-authored-by: stefanstokic-oai <stefanstokic@openai.com> Co-authored-by: bryane-oai <bryane@openai.com>
…on (openai#234) (openai#235) * fix: route /codex:rescue through the Agent tool to stop Skill recursion (openai#234) `/codex:rescue` previously combined two things that together caused a hang: - `context: fork` in the frontmatter, which spawns a `general-purpose` subagent for the command body. - Body prose "Route this request to the `codex:codex-rescue` subagent." without naming the transport. When the main agent called `Skill(codex:rescue)` programmatically, the fork resolved the ambiguous prose by trying `Skill(codex:codex-rescue)` (unknown skill) and then falling back to `Skill(codex:rescue)`, which re-entered this command and hung the session until the user cancelled. No Codex job was ever created. Naming the transport as `Agent(codex:codex-rescue)` alone is not enough: forked general-purpose subagents do not expose the `Agent` tool, so the forked runner cannot reach the subagent that way either. The minimal fix is therefore two coordinated changes: - Drop `context: fork` so the command body runs inline in the calling agent's context, where `Agent` is in scope. - Say explicitly "use the `Agent` tool with `subagent_type: "codex:codex-rescue"`", and call out that `Skill(codex:codex-rescue)` and `Skill(codex:rescue)` are not valid routing paths. Add `Agent` to `allowed-tools` so the call does not prompt for permission. Everything else in rescue.md (resume-candidate check, flag handling, background/foreground semantics, operating rules) is unchanged. The `codex:codex-rescue` subagent itself is unchanged. Tests pin the new allow-list, the explicit `subagent_type`, the ban on `Skill(codex:codex-rescue)`, and the absence of `context: fork`. The existing "run the `codex:codex-rescue` subagent in the background" assertion continues to hold since that sentence still reads correctly with the Agent-tool transport. Fixes openai#234 * test: match quoted result and cancel command arguments --------- Co-authored-by: Dominik Kundel <dkundel@openai.com>
Fixes #234.
What was broken
/codex:rescuewas markedcontext: fork, so the command body ran inside a forkedgeneral-purposesubagent. The body said:When the main Claude Code agent called
Skill(codex:rescue)programmatically (as opposed to the user typing the slash command), the forked runner read that ambiguous prose and guessed the transport. It triedSkill(codex:codex-rescue)— unknown skill — then fell back toSkill(codex:rescue), which re-entered this same command and recursed until the user cancelled. No Codex job was ever created.The user-facing symptom, quoted from the issue:
Why the fix is two coordinated changes and not one
The issue's "Suggested fix" section lists two options:
Agenttool).allowed-toolsto excludeSkill.In practice, neither alone is sufficient. Forked
general-purposesubagents do not expose theAgenttool, so instructing the fork to "useAgent" just trades the hang for a clean failure — the main agent then retries on its own. The minimum fix is therefore two coordinated changes:context: fork. The command body now runs inline in the calling agent's context, where theAgenttool is in scope.codex:codex-rescueviaAgent(subagent_type: "codex:codex-rescue"), addAgenttoallowed-tools, and call out thatSkill(codex:codex-rescue)/Skill(codex:rescue)are not valid paths.Everything else in
rescue.md— resume-candidate check,--background/--wait/--resume/--freshhandling, operating rules, model/effort pass-through — is unchanged.agents/codex-rescue.mdandskills/codex-cli-runtime/SKILL.mdare not touched.Diff
plugins/codex/commands/rescue.md: 4 lines changed (one frontmatter swap, two body sentences rewritten).tests/commands.test.mjs: pin the newallowed-tools, pin the explicitsubagent_type: "codex:codex-rescue"prose, pin theSkill(codex:codex-rescue)ban, and assertcontext: forkis absent. Total +9 / −1.Not in scope
During iteration on this fix I observed adjacent duplicate-
task-call patterns insidecodex:codex-rescue(model-name drift retried on Codex 4xx, long prompts failing on zsh shell-quoting, main-agent retry on upstream failure). Those are separate bugs and deserve separate issues and separate PRs where they can be discussed on their own merits. I deliberately kept this PR scoped to #234's routing hang to keep the diff small, reviewable, and directly tied to the filed issue.Test plan
node --test tests/commands.test.mjs -t "rescue command absorbs continue semantics"passes.rescue.mdrestored, the new assertions fail onallowed-tools→Agentmissing, then would also fail on thesubagent_typeand no-context: forkassertions.node --test tests/commands.test.mjsas a whole: 7/8 — the one failure is a pre-existingresult "$ARGUMENTS"quoting assertion unrelated to rescue (introduced in fix: quote $ARGUMENTS in cancel, result, and status commands #168, also fails on pristine upstream/main).Commits
One commit:
cd4cb60 fix: route /codex:rescue through the Agent tool to stop Skill recursion (#234).