Skip to content

feat: elicitation consent channel for policy "prompt" (issue #10 channel 1) - #24

Merged
V3RON merged 3 commits into
mainfrom
claude/human-loop-sensitive-tools-rioudz
Sep 6, 2026
Merged

V3RON merged 3 commits into
mainfrom
claude/human-loop-sensitive-tools-rioudz

Conversation

@V3RON

@V3RON V3RON commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

Implements channel 1 of #10: MCP elicitation as a second consent channel for policy: "prompt" tool calls, alongside the _meta["anthropic/requiresUserInteraction"] gate from #14. This makes human-in-the-loop work on any MCP client that declares the elicitation capability (Claude Code CLI, recent Codex), not just Claude Code ≥ v2.1.199.

Behavior

  • Negotiation, per connection (mcp/server.ts): when the client's initialize declares the elicitation capability, the elicitation channel is preferred and tools/list never emits _meta["anthropic/requiresUserInteraction"] for "prompt" tools — the two channels never both arm for one call, so the human never sees two prompts. Clients without the capability keep the existing flag-based fallback; everything else keeps failing closed exactly as today.
  • Call time: a "prompt"-policy call sends one elicitation/create request naming the tool, session alias, and the call's arguments (JSON, truncated past 2 KB), bounded server-side at 10 minutes — comfortably under the 30-minute stdio idle window, so an unanswered prompt resolves as our own clean result rather than an opaque client abort.
    • accept → the call proceeds with consent: "elicitation".
    • decline / cancel / timeout → the daemon is never called; the agent gets an ordinary tool result with isError: true naming what happened, so it can adapt.
    • A failed elicitation request (declared-but-unsupported, transport error) is treated as no consent obtained, never as approval — it falls through to the existing policy_denied / no_consent_channel path.
  • Daemon (daemon.ts, audit.ts, @cordierite/shared rpc.ts): the tools.call consent param widens to "client" | "elicitation". Either satisfies the "prompt" gate, but the audit log distinguishes them: "elicitation" is an observed decision (the client replied accept to a live request), where "client" is only evidence the flag was armed. The existing trust caveat is unchanged and documented — the daemon can't verify either marker against a process that already has socket access.
  • Consent resolution is factored into one place (resolveToolCallConsent), shared by the plain and progress-tracked call paths, recomputed at call time so nothing rides on a stale listing.

Docs

docs/ARCHITECTURE.md §1/§5/§9/§12/§14 updated to describe both channels, the preference order, and the known limitation that a client can declare elicitation and auto-decline (older Codex behavior) — which fails closed, the acceptable outcome.

Tests

  • policy-and-audit.integration.test.ts — new policy: prompt via MCP elicitation block: flag suppression for elicitation-capable clients; accept → consent: "elicitation" audited and the prompt message carries tool/alias/args; decline and cancel (parameterized) → app never contacted, agent gets an isError result, no audit record fabricated; capability-declared-but-request-fails → policy_denied / no_consent_channel; and a fast injectable-timeout case (elicitationTimeoutMs test-only seam).
  • tool-mapping.test.ts (new) — pure unit test for the flag-suppression seam in toMcpTool.

Validation

  • pnpm typecheck / pnpm build — clean across all 3 packages.
  • pnpm lint — 0 errors (194 pre-existing warnings, all in @cordierite/react-native, untouched).
  • cordierite: 34 files / 309 tests passed (non-e2e; the one e2e failure in daemon-restart.e2e.test.ts reproduces identically on unmodified main — environment timing, unrelated). shared: 4 files / 122 tests passed.

Related: #10 (parent design), #14 (flag-based channel, merged as #16), #12 (local approval channel, still parked).

🤖 Generated with Claude Code

https://claude.ai/code/session_01Bry2KSYdEDaHrkgjGvmCJw


Generated by Claude Code

…nel 1)

Adds a second MCP consent channel for tools.call's "prompt" policy gate,
alongside the existing requiresUserInteraction flag (issue #14):

- packages/cordierite/src/mcp/server.ts: negotiates the channel per
  connection from the client's declared MCP capabilities
  (Server.getClientCapabilities()). When a client declares `elicitation`,
  tools/list never emits _meta["anthropic/requiresUserInteraction"] for
  "prompt" tools (the two channels never both arm for one call), and
  callProxiedTool's consent resolution is factored into a single
  resolveToolCallConsent helper shared by both the plain and
  progress-tracked call paths. On a "prompt" call it sends one
  elicitation/create request naming the tool, session alias, and the call's
  arguments (JSON, truncated past 2KB), bounded server-side at 10 minutes.
  Accept forwards consent: "elicitation" to the daemon; decline/cancel/
  timeout short-circuit to an MCP tool result with isError: true naming
  what happened, without ever calling the daemon; a failed elicitation
  request (declared but unsupported, transport error) is treated as no
  consent obtained and falls through to the daemon's existing
  policy_denied/no_consent_channel path — never as approval.
- packages/cordierite/src/mcp/tool-mapping.ts: toMcpTool's flag-emission
  boolean is now handed the already-arbitrated decision from server.ts
  rather than only the client-compliance check.
- packages/cordierite/src/daemon/daemon.ts, packages/cordierite/src/daemon/audit.ts,
  packages/shared/src/domains/rpc.ts: widen tools.call's consent param and
  the audit record's consent field from "client" to "client" | "elicitation";
  the daemon treats either value as satisfying the "prompt" gate the same
  way, and the audit log now distinguishes an observed elicitation decision
  from the upstream flag-based gate. packages/cordierite/src/daemon/config.ts's
  doc comment updated to match.
- docs/ARCHITECTURE.md: §12 documents both channels and their preference
  order, plus the known limitation that a client can declare elicitation
  and always auto-decline (older Codex behavior) — an acceptable fail-closed
  outcome. §1, §9, and §14 updated so they no longer describe elicitation as
  unimplemented.
- Tests: packages/cordierite/src/__tests__/policy-and-audit.integration.test.ts
  gains a "policy: prompt via MCP elicitation" describe block (flag
  suppression, accept, decline/cancel, unsupported-despite-declared, and a
  fast injectable-timeout case); packages/cordierite/src/__tests__/tool-mapping.test.ts
  is a new pure unit test for the flag-suppression seam in toMcpTool.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bry2KSYdEDaHrkgjGvmCJw
@V3RON
V3RON force-pushed the claude/human-loop-sensitive-tools-rioudz branch from 9e8e656 to 8393202 Compare September 5, 2026 21:20
- `mcp/tool-mapping.ts`: main's per-server mapper factory (#42's schema gate)
  keeps its structure; this branch's rename of the flag argument to
  `emitRequiresUserInteractionFlag` carries over, since `mcp/server.ts` folds
  the elicitation preference into it before calling.
- `mcp/server.ts`: keeps both the consent resolver and the mapper instance, and
  the `tools/list` handler maps through the mapper with this branch's
  suppress-the-flag-when-elicitation-is-preferred decision.
- `tool-mapping.test.ts`: main's schema-gate suite plus this branch's flag-gate
  cases, rewritten against `createMcpToolMapper`.
- ARCHITECTURE §12/audit: one record entry describing both consent channels and
  the retention schedule from #43.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JjqwCbdEfwSQXwHnNth2ge
@V3RON
V3RON merged commit f701052 into main Sep 6, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants