Skip to content

agent: MCP Skills extension (SEP-2640) client support - #131

Merged
jaredLunde merged 1 commit into
mainfrom
jared/mcp-skills
Oct 7, 2026
Merged

jaredLunde merged 1 commit into
mainfrom
jared/mcp-skills

Conversation

@jaredLunde

@jaredLunde jaredLunde commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

What

This PR adds client-side support for the MCP Skills extension (SEP-2640, io.modelcontextprotocol/skills, Final) to crates/agent.

When a connected server declares the extension, its skills now work like this:

  • They are listed for the model and in serve's get_commands.
  • The model can load them and read their files, with every read checked against the server's manifest.
  • They can be invoked with /skill:<server>:<name>.
  • set_mcp_enabled gates them.
  • Loading a skill the model chose needs the user's approval. So does running code, or reading from another server, while the session is acting on a skill.
  • Their state belongs to the session that loaded them, and survives resume, switch, fork and restart.
  • The listing is re-fetched when it is due, and cached in the idle-reap manifest.

Servers that don't declare the extension behave exactly as before.

The code lives mostly in tools/mcp_skills.rs. It is rebased onto MCP Events (#133), MCP Tasks (#130) and MCP Apps (#132): stdio servers go through their one stdio transport (tools/mcp_stdio.rs), the rmcp _meta bug has one workaround (mcp_stdio::rescue), and nested requests are attributed with track_call. tools/mcp_wire.rs keeps only the streamable-HTTP client that applies that same rescue to skills/* responses.

Design

Server state vs. session state.

  • ServerSkills exists once per connection. It holds what the server published: the listing, and every entry learned since.
  • SkillSession exists once per session, on McpEnabledSet. It holds what that session loaded, what its user approved, and who to ask.
  • Each session's registry rebinds the skill tools to its own SkillSession. Subagents share their parent's.
  • The acting window is rebuilt from the transcript before every run. A loaded skill enters context under a <skill name server location manifest> tag. As long as that tag is in context, the session counts as acting on that skill, whether it was resumed, switched to, forked or reopened after a restart. The spec allows the window to run longer than the skill's time in context, never shorter.
    • Only host-authored records count: a skill-tool result matched to its mcp__<server>__skill__read call by tool_use_id, or a user turn that is itself a /skill: expansion.
    • A tag that some other tool, a file or the model reproduced activates nothing.

Approval (the spec's "no implicit local execution").

  • Activation. A skill the model chose loads only after the user approves it. The question comes before SKILL.md is fetched, and the answer is bound to the manifest fingerprint (full SHA-256 over every {uri, digest}), so a changed skill is asked about again.
    • A /skill: the user types counts as consent.
    • A nested skill always needs its own approval.
    • The question carries the entry's frontmatter and file manifest, so the user can inspect the skill before it is loaded.
  • Code execution. While the session is acting on an MCP skill, every bash/execute call needs approval, asked per call and naming the active skills.
    • A "session" answer covers only that exact call (the verbatim command, or a digest of the whole input for execute), and only under that exact set of manifests.
  • Cross-origin reads. While acting on a skill from server A, any resource read from server B needs approval per call, naming both servers. These answers are never remembered.
  • Who answers.
    • serve sends approval_request frames with an mcp_skill object, through its existing approval gate (now always present).
    • run has no one to ask, so it denies.
    • --approve-mcp-skills (on both run and serve) approves everything in advance.
    • The gate is installed in all three hook paths: ServeHooks, RunHooks and ChildHooks.
  • allowed-tools is not honoured for any skill.

Integrity.

  • Every read is verified: byte size, then SHA-256, then for SKILL.md a field-by-field frontmatter comparison.
  • Frontmatter is compared as YAML values, not as JSON text: 1 equals 1.0, YAML 1.1's yes equals true, and the null spellings are equal. This deliberately loosens the spec's "identical in content" from one JSON rendering to the values. The fields the host acts on (disable-model-invocation) are read with the same normalization, so anything that verifies as true is treated as true.
  • If a read fails verification, the entry is refreshed once with skills/get; after that the read is refused.
  • A file that isn't listed in the manifest is refused, as is a content block labelled as a different resource.

No early fetch. Discovery reads only skills/list.

Freshness.

  • The listing is re-fetched when it is needed and due: its ttlMs ran out, or the server sent notifications/resources/list_changed. SEP-2640 has no skills-specific notification; notifications/skills/list_changed is accepted as well.
  • A re-list that fails keeps the last good listing. A notification stays pending until a re-list succeeds, so a re-list cut short doesn't lose it.
  • cacheScope: "private" listings are never cached on disk.

Limits.

  • Skills over 512 files or 16 MiB are declined, and so are dynamic skills.
  • None of these are offered as generic resource tools: the declined skills' files, anything under skill://, and any .../SKILL.md under any scheme together with everything under its directory.

Diagnostics. These appear in get_commands' collisions and as run warnings, and survive a boot from the cache:

  • Declined or invalid entries.
  • A failed listing.
  • A listing cut off at the 64-page cap.
  • Duplicate names.
  • Clashes with on-disk skill names.

Names. Skills are named <server>:<name>. Two same-named entries on one server become <server>:<skill-path>. They never shadow local skills. --no-skills suppresses MCP skills too.

/skill: during a run. Steers and follow-ups received during a run go through PendingSteers:

  • Each takes its place at receipt and is queued and acked strictly in arrival order, so a plain steer never overtakes an earlier /skill: one.
  • A /skill: expansion runs on its own task, so the busy loop keeps polling the run and reading abort/approve.
  • Expansions are bound to their run. When the run ends (finished, aborted, or cancelled by new_session/switch_session/fork/clone), any expansion still in flight is cancelled and awaited, and acked as not queued with the reason.
  • An expansion opens no acting window itself. The skill is activated only when its text is actually queued, so nothing reaches a later prompt or another session.

The rmcp bug: one workaround, shared with Events

The bug. rmcp 3.2–3.5.1 decodes a result that carries _meta (which every 2026-07-28 result does) as an empty CallToolResult, so skills results disappear (rust-sdk#1197, open). The Python SDK's reference server exposed it.

The workaround is mcp_stdio::rescue / unwrap_rescued, merged with #133. A lossy result is wrapped as {"x-beyond-raw-result": …} before rmcp parses it, and each custom-request caller unwraps it. This PR's own stdio transport and its _meta stripping are gone.

  • stdio: main's tools/mcp_stdio.rs, unchanged. It applies the rescue to every line, skips non-UTF-8 lines, and keeps the exit grace and process-group sweep.
  • HTTP: mcp_wire::HttpClient answers skills/* POSTs itself and applies the same rescue. It builds requests and handles statuses the way rmcp does:
    • reserved headers are refused;
    • 401 maps to AuthRequired, 403 to InsufficientScope, and 404 with a session to SessionExpired;
    • SSE responses go back to rmcp as a bounded stream, rescued event by event, so rmcp still routes everything else on that stream. Event boundaries are found byte by byte across chunks, for \n, \r\n and \r alike.
    • JSON bodies are read up to the same cap.

Nested-request attribution (rebased onto #130)

Every skills request (skills/list, skills/get, resources/read) now goes through mcp::request_tracked. It registers the request with track_call under the calling session's host for as long as the request is outstanding. A nested request the server raises meanwhile is attributed to that session or refused. It is never sent to a shared connection's own host.

  • A skill-tool call is scoped by its run.
  • A /skill: steered into a running prompt expands as its session (McpSkills::in_run). That session's command loop is live, so it can answer.
  • A re-list, or an expansion between runs, runs under a host with no client (McpSkills::outside_run), so a nested request is declined at once. Nothing could answer it: serve's command loop is waiting on that very read, so routing the request to the session only stalled the prompt until the 30s expansion timeout. A first version did exactly that, and the new test caught it.

Composition with MCP Apps (rebased onto #132)

  • Resource filtering: tools_from_client drops skill resources from the generic resource tools, and Apps' ui:// filter runs in the same function. Both stay. mcp_skills::run_lists_mcp_skills_without_fetching_any_skill_file now also serves a ui:// resource and asserts it is not a resource tool either.
  • Notifications: a single on_resource_list_changed handler clears the Apps views and also invalidates the skills listing.
  • Dial: the skills invalidation flag rides the Dial into McpHandler::for_dial, next to apps.
  • Session switch: reset_session (enablement and skill state) runs alongside the Apps close.
  • Apps sessions: a skill tool coming from the apps view is rebound to the session's skill state by name, the same way a plain one is.

Audit findings: all fixed, each with its proving test

Every test below fails without its fix. Each fix was reverted in turn by a scripted mutation run, and all 36 mutations were killed:

  • Round 1 (23): F1–F12 plus the surviving mutants M3, M5+M5b, M12b, M16 and M24.
  • Round 2 (7): N-A, N-B, N-C, N-D, N-E, and N-F (two mutations).
  • Round 3 (3): plain steers surviving an abort, R2 (activating on expansion), and an expanded steer queued after an abort.
  • Rebase (3): a steered expansion left unscoped, an expansion between runs routed to the session, and a skills request left untracked.
Finding Fix Proving test(s)
F1 gate lost on resume/switch/fork/restart transcript-rebuilt acting window mcp_skills_resume: a_resumed_run_is_still_gated, switching_back_to_a_session_that_loaded_a_skill_restores_the_gate, a_fork_of_a_session_that_loaded_a_skill_is_gated (clone + fork), a_restarted_serve_reopening_the_session_is_gated
F2 /skill: steer freezes the busy loop expanded on its own task mcp_skills_steer::a_skill_steered_mid_run_does_not_stall_the_command_loop (slow fixture read, get_state answered in under 3s)
F3 non-UTF-8 line kills the stdio pump now main's mcp_stdio transport, which skips bad lines mcp_skills_integrity::a_stdio_server_writing_non_utf8_lines_keeps_working (bad line before the handshake and before every reply), against main's transport
F4 HTTP SSE read waits for stream close streamed, bounded, routed by rmcp mcp_wire::an_sse_response_is_streamed_rescued_and_does_not_wait_for_the_server_to_close, an_oversized_sse_event_ends_the_stream_with_an_error
F5 HTTP status ignored rmcp-equivalent mapping mcp_wire::statuses_map_to_the_errors_rmcp_acts_on, a_reserved_custom_header_is_refused
F6 execute keyed as tool:execute key on the full input mcp_skills::execute_is_gated_and_remembered_per_program_not_per_tool (also kills M12b)
F7 bypass via generic resource tools / cross-origin declined and skill:// files hidden; per-call cross-origin gate mcp_skills::run_lists_mcp_skills_without_fetching_any_skill_file; mcp_skills_approval: run_denies_a_cross_origin_read_while_acting_on_another_servers_skill, serve_asks_per_call_before_a_cross_origin_read
F8 frontmatter check unproven (M4) e2e tampered frontmatter mcp_skills_integrity::tampered_frontmatter_is_refused_on_the_real_load_path
F9 tests hang on regression reader thread plus 60s deadlines every Serve wait in tests/common/skills_env.rs (exercised by the mutation run: regressions fail, not hang)
F10 type-strict frontmatter YAML-value equality mcp_skills_integrity::an_honest_skill_is_not_refused_for_yaml_versus_json_representation, mcp_skills::frontmatter_equality_is_of_yaml_values_not_their_rendering
F11 failed/cut-short re-list keep good listing; generation-counted notices mcp_skills_integrity: a_failed_re_list_keeps_the_last_good_listing, a_re_list_cut_short_keeps_the_notice_pending
F12 unlisted nested by URI load first, file read only on -32602 mcp_skills_integrity::an_unlisted_nested_skill_inside_a_loaded_one_can_be_activated_by_uri
F12 --no-skills suppresses MCP skills mcp_skills_integrity::no_skills_suppresses_mcp_skills_too
F12 page cap diagnostic mcp_skills_integrity::a_listing_that_never_ends_is_cut_off_and_said_so
F12 content uri exact match required mcp_skills::a_content_block_labelled_as_another_resource_is_refused
F12 fingerprint width full SHA-256 mcp_skills::the_fingerprint_is_a_full_sha256
F12 dead code handle_approve's no-gate branch removed (the gate always exists) n/a (compiler)
F12 inspect before load frontmatter plus manifest_files in the question mcp_skills_approval::serve_asks_to_activate_a_skill_and_to_run_code_while_it_is_active
M3 size independent of digest — mcp_skills_integrity::a_wrong_size_is_refused_even_when_the_digest_matches
M5 unlisted file read — mcp_skills_integrity::a_file_the_loaded_skills_manifest_does_not_list_is_never_read
M16 approval bound to manifest — mcp_skills_approval::a_skill_whose_manifest_changed_is_asked_about_again
M24 verify-fail refresh retry — mcp_skills_integrity::a_held_entry_that_goes_stale_is_refreshed_and_the_current_body_served
N-A /skill: steer leaked across abort and sessions, acked after its run ended expansion bound to the run: cancelled and awaited at run end, acked not queued; activation only when queued mcp_skills_steer: a_skill_steer_from_an_aborted_run_never_reaches_the_next_session (the audit's steer_leak sequence: no body, and no active skill gating bash), a_skill_steer_from_an_aborted_run_never_reaches_the_next_prompt, a_skill_steered_mid_run_does_not_stall_the_command_loop (late ack is success: false with the reason)
N-B ordering inversion FIFO slots reserved at receipt, acks in order mcp_skills_steer::a_plain_steer_does_not_overtake_an_earlier_skill_steer (the audit's steer_order sequence)
N-C disable-model-invocation: yes read with the verifier's normalization mcp_skills_integrity::a_yaml_1_1_disable_model_invocation_is_honoured
N-D phantom skill from a forged tag on resume restore only from host-authored records mcp_skills_integrity::a_forged_skill_tag_in_other_content_activates_nothing_on_resume
N-E unlisted SKILL.md under another scheme stayed a generic resource any .../SKILL.md and its directory hidden; loads go through the skill tool mcp_skills_integrity::an_unlisted_skill_md_under_another_scheme_is_not_a_generic_resource
N-F SSE bounding saw \n\n only within one chunk; JSON body uncapped byte-level boundary tracking across chunks (\n/\r\n/\r); capped body mcp_wire: event_boundaries_are_found_across_chunks_and_line_endings, a_small_events_stream_split_with_crlf_is_not_mistaken_for_an_oversized_one, a_json_body_over_the_cap_is_refused
R3 plain steers queued behind a cancelled /skill: survived into the next prompt on a cancelled run, everything still waiting is acked not queued (the run already cleared its lanes) mcp_skills_steer::a_plain_steer_behind_a_skill_steer_dies_with_the_aborted_run
R2 activating on expansion instead of when queued survived pinned: an expansion that finished but never reached the model activates nothing mcp_skills_steer::a_finished_expansion_that_was_never_queued_does_not_activate_its_skill (a fast skill done behind a slow one at abort; the next bash runs ungated)
Rebase: skills requests unattributed request_tracked + in_run / outside_run mcp_skills_steer: a_nested_request_during_a_steered_skill_expansion_reaches_the_session, a_nested_request_during_an_expansion_between_runs_is_refused_at_once
Rebase: one rescue boundary mcp_stdio::rescue for stdio and skills HTTP mcp_wire::rmcp_still_shadows_a_skills_result_that_carries_meta, the SSE test above, the Python SDK runs below

Proof

The end-to-end tests drive the real run/serve binaries against a real subprocess fixture, src/bin/mcp_skills_fixture_server.rs (stdio or --http, _meta on every result, request log). Every child is started with spawn_guarded.

Local results (f514ebf, rebased on 7b7e9ef):

  • Lib plus all 35 tests/mcp_*.rs targets: 1730 passed, 0 failed. That covers skills, apps, tasks, events and mcp_stdio_transport.
    • Skills suites: skills 9, skills_approval 7, skills_integrity 14, skills_listing 4, skills_resume 4, skills_sessions 2, skills_steer 8.
  • clippy (--all-targets --features code-mode -D warnings), cargo fmt --check and dprint check: clean.
  • Mutations re-run on this commit: the 6 from round 3 and the rebase, all killed.

CI (f514ebf): green on the first attempt: all 20 jobs passed (run 37557496219), including the three agent test shards: agent-lib, agent-rest and agent-code-mode, which finished in under 5 minutes each. This is the first CI run on this branch in which the e2e tests completed.

Independent checks (run locally, not in CI)

  1. Official conformance suite (conformance @ c37eec8, the SEP-2640 client scenarios): 4/4 passed. The driver uses run --approve-mcp-skills; without the flag, a load is correctly denied before any read.
    • no-prefetch
    • verify-digest
    • verify-size
    • verify-frontmatter
  2. Python SDK reference server (python-sdk#3485): passes over stdio, and over streamable HTTP in both JSON and SSE modes. The SSE mode exercises the new streaming path.

Remaining limits

  • Approvals are not persisted across restarts. After a restart, the restored acting window asks again, which is the conservative direction.
  • Activation questions report origin: "main". The skill tool has no per-call origin. Code-execution and cross-origin questions do name the subagent.
  • A refreshed listing isn't written back to the on-disk cache until the next live connect.

🤖 Generated with Claude Code

https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk

jaredLunde added a commit that referenced this pull request Oct 6, 2026
…agnostics

Closes the gaps PR #131 listed against SEP-2640:

- Approval (spec MUST, no implicit local execution). Activating a
  model-chosen MCP skill needs the user's approval, bound to the entry's
  manifest fingerprint and asked before SKILL.md is fetched; a `/skill:`
  the user typed is that consent. A nested skill's approval is its own and
  the question names the enclosing skill. While a session is acting on an
  MCP skill, every `bash`/`execute` call needs approval too, in serve,
  run and subagent hook paths. serve asks through its approval gate (now
  always present) as `approval_request` frames with an `mcp_skill`
  object; run has no one to ask and denies; `--approve-mcp-skills`
  approves in advance.
- Loaded-skill state is per session (`SkillSession` on `McpEnabledSet`),
  not per connection: a daemon's sessions share connections, so each
  session's registry rebinds the loading tools to its own state, reset on
  a session switch and shared with its subagents.
- Listings refresh when needed and due: `ttlMs` expiry (absent/0 = stale)
  or `notifications/resources/list_changed`; held entries past `ttlMs` are
  re-fetched before a load; `cacheScope: "private"` listings are never
  written to the on-disk manifest.
- Subagents get their parent session's MCP skills listing.
- The 512-file / 16 MiB per-skill limits are enforced from the entry.
- Declined/invalid entries, failed listings, same-name entries and
  on-disk name clashes are surfaced via get_commands' `collisions` and
  run's warnings, and survive a boot from the manifest cache.

New e2e suites: mcp_skills_approval, mcp_skills_sessions,
mcp_skills_listing (shared harness in tests/common/skills_env.rs).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
@jaredLunde

Copy link
Copy Markdown
Contributor Author

CI status for 8e89010: build, clippy, fmt, all mutants shards, core-providers, exec-live, gateway-unit, gateway-e2e and gateway stress pass. The three agent test shards (agent-lib, agent-rest, agent-code-mode) have not completed in CI: they wedged and were cancelled at 30:00 twice on this commit (22:00 and 22:31 UTC), as they did on fcd72c2. In the same 22:16–22:54 window the agent shards of the Apps, Events and Tasks PRs were cancelled the same way, so this looks environmental, but none of this PR's e2e tests have completed in CI yet. Locally on 8e89010: cargo nextest --profile ci with the agent-rest filter passed 424/424; all tests/mcp_*.rs passed 81/81; lib mcp_ unit tests passed 38/38. I've requested another rerun.

Client side of `io.modelcontextprotocol/skills`: a declaring server's
`skills/list` entries become `<server>:<name>` skills in a separate
`<available_skills origin="mcp">` block, loaded (size, digest and
frontmatter verified, never prefetched) through
`mcp__<server>__skill__read`, invocable as `/skill:<server>:<name>`, listed
by `get_commands`, re-listed on `ttlMs` expiry or list-changed, and cached
in the manifest. Activation of a model-chosen skill and code execution
while one is active need the user's approval (`run` denies unless
`--approve-mcp-skills`); state is per session and restored only from
host-authored transcript records.

Rebased onto MCP Events (#133): `tools/mcp_stdio.rs` is the one stdio
transport, and its `rescue`/`unwrap_rescued` is the one workaround for
rmcp's untagged-result bug (rust-sdk#1197), for events and skills results
alike. `tools/mcp_wire.rs` keeps only the streamable-HTTP client, which
applies the same `rescue` to `skills/*` JSON bodies and SSE events.

Rebased onto MCP Tasks (#130) too: every skills request (skills/list,
skills/get, resources/read) goes through `mcp::request_tracked`
(`track_call` under the calling session's host), so a nested request is
attributed or refused, never sent to a shared connection's own host. A
steered `/skill:` expands as its session (its command loop can answer);
a re-list or expansion between runs runs under a client-less host, so a
nested request is declined at once instead of stalling the prompt.

Rebased onto MCP Apps (#132): the skills resource hiding and the Apps
`ui://` filter both run in `tools_from_client`; one resource-list-changed
handler clears the Apps views and invalidates the skills listing; and the
skills invalidation flag rides the dial into `McpHandler::for_dial`.

Round-3 fixes: a steer waiting behind a `/skill:` when its run is
cancelled is acked as not queued (the cancellation already dropped the
run's lanes), and an expansion that finished but was never queued
activates nothing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
@jaredLunde
jaredLunde merged commit 7271b83 into main Oct 7, 2026
20 checks passed
@jaredLunde
jaredLunde deleted the jared/mcp-skills branch October 7, 2026 01:41
jaredLunde added a commit that referenced this pull request Oct 7, 2026
* agent: MCP skills follow-ups, and the serve suite back in CI

MCP skills (SEP-2640) gaps from #131:
- A subagent's skill-load question names that subagent (`origin`), the
  same provenance its code-execution questions carry: its registry binds
  the loading tools in its own name (`filter_by_enabled_as`).
- A listing re-fetched after connect is written back to the on-disk
  manifest cache at once (`mcp_manifest::store_skills`), or forgotten if
  it is now `cacheScope: private`.
- Remembered approvals persist with the session (`mcp_skill_approval`
  custom entries, journaled at run end) and are restored before each
  prompt. Keys embed the content fingerprint, so an unchanged skill is
  not asked about again after a restart and a changed one is.

The serve test suite:
- `serve_get_tree_since_works_from_the_busy_loop_mid_prompt` hung on main
  too, deterministically: the session-title request serve makes after a
  first run took a reply from the in-order mock model's script, shifting
  every later turn. The same cause failed or hung 48 serve tests, unseen
  because CI has excluded `binary(~serve)` since #39. The scripted mock
  servers now answer title requests themselves, off the script and out of
  the record; title tests match `SESSION_TITLE_MARKER` on a routed server.
- Two real bugs it had been hiding: session listings tied on `updated_at`
  (one-second resolution) had no stable order, so `offset` paging could
  repeat a session (tie broken by id); and since #131 an `approve` in a
  session without `--approve` was acknowledged even when it matched no
  question (refused again, pointing at `--approve`).
- CI runs the serve tests again, as an `agent-serve` shard.

Lib tests that assert "no enclosing repo / context file" now use a temp
root with no project above it (`test_support::isolated_tempdir`), so they
pass with TMPDIR inside a checkout or under a CLAUDE.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk

* agent: approvals never persisted, host custom kinds reserved, race-free test ports

Audit of 237d9a5:
- F1: persisted MCP-skill approvals were forgeable. A model with
  write/edit tools can append an `mcp_skill_approval` entry to the
  user's session file (it knows the fingerprint from the loaded tag),
  and a client could `append_custom` one; an `execute` approval would
  then run unasked after a restart. Approvals are no longer persisted
  or restored: session-lifetime, in memory. A client's `append_custom`
  of a kind the agent writes itself (`mcp_task`, `mcp_task_result`,
  `mcp_skill_approval`: `session_store::HOST_CUSTOM_KINDS`) is refused.
- F4: a typed `/skill:` is consent for that invocation only; it no
  longer approves the model's later loads of the skill.
- F3: same-second session listings tie-break by id descending, so the
  newest (generated ids lead with their creation time) comes first.
- F5: manifest-cache writes are read-modify-writes of one shared file;
  they now run under a lockfile with a per-writer temporary name.
- F6: `serve_state_reporting` reads stdout with a deadline instead of
  hanging on a stall, and the scripted servers count title calls, so a
  change in title frequency shows (one per session, pinned).

Test ports: `free_port()` released the port it picked and hoped the
child bound it first; under parallel suites another process could take
it. `serve` children now bind `--listen 127.0.0.1:0` and the test reads
the port from serve's announcement (`spawn_listening`); a daemon whose
own arguments name its port (an MCP Events callback URL) is handed a
listener the test bound, by socket activation (`HeldPort`), kept across
a restart; a "nothing listens here" port is held bound, never listening
(`DeadPort`). `free_port` remains only for the gateway and nats-server.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
jaredLunde added a commit that referenced this pull request Oct 7, 2026
…title-proof test mock

- run journals the MCP tasks it starts (the same mcp_task entries serve
  writes) and `run --continue` resumes the session's pending tasks before its
  turn, journaling each result and splicing it into the turn it sends.
- A session file's mode is tightened to 0600 by the next append if it was
  looser (create already set 0600; an existing file kept its mode). Task
  journal entries are confirmed sealed with the tenant key in service
  storage and read back intact.
- The in-order test mock answers serve's session-title call outside its
  script and does not record it: the title call raced the next prompt and
  took its scripted reply under load (serve_session_tree, serve_compaction_
  retry, serve_stop_after_turn and others). The title tests route titles.
- The sessionId ownership rule now has an e2e test that fails without it (a
  session file copied under a new id carries the journal; no fork does).
- serve_approval: a session without --approve has had a gate since #131
  (MCP skills ask through it); the stale "no gate" expectation is updated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant