Repository navigation
BrokenResourceError race condition in stdio_client cleanup when context exits quickly #1960
Description
Activity
- addedbugSomething isn't workingSomething isn't workingneeds reproneeds additional information to be able to reproduce bugneeds additional information to be able to reproduce bug
on Feb 10, 2026 Thanks for the detailed analysis! Could you provide a single-file runnable script that reproduces this? Would help us validate and prioritize the fix.
I was able to reproduce this locally on current
main(b33c811).I used this minimal script from the repo root:
import sys import textwrap import anyio from mcp.client.stdio import StdioServerParameters, stdio_client SERVER_SCRIPT = textwrap.dedent(""" import sys import time sys.stdout.write('{"jsonrpc":"2.0","id":1,"result":{}}\n') sys.stdout.flush() time.sleep(2.0) """) async def main() -> None: server_params = StdioServerParameters( command=sys.executable, args=["-c", SERVER_SCRIPT], ) with anyio.fail_after(5.0): async with stdio_client(server_params) as (_read_stream, _write_stream): await anyio.sleep(0.2) anyio.run(main)
Run command:
python3 -m uv run python repro_1960.py
For me this fails consistently on
mainwhen the client exits without consumingread_stream.Traceback excerpt:
ExceptionGroup: unhandled errors in a TaskGroup (1 sub-exception) ... File ".../src/mcp/client/stdio.py", line 160, in stdout_reader await read_stream_writer.send(session_message) ... anyio.BrokenResourceErrorI also ran the same script against the branch from #2219, and it exited cleanly there.
Root Cause Analysis
I traced through the code in
src/mcp/client/stdio/__init__.pyon the currentmainbranch. The race condition is straightforward:Timeline of the bug:
stdio_clientyields(read_stream, write_stream)to the caller (line 189)stdout_readerruns in the task group, reading process stdout and callingawait read_stream_writer.send(session_message)(line 162) — this is a zero-buffer memory stream, sosend()blocks until a consumer callsreceive()- The caller exits the context without consuming
read_stream(or exits quickly) - The
finallyblock executes (line 190). It terminates the process, then callsawait read_stream_writer.aclose()(line 215) stdout_readeris still alive inside the task group, blocked onsend(). The stream it's sending into just got closed from the other end →BrokenResourceError- The existing
excepton line 163 only catchesClosedResourceError, notBrokenResourceError, so the exception propagates into theExceptionGroup
Key distinction:
ClosedResourceErroris raised when you call.send()on a stream handle that you already closed.BrokenResourceErroris raised when the receiving end of the stream is closed. Thefinallyblock closesread_stream(the receive end, line 213) andread_stream_writer(the send end, line 215). Ifstdout_readeris mid-send()when the receive end closes, it getsBrokenResourceError, notClosedResourceError.Proposed Fix
The proper fix requires both a structural change and a defensive catch. Either alone is incomplete:
1. Cancel the task group scope before closing streams (structural fix)
finally: # Cancel background tasks FIRST so they're not racing against stream teardown tg.cancel_scope.cancel() # MCP spec: stdio shutdown sequence (existing code follows) if process.stdin: ...
Adding
tg.cancel_scope.cancel()at the top of thefinallyblock causes anyio to deliverCancelledtostdout_readerandstdin_writerat their next checkpoint (i.e., the blockedsend()call). This ensures the tasks are winding down before the streams are closed.2. Catch
BrokenResourceErroralongsideClosedResourceError(defensive fix)async def stdout_reader(): ... except (anyio.BrokenResourceError, anyio.ClosedResourceError): await anyio.lowlevel.checkpoint()
Same change for
stdin_writer. This is defense-in-depth: even with the cancel-first approach, there's a narrow window where the task could be between checkpoints when the stream closes. Catching both error types makes the cleanup fully robust.Why both are needed
- Cancel-first alone: there's still a theoretical window between the cancel signal and the task actually reaching a checkpoint. If a stream close happens in that window,
BrokenResourceErrorstill escapes. - Catch-only alone (what PR fix: avoid stdio cleanup BrokenResourceError race #2219 did): it suppresses the symptom but doesn't address the root ordering problem. The tasks still run concurrently with stream teardown, which could cause other subtle issues if anyio's internal invariants change.
How This Prevents the Race Condition
With both changes applied:
finallyfires →tg.cancel_scope.cancel()marks all tasks for cancellationstdout_readeris blocked onsend()→ anyio deliversCancelled, the task exits itsasync with read_stream_writer:block cleanly- Stream
.aclose()calls now operate on streams that no active task is using - Even if timing is unlucky, the
BrokenResourceErrorcatch prevents any exception from leaking into anExceptionGroup
Note: PR #2219 was closed because the diff only contained the defensive catch while the description promised the structural reordering. A complete fix should include both changes.
- added a commit that references this issue
on Mar 11, 2026 Paired PRs with a fix + regression test:
main(V2): fix(stdio): handle BrokenResourceError in stdout_reader race (#1960) #2450v1.x: [v1.x] fix(stdio): handle BrokenResourceError in stdout_reader race (#1960) #2449
Fix: wrap both
read_stream_writer.send(...)call sites instdout_readerwithtry/except (ClosedResourceError, BrokenResourceError): return, and widen the outerexceptto the same union.ClosedResourceErroris the sibling class (raised on already-closed streams);BrokenResourceErroris the one raised when the receiver is closed during an in-flight send — the exact shape of this shutdown race. No API changes.Tests:
test_stdio_client_exits_cleanly_while_server_still_writingintests/client/test_stdio.py— spawns a subprocess that emits a burst of JSONRPC notifications, exits thestdio_clientcontext immediately, asserts noExceptionGrouppropagates. Wrapped inanyio.fail_after(5.0)per AGENTS.md. Fails before the patch, passes after.Repro trigger used to verify: an MCP stdio server that emits a few
notifications/messageframes on startup (jules-mcp-serverin our case), driven by a caller that opens a freshstdio_clientper invocation (mcp2cli). Every tool call used to surface asExceptionGroup→anyio.BrokenResourceError. After the patch, tool calls return cleanly.
GitHub Issue: MCP SDK stdio_client Race Condition
Repository
modelcontextprotocol/python-sdkTitle
BrokenResourceErrorrace condition instdio_clientcleanup when context exits quicklyBody
Description
When the
stdio_clientasync context manager exits quickly (before the subprocess has finished outputting data), a race condition occurs between thestdout_readertask and the cleanup code in thefinallyblock, resulting inBrokenResourceError.Reproduction
This occurs in scenarios where:
stdio_clientfinallyblock closesread_stream_writerwhilestdout_readeris mid-sendError
Root Cause
In
mcp/client/stdio/__init__.py, thestdio_clientfunction:The
finallyblock closesread_stream_writerbefore the TaskGroup has cancelled its tasks. Ifstdout_readeris in the middle ofawait read_stream_writer.send(...), it receivesBrokenResourceError.Suggested Fix
Cancel the TaskGroup's scope before closing the streams:
Alternatively, wrap the stream operations in
stdout_readerwithBrokenResourceErrorhandling:Environment
Workaround
We're currently serializing MCP lifecycle operations with an asyncio.Lock to prevent overlapping enter/exit operations, which avoids triggering the race condition.