Skip to content

BrokenResourceError race condition in stdio_client cleanup when context exits quickly #1960

Description

@newsbubbles

GitHub Issue: MCP SDK stdio_client Race Condition

Repository

modelcontextprotocol/python-sdk

Title

BrokenResourceError race condition in stdio_client cleanup when context exits quickly

Body

Description

When the stdio_client async context manager exits quickly (before the subprocess has finished outputting data), a race condition occurs between the stdout_reader task and the cleanup code in the finally block, resulting in BrokenResourceError.

Reproduction

This occurs in scenarios where:

  1. A subprocess is spawned via stdio_client
  2. The subprocess starts outputting data (e.g., initialization messages)
  3. The calling code exits the context quickly (e.g., user disconnects)
  4. The finally block closes read_stream_writer while stdout_reader is mid-send

Error

ExceptionGroup: unhandled errors in a TaskGroup (1 sub-exception)
  +-+---------------- 1 ----------------
    | Traceback (most recent call last):
    |   File "mcp/client/stdio/__init__.py", line 162, in stdout_reader
    |     await read_stream_writer.send(session_message)
    |   File "anyio/streams/memory.py", line 242, in send
    |     self.send_nowait(item)
    |   File "anyio/streams/memory.py", line 213, in send_nowait
    |     raise BrokenResourceError
    | anyio.BrokenResourceError
    +------------------------------------

Root Cause

In mcp/client/stdio/__init__.py, the stdio_client function:

async with (
    anyio.create_task_group() as tg,
    process,
):
    tg.start_soon(stdout_reader)
    tg.start_soon(stdin_writer)
    try:
        yield read_stream, write_stream
    finally:
        # ... process cleanup ...
        await read_stream_writer.aclose()  # <-- Closes stream
        await write_stream_reader.aclose()

The finally block closes read_stream_writer before the TaskGroup has cancelled its tasks. If stdout_reader is in the middle of await read_stream_writer.send(...), it receives BrokenResourceError.

Suggested Fix

Cancel the TaskGroup's scope before closing the streams:

finally:
    # Cancel background tasks before closing streams
    tg.cancel_scope.cancel()
    
    # ... existing cleanup ...
    await read_stream_writer.aclose()
    await write_stream_reader.aclose()

Alternatively, wrap the stream operations in stdout_reader with BrokenResourceError handling:

async def stdout_reader():
    try:
        async with read_stream_writer:
            # ... existing code ...
    except anyio.BrokenResourceError:
        # Context is closing, exit gracefully
        pass

Environment

  • Python: 3.12.1
  • mcp: (version from pip)
  • anyio: (version from pip)
  • OS: Linux

Workaround

We're currently serializing MCP lifecycle operations with an asyncio.Lock to prevent overlapping enter/exit operations, which avoids triggering the race condition.

Activity

  1. added
    bugSomething isn't working
    needs reproneeds additional information to be able to reproduce bug
    on Feb 10, 2026
  2. maxisbey commented on Feb 10, 2026

    @maxisbey
    Contributor

    Thanks for the detailed analysis! Could you provide a single-file runnable script that reproduces this? Would help us validate and prioritize the fix.

    AI Disclaimer

  3. lavish0000 commented on Mar 6, 2026

    @lavish0000

    I was able to reproduce this locally on current main (b33c811).

    I used this minimal script from the repo root:

    import sys
    import textwrap
    import anyio
    from mcp.client.stdio import StdioServerParameters, stdio_client
    
    SERVER_SCRIPT = textwrap.dedent("""
    import sys
    import time
    
    sys.stdout.write('{"jsonrpc":"2.0","id":1,"result":{}}\n')
    sys.stdout.flush()
    time.sleep(2.0)
    """)
    
    async def main() -> None:
        server_params = StdioServerParameters(
            command=sys.executable,
            args=["-c", SERVER_SCRIPT],
        )
        with anyio.fail_after(5.0):
            async with stdio_client(server_params) as (_read_stream, _write_stream):
                await anyio.sleep(0.2)
    
    anyio.run(main)

    Run command:

    python3 -m uv run python repro_1960.py

    For me this fails consistently on main when the client exits without consuming read_stream.

    Traceback excerpt:

    ExceptionGroup: unhandled errors in a TaskGroup (1 sub-exception)
      ...
      File ".../src/mcp/client/stdio.py", line 160, in stdout_reader
        await read_stream_writer.send(session_message)
      ...
      anyio.BrokenResourceError
    

    I also ran the same script against the branch from #2219, and it exited cleanly there.

  4. weiguangli-io commented on Mar 9, 2026

    @weiguangli-io

    Root Cause Analysis

    I traced through the code in src/mcp/client/stdio/__init__.py on the current main branch. The race condition is straightforward:

    Timeline of the bug:

    1. stdio_client yields (read_stream, write_stream) to the caller (line 189)
    2. stdout_reader runs in the task group, reading process stdout and calling await read_stream_writer.send(session_message) (line 162) — this is a zero-buffer memory stream, so send() blocks until a consumer calls receive()
    3. The caller exits the context without consuming read_stream (or exits quickly)
    4. The finally block executes (line 190). It terminates the process, then calls await read_stream_writer.aclose() (line 215)
    5. stdout_reader is still alive inside the task group, blocked on send(). The stream it's sending into just got closed from the other end → BrokenResourceError
    6. The existing except on line 163 only catches ClosedResourceError, not BrokenResourceError, so the exception propagates into the ExceptionGroup

    Key distinction: ClosedResourceError is raised when you call .send() on a stream handle that you already closed. BrokenResourceError is raised when the receiving end of the stream is closed. The finally block closes read_stream (the receive end, line 213) and read_stream_writer (the send end, line 215). If stdout_reader is mid-send() when the receive end closes, it gets BrokenResourceError, not ClosedResourceError.

    Proposed Fix

    The proper fix requires both a structural change and a defensive catch. Either alone is incomplete:

    1. Cancel the task group scope before closing streams (structural fix)

            finally:
                # Cancel background tasks FIRST so they're not racing against stream teardown
                tg.cancel_scope.cancel()
    
                # MCP spec: stdio shutdown sequence (existing code follows)
                if process.stdin:
                    ...

    Adding tg.cancel_scope.cancel() at the top of the finally block causes anyio to deliver Cancelled to stdout_reader and stdin_writer at their next checkpoint (i.e., the blocked send() call). This ensures the tasks are winding down before the streams are closed.

    2. Catch BrokenResourceError alongside ClosedResourceError (defensive fix)

        async def stdout_reader():
            ...
            except (anyio.BrokenResourceError, anyio.ClosedResourceError):
                await anyio.lowlevel.checkpoint()

    Same change for stdin_writer. This is defense-in-depth: even with the cancel-first approach, there's a narrow window where the task could be between checkpoints when the stream closes. Catching both error types makes the cleanup fully robust.

    Why both are needed

    • Cancel-first alone: there's still a theoretical window between the cancel signal and the task actually reaching a checkpoint. If a stream close happens in that window, BrokenResourceError still escapes.
    • Catch-only alone (what PR fix: avoid stdio cleanup BrokenResourceError race #2219 did): it suppresses the symptom but doesn't address the root ordering problem. The tasks still run concurrently with stream teardown, which could cause other subtle issues if anyio's internal invariants change.

    How This Prevents the Race Condition

    With both changes applied:

    1. finally fires → tg.cancel_scope.cancel() marks all tasks for cancellation
    2. stdout_reader is blocked on send() → anyio delivers Cancelled, the task exits its async with read_stream_writer: block cleanly
    3. Stream .aclose() calls now operate on streams that no active task is using
    4. Even if timing is unlucky, the BrokenResourceError catch prevents any exception from leaking into an ExceptionGroup

    Note: PR #2219 was closed because the diff only contained the defensive catch while the description promised the structural reordering. A complete fix should include both changes.

  5. added a commit that references this issue on Mar 11, 2026
    f4090bd
  6. ddullah commented on Apr 15, 2026

    @ddullah

    Paired PRs with a fix + regression test:

    Fix: wrap both read_stream_writer.send(...) call sites in stdout_reader with try/except (ClosedResourceError, BrokenResourceError): return, and widen the outer except to the same union. ClosedResourceError is the sibling class (raised on already-closed streams); BrokenResourceError is the one raised when the receiver is closed during an in-flight send — the exact shape of this shutdown race. No API changes.

    Tests: test_stdio_client_exits_cleanly_while_server_still_writing in tests/client/test_stdio.py — spawns a subprocess that emits a burst of JSONRPC notifications, exits the stdio_client context immediately, asserts no ExceptionGroup propagates. Wrapped in anyio.fail_after(5.0) per AGENTS.md. Fails before the patch, passes after.

    Repro trigger used to verify: an MCP stdio server that emits a few notifications/message frames on startup (jules-mcp-server in our case), driven by a caller that opens a fresh stdio_client per invocation (mcp2cli). Every tool call used to surface as ExceptionGroup → anyio.BrokenResourceError. After the patch, tool calls return cleanly.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingneeds reproneeds additional information to be able to reproduce bug

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions