Skip to content

streamable-http client: no size bound before JSONRPCMessage.model_validate_json — one large server message can OOM the client #3330

Description

@aartaria

Summary

StreamableHTTPTransport parses every inbound SSE event and JSON response body with JSONRPCMessage.model_validate_json / model_validate_json(content) and no size bound. Because pydantic validation of a large JSON document allocates several times the wire size in live Python objects, a single oversized message from a server can exhaust the client's memory. There is no hook to inspect or reject a message before it is parsed.

We hit this in production: a client connected to ~30 MCP servers, one of which answered tools/list with a 21.9 MB single SSE event. Measured with tracemalloc at the parse site:

#1  663.4 MB in 9,957,044 blocks
    mcp/client/streamable_http.py:217
      message = JSONRPCMessage.model_validate_json(sse.data)
#2  46.9 MB in 3 blocks
    httpx_sse/_decoders.py:61   (whole event buffered as one string)
#3  46.9 MB in 2 blocks
    httpx_sse/_decoders.py:120  (copy from slicing)

Roughly a 7× amplification of the wire size in live objects, transient but concurrent with other parses. A trivial request that triggered only catalogue loading peaked at 2.35 GB RSS from a 56 MB baseline; heavier concurrent work reached 4.2 GB. Removing that one server dropped the same request's peak to 456 MB. Versions: mcp 1.27.1, Python 3.14.

Why a client-side bound is needed

The client cannot know in advance that a server will return a huge payload, and a misbehaving or misconfigured server should not be able to OOM its client. Today the only outcome is process death with nothing identifying the responsible server — the failure surfaces as an unexplained kill rather than an actionable error.

What we did as a workaround, and why it was awkward

We wrapped StreamableHTTPTransport._handle_sse_event and _handle_json_response to measure sse.data / the response body and reject anything over a configured cap before parsing. Two things made this harder than expected, and both seem worth addressing upstream:

  1. Raising from the handler is not viable. Every call site wraps it in except Exception: logger.debug(...) and then reconnects with Last-Event-ID, which replays the same oversized payload. The raise is swallowed and the pending request hangs.
  2. Delivering a bare Exception on the read stream does not fail the request either. In BaseSession._receive_loop, isinstance(message, Exception) routes to _handle_incoming, which for the default message handler is a no-op; only a JSONRPCResponse/JSONRPCError matching _response_streams[request_id] completes send_request, and that awaits with timeout=None. To fail the request we had to synthesize a JSONRPCError and recover the request id from the raw payload with a regex, since parsing it is exactly what we were trying to avoid.

Suggested improvements

  • An optional client-side maximum message size (constructor argument and/or environment variable) enforced before model_validate_json, on both the SSE and application/json paths.
  • When it trips, fail the corresponding pending request with a distinct error identifying the endpoint and the observed size, rather than letting the reconnect path replay the payload.
  • Failing that, a documented hook to inspect a raw message before parsing would let clients implement this without patching private methods.
  • Independently: _handle_sse_event returning False for a parse failure causes the caller to treat the stream as ended and reconnect with Last-Event-ID; for a deterministic failure (such as a payload that will always be too large, or malformed JSON) this retries something that cannot succeed.

Happy to open a PR if a maintainer indicates the preferred shape (constructor arg vs. env var vs. pre-parse hook).

Activity

  1. added
    v1Affects the v1.x maintenance line
    v2Affects the v2 line (2.x on main)
    on Aug 18, 2026
  2. tarunag10 commented on Aug 20, 2026

    @tarunag10

    I checked this against draft #3338. That PR intentionally restores unbounded SSE compatibility by passing max_event_size=None; it is useful groundwork for consistent decoder configuration, but it does not solve this OOM issue and leaves JSON unbounded too.

    Before implementing a public API, I propose this contract:

    • expose one transport-neutral client option such as max_inbound_message_size, rather than an SSE-only max_event_size;
    • apply it before JSON-RPC parsing on both paths: configure the httpx2 SSE decoder with the same cap, and stream/count application/json response bytes instead of materializing an unlimited body first;
    • treat limit breaches as terminal and non-retryable so Last-Event-ID does not replay the same oversized event;
    • when a request ID can be recovered safely, fail that pending request with a distinct error containing endpoint, configured limit, and observed size; for standalone notifications/requests where correlation is unavailable, surface a transport/session error rather than silently dropping it;
    • retain litellm_trace_id-style session usability after a request-scoped rejection where the underlying stream remains safe.

    Suggested acceptance matrix: exactly-at-limit and limit+1 for SSE and JSON; multibyte payloads counted as wire bytes; assert model_validate_json is never called after rejection; no reconnect/replay for deterministic oversize; correlated request completion; standalone GET-stream behavior; and a subsequent normal request proving session recovery. A default and v1/v2 exposure policy are the remaining maintainer decisions. If this shape is acceptable, I can implement it after assignment in line with the repository contribution policy.

  3. Vish12345678673 commented on Sep 18, 2026

    @Vish12345678673

    Hi! I would like to work on this issue.

    I reviewed the Streamable HTTP receive path and the reported memory-amplification failure. I would like to reproduce the behavior locally and implement a focused pre-parse message-size guard with regression coverage, while preserving the existing reconnect and pending-request semantics.

    I have not started a PR yet and will wait for maintainer confirmation before proceeding.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingv1Affects the v1.x maintenance linev2Affects the v2 line (2.x on main)

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions