Skip to content

L7 inference proxy silently drops tool_calls chunks on large streaming responses #829

Description

@MitchFuchs

Agent Diagnostic

The OpenShell L7 inference proxy passes reasoning tokens through correctly but silently drops tool_calls delta fields from SSE streaming responses when the tool call payload is large (>~5KB). Small tool calls (~1-2KB) pass through. The model generates valid responses — confirmed by bypassing the proxy and connecting directly to the inference backend.

Description

When an LLM streams a response with tool_calls via /v1/chat/completions (OpenAI-compatible, stream: true, tools parameter), the OpenShell inference routing proxy (inference.local) strips the tool_calls fields from the SSE chunks. The stream completes with reasoning content only — no tool call, no content, no finish_reason. The sandbox OCSF log shows NET:FAIL [LOW] inference.local:443 after the stream ends.

This affects all models — tested with both nemotron-3-super (120B MoE) and gemma4:e4b via Ollama. Both produce valid tool call JSON when accessed directly, but fail identically through the proxy.

Reproduction Steps

  1. Configure an OpenAI-compatible inference provider (e.g., Ollama) with tool-calling-capable models
  2. Set up inference routing: openshell inference set --provider <name> --model <model> --timeout 1800
  3. From inside the sandbox, send a streaming request with tools that requires a large response:
# Through proxy (FAILS - tool_calls dropped):
# URL: https://inference.local/v1/chat/completions
#
# Direct to backend (WORKS - 23KB valid JSON tool call):
# URL: http://<backend-host>:<port>/v1/chat/completions

payload = {
"model": "<model>",
"messages": [{"role": "user", "content": "Write a comprehensive 2000-word project plan. Save it to project_plan.md"}],
"tools": [{"type": "function", "function": {"name": "write", "description": "Write content to a file.", "parameters": {"type": "object", "properties": {"file_path": {"type": "string"}, "content": {"type": "string"}}, "required": ["file_path", "content"]}}}],
"stream": True,
"max_tokens": 16384
}

  1. Expected: SSE stream includes delta.tool_calls chunks, finishes with finish_reason: "tool_calls"
  2. Actual: SSE stream includes only delta.reasoning chunks, then [DONE] with no tool call data. OCSF log shows NET:FAIL [LOW] inference.local:443

Test results summary:

Test Direct to backend Through proxy
nemotron-3-super short tool call (~1.8KB) works works
nemotron-3-super long tool call (~23KB) 324s, valid JSON reasoning only, tool call dropped
gemma4:e4b long tool call (~13KB) 101s, valid JSON reasoning only, tool call dropped
Plain chat (no tools) works works

Environment

OpenShell: 0.0.25 (CLI and gateway)
NemoClaw: 0.0.10
Host: Raspberry Pi 5 (8GB), Ubuntu Server 24.04, aarch64
Inference backend: Ollama (remote via SSH tunnel), models: nemotron-3-super, gemma4:e4b
Inference route configured with protocols=openai_chat_completions,openai_completions,openai_responses,model_discovery
Proxy path: sandbox → inference.local:443 → OpenShell L7 proxy → backend endpoint

Logs

Sandbox OCSF log around failure:


[timestamp] NET:OPEN [INFO] ALLOWED inference.local:443
[timestamp] [openshell_router] routing proxy inference request (streaming) endpoint=http://<backend>:11435/v1 method=POST path=/v1/chat/completions protocols=openai_chat_completions,openai_completions,openai_responses,model_discovery
# ... sendChatAction calls every ~3s (typing indicator) ...
[timestamp] NET:FAIL [LOW] inference.local:443
[timestamp] HTTP:POST [INFO] ALLOWED POST http://api.telegram.org/bot[CREDENTIAL]/deleteMessage [policy:telegram]
The deleteMessage confirms OpenClaw received no usable response and cleaned up the partial reasoning message from the chat channel.

Agent-First Checklist

  • I pointed my agent at the repo and had it investigate this issue
  • I loaded relevant skills (e.g., debug-openshell-cluster, debug-inference, openshell-cli)
  • My agent could not resolve this — the diagnostic above explains why

Activity

  1. added theissue type on Apr 14, 2026
  2. self-assigned this
    on Apr 14, 2026
  3. johntmyers commented on Apr 14, 2026

    @johntmyers
    Collaborator

    Thanks @MitchFuchs - I was able to reproduce. Will dig further.

  4. johntmyers commented on Apr 14, 2026

    @johntmyers
    Collaborator

    🏗️ build-plan

    Implementation Plan

    Issue type: fix
    Complexity: Medium
    Confidence: High — root cause confirmed and reproduced

    Summary

    The inference proxy silently drops large streaming tool_calls due to three interacting bugs: an aggressive 30s per-chunk idle timeout that kills reasoning model "think" pauses, a reqwest total-request timeout (60s default) that caps the entire body stream, and silent truncation that writes a valid HTTP terminator on error paths. Additionally, per-chunk flush() causes ~5x latency overhead. The fix touches two crates across four independently testable concerns.

    Scope

    • crates/openshell-router/src/lib.rs: Add connect_timeout(30s) to reqwest client builder
    • crates/openshell-router/src/backend.rs: Extract prepare_backend_request() helper, create send_backend_request_streaming() without total timeout, update proxy_to_backend_streaming to use it
    • crates/openshell-sandbox/src/proxy.rs: Increase CHUNK_IDLE_TIMEOUT from 30s to 120s, inject SSE error events before chunked terminator on all truncation paths, wrap streaming relay in BufWriter
    • crates/openshell-sandbox/src/l7/inference.rs: Add format_sse_error() helper

    Implementation Steps

    Step 1: Remove reqwest total timeout from streaming path

    • Extract shared request-building logic into prepare_backend_request() helper in backend.rs
    • Create send_backend_request_streaming() that omits .timeout(route.timeout) — body stream lifetime is governed by the sandbox idle timeout instead
    • Add connect_timeout(30s) to the reqwest::Client::builder() in Router::new()
    • Update proxy_to_backend_streaming to call the new streaming variant
    • Non-streaming proxy_to_backend retains the total timeout (correct for buffered responses)

    Step 2: Increase CHUNK_IDLE_TIMEOUT for reasoning models

    • Change CHUNK_IDLE_TIMEOUT from 30s to 120s in proxy.rs
    • 120s provides 4x headroom over observed 32s pauses from reasoning models

    Step 3: Signal truncation to the client instead of silent corruption

    • Add format_sse_error(reason) in l7/inference.rs — produces a parseable SSE error event
    • On all three truncation paths (idle timeout, upstream error, byte limit), inject the error event before the chunked terminator
    • Error messages must NOT leak internal URLs/hostnames — OCSF log captures full detail server-side
    • Bump idle timeout OCSF severity from Low to Medium (data loss)

    Step 4: Reduce per-chunk flush overhead

    • Wrap the TLS writer in tokio::io::BufWriter::with_capacity(16384) for the streaming relay
    • Write chunks through BufWriter (auto-flushes at capacity), single explicit flush() at loop exit
    • Error-path SSE events also go through the BufWriter so the final flush delivers them

    Test Plan

    • Unit tests:
      • format_sse_error produces valid parseable SSE (in l7/inference.rs)
      • prepare_backend_request shares logic correctly for both paths (in backend.rs)
    • Integration tests:
      • Streaming proxy with slow chunks (3s delay × 30 chunks = 90s) completes without timeout (in backend_integration.rs)
      • Idle timeout emits SSE error event in stream
      • Byte limit emits SSE error event in stream
    • E2E tests: N/A for this PR — existing E2E tests validate inference proxy path

    Risks & Open Questions

    • 120s idle timeout too generous? — MAX_STREAMING_BODY (32 MiB) still bounds total data. Acceptable tradeoff vs. silently dropping tool_calls.
    • Removing total timeout from streaming — worst case bounded by idle timeout (120s) + byte cap (32 MiB). Better than current silent truncation.
    • SSE error format compatibility — uses standard data: SSE format with JSON matching OpenAI error schema. SDKs that don't parse it are no worse off than today.
    • Refactoring send_backend_request — existing backend_integration.rs tests cover auth injection, header stripping, model rewrite.

    Documentation Impact

    • Update architecture/inference-routing.md with idle timeout behavior and SSE error signaling
    • Update OCSF severity for idle timeout from Low to Medium

    Revision 1 — initial plan

  5. added a commit that references this issue on Apr 14, 2026
    4f7d0c4
  6. johntmyers commented on Apr 14, 2026

    @johntmyers
    Collaborator

    🏗️ build-from-issue-agent

    Implementation Complete

    PR: #834

    What was built

    Fixed the L7 inference proxy silently truncating large streaming tool_calls by removing the reqwest total-request timeout from the streaming path, increasing the per-chunk idle timeout from 30s to 120s for reasoning models, injecting SSE error events on all truncation paths so clients can detect data loss, and wrapping the streaming relay in BufWriter to reduce per-chunk TLS flush overhead.

    Tests

    • Unit: 2 tests added (format_sse_error output format and JSON escaping)
    • Integration: 2 tests added (streaming proxy completes without total timeout, buffered proxy enforces timeout)
    • E2E: N/A

    Docs updated

    • architecture/inference-routing.md: Updated timeout model, response streaming, and truncation signaling sections

    The issue will auto-close when the PR is merged.

  7. added a commit that references this issue on Apr 14, 2026
    355d845
  8. vnicolici commented on Apr 24, 2026

    @vnicolici

    Follow-up: the 120s ceiling introduced here is still insufficient for tool-call streams on the same nemotron-3-super + qwen3_coder setup that motivated PR #834. Reproduction and a relatively simple patch that makes the timeout configurable in #866 (reproduction comment).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

state:agent-readyApproved for agent implementationstate:pr-openedPR has been opened for this issue

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions