Repository navigation
L7 inference proxy silently drops tool_calls chunks on large streaming responses #829
Description
Activity
Thanks @MitchFuchs - I was able to reproduce. Will dig further.
🏗️ build-plan
Implementation Plan
Issue type:
fix
Complexity: Medium
Confidence: High — root cause confirmed and reproducedSummary
The inference proxy silently drops large streaming tool_calls due to three interacting bugs: an aggressive 30s per-chunk idle timeout that kills reasoning model "think" pauses, a reqwest total-request timeout (60s default) that caps the entire body stream, and silent truncation that writes a valid HTTP terminator on error paths. Additionally, per-chunk
flush()causes ~5x latency overhead. The fix touches two crates across four independently testable concerns.Scope
crates/openshell-router/src/lib.rs: Addconnect_timeout(30s)to reqwest client buildercrates/openshell-router/src/backend.rs: Extractprepare_backend_request()helper, createsend_backend_request_streaming()without total timeout, updateproxy_to_backend_streamingto use itcrates/openshell-sandbox/src/proxy.rs: IncreaseCHUNK_IDLE_TIMEOUTfrom 30s to 120s, inject SSE error events before chunked terminator on all truncation paths, wrap streaming relay inBufWritercrates/openshell-sandbox/src/l7/inference.rs: Addformat_sse_error()helper
Implementation Steps
Step 1: Remove reqwest total timeout from streaming path
- Extract shared request-building logic into
prepare_backend_request()helper inbackend.rs - Create
send_backend_request_streaming()that omits.timeout(route.timeout)— body stream lifetime is governed by the sandbox idle timeout instead - Add
connect_timeout(30s)to thereqwest::Client::builder()inRouter::new() - Update
proxy_to_backend_streamingto call the new streaming variant - Non-streaming
proxy_to_backendretains the total timeout (correct for buffered responses)
Step 2: Increase CHUNK_IDLE_TIMEOUT for reasoning models
- Change
CHUNK_IDLE_TIMEOUTfrom 30s to 120s inproxy.rs - 120s provides 4x headroom over observed 32s pauses from reasoning models
Step 3: Signal truncation to the client instead of silent corruption
- Add
format_sse_error(reason)inl7/inference.rs— produces a parseable SSE error event - On all three truncation paths (idle timeout, upstream error, byte limit), inject the error event before the chunked terminator
- Error messages must NOT leak internal URLs/hostnames — OCSF log captures full detail server-side
- Bump idle timeout OCSF severity from
LowtoMedium(data loss)
Step 4: Reduce per-chunk flush overhead
- Wrap the TLS writer in
tokio::io::BufWriter::with_capacity(16384)for the streaming relay - Write chunks through
BufWriter(auto-flushes at capacity), single explicitflush()at loop exit - Error-path SSE events also go through the
BufWriterso the final flush delivers them
Test Plan
- Unit tests:
format_sse_errorproduces valid parseable SSE (inl7/inference.rs)prepare_backend_requestshares logic correctly for both paths (inbackend.rs)
- Integration tests:
- Streaming proxy with slow chunks (3s delay × 30 chunks = 90s) completes without timeout (in
backend_integration.rs) - Idle timeout emits SSE error event in stream
- Byte limit emits SSE error event in stream
- Streaming proxy with slow chunks (3s delay × 30 chunks = 90s) completes without timeout (in
- E2E tests: N/A for this PR — existing E2E tests validate inference proxy path
Risks & Open Questions
- 120s idle timeout too generous? —
MAX_STREAMING_BODY(32 MiB) still bounds total data. Acceptable tradeoff vs. silently dropping tool_calls. - Removing total timeout from streaming — worst case bounded by idle timeout (120s) + byte cap (32 MiB). Better than current silent truncation.
- SSE error format compatibility — uses standard
data:SSE format with JSON matching OpenAI error schema. SDKs that don't parse it are no worse off than today. - Refactoring
send_backend_request— existingbackend_integration.rstests cover auth injection, header stripping, model rewrite.
Documentation Impact
- Update
architecture/inference-routing.mdwith idle timeout behavior and SSE error signaling - Update OCSF severity for idle timeout from Low to Medium
Revision 1 — initial plan
- addedstate:review-readyReady for human reviewReady for human reviewstate:agent-readyApproved for agent implementationApproved for agent implementationstate:in-progressWork is currently in progressWork is currently in progress
on Apr 14, 2026 - added a commit that references this issue
on Apr 14, 2026 🏗️ build-from-issue-agent
Implementation Complete
PR: #834
What was built
Fixed the L7 inference proxy silently truncating large streaming tool_calls by removing the reqwest total-request timeout from the streaming path, increasing the per-chunk idle timeout from 30s to 120s for reasoning models, injecting SSE error events on all truncation paths so clients can detect data loss, and wrapping the streaming relay in BufWriter to reduce per-chunk TLS flush overhead.
Tests
- Unit: 2 tests added (format_sse_error output format and JSON escaping)
- Integration: 2 tests added (streaming proxy completes without total timeout, buffered proxy enforces timeout)
- E2E: N/A
Docs updated
architecture/inference-routing.md: Updated timeout model, response streaming, and truncation signaling sections
The issue will auto-close when the PR is merged.
- addedstate:pr-openedPR has been opened for this issuePR has been opened for this issueand removedstate:in-progressWork is currently in progressWork is currently in progressstate:review-readyReady for human reviewReady for human review
on Apr 14, 2026 - added a commit that references this issue
on Apr 14, 2026 Follow-up: the 120s ceiling introduced here is still insufficient for tool-call streams on the same
nemotron-3-super+qwen3_codersetup that motivated PR #834. Reproduction and a relatively simple patch that makes the timeout configurable in #866 (reproduction comment).
Agent Diagnostic
The OpenShell L7 inference proxy passes reasoning tokens through correctly but silently drops tool_calls delta fields from SSE streaming responses when the tool call payload is large (>~5KB). Small tool calls (~1-2KB) pass through. The model generates valid responses — confirmed by bypassing the proxy and connecting directly to the inference backend.
Description
When an LLM streams a response with tool_calls via /v1/chat/completions (OpenAI-compatible, stream: true, tools parameter), the OpenShell inference routing proxy (inference.local) strips the tool_calls fields from the SSE chunks. The stream completes with reasoning content only — no tool call, no content, no finish_reason. The sandbox OCSF log shows NET:FAIL [LOW] inference.local:443 after the stream ends.
This affects all models — tested with both nemotron-3-super (120B MoE) and gemma4:e4b via Ollama. Both produce valid tool call JSON when accessed directly, but fail identically through the proxy.
Reproduction Steps
openshell inference set --provider <name> --model <model> --timeout 1800delta.tool_callschunks, finishes withfinish_reason: "tool_calls"delta.reasoningchunks, then[DONE]with no tool call data. OCSF log showsNET:FAIL [LOW] inference.local:443Test results summary:
Environment
OpenShell: 0.0.25 (CLI and gateway)
NemoClaw: 0.0.10
Host: Raspberry Pi 5 (8GB), Ubuntu Server 24.04, aarch64
Inference backend: Ollama (remote via SSH tunnel), models: nemotron-3-super, gemma4:e4b
Inference route configured with protocols=openai_chat_completions,openai_completions,openai_responses,model_discovery
Proxy path: sandbox → inference.local:443 → OpenShell L7 proxy → backend endpoint
Logs
Agent-First Checklist
debug-openshell-cluster,debug-inference,openshell-cli)