Skip to content

feat(llm): add shared request pacing and cooldown - #418

Open
cpakkamisaac-sae wants to merge 1 commit into
NVIDIA-NeMo:mainfrom
cpakkamisaac-sae:feature/shared-llm-pacing
Open

cpakkamisaac-sae wants to merge 1 commit into
NVIDIA-NeMo:mainfrom
cpakkamisaac-sae:feature/shared-llm-pacing

Conversation

@cpakkamisaac-sae

@cpakkamisaac-sae cpakkamisaac-sae commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Closes #417

Builds on #349 without changing its default concurrency-only behavior.

Why

An active-call ceiling does not necessarily control attempt frequency. Fast requests and independent retries can exceed a provider's request quota even while respecting max_in_flight. Applications can now opt into shared timing policy at the existing provider-attempt boundary.

Changes

  • Add requests_per_second minimum-spacing admission and max_cooldown bounded, shared Retry-After feedback to process-local groups and the application-owned host-local broker. Both default to None.
  • Jointly schedule concurrency, pacing, and cooldown through one FIFO queue. Waiting for time reserves no concurrency; queue deadlines and cancellations still apply. Idle time does not accumulate burst credits.
  • Publish structured 429/503/529 feedback before releasing the failed attempt's permit, including terminal attempts and shielded provider tasks that finish after caller cancellation. Existing retry ownership and synchronous calls remain unchanged.
  • Preserve external acquire/release-only controllers through an optional CooldownAdmissionPermit protocol.
  • Parse RFC integer seconds and HTTP dates without inspecting error messages. Follow bounded, cycle-safe SDK exception chains because real LiteLLM transports can retain headers only on an underlying exception.
  • Acknowledge broker feedback before release, bound the exchange, close dirty sockets on failure, and reject incompatible protocol peers. Delayed handshakes cannot accumulate paced grants into a burst.
  • Document configuration, lifecycle, and explicit scope boundaries.

Validation

Completed locally on current main (564a3401) before opening the issue or PR:

  • Full default non-live suite: 9,454 passed, 9 skipped, 3 expected failures; integration/live-provider, stress, and sandbox markers excluded by the repository's default selection.
  • 180 new tests passed together, with ResourceWarning and PytestUnraisableExceptionWarning treated as errors. These include 23 real loopback HTTP end-to-end cases using existing Chat and Responses transports, not mocked provider calls.
  • Header parsing, FIFO/rate/concurrency sharing, group consistency, queue timeout, cancellation before and after dispatch, retry accounting, terminal overload feedback, broker authentication/protocol errors, foreign event loops, delayed handshakes, child crashes, and restart behavior covered.
  • Additional standalone local soak: 48,000 synthetic attempts across four 12-process batches, 384 acknowledged cooldown reports, exact attempt-cap accounting, final active/queued counts zero, and flat descriptor/thread counts across repeated batches. Mixed failures include timeouts, cancellations, retries, call-cap rejections, and queued broker shutdown.
  • Project-wide Ruff lint/format, license headers, diff hygiene, and non-type submission hooks passed. Scoped admission/broker/cooldown/export Pyright check: zero errors.

The complete unifiedllm.py type check retains an existing cache-mapping annotation error, reproduced on unchanged main. The shared pacing change does not touch that code; the generic Pyright pre-commit hook was checked separately rather than claimed green.

Focused reproduction:

uv run pytest tests/unifiedllm/test_pacing.py \
  tests/unifiedllm/test_broker_pacing.py \
  tests/unifiedllm/test_cooldown.py \
  tests/unifiedllm/test_pacing_end_to_end.py \
  -W error::ResourceWarning -W error::pytest.PytestUnraisableExceptionWarning

Against the synthetic rate quota, six logical calls through either HTTP transport produced:

Policy Completed calls Overload responses Wire attempts
Concurrency only 1 / 6 10 11
Configured pacing 6 / 6 0 6

The paced run takes about 1.26 seconds; the baseline stops quickly after exhausted retries. This demonstrates completion/retry behavior for this configured fixture, not a general production performance claim. SDK retries are disabled in these comparisons.

Boundaries

Pacing controls NOOA admission grants, not exact remote wire counts or tokens. Already granted/offered leases are not revoked by later feedback. A per-report cooldown cap can resume before the provider's requested delay; this is documented as an application choice. The broker remains host-local; multi-host coordination, adaptive concurrency, queue-size limits, and production gateway validation are separate work. No new dependencies.

Signed-off-by: Clement Pakkam Isaac <cpakkamisaac@nvidia.com>
@coderabbitai

coderabbitai Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: NVIDIA-NeMo/labs-OO-Agents/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Enterprise
  • Run ID: 6ea5220f-e776-49de-b776-8b4cd438d6a0
📥 Commits

Reviewing files that changed from the base of the PR and between 564a340 and 6c2ef03.

📒 Files selected for processing (10)
  • docs/concepts/llm-admission-control.md
  • src/nooa/unifiedllm/__init__.py
  • src/nooa/unifiedllm/admission.py
  • src/nooa/unifiedllm/broker_admission.py
  • src/nooa/unifiedllm/cooldown.py
  • src/nooa/unifiedllm/unifiedllm.py
  • tests/unifiedllm/test_broker_pacing.py
  • tests/unifiedllm/test_cooldown.py
  • tests/unifiedllm/test_pacing.py
  • tests/unifiedllm/test_pacing_end_to_end.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.


📝 Walkthrough

Walkthrough

Adds optional shared request pacing and capped overload cooldown to local admission groups and the broker. Structured Retry-After values from supported provider errors can update cooldown deadlines before permit release. The change also adds configuration validation, protocol updates, documentation, and tests.

Changes

Shared admission control

Layer / File(s) Summary
Parse overload feedback and publish cooldown
src/nooa/unifiedllm/cooldown.py, src/nooa/unifiedllm/unifiedllm.py, src/nooa/unifiedllm/admission.py, src/nooa/unifiedllm/__init__.py, tests/unifiedllm/test_cooldown.py, tests/unifiedllm/test_pacing.py
Adds bounded parsing for structured Retry-After headers on supported overload errors. Provider calls publish parsed delays through permits that support cooldown feedback before releasing them.
Add process-local pacing and cooldown
src/nooa/unifiedllm/admission.py, tests/unifiedllm/test_pacing.py, docs/concepts/llm-admission-control.md
Local groups combine concurrency limits with optional FIFO pacing and shared cooldown deadlines. Timed waits do not reserve concurrency capacity. Configuration validation and same-group policy consistency checks cover both new settings.
Share pacing and cooldown through the broker
src/nooa/unifiedllm/broker_admission.py, tests/unifiedllm/test_broker_pacing.py, docs/concepts/llm-admission-control.md
The broker adds paced offers and bounded cooldown commands, and advances its protocol version to 3. Tests cover connection handling, shutdown, timing, and independent processes.
Verify pacing and cooldown in provider transports
tests/unifiedllm/test_pacing_end_to_end.py
Loopback tests exercise Chat and Responses calls with local and broker admission, including retries, cancellation, queue timeouts, and cross-process cooldown sharing.

Priority: ➖ Normal

Estimated code review effort: 4 (Complex) | ~60 minutes

Change: Feature

Sequence Diagram(s)

sequenceDiagram
  participant Provider
  participant UnifiedLLM
  participant retry_after_delay
  participant CooldownAdmissionPermit
  participant AdmissionGroup
  participant QueuedRequest
  Provider->>UnifiedLLM: Return provider error
  UnifiedLLM->>retry_after_delay: Parse structured Retry-After
  retry_after_delay-->>UnifiedLLM: Return delay or None
  UnifiedLLM->>CooldownAdmissionPermit: Publish cooldown delay
  CooldownAdmissionPermit->>AdmissionGroup: Extend shared admission deadline
  AdmissionGroup-->>QueuedRequest: Grant after deadline and available capacity
Loading

Merge Risk: ⚪ Minimal · up to 6c2ef

No actionable issue remains in the reviewed change; it is ready for normal merge checks.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Issue #417 requires optional, non-bursty pacing and bounded shared Retry-After cooldown, with both settings defaulting to None. The changes add these policies to local admission groups and the host-lo…
Out of Scope Changes check ✅ Passed The changed admission, broker, cooldown, provider-call, export, documentation, and test files all support issue #417. The PR summary identifies no unrelated changes.
Docstring Coverage ✅ Passed Docstring coverage is 92.18% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 179 functions across 9 files. (1 skipped: 1…
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: shared request pacing and cooldown for LLM admission control.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(llm): coordinate shared request pacing and cooldown

1 participant