Repository navigation
[SGLang follow-up] Hybrid eviction fast path and progressive layer delivery - #44
Closed
ziqifan617 wants to merge 5 commits into
Closed
ziqifan617 wants to merge 5 commits into
ziqifan617 wants to merge 5 commits into
Conversation
Collaborator
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Scope
Follow-up to #39, intentionally targeting
idhanani/framework-gpu-regions(
f3cdff142a2e601a62fab5d2fb83b5eae7d91539), not main. Ports the reusableKVCR-side changes from our SGLang hybrid-model linker experiments. Companion PR:
ishandhanani/sglang#11. The companion
SGLang branch is
ziqifan617/sglang:codex/kvcr-linker-followup, based onishandhanani/sglang:idhanani/kvcr-direct-linker.Commits
scanning unrelated pool objects under hybrid cache pressure. Keep the global
policy queue whenever an overlapping layout can satisfy the deficit. Claims,
releases and removals update both indexes. Regression tests cover repeated
eviction, claim/release, overlapping-layout ordering and existing FIFO/LRU
behavior.
deliver()'s signature and allowan ordered subset of uniquely named stored spans, with size validation. Use
the same projection for local delivery. Fetch still requires the exact stored
object layout: only remote delivery sets the new control-message flag. Empty,
duplicate, unknown and reordered subset layouts fail closed. The bounded
layout cache contains indices, never memory addresses or source claims.
transfer while processing peer control, avoiding an additional queue turn
before the early layer can move. Keep normal per-operation claim/dependency
cleanup and completion ownership.
Safety / deliberately excluded
layer deliveries. Per-operation source claims protect in-flight reads only.
The SGLang direct-HBM path is opt-in and fails its layer counter if a selected
source disappears after admission. Production reservation semantics remain
follow-up work.
deliver_manypublic API, batched-control prototype, experimental requestleases, or process-global polling/GIL tuning is included.
Validation
Fresh checks of this branch in an isolated Linux/aarch64 GB300 container:
PYTHONPATH=src python3 -m pytest tests/unit -q \ -k 'not promoted_guard_serves_real_nixl_transfers and not g3-invalid'406 passed, 3 deselected. Ruff and
git diff --checkpass.The full run before the final inline-submit commit had 404 passed / 3 failed;
all three failures were reproduced on the unchanged #39 head in the same
container: two real-NIXL Guard shared-memory registration tests fail with UCX
ibv_reg_mr/ucp_mem_maperrors, andg3-invalidexpects a 4-KiB-alignment errorbut this ARM host reports the local-pool capacity error first. They are not
silently counted as passes. A CPU-only UCX transport override was also tried as
a diagnostic, but it is not the validation configuration above and disabled GPU
registration; no passing RDMA claim is made from that diagnostic.
The companion SGLang CPU/control suites pass (92 tests; 7 Rust-backed tests
excluded because no matching extension was available in the isolated checkout).
Earlier DeepSeek-V4.1/Kimi K3 GPU benchmarks motivated these changes but used older
runtime overlays. No new full-model throughput/TTFT validation is claimed for
these rebased commits. Draft pending review of the projection protocol and
direct-HBM lifecycle.