Skip to content

[SGLang follow-up] Hybrid eviction fast path and progressive layer delivery - #44

Closed
ziqifan617 wants to merge 5 commits into
idhanani/framework-gpu-regionsfrom
codex/linker-hybrid-followup
Closed

ziqifan617 wants to merge 5 commits into
idhanani/framework-gpu-regionsfrom
codex/linker-hybrid-followup

Conversation

@ziqifan617

@ziqifan617 ziqifan617 commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Scope

Follow-up to #39, intentionally targeting idhanani/framework-gpu-regions
(f3cdff142a2e601a62fab5d2fb83b5eae7d91539), not main. Ports the reusable
KVCR-side changes from our SGLang hybrid-model linker experiments. Companion PR:
ishandhanani/sglang#11. The companion
SGLang branch is ziqifan617/sglang:codex/kvcr-linker-followup, based on
ishandhanani/sglang:idhanani/kvcr-direct-linker.

Commits

  1. Index evictable DRAM objects by disjoint pool layout. Avoid repeatedly
    scanning unrelated pool objects under hybrid cache pressure. Keep the global
    policy queue whenever an overlapping layout can satisfy the deficit. Claims,
    releases and removals update both indexes. Regression tests cover repeated
    eviction, claim/release, overlapping-layout ordering and existing FIFO/LRU
    behavior.
  2. Deliver ordered named layer spans. Keep deliver()'s signature and allow
    an ordered subset of uniquely named stored spans, with size validation. Use
    the same projection for local delivery. Fetch still requires the exact stored
    object layout: only remote delivery sets the new control-message flag. Empty,
    duplicate, unknown and reordered subset layouts fail closed. The bounded
    layout cache contains indices, never memory addresses or source claims.
  3. Start ready source writes inline on the progress thread. Submit the
    transfer while processing peer control, avoiding an additional queue turn
    before the early layer can move. Keep normal per-operation claim/dependency
    cleanup and completion ownership.

Safety / deliberately excluded

  • Query is advisory; this is not a reservation/lease across lookup and all
    layer deliveries. Per-operation source claims protect in-flight reads only.
    The SGLang direct-HBM path is opt-in and fails its layer counter if a selected
    source disappears after admission. Production reservation semantics remain
    follow-up work.
  • No deliver_many public API, batched-control prototype, experimental request
    leases, or process-global polling/GIL tuning is included.
  • Old sources reject partial layouts; both endpoints need this support.
  • No changes to Lin's separate HiCacheStorage adapter.

Validation

Fresh checks of this branch in an isolated Linux/aarch64 GB300 container:

PYTHONPATH=src python3 -m pytest tests/unit -q \
  -k 'not promoted_guard_serves_real_nixl_transfers and not g3-invalid'

406 passed, 3 deselected. Ruff and git diff --check pass.

The full run before the final inline-submit commit had 404 passed / 3 failed;
all three failures were reproduced on the unchanged #39 head in the same
container: two real-NIXL Guard shared-memory registration tests fail with UCX
ibv_reg_mr/ucp_mem_map errors, and g3-invalid expects a 4-KiB-alignment error
but this ARM host reports the local-pool capacity error first. They are not
silently counted as passes. A CPU-only UCX transport override was also tried as
a diagnostic, but it is not the validation configuration above and disabled GPU
registration; no passing RDMA claim is made from that diagnostic.

The companion SGLang CPU/control suites pass (92 tests; 7 Rust-backed tests
excluded because no matching extension was available in the isolated checkout).
Earlier DeepSeek-V4.1/Kimi K3 GPU benchmarks motivated these changes but used older
runtime overlays. No new full-model throughput/TTFT validation is claimed for
these rebased commits. Draft pending review of the projection protocol and
direct-HBM lifecycle.

@mkhazraee

Copy link
Copy Markdown
Collaborator

Main components were already merged in main, few remaining being tracked in #62 / #63 / #64 .

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants