Skip to content

feat: trace KVCR remote delivery from hint to caller completion - #69

Draft
aknvda wants to merge 4 commits into
aniket/nvtx-instrumentationfrom
aniket/nvtx-target-lifecycle
Draft

aknvda wants to merge 4 commits into
aniket/nvtx-instrumentationfrom
aniket/nvtx-target-lifecycle

Conversation

@aknvda

@aknvda aknvda commented Oct 5, 2026 •

Copy link
Copy Markdown

This extends #67 with correlation from remote hint submission through target receipt and caller completion. A successful native write, target receipt, completion of the mixed-tier join, and the caller observing the result remain separate events. Cancellation, quarantine/quiescence, stale routes and unresolved shutdown retain explicit causes without changing ownership or wire contracts.

Add a two-process capture example for real KVCR/NIXL DRAM or VRAM transfers, a SQLite payload decoder, and profiling/schema documentation. The decoder validates request context, worker joins and ordering, and rejects missing or mismatched context streams. One KVCR domain and bounded event names use fresh structured payloads; unavailable session/parent IDs and framework utilization stay unknown.

This draft is stacked on aniket/nvtx-instrumentation (#67). The unrelated Guard test allocation fix remains in #68.

Validation:

  • Linux full suite: 500 passed, 1 skipped, with test: align Guard recovery buffers for direct I/O #68 applied. The NVTX branch alone retains main's pre-existing unaligned Guard fixture failure. The source branch separately passed 183 focused tests.
  • Ruff lint/format, diff checks, lockfile check, wheel and source-distribution builds passed. All 67 tracked files in the tested Linux tree matched the final branch plus the separately identified Guard test file.
  • Two-process DRAM success, partial, failed-pin and deadline captures decoded; off mode produced zero KVCR payload events. Exact UTF-8 request mapping, bytes and pin release counts verified.
  • Two-process VRAM capture on one H100 verified bytes and decoded 72 KVCR events across three operations. This covers same-host GPU IPC/UCX, not cross-node RDMA.
  • A real Qwen2.5-7B vLLM run served three requests and deposited 83 blocks / 76,152,832 bytes into local G2. That run needed the public vLLM descriptor/backpressure integration fixes; it does not validate Dynamo remote hint routing.
  • Independent review found lost failure causes and a context-decoder false-positive gap. Both were reproduced before fixes; focused regressions and the full suite above passed afterward. Native error/release/quarantine cases use controlled native-state doubles, not hardware fault injection.

Measured delivery overhead:

Collection Level Delivery latency (ms) Run-mean range (ms) Change vs off
None off 4.500 4.134–4.981 +0.0%
None low 4.803 4.552–5.375 +6.7%
None medium 4.746 4.334–5.748 +5.5%
Nsight off 4.653 4.184–5.097 +0.0%
Nsight low 6.461 5.902–7.241 +38.9%
Nsight medium 6.058 5.536–6.716 +30.2%

Latency is the median of five run means, measured from deliver dispatch through
caller completion. Each run uses 20 warmups and 200 measured deliveries of
16 blocks × 4 KiB between two real KVCR/NIXL processes. Thirty runs cover off,
low and medium with and without Nsight: 6,000 measured deliveries, plus 600
warmups. Level order rotates by repetition; runs are sequential. Every run
verified received bytes and 220 pin requests/releases, with zero cancellations.
Nsight uses NVTX/CUDA tracing with CPU sampling/context-switch collection off.
Hint construction, byte verification, process startup and teardown are outside
the timed delivery interval. This is a DRAM microbenchmark, not model throughput
or a GPU/RDMA performance claim. The ranges overlap; the apparent medium/low
ordering is not evidence that medium is cheaper. The successful path exercises
few medium-only events. No production overhead acceptance threshold was supplied.

Median delivery rates and raw Nsight report sizes across the five runs:

Collection Level Deliveries/s Nsight report KiB
None off 222.2 —
None low 208.2 —
None medium 210.7 —
Nsight off 214.9 307.7
Nsight low 154.8 460.5
Nsight medium 165.1 457.8

The delivery rate is the reciprocal of each run’s mean timed delivery latency;
it excludes setup/verification and is not full application throughput. Report
size includes the warmups and Nsight/NIXL data as well as KVCR payloads.

Review boundaries: full Dynamo remote reuse still needs a compatible hint carrier between the producer and adapter. Session/parent correlation and framework utilization need explicit integration seams. Full public-fetch fan-out and other KVCR operations are outside this contribution. Human schema/capture approval and an overhead threshold remain open.

Known minor: when the source selects zero descriptors it reports PARTIAL with an explicit zero count while the target/caller reports FAILED; transfer behavior and ownership are unchanged. Use at least two blocks for the capture example's partial scenario; a one-block partial scenario produces zero successes and does not match the decoder's PARTIAL expectation.

aknvda added 4 commits October 5, 2026 16:20
Signed-off-by: aknvda <anikkulkarni@nvidia.com>
Signed-off-by: aknvda <anikkulkarni@nvidia.com>
Signed-off-by: aknvda <anikkulkarni@nvidia.com>
Signed-off-by: aknvda <anikkulkarni@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant