Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
15 commits
Select commit Hold shift + click to select a range
803bd73
config(dsv41flash): sweep MI355X at TP=2 and TP=4 through concurrency…
Fangzhou-Ai Sep 21, 2026
0de2d6e
config(dsv41flash): record the sweep PR link in the changelog
Fangzhou-Ai Sep 21, 2026
99f0859
config(dsv41flash): let vLLM pick max_num_seqs for the MI355X AgentX …
Fangzhou-Ai Sep 21, 2026
039194e
config(dsv41flash): keep the placeholder on the existing ROCm nightly…
Fangzhou-Ai Sep 21, 2026
a1ddae9
Revert "config(dsv41flash): keep the placeholder on the existing ROCm…
Fangzhou-Ai Sep 21, 2026
4db75a6
config(dsv41flash): pin the MI355X AgentX sweep to the ROCm 10 nightly
Fangzhou-Ai Sep 21, 2026
beec2ef
config(dsv41flash): offload Engram to host memory at TP=2 on MI355X
Fangzhou-Ai Sep 21, 2026
a9ab48e
config(dsv41flash): disable SWA bounded replay on MI355X
Fangzhou-Ai Sep 21, 2026
42bb4f5
config(dsv41flash): fix the MI355X KV split at high concurrency
Fangzhou-Ai Sep 21, 2026
0b1bebc
Merge remote-tracking branch 'origin/main' into config/dsv41flash-mi3…
Fangzhou-Ai Sep 21, 2026
64dab29
config(dsv41flash): change the MI355X memory split only where KV ran …
Fangzhou-Ai Sep 21, 2026
baaa4fc
Merge remote-tracking branch 'origin/main' into config/dsv41flash-mi3…
functionstackx Sep 22, 2026
e913c5c
config(dsv41flash): measure the full MI355X concurrency range on the …
chunfangamd Sep 22, 2026
da6d453
docs(dsv41flash): retire the MI355X draft status and repoint GPU vali…
chunfangamd Sep 22, 2026
f5249b3
Merge branch 'main' into config/dsv41flash-mi355x-tp2-tp4-c128
chunfangamd Sep 22, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 0 additions & 1 deletion MODELS.md
Original file line number Diff line number Diff line change
Expand Up @@ -159,7 +159,6 @@ Other offloading tiers, including NVMe KV cache offloading, are outside the init
| Kimi-K3 | `kimik3` | 2026-07-27 ([#2391](https://github.com/SemiAnalysisAI/InferenceX/pull/2391)) | Agentic coding (DSpark may be disabled for better Pareto points) | Standalone non-DSpark A/B baseline (not required from day 0) |
| GLM-5.2 | `glm5.2` | 2026-07-18 ([#2268](https://github.com/SemiAnalysisAI/InferenceX/pull/2268)) | Agentic coding (non-MTP points remain eligible under the Pareto policy; see Deprecation Notice) | |
| MiniMax-M3 | `minimaxm3` | 2026-06-12 ([#1724](https://github.com/SemiAnalysisAI/InferenceX/pull/1724)) | Agentic coding | Single-turn 1k1k, Single-turn 8k1k (removed 2026-08-04, [#2493](https://github.com/SemiAnalysisAI/InferenceX/pull/2493)) |
| DeepSeek-V4.1-Flash | `dsv41flash` | Pending | Agentic coding on MI355X (draft; GPU validation pending) | — |
| DeepSeek-V4-Pro | `dsv4` | 2026-04-24 ([#1130](https://github.com/SemiAnalysisAI/InferenceX/pull/1130)) | Agentic coding (non-spec-decode points remain eligible under the Pareto policy) | Single-turn 1k1k, Single-turn 8k1k (removed 2026-09-09, [#2921](https://github.com/SemiAnalysisAI/InferenceX/pull/2921)) |
| GLM-5 / GLM-5.1 | `glm5`, `glm5.1` | 2026-03-06 ([#762](https://github.com/SemiAnalysisAI/InferenceX/pull/762)), with GLM-5.1 added 2026-04-21 ([#1098](https://github.com/SemiAnalysisAI/InferenceX/pull/1098)) | GLM-5.1 B200 TileRT only: 1k1k and 8k1k added 2026-08-09 ([#2533](https://github.com/SemiAnalysisAI/InferenceX/pull/2533)); Agentic coding added in [#2650](https://github.com/SemiAnalysisAI/InferenceX/pull/2650) | The earlier GLM-5 / GLM-5.1 recipes were retired 2026-07-18 ([#2276](https://github.com/SemiAnalysisAI/InferenceX/pull/2276)) |
| MiniMax-M2.5/2.7 | `minimaxm2.5` | 2026-02-18 ([#755](https://github.com/SemiAnalysisAI/InferenceX/pull/755)) | None (retired 2026-06-20, [#1874](https://github.com/SemiAnalysisAI/InferenceX/pull/1874)) | Single-turn 1k1k, Single-turn 1k8k, Single-turn 8k1k |
Expand Down
1 change: 0 additions & 1 deletion MODELS_zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -159,7 +159,6 @@ InferenceX 支持 SGLang 和 vLLM 双方的维护者,并响应 AI 实验室和
| Kimi-K3 | `kimik3` | 2026-07-27 ([#2391](https://github.com/SemiAnalysisAI/InferenceX/pull/2391)) | 智能体编码(可关闭 DSpark 以获得更优帕累托点) | 独立非 DSpark A/B 基线(自第 0 天起即不要求) |
| GLM-5.2 | `glm5.2` | 2026-07-18([#2268](https://github.com/SemiAnalysisAI/InferenceX/pull/2268)) | 智能体编码(非 MTP 数据点仍可按帕累托策略参与发布;见弃用公告) | |
| MiniMax-M3 | `minimaxm3` | 2026-06-12([#1724](https://github.com/SemiAnalysisAI/InferenceX/pull/1724)) | 智能体编码 | 单轮 1k1k、单轮 8k1k(2026-08-04 移除,[#2493](https://github.com/SemiAnalysisAI/InferenceX/pull/2493)) |
| DeepSeek-V4.1-Flash | `dsv41flash` | 待验证 | MI355X 上的 Agentic coding(草案;等待 GPU 验证) | — |
| DeepSeek-V4-Pro | `dsv4` | 2026-04-24([#1130](https://github.com/SemiAnalysisAI/InferenceX/pull/1130)) | 智能体编码(非投机解码数据点仍可按帕累托策略参与发布) | 单轮 1k1k、单轮 8k1k(已于 2026-09-09 移除,[#2921](https://github.com/SemiAnalysisAI/InferenceX/pull/2921)) |
| GLM-5 / GLM-5.1 | `glm5`、`glm5.1` | 2026-03-06([#762](https://github.com/SemiAnalysisAI/InferenceX/pull/762)),GLM-5.1 于 2026-04-21 加入([#1098](https://github.com/SemiAnalysisAI/InferenceX/pull/1098)) | 仅 GLM-5.1 B200 TileRT:1k1k 和 8k1k 于 2026-08-09 加入([#2533](https://github.com/SemiAnalysisAI/InferenceX/pull/2533));智能体编码由 [#2650](https://github.com/SemiAnalysisAI/InferenceX/pull/2650) 加入 | 此前的 GLM-5 / GLM-5.1 配方于 2026-07-18 退役([#2276](https://github.com/SemiAnalysisAI/InferenceX/pull/2276)) |
| MiniMax-M2.5/2.7 | `minimaxm2.5` | 2026-02-18([#755](https://github.com/SemiAnalysisAI/InferenceX/pull/755)) | 无(2026-06-20 退役,[#1874](https://github.com/SemiAnalysisAI/InferenceX/pull/1874)) | 单轮 1k1k、单轮 1k8k、单轮 8k1k |
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -41,15 +41,67 @@ export VLLM_ENGINE_READY_TIMEOUT_S=3600
export VLLM_USE_RUST_FRONTEND=1
export PYTHONUNBUFFERED=1

# Upstream picks 1024 on GPUs with >= 160 GiB, and 2*CONC starves AgentX
# subagent fan-out at low CONC. 128 also keeps CAPTURE_SIZE deterministic.
MAX_NUM_SEQS=128
# vllm-project/vllm#57491 widened the two is_cuda() gates to is_cuda_alike(), so
# on gfx950 this image resolves an Engram config and an explicit value is needed
# rather than the VLLM_PLE_CPU_OFFLOAD default.
#
# The tables cost 47.2 GiB per rank at TP=4, so 94.4 GiB at TP=2, which does not
# fit beside half of the 511 GB checkpoint on a 288 GiB card: TP=2 always
# offloads. TP=4 keeps them resident, as the validated concurrency 1-32 run
# measured, because offloaded lookups go to pinned host memory over UVA and
# nothing below c128 is short of KV. Resident leaves 14.83M KV tokens, which is
# 232K per request at c64 and healthy, but only 116K at c128, under the 122K at
# which TP=2 c64 collapsed. Offloading lifts it to 33.06M, so 258K at c128.
if (( TP == 2 || CONC >= 128 )); then
ENGRAM_CONFIG='{"cpu_offload":true}'
else
ENGRAM_CONFIG='{"cpu_offload":false}'
fi

# Graph capture covers twice the outer concurrency, floored at the #3058 size of
# 128 sequences, across the 1+5 DSpark token shape. Twice leaves headroom for
# AgentX subagent fan-out above the outer concurrency.
NUM_SPEC_TOKENS=5
GRAPH_NUM_SEQS=$((2 * CONC))
if (( GRAPH_NUM_SEQS < 128 )); then
GRAPH_NUM_SEQS=128
fi
CAPTURE_SIZE=1
while (( CAPTURE_SIZE < MAX_NUM_SEQS * (1 + NUM_SPEC_TOKENS) && CAPTURE_SIZE < 2048 )); do
while (( CAPTURE_SIZE < GRAPH_NUM_SEQS * (1 + NUM_SPEC_TOKENS) && CAPTURE_SIZE < 2048 )); do
CAPTURE_SIZE=$((CAPTURE_SIZE * 2))
done

# The sparse-attention indexer and its companion per-rank buffers scale with
# --max-num-batched-tokens at roughly 4.4 MiB per token, measured on gfx950, so
# a smaller prefill chunk buys KV room. TP=2 starts from half the per-rank space
# and is the arm that runs short: at the upstream 16384 it holds 7.84M KV
# tokens, 122K per request at c64, where run 35574132719 fell to a 17.6% prefix
# cache hit rate, 187 s TTFT and 150 tok/s against 94.8%, 1.3 s and 957 tok/s at
# c32. Every point that held had 232K per request or more, so keep the upstream
# chunk through c32 (245K at TP=2) and trade it away only above that. B300 runs
# 8192 at TP=4 and the Blackwell TP=2 arms run 4096 (#3320, #3321).
if (( CONC >= 128 )); then
(( TP == 2 )) && BATCHED_TOKENS=4096 || BATCHED_TOKENS=8192
elif (( TP == 2 && CONC >= 64 )); then
BATCHED_TOKENS=8192
else
BATCHED_TOKENS=16384
fi

# DSpark verifies 1+5 tokens per sequence, so a decode batch of max_num_seqs
# needs six times that many token slots. The MI355X API-server default of 1024
# sequences therefore wants 6144, and below that the engram projection faults
# during profiling: TP=2 c128 at 4096 leaves four slots per sequence and dies
# with HSA_STATUS_ERROR_EXCEPTION at M=1024, N=129280, K=256, reproduced on two
# separate GPU pairs. Where the chunk is that small, cap in-flight sequences at
# the shape graph capture already covers, which also keeps the largest decode
# batch on a captured graph. Leave the default alone everywhere else.
DEFAULT_MAX_NUM_SEQS=1024
MAX_NUM_SEQS=""
if (( BATCHED_TOKENS < DEFAULT_MAX_NUM_SEQS * (1 + NUM_SPEC_TOKENS) )); then
MAX_NUM_SEQS="$GRAPH_NUM_SEQS"
fi

# Use the runner-specific port assigned by launch_mi355x-amds.sh.
export AIPERF_SERVER_URL="http://localhost:${PORT}"
export AIPERF_SERVER_METRICS_URLS="${AIPERF_SERVER_URL}/metrics"
Expand All @@ -73,6 +125,7 @@ VLLM_CMD=(
--tokenizer-mode deepseek_v41
--tool-call-parser deepseek_v41 --enable-auto-tool-choice
--reasoning-parser deepseek_v41
--engram-config "$ENGRAM_CONFIG"
# aiter, not aiter_triton_mxfp4_bf16: the plain name opens vLLM's full
# priority list and the CK kernel at its head wins. CK quantizes
# activations to FP8 internally and dispatches the a8w4 experts
Expand All @@ -82,11 +135,21 @@ VLLM_CMD=(
--gpu-memory-utilization 0.9
--speculative-config "$SPEC_CONFIG"
--max-model-len 1048576
--max-num-seqs "$MAX_NUM_SEQS"
--max-cudagraph-capture-size "$CAPTURE_SIZE"
--max-num-batched-tokens 16384
--max-num-batched-tokens "$BATCHED_TOKENS"
# vllm-project/vllm#56227 added SWA bounded replay (default on) after the
# eed1f3d0 pin and before this one. It pads the replayed tokens' slots in the
# prefix-cacheable groups, but the window clamp it relies on landed in the
# FlashInfer and FlashMLA kernels; the ROCm sparse SWA path only gained the
# replay_start kwarg. On gfx950 every TP=2 and TP=4 point of run 35567570539
# died with HSA_STATUS_ERROR_MEMORY_FAULT at the first prefix hit carrying a
# replay start. Drop this once ROCm clamps too; prefix caching stays on.
--no-swa-bounded-replay
--disable-uvicorn-access-log
)
if [[ -n "$MAX_NUM_SEQS" ]]; then
VLLM_CMD+=(--max-num-seqs "$MAX_NUM_SEQS")
fi
printf '%q ' "${VLLM_CMD[@]}" | tee "$RESULT_DIR/vllm_command.txt"
printf '\n' | tee -a "$RESULT_DIR/vllm_command.txt"
SERVER_PID=""
Expand Down
17 changes: 14 additions & 3 deletions configs/amd-master.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -1466,8 +1466,14 @@ dsv4-fp4-mi355x-sglang-agentic-mtp:
# tag predates vllm-project/vllm#56503, which moves the mHC delayed pre block off
# the eager Torch reference and onto AITER. MI355X run 34710937012 passed the
# full concurrency 1-32 sweep and eval-only concurrency 32.
# TP=2 joins TP=4, and both extend to concurrency 128. TP=2 became feasible on
# this SKU once the Engram tables moved to host memory: they cost 47.2 GiB of
# device memory per rank, which is what previously forced four GPUs per server.
# Two GPUs per server doubles the servers per node and is the layout that
# decides whether DSv4.1-Flash is throughput- or interactivity-bound here.
dsv41flash-fp4-mi355x-vllm-agentic-dspark:
image: vllm/vllm-openai-rocm:nightly-eed1f3d0c6043bd494424a22443ee198dd56f657
# ROCm 10.0 nightly channel, shared with kimik3-fp4-mi355x-vllm-agentic-mtp.
image: vllm/vllm-openai-rocm:nightly-rocm100-3df4ae153eb385e27b52f26c81f8edb9e20b9984
model: deepseek-ai/DeepSeek-V4.1-Flash
model-prefix: dsv41flash
runner: cluster:mi355x-amds
Expand All @@ -1478,5 +1484,10 @@ dsv41flash-fp4-mi355x-vllm-agentic-dspark:
agentic-coding:
- dram-utilization: 0.60
search-space:
# Follow upstream AMD Engram defaults; omit the CUDA-only config flag.
- { tp: 4, kv-offloading: none, spec-decoding: mtp, conc-list: [1, 2, 4, 8, 16, 32] }
# The recipe only changes the memory split where a point ran short of KV:
# TP=4 offloads the engram tables at c128, and the prefill chunk shrinks at
# TP=2 c64, TP=4 c128 and TP=2 c128. Every point is measured here rather
# than combined from the earlier run: the image move leaves concurrency
# 1-32 measured only on the superseded nightly-eed1f3d0 pin.
- { tp: 4, kv-offloading: none, spec-decoding: mtp, conc-list: [1, 2, 4, 8, 16, 32, 64, 128] }
- { tp: 2, kv-offloading: none, spec-decoding: mtp, conc-list: [1, 2, 4, 8, 16, 32, 64, 128] }
9 changes: 5 additions & 4 deletions docs/configuration-procedures.md
Original file line number Diff line number Diff line change
Expand Up @@ -555,16 +555,17 @@ A configuration is ready for sweep only when the executable files agree, the exa

## DeepSeek-V4.1-Flash on MI355X

The draft `dsv41flash-fp4-mi355x-vllm-agentic-dspark` recipe extends [#2958](https://github.com/SemiAnalysisAI/InferenceX/pull/2958) to MI355X AgentX: TP4, concurrency 1–32, native five-token DSpark. Throughput uses the [committed golden AL](../golden_al_distribution/dsv41flash_dspark.yaml) of 3.51 for thinking on and five draft tokens, with synthetic rejection sampling and adaptive verification disabled. Accuracy evals retain real block rejection but, unlike the CUDA arms, also keep adaptive verification disabled: it trims verification requests on device, which the ROCm `DeepseekV4IndexerBackend` does not support, and the engine refused to start with it enabled ([run 34651830283](https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34651830283)). FP4 describes the MXFP4 experts; the checkpoint also contains MXFP8 weights.
The `dsv41flash-fp4-mi355x-vllm-agentic-dspark` recipe extends [#2958](https://github.com/SemiAnalysisAI/InferenceX/pull/2958) to MI355X AgentX: TP4 and TP2, concurrency 1–128, native five-token DSpark. Throughput uses the [committed golden AL](../golden_al_distribution/dsv41flash_dspark.yaml) of 3.51 for thinking on and five draft tokens, with synthetic rejection sampling and adaptive verification disabled. Accuracy evals retain real block rejection but, unlike the CUDA arms, also keep adaptive verification disabled: it trims verification requests on device, which the ROCm `DeepseekV4IndexerBackend` does not support, and the engine refused to start with it enabled ([run 34651830283](https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34651830283)). FP4 describes the MXFP4 experts; the checkpoint also contains MXFP8 weights.

Follow the AMD overrides in the merged [upstream recipe #968](https://github.com/vllm-project/recipes/pull/968): `VLLM_ROCM_USE_AITER=1`, `VLLM_ROCM_USE_AITER_MOE=1`, and `--moe-backend aiter`. The generic AITER selector lets vLLM pick the CK a8w4 experts, matching the DSV4-Pro MI355X recipe. The recipe pins `semianalysis_cc_traces_weka_062126` (the unfiltered corpus) via `WEKA_LOADER_OVERRIDE`. KV stays GPU-resident; Engram follows upstream AMD defaults. Do not copy the NVIDIA `--engram-config` option: upstream currently rejects it on ROCm. The MI355X launcher uses the shared HF cache and mounts this model's repository at `/ix`, and exports `INFMAX_CONTAINER_WORKSPACE=/ix` so AgentX dependencies and outputs resolve inside that mount.
Follow the AMD overrides in the merged [upstream recipe #968](https://github.com/vllm-project/recipes/pull/968): `VLLM_ROCM_USE_AITER=1`, `VLLM_ROCM_USE_AITER_MOE=1`, and `--moe-backend aiter`. The generic AITER selector lets vLLM pick the CK a8w4 experts, matching the DSV4-Pro MI355X recipe. The recipe pins `semianalysis_cc_traces_weka_062126` (the unfiltered corpus) via `WEKA_LOADER_OVERRIDE`. KV stays GPU-resident. Engram stayed on GPU under the upstream AMD defaults until [vllm-project/vllm#57491](https://github.com/vllm-project/vllm/pull/57491) widened the two `is_cuda()` gates to `is_cuda_alike()`. From that commit on, ROCm resolves an `EngramConfig` and `cpu_offload` defaults to on through `VLLM_PLE_CPU_OFFLOAD`, so the recipe sets `--engram-config` explicitly rather than leaning on that default. TP=2 always offloads, since the tables need 94.4 GiB per rank there; TP=4 keeps them resident through concurrency 64, where the KV pool is not the constraint, and offloads only at 128. The recipe likewise trims `--max-num-batched-tokens` only above concurrency 32, to 8192 at TP=2 c64 and TP=4 c128 and to 4096 at TP=2 c128, because the sparse-attention indexer and its companion per-rank buffers grow at roughly 4.4 MiB per batched token. Where that chunk falls below six times the API-server default of 1024 sequences, `--max-num-seqs` is capped at the graph-capture shape: DSpark verifies 1+5 tokens per sequence, and at 4096 against 1024 sequences the engram projection faults during profiling. The rule in every case is to spend device memory on KV only at the concurrencies that ran short of it, leaving the validated low-concurrency settings alone. Images built before that merge still reject the option on ROCm. The MI355X launcher uses the shared HF cache and mounts this model's repository at `/ix`, and exports `INFMAX_CONTAINER_WORKSPACE=/ix` so AgentX dependencies and outputs resolve inside that mount.

**GPU validation:** [Run 34710937012](https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34710937012) passed the exact pinned image for throughput at concurrency 1, 2, 4, 8, 16, and 32, plus eval-only concurrency 32. The recipe uses `vllm/vllm-openai-rocm:nightly-eed1f3d0c6043bd494424a22443ee198dd56f657` (digest `sha256:960228cf…`, published 2026-09-12). The earlier `deepseekv41-flash-0909` tag predates [vllm-project/vllm#56503](https://github.com/vllm-project/vllm/pull/56503), which moves the mHC delayed pre block off the eager Torch reference and onto AITER; the merged [upstream recipe #968](https://github.com/vllm-project/recipes/pull/968) pins the same nightly and records the complete InferenceX command. Follow the [AgentX procedure](./eval-agentx-procedures.md#7-run-agentx-fast-feedback-versus-canonical-evidence) for future runtime evidence; local generation and registry metadata alone are not GPU proof.
**GPU validation:** The recipe uses `vllm/vllm-openai-rocm:nightly-rocm100-3df4ae153eb385e27b52f26c81f8edb9e20b9984` (digest `sha256:eccb72b7…`, published 2026-09-21) on the ROCm 10.0 nightly channel that `kimik3-fp4-mi355x-vllm-agentic-mtp` already runs on this cluster. The sweep in [#3326](https://github.com/SemiAnalysisAI/InferenceX/pull/3326) qualifies that pin across TP4 and TP2 at concurrency 1–128, and is the only evidence for it: [run 34710937012](https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34710937012) covered TP4 concurrency 1–32 plus eval-only concurrency 32 on the superseded `nightly-eed1f3d0c6043bd494424a22443ee198dd56f657`, so its points do not carry onto this image. The merged [upstream recipe #1006](https://github.com/vllm-project/recipes/pull/1006) documents the MI355X TP2 Engram offload and `--no-swa-bounded-replay`, and the merged [#968](https://github.com/vllm-project/recipes/pull/968) records the original AMD overrides and the complete InferenceX command. Follow the [AgentX procedure](./eval-agentx-procedures.md#7-run-agentx-fast-feedback-versus-canonical-evidence) for future runtime evidence; local generation and registry metadata alone are not GPU proof.

## DeepSeek-V4.1-Flash on MI300X and MI325X

`dsv41flash-fp4-mi300x-vllm-agentic-dspark` and `dsv41flash-fp4-mi325x-vllm-agentic-dspark`
copy the validated MI355X vLLM arm onto gfx942, on the same ROCm nightly and with the same
copy the validated MI355X vLLM arm onto gfx942, on the `nightly-eed1f3d0` ROCm nightly the
MI355X arm used before it moved to the `nightly-rocm100` channel, and with the same
AMD overrides (`VLLM_ROCM_USE_AITER=1`, `VLLM_ROCM_USE_AITER_MOE=1`,
`VLLM_USE_BREAKABLE_CUDAGRAPH=1`, `--moe-backend aiter`, adaptive verification off). gfx942
is not in the upstream hardware table, and it has no FP4 MFMA: the plain `aiter` MoE
Expand Down
Loading
Loading