Skip to content

Commit fbf1f40

Browse files
authored
config(dsv41flash): re-sweep MI355X on the first ROCm 10.0 nightly carrying vllm#58510 / 在首个包含 vllm#58510 的 ROCm 10.0 nightly 上重新扫描 MI355X (#3420)
1 parent cd531ee commit fbf1f40

2 files changed

Lines changed: 12 additions & 1 deletion

File tree

‎configs/amd-master.yaml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1410,7 +1410,7 @@ dsv4-fp4-mi355x-sglang-agentic-mtp:
14101410
# decides whether DSv4.1-Flash is throughput- or interactivity-bound here.
14111411
dsv41flash-fp4-mi355x-vllm-agentic-dspark:
14121412
# ROCm 10.0 nightly channel, shared with kimik3-fp4-mi355x-vllm-agentic-mtp.
1413-
image: vllm/vllm-openai-rocm:nightly-rocm100-7f1a5398e9610d96c473931a26c0e12bbe0d0423
1413+
image: vllm/vllm-openai-rocm:nightly-rocm100-29468dde8b515031dc6d4d9d06bf0a2fa0442098
14141414
model: deepseek-ai/DeepSeek-V4.1-Flash
14151415
model-prefix: dsv41flash
14161416
runner: cluster:mi355x-amds

‎perf-changelog.yaml‎

Lines changed: 11 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -8909,3 +8909,14 @@
89098909
description:
89108910
- "Update SGLang ROCm image from v0.5.18-rocm720-mi35x-20260828 to v0.5.20-rocm720-mi35x-20260924 (latest nightly)"
89118911
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3423
8912+
8913+
- config-keys:
8914+
- dsv41flash-fp4-mi355x-vllm-agentic-dspark
8915+
scenario-type:
8916+
- agentic-coding
8917+
description:
8918+
- "Re-sweep the MI355X DeepSeek-V4.1-Flash vLLM AgentX arm on the next ROCm 10.0 nightly, replacing nightly-rocm100-7f1a5398. Only the image changes: the TP=4 and TP=2 rows keep concurrency 1, 2, 4, 8, 16, 32, 64 and 128, and every recipe setting is unchanged, including Engram placement per TP and concurrency, the prefill chunk and sequence cap ladder, graph capture, and five-token DSpark with golden acceptance 3.51 for throughput and real block verification for evals."
8919+
- "The pin is mainly for vllm-project/vllm#58510, which computes the MXFP8 GEMM on native 32x32 block scales for gfx950. This arm serves a checkpoint whose non-routed weights are MXFP8 on exactly that architecture, so the target image is the first nightly whose commit contains that merge rather than simply the next one published. The exact tag, its publication time and digest are filled in once that build exists."
8920+
- "在下一个 ROCm 10.0 nightly 上重新扫描 MI355X DeepSeek-V4.1-Flash vLLM AgentX 臂,替换 nightly-rocm100-7f1a5398。本次仅更换镜像:TP=4 与 TP=2 两行保持并发 1、2、4、8、16、32、64 与 128,其余 recipe 设置全部不变,包括按 TP 与并发决定的 Engram 放置、prefill 分块与序列上限阶梯、图捕获,以及吞吐用黄金接受长度 3.51、评测用真实分块验证的五 token DSpark。"
8921+
- "此次固定主要是为了 vllm-project/vllm#58510:它在 gfx950 上使用原生 32x32 block scale 计算 MXFP8 GEMM。本臂所服务的 checkpoint 其非 routed 权重正是该架构上的 MXFP8,因此目标镜像是第一个其提交包含该合并的 nightly,而非简单地取下一个发布的 nightly。具体标签、发布时间与摘要将在该构建出现后补充。"
8922+
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3420

0 commit comments

Comments
 (0)