Repository navigation
Conversation
将 AgentX SRT 配置迁移至 schema 2,并更新共用的 srt-slurm 版本。 Signed-off-by: Rohit Pujar Nagraj <rpujarnagraj@nvidia.com>
将 schema 迁移的变更记录条目关联到此 PR。 Signed-off-by: Rohit Pujar Nagraj <rpujarnagraj@nvidia.com>
为归档校验测试指定固定参数 ID,避免并行收集时因 gzip 时间戳不同而失败。 Signed-off-by: Rohit Pujar Nagraj <rpujarnagraj@nvidia.com>
合并最新 main,以更新 PR 分支并保留现有配置及变更记录。 Signed-off-by: Rohit Pujar Nagraj <rpujarnagraj@nvidia.com>
合并最新 main,并让 AgentX 配置使用自动注入的黄金接受长度。
…v2.43.4 Replace benchmark.client_placement with benchmark.placement.node in eleven Kimi-K3 vLLM GB200/GB300 AgentX recipes, which the v2.43.4 pin rejects as a v1 key. 将十一个 Kimi-K3 vLLM GB200/GB300 AgentX 配置中的 benchmark.client_placement 替换为 benchmark.placement.node;v2.43.4 固定版本会将前者视为 v1 字段并拒绝加载。
There was a problem hiding this comment.
Looks good, straightforward mechanical config migration. I reviewed all 11 recipe diffs and confirmed each is structurally identical: removes the deprecated benchmark.client_placement: head field and adds the equivalent benchmark.placement: { node: head } nested field, with no other content (concurrencies, env, telemetry, topology) changed. The perf-changelog.yaml change is a single new English-only bullet appended at the physical end of the existing PR #3755 entry, consistent with the append-only convention.
Extended reasoning...
The diff touches 11 Kimi-K3 vLLM AgentX srt-slurm recipe YAMLs plus the changelog; verified via git diff that all 11 changes are byte-for-byte structurally identical (one field removed, one nested field added) with no other recipe content altered, and the changelog edit only appends a line at the tail. No security-sensitive surface (auth, crypto, permissions) is touched. No bugs were found by the bug hunting system and no third-party objections appear in the timeline, so this small, mechanical, schema-conformant change is safe to approve without further human review.
Description
Follow-up to #3755, targeting its branch. srt-slurm v2.43.4 rejects
benchmark.client_placementas a pre-2.0 key, which breaks 11 Kimi-K3 vLLM AgentX recipes that #3755 did not touch (flagged by Alec Ibarra in review):kimik3/vllm/gb200-fp4/agentx/:agg-dcp16-dspark4-maxseq2-mooncake,agg-dcp16-nospec-mooncake,agg-tp8pp2-mooncake-c{16,32,48,72,96}kimik3/vllm/gb300-fp4/agentx/:disagg-1p1d-dcp8-dcp8-dspark4-mooncake,disagg-1p2d-dcp8-dcp8-dspark4-mooncake,disagg-1p3d-dcp8-dcp8-dspark4-mooncake,disagg-1p3d-dcp8-dcp8-dspark7-mooncakeEach file was rewritten with
srtctl migrate --in-placefrom the v2.43.4 pin. The only change isclient_placement: head->placement: {node: head}underbenchmark; placement is unchanged. Also adds a line to the #3755 changelog entry.Local validation (v2.43.4 pin): all 11 recipes pass
srtctl dry-runwith their ownbenchmark.concurrencies;infx/tests/launchand changelog tests pass (227 tests);kimik3full-sweep matrix generates. GPU execution remains unverified.Note: a dry-run over every
schema: 2recipe also flags three override bundles unrelated to this change (glm5.2/sglang/h200-fp8/agentx/disagg-2p2d-pcp8-tp8-dp8-mtp.yaml,glm5.2/sglang/h200-fp8/agentx/disagg-1p1d-pcp8-tp8-dp8-mtp6-hicache.yaml,dsv4/sglang/h200-fp8/agentx/agg-tp8-mtp-kvoffload.yaml): "telemetry requires the discovery plane on the infra node". Not addressed here.AI model disclosure
Type of Change
Checklist
inferencex-e2e/perf-changelog.yamland have not edited historical entries (extends the unmerged Update srt-slurm and migrate AgentX recipes to schema 2 / 更新 srt-slurm 并将 AgentX 配置迁移至 schema 2 #3755 entry)中文
改动说明
本 PR 是 #3755 的后续,目标分支为 #3755 的分支。srt-slurm v2.43.4 将
benchmark.client_placement视为 2.0 之前的字段并拒绝加载,导致 #3755 未涉及的 11 个 Kimi-K3 vLLM AgentX 配置无法使用(由 Alec Ibarra 在评审中指出)。文件列表见上文英文部分。每个文件均使用 v2.43.4 固定版本中的
srtctl migrate --in-place重写。唯一改动是将benchmark下的client_placement: head改为placement: {node: head},实际放置位置不变。同时在 #3755 的变更记录条目中追加一行说明。本地验证(v2.43.4):11 个配置在使用各自的
benchmark.concurrencies时均通过srtctl dry-run;infx/tests/launch与变更记录测试通过(227 项);kimik3full-sweep 矩阵可正常生成。尚未验证 GPU 运行。另外,对所有
schema: 2配置执行 dry-run 时,还发现三个与本改动无关的覆盖配置报错("telemetry requires the discovery plane on the infra node"),本 PR 未处理。AI 模型使用说明
改动类型与检查状态
配置改动已完成;本地测试已完成。尚未添加扫描标签,也未选择可复用的最终扫描。