Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@

InferenceX owns the recipes in this directory. For every NVIDIA srt-slurm launch, the srt driver ([`infx/launch/drivers/srt/checkout.py`](../../../infx/launch/drivers/srt/checkout.py)) makes a job-local Git clone of the pinned submodule and copies this entire tree into `recipes/`. It records the actual revision in `srt-slurm-sha.txt`; power lanes copy that revision into `power-producer-sha.txt` for result validation.

The shared version is the Git submodule pointer at [`utils/srt-slurm`](../../../utils/srt-slurm), currently [v2.36.0](https://github.com/NVIDIA/srt-slurm/releases/tag/v2.36.0) (`7b5863a7837673d81403b076be219bbf18a7700f`). Update that submodule pointer when upgrading, then run the recipe and integration checks. Do not add model-specific checkout branches to launchers.
The shared version is the Git submodule pointer at [`utils/srt-slurm`](../../../utils/srt-slurm), currently [v2.43.4](https://github.com/NVIDIA/srt-slurm/releases/tag/v2.43.4) (`848f72d45b05af0fc082a9d1df49a8b4e7e61507`). Update that submodule pointer when upgrading, then run the recipe and integration checks. Do not add model-specific checkout branches to launchers.

InferenceX requires srt-slurm 2.0 or newer and `schema: 2` recipes. Legacy recipe layouts are unsupported; migrate them before adding them to this tree.

Expand Down Expand Up @@ -54,8 +54,6 @@ All referenced recipes must be checked in: srt-slurm 2 ships curated examples in
Install the shared pin in an isolated environment, then use its CLI:

```bash
# Verify each supported recipe directory before rewriting it.
srtctl migrate --verify -f benchmarks/multi_node/srt-slurm-recipes/dsr1/sglang
srtctl migrate --in-place -f benchmarks/multi_node/srt-slurm-recipes/dsr1/sglang
# Repeat for the other model/engine directories.
python -m pytest infx/tests/matrix/ -q
Expand All @@ -64,7 +62,7 @@ python -m infx.matrix.generate full-sweep \
--framework dynamo-sglang dynamo-trt dynamo-vllm --multi-node
```

Validate recipes with the exact launcher pin, including all override variants. For a path-only reorganization, compare generated matrices before and after with the path mapping applied; all other fields, including eval selection and node counts, must match. A passing local schema check does not replace the full hardware sweep and evals.
Validate recipes with the exact launcher pin, including all override variants. The current migration CLI has no `--verify`; use `srtctl dry-run` for each selected recipe after migration, supplying launcher-provided values such as `benchmark.concurrencies` through `--set` when needed. For a path-only reorganization, compare generated matrices before and after with the path mapping applied; all other fields, including eval selection and node counts, must match. A passing local schema check does not replace the full hardware sweep and evals.

The initial migration also resolves compatibility issues that `srtctl migrate` cannot fix itself:

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@

InferenceX 负责维护本目录中的配置。每次 NVIDIA srt-slurm 启动时,srt 驱动([`infx/launch/drivers/srt/checkout.py`](../../../infx/launch/drivers/srt/checkout.py))都会为作业创建固定版本子模块的本地 Git 克隆,并将整个目录复制到 `recipes/`。它将实际提交记录到 `srt-slurm-sha.txt`;功耗测试路径还会将其复制到 `power-producer-sha.txt`,供结果校验使用。

统一版本由 [`utils/srt-slurm`](../../../utils/srt-slurm) 的 Git 子模块指针指定,目前为 [v2.36.0](https://github.com/NVIDIA/srt-slurm/releases/tag/v2.36.0)(`7b5863a7837673d81403b076be219bbf18a7700f`)。升级时更新该子模块指针,然后运行配置和集成检查。不要在启动器中新增按模型选择检出版本的分支。
统一版本由 [`utils/srt-slurm`](../../../utils/srt-slurm) 的 Git 子模块指针指定,目前为 [v2.43.4](https://github.com/NVIDIA/srt-slurm/releases/tag/v2.43.4)(`848f72d45b05af0fc082a9d1df49a8b4e7e61507`)。升级时更新该子模块指针,然后运行配置和集成检查。不要在启动器中新增按模型选择检出版本的分支。

InferenceX 要求 srt-slurm 2.0 或更新版本,且配置必须声明 `schema: 2`。不支持旧版配置结构;加入本目录前必须先完成迁移。

Expand Down Expand Up @@ -54,8 +54,6 @@ TileRT 使用固定版本的上游 srt-slurm 子模块。配置指定 `roles.pre
在隔离环境中安装统一版本,然后使用其 CLI:

```bash
# 重写前先验证每个受支持的配置目录。
srtctl migrate --verify -f benchmarks/multi_node/srt-slurm-recipes/dsr1/sglang
srtctl migrate --in-place -f benchmarks/multi_node/srt-slurm-recipes/dsr1/sglang
# 对其他模型/引擎目录重复执行。
python -m pytest infx/tests/matrix/ -q
Expand All @@ -64,7 +62,7 @@ python -m infx.matrix.generate full-sweep \
--framework dynamo-sglang dynamo-trt dynamo-vllm --multi-node
```

使用启动器指定的确切提交验证配置,包括全部覆盖变体。仅调整路径时,应按路径映射比较变更前后的生成矩阵;其他字段(包括评估选择和节点数)必须完全一致。本地配置校验通过不能替代完整硬件扫描和准确性评估。
使用启动器指定的确切提交验证配置,包括全部覆盖变体。当前迁移 CLI 已不提供 `--verify`;迁移后对每个选中的配置运行 `srtctl dry-run`,必要时通过 `--set` 提供启动器注入的值,例如 `benchmark.concurrencies`。仅调整路径时,应按路径映射比较变更前后的生成矩阵;其他字段(包括评估选择和节点数)必须完全一致。本地配置校验通过不能替代完整硬件扫描和准确性评估。

本次迁移还修复了 `srtctl migrate` 无法自动处理的兼容性问题:

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,7 @@
# with DSpark. Every arm offloads prefill KV to UMBP as a direct external store
# for the unified radix tree (no host cache tier), backed by a 1.5 TB
# hugepage DRAM tier on the prefill node.
schema: 2
base:
name: mi355x-dsv4-pro-0813-agentx-umbp
model:
Expand Down
Original file line number Diff line number Diff line change
@@ -1,3 +1,4 @@
schema: 2
name: agg-h200-tp8-mtp-c12

# Aggregated TP8 GLM-5.2 AgentX topology on one 8-GPU H200 node. MTP uses the
Expand Down Expand Up @@ -27,14 +28,17 @@ dynamo:
resources:
gpu_type: h200
gpus_per_node: 8
gpus_per_agg: 8
agg_nodes: 1
agg_workers: 1

infra:
etcd_nats_dedicated_node: false
nats_max_payload_mb: 64

services:
- name: etcd
type: etcd
placement:
node: infra
- name: nats
type: nats
placement:
node: infra
options:
max_payload_mb: 64
frontend:
type: dynamo
nginx_session_affinity: true
Expand All @@ -49,10 +53,27 @@ frontend:
active-prefill-tokens-threshold: None
active-prefill-tokens-threshold-frac: None

backend:
type: sglang
sglang_config:
aggregated:
engine: sglang
roles:
agg:
nodes: 1
workers: 1

gpus: 8
env:
HF_HOME: /hf_hub_cache
HF_HUB_CACHE: /hf_hub_cache
PIP_BREAK_SYSTEM_PACKAGES: "1"
PYTHONUNBUFFERED: "1"
SGLANG_OPT_USE_TOPK_V2: "1"
SGLANG_TIMEOUT_KEEP_ALIVE: "900"
SGLANG_ENABLE_THINKING: "1"
SGLANG_REASONING_EFFORT: max
SGLANG_ENABLE_UNIFIED_RADIX_TREE: "1"
SGLANG_HICACHE_DEBUG_LOG: "1"
SGLANG_HICACHE_DEBUG_SAMPLE_RATE: "16384"

args:
served-model-name: zai-org/GLM-5.2-FP8
trust-remote-code: true
tool-call-parser: glm47
Expand Down Expand Up @@ -83,22 +104,6 @@ backend:
stream-interval: 60
enable-metrics: true
enable-cache-report: true
aggregated_environment:
HF_HOME: /hf_hub_cache
HF_HUB_CACHE: /hf_hub_cache
PIP_BREAK_SYSTEM_PACKAGES: "1"
PYTHONUNBUFFERED: "1"
SGLANG_OPT_USE_TOPK_V2: "1"
SGLANG_TIMEOUT_KEEP_ALIVE: "900"
SGLANG_ENABLE_THINKING: "1"
SGLANG_REASONING_EFFORT: max
SGLANG_ENABLE_UNIFIED_RADIX_TREE: "1"
SGLANG_HICACHE_DEBUG_LOG: "1"
SGLANG_HICACHE_DEBUG_SAMPLE_RATE: "16384"
SGLANG_SIMULATE_ACC_LEN: "2.99"
SGLANG_SIMULATE_ACC_METHOD: match-expected
SGLANG_SIMULATE_ACC_TOKEN_MODE: real-draft-token

health_check:
max_attempts: 1440
interval_seconds: 10
Expand All @@ -118,7 +123,6 @@ telemetry:

benchmark:
type: custom
use_chat_template: true
command: bash /infmax-workspace/benchmarks/srt_agentic.sh
env:
INFMAX_CONTAINER_WORKSPACE: /infmax-workspace
Expand Down
Original file line number Diff line number Diff line change
@@ -1,3 +1,4 @@
schema: 2
name: agg-h200-tp8-mtp-c16

# Aggregated TP8 GLM-5.2 AgentX topology on one 8-GPU H200 node. MTP uses the
Expand Down Expand Up @@ -27,14 +28,17 @@ dynamo:
resources:
gpu_type: h200
gpus_per_node: 8
gpus_per_agg: 8
agg_nodes: 1
agg_workers: 1

infra:
etcd_nats_dedicated_node: false
nats_max_payload_mb: 64

services:
- name: etcd
type: etcd
placement:
node: infra
- name: nats
type: nats
placement:
node: infra
options:
max_payload_mb: 64
frontend:
type: dynamo
nginx_session_affinity: true
Expand All @@ -49,10 +53,27 @@ frontend:
active-prefill-tokens-threshold: None
active-prefill-tokens-threshold-frac: None

backend:
type: sglang
sglang_config:
aggregated:
engine: sglang
roles:
agg:
nodes: 1
workers: 1

gpus: 8
env:
HF_HOME: /hf_hub_cache
HF_HUB_CACHE: /hf_hub_cache
PIP_BREAK_SYSTEM_PACKAGES: "1"
PYTHONUNBUFFERED: "1"
SGLANG_OPT_USE_TOPK_V2: "1"
SGLANG_TIMEOUT_KEEP_ALIVE: "900"
SGLANG_ENABLE_THINKING: "1"
SGLANG_REASONING_EFFORT: max
SGLANG_ENABLE_UNIFIED_RADIX_TREE: "1"
SGLANG_HICACHE_DEBUG_LOG: "1"
SGLANG_HICACHE_DEBUG_SAMPLE_RATE: "16384"

args:
served-model-name: zai-org/GLM-5.2-FP8
trust-remote-code: true
tool-call-parser: glm47
Expand Down Expand Up @@ -83,22 +104,6 @@ backend:
stream-interval: 60
enable-metrics: true
enable-cache-report: true
aggregated_environment:
HF_HOME: /hf_hub_cache
HF_HUB_CACHE: /hf_hub_cache
PIP_BREAK_SYSTEM_PACKAGES: "1"
PYTHONUNBUFFERED: "1"
SGLANG_OPT_USE_TOPK_V2: "1"
SGLANG_TIMEOUT_KEEP_ALIVE: "900"
SGLANG_ENABLE_THINKING: "1"
SGLANG_REASONING_EFFORT: max
SGLANG_ENABLE_UNIFIED_RADIX_TREE: "1"
SGLANG_HICACHE_DEBUG_LOG: "1"
SGLANG_HICACHE_DEBUG_SAMPLE_RATE: "16384"
SGLANG_SIMULATE_ACC_LEN: "2.99"
SGLANG_SIMULATE_ACC_METHOD: match-expected
SGLANG_SIMULATE_ACC_TOKEN_MODE: real-draft-token

health_check:
max_attempts: 1440
interval_seconds: 10
Expand All @@ -118,7 +123,6 @@ telemetry:

benchmark:
type: custom
use_chat_template: true
command: bash /infmax-workspace/benchmarks/srt_agentic.sh
env:
INFMAX_CONTAINER_WORKSPACE: /infmax-workspace
Expand Down
Original file line number Diff line number Diff line change
@@ -1,3 +1,4 @@
schema: 2
name: agg-h200-tp8-mtp-c2

# Aggregated TP8 GLM-5.2 AgentX topology on one 8-GPU H200 node. MTP uses the
Expand Down Expand Up @@ -27,14 +28,17 @@ dynamo:
resources:
gpu_type: h200
gpus_per_node: 8
gpus_per_agg: 8
agg_nodes: 1
agg_workers: 1

infra:
etcd_nats_dedicated_node: false
nats_max_payload_mb: 64

services:
- name: etcd
type: etcd
placement:
node: infra
- name: nats
type: nats
placement:
node: infra
options:
max_payload_mb: 64
frontend:
type: dynamo
nginx_session_affinity: true
Expand All @@ -49,10 +53,27 @@ frontend:
active-prefill-tokens-threshold: None
active-prefill-tokens-threshold-frac: None

backend:
type: sglang
sglang_config:
aggregated:
engine: sglang
roles:
agg:
nodes: 1
workers: 1

gpus: 8
env:
HF_HOME: /hf_hub_cache
HF_HUB_CACHE: /hf_hub_cache
PIP_BREAK_SYSTEM_PACKAGES: "1"
PYTHONUNBUFFERED: "1"
SGLANG_OPT_USE_TOPK_V2: "1"
SGLANG_TIMEOUT_KEEP_ALIVE: "900"
SGLANG_ENABLE_THINKING: "1"
SGLANG_REASONING_EFFORT: max
SGLANG_ENABLE_UNIFIED_RADIX_TREE: "1"
SGLANG_HICACHE_DEBUG_LOG: "1"
SGLANG_HICACHE_DEBUG_SAMPLE_RATE: "16384"

args:
served-model-name: zai-org/GLM-5.2-FP8
trust-remote-code: true
tool-call-parser: glm47
Expand Down Expand Up @@ -83,22 +104,6 @@ backend:
stream-interval: 60
enable-metrics: true
enable-cache-report: true
aggregated_environment:
HF_HOME: /hf_hub_cache
HF_HUB_CACHE: /hf_hub_cache
PIP_BREAK_SYSTEM_PACKAGES: "1"
PYTHONUNBUFFERED: "1"
SGLANG_OPT_USE_TOPK_V2: "1"
SGLANG_TIMEOUT_KEEP_ALIVE: "900"
SGLANG_ENABLE_THINKING: "1"
SGLANG_REASONING_EFFORT: max
SGLANG_ENABLE_UNIFIED_RADIX_TREE: "1"
SGLANG_HICACHE_DEBUG_LOG: "1"
SGLANG_HICACHE_DEBUG_SAMPLE_RATE: "16384"
SGLANG_SIMULATE_ACC_LEN: "2.99"
SGLANG_SIMULATE_ACC_METHOD: match-expected
SGLANG_SIMULATE_ACC_TOKEN_MODE: real-draft-token

health_check:
max_attempts: 1440
interval_seconds: 10
Expand All @@ -118,7 +123,6 @@ telemetry:

benchmark:
type: custom
use_chat_template: true
command: bash /infmax-workspace/benchmarks/srt_agentic.sh
env:
INFMAX_CONTAINER_WORKSPACE: /infmax-workspace
Expand Down
Loading