Skip to content

Add Narwhal as a multi-node disaggregated framework / 添加 Narwhal 作为多节点分离式推理框架 #3768

Description

@athrael-soju

Is your feature request related to a problem? Please describe.

InferenceX benchmarks disaggregated prefill/decode serving with a fixed split of prefill and decode workers per run.

Narwhal (Apache-2.0, narwhal-inference on PyPI) is a new adaptive disaggregated inference framework, which serves a model on a fleet of vLLM engines with NIXL KV transfer, and a role controller moves engines between prefill and decode as demand changes, without reloading weights. Adding Narwhal as a framework would let InferenceX measure adaptive prefill/decode splits alongside its fixed-split results.

Describe the solution you'd like

A new multi-node framework, narwhal-vllm, with its own launch path modelled on the llmd-vllm driver and Slurm scripts. No current master entry uses the llm-d path, and it has no AgentX step. Paths are under inferencex-e2e/:

  • benchmarks/multi_node/narwhal/{submit.sh,job.slurm,server.sh}: InferenceX owns the Slurm allocation and runs one container per node on an unmodified upstream vllm/vllm-openai image. Each node starts its vLLM engines directly with python3 -m vllm.entrypoints.openai.api_server and NIXL kv_both. The router node writes the Narwhal fleet file from the allocated addresses, profiles the engines with narwhal-profile, starts narwhal-serve from its own uv environment and waits on /ready. It then runs the fixed-sequence points, the AgentX replay or the eval-only job that the matrix row selects.
  • benchmarks/multi_node/narwhal-recipes/: the Narwhal fleet settings and engine arguments for each point, selected through CONFIG_FILE in additional-settings.
  • A LaunchPath, its route, a request type and infx/launch/drivers/narwhal.py modelled on llmd.py, plus a narwhal-vllm row in benchmarks/multi_node/runtime_settings.sh.
  • Master-config entries with framework: narwhal-vllm, multinode: true, disagg: true, kv-p2p-transfer: nixl and router: { name: narwhal, version: <narwhal-inference release> }, plus a perf-changelog.yaml entry, behavioural tests and synchronized _zh.md docs.

Most active models run AgentX only, so the first PR would cover AgentX, fixed-sequence and evals.

Questions for maintainers before I start

  1. Image. Only the router node needs Narwhal. It can run from its own uv environment, as AIPerf does, so the upstream vllm/vllm-openai image runs as shipped, and Narwhal's own engine launcher isn't used. Is that acceptable?
  2. Engine-first ordering. Would narwhal-vllm count as an open-source vLLM deployment, as dynamo-vllm does, or as a vendor-specific framework under Check 6(b)?
  3. Upstream recipe. For multi-node disaggregated vLLM, the review checklist asks for the prefill, decode and router/frontend commands in vLLM recipes. Would you expect vLLM recipes to cover the Narwhal router commands?
  4. Result labelling. Master entries record fixed prefill/decode worker counts, but Narwhal's split changes during a run. Narwhal can hold its starting roles (controller.advisory: true), which would give a fixed-split baseline with the current labels. For the adaptive runs, which option do you prefer: label each run with its starting split and record role changes in the artifacts, or add a schema field for the adaptive mode? The second option probably also needs an InferenceX-app change.
  5. Engine profiling. The Narwhal router doesn't start until every engine has a measured profile from narwhal-profile. In this launch, profiles bind to each engine process, so every job, including each eval-only job, profiles its engines before the benchmark runs. Is a profiling phase inside each job acceptable?

Describe alternatives you've considered

  • Publishing results outside InferenceX. Those results would be unofficial and harder to compare.

Additional context

The PR will follow the review flow in CONTRIBUTING.md, including the AI model disclosure.

中文

功能请求是否与某个问题相关?请描述。

InferenceX 对分离式预填充/解码推理进行基准测试时,每次运行都使用固定数量的预填充和解码 worker。

Narwhal(Apache-2.0,PyPI 包名 narwhal-inference)是一个新的自适应分离式推理框架,在一组 vLLM 引擎上服务一个模型,通过 NIXL 传输 KV。其角色控制器会根据负载变化在预填充和解码之间调整引擎,且无需重新加载权重。将 Narwhal 添加为框架后,InferenceX 可以在固定划分结果之外,测量自适应的预填充/解码划分。

期望的方案

新增多节点框架 narwhal-vllm,使用参照 llmd-vllm driver 和 Slurm 脚本编写的独立启动路径。目前没有主配置条目使用 llm-d 路径,且该路径没有 AgentX 步骤。以下路径均位于 inferencex-e2e/ 下:

  • benchmarks/multi_node/narwhal/{submit.sh,job.slurm,server.sh}:由 InferenceX 负责 Slurm 分配,每个节点在未经修改的上游 vllm/vllm-openai 镜像上运行一个容器。每个节点直接使用 python3 -m vllm.entrypoints.openai.api_server 和 NIXL kv_both 启动各自的 vLLM 引擎。router 节点根据分配到的地址写入 Narwhal fleet 文件,用 narwhal-profile 为引擎采集性能画像,从独立的 uv 环境启动 narwhal-serve 并等待 /ready。随后按矩阵行的选择,执行固定序列测试点、AgentX 回放或仅评估作业。
  • benchmarks/multi_node/narwhal-recipes/:每个测试点的 Narwhal fleet 设置和引擎参数,通过 additional-settings 中的 CONFIG_FILE 选择。
  • 一个 LaunchPath 及其路由、一个请求类型,以及参照 llmd.py 编写的 infx/launch/drivers/narwhal.py,并在 benchmarks/multi_node/runtime_settings.sh 中添加 narwhal-vllm 行。
  • 主配置条目设置 framework: narwhal-vllm、multinode: true、disagg: true、kv-p2p-transfer: nixl 和 router: { name: narwhal, version: <narwhal-inference 版本> },并添加 perf-changelog.yaml 条目、行为测试和同步的 _zh.md 文档。

大多数活跃模型只运行 AgentX,因此第一个 PR 将覆盖 AgentX、固定序列测试和评估。

开始前需要维护者确认的问题

  1. 镜像。 只有 router 节点需要 Narwhal。它可以像 AIPerf 一样运行在独立的 uv 环境中,因此上游 vllm/vllm-openai 镜像可以按原样运行,也不会使用 Narwhal 自带的引擎启动器。这种方式是否可以接受?
  2. 引擎优先顺序。 在 Check 6(b) 中,narwhal-vllm 会像 dynamo-vllm 一样被视为开源 vLLM 部署,还是被视为特定厂商框架?
  3. 上游 recipe。 对于多节点分离式 vLLM,评审清单要求在 vLLM recipes 中记录预填充、解码和 router/frontend 命令。你们是否期望 vLLM recipes 覆盖 Narwhal router 的命令?
  4. 结果标注。 主配置条目记录固定的预填充/解码 worker 数量,但 Narwhal 的划分会在运行中变化。Narwhal 可以保持初始角色不变(controller.advisory: true),从而在现有标注下提供固定划分的基线。对于自适应运行,你们倾向于哪种方式:用初始划分标注每次运行并在产物中记录角色变化,还是为自适应模式新增 schema 字段?后者可能还需要修改 InferenceX-app。
  5. 引擎性能画像。 只有每个引擎都具备 narwhal-profile 测得的性能画像后,Narwhal router 才会启动。在这种启动方式下,性能画像绑定到每个引擎进程,因此每个作业(包括每个仅评估作业)都需要在基准测试前为引擎采集性能画像。是否可以在每个作业内加入这一画像采集阶段?

考虑过的替代方案

  • 在 InferenceX 之外发布结果。这类结果属于非官方结果,也更难比较。

补充说明

该 PR 将遵循 CONTRIBUTING.md 中的评审流程,包括 AI 模型披露。

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions