Skip to content

About

Pinned, gate-verified vLLM deployment stack for AMD MI210 (gfx90a/CDNA2)

Resources

Contributing

Stars

4 stars

Watchers

0 watching

Forks

Repository files navigation

vLLM on AMD MI210 (gfx90a / CDNA2)

Deployment stack for running vLLM on 2x AMD Instinct MI210. It carries the patches this hardware needs — several upstream, several local — pinned, verified at build time, and documented with the measurements that justify them.

Status: builds and verifies. Built on 2x MI210 on 2026-08-04; all three verification tiers pass with hardware_validated=true (23 numeric tests). DERIVED_IMAGE in VERSIONS is still empty because the image has not been pushed to a registry — run build/build.sh to produce your own, or publish one and record its digest.


What this carries

patch origin what it does
multi-pass paged attention vllm#39001 (Eugene Kuznetsov) removes the 131,072-token context ceiling
gfx90a free-kernel guards local declines shapes that compute silently wrong attention on CDNA2
single-pass 131k–262k local keeps the common long-context range off the multi-pass path
int4 interleave packing vllm#43389 (amd-xavierwang), gate widened 1.45–4.8x on the int4 MoE kernel, bit-identical output
wvSplitK stride guard vllm#50618 (John Qin / Yanyuan Qin) fixes an out-of-bounds read on strided activations
sharded_state TP guard local rejects checkpoints saved at a different TP size instead of half-loading them
benchmark_moe int8_w8a16 local, grouping from vllm#31011 the tuner could not run at all before this
W4A16 gfx90a tile table local gfx90a inherited the MI300 tiles, which assume 304 CUs against MI210's 104
W4A16 magic-bias dequant local bit-trick dequant + scale hoist in the dense W4A16 GEMM inner loop; gfx90a has no bf16 VALU arithmetic
W4A16 narrow rung to M≤16 local 1.37–1.45x GB/s on M=9..16, by keeping the narrow tile active over that range instead of dropping to BLOCK_M=64
W4A16 magic-bias gate on BLOCK_M local keys the gate on the tile rather than on M, where the cost crossover actually lives, so a future ladder change cannot silently mis-enable it
fp8 W8A16 Triton kernel local first ROCm entry in _POSSIBLE_WFP8A16_KERNELS (upstream ships that list empty), plus the CDNA2 dispatcher fix that was routing fp8 checkpoints into a torch._scaled_mm crash
NVFP4 W4A16 Triton kernel local packed e2m1 decode + per-16-group scales for gfx90a — not in the pinned tag: this one is post-VLLM_REF, so it is not in an image built from VERSIONS as it stands (see patches/registry.yaml)
fused GDN decode on CDNA local builds vLLM's CUDA-only fused gated-delta-net decode kernel for gfx90a, so GDN layers stop falling back to unfused Triton — not in the pinned tag either (post-VLLM_REF, see patches/registry.yaml)

Every row except the last two is in the tag VERSIONS pins and therefore in any image built from it. The NVFP4 and fused-GDN rows are not: both are marked status: post-tag in patches/registry.yaml and land in the next pin.

Measurements behind each are in the commit messages on davetha/vllm, and the investigation history is in mi210-llm-stack.

Prior art

The gfx90a work here builds on people who got there first:

  • Eugene Kuznetsov (vllm#39001) — the multi-pass reduction. Developed on gfx942; this repo validated it on gfx90a and added the CDNA2 guards.
  • amd-xavierwang (vllm#43389) — int4 interleave packing, gated to RDNA upstream. Measured here on CDNA2 and widened.
  • rlrs (vllm#49888 and others) — upstreaming AITER attention for gfx90a on MI250X.
  • The packaging pattern — digest-pinned base, overlays as source of truth, diffs as documentation — is taken from ryanzhou/deepseek-v4-flash-mi300x.

Layout

VERSIONS              every pin; base image by DIGEST, never a tag
run.sh                serve any model in one command -- docs/RUNNING.md
docs/RECIPE-QWEN38-27B-INT8.md
                      end-to-end: the fastest known config for this model on
                      one MI210, copy-pasteable
docs/SPEC-DECODE.md   speculation: dflash N=12, the draft, and why depth is not
                      a fixed property of the model
docs/INT8-GFX90A.md   INT8 W8A8: confirming the AITER kernel, and what a stock
                      W8A8 recipe leaves in BF16
gpu-nodes.sh          picks the gfx90a cards on a mixed-GPU host; sourced by
                      run.sh and build/add-aiter.sh
compose.yaml          the stack consumers run
build/                the ONE compiled layer + its gates
  Dockerfile          rebuilds _rocm_C for gfx90a. NEEDS NO GPU.
  entrypoint.sh       the image's own dispatch + arch preamble
  build.sh            build, verify on real cards, record the digest
  add-aiter.sh        OPTIONAL, needs the cards: AITER + gfx90a ASM kernels
  verify.sh           static markers -> runtime gates -> numeric tests
  quant_nemotron_heretic_vast.sh  W4A16 on a rented box; the ignore list is
                      NVIDIA's own, because quantising a Mamba-2 hybrid's SSM
                      projections yields fluent WRONG text rather than an error
  watch_vast_quant.sh watchdog that kills the rental when it stops progressing
patches/              NO diffs here -- the patches are branches on the fork,
  registry.yaml       already merged into VLLM_REF. This is their index, with
  README.md           an obsolete_when predicate per patch. Read patches/README.md.
tuning/               tuned fused_moe configs, and why there are none
probe/                probe_image_patches.sh, also shipped inside the image
upgrade.sh            triage every patch against a new upstream tag

Why the build gates exist

A half-patched image fails by being slow, not by erroring. This project has already published a throughput ratio it had not earned, because a benchmark client image was patched and the server image was not — nothing errored, and the number looked reasonable.

So build/verify.sh runs during the build and refuses to produce an image unless it passes, in three tiers that are not substitutes for one another:

  1. static markers — could this image possibly do X
  2. runtime gates — does the gate actually select it
  3. numeric acceptance — 11 + 4 + 8 tests against reference implementations

Two runtime gates encode findings specific to this hardware:

gate ACCEPTS  head_size 256    Qwen3-Next needs it; stock vLLM refuses it
gate DECLINES block_size 544   Qwen3-Next safety; silently wrong otherwise

Running a model

./run.sh /path/to/your/model        # or an HF repo id

Serves an OpenAI-compatible API on :8000. Needs only docker run — no compose plugin, no config file.

For Qwen3.8-27B specifically there is a complete recipe — image, checkpoint, draft model, device nodes, flags and the three greps that confirm it took the fast paths: docs/RECIPE-QWEN38-27B-INT8.md. It is worth 2.4x over a naive launch on a mixed-GPU host.

The rest of the CLI is there too, with the GPUs, mounts and ROCm environment already correct:

./run.sh bench latency --model /path/to/model
./run.sh complete --url http://localhost:8000/v1 --model /path/to/model --quick "hello"
./run.sh exec probe-image-patches      # or shell, or any command

Before downloading something large, ask what it will actually hit — a local path, an HF repo id or a pasted Hub URL, reading config.json only:

./run.sh exec model-fastpath https://huggingface.co/Qwen/Qwen3-8B --tp 2
./run.sh exec model-convert /path/to/model --to W4A16 --out /models/out

model-fastpath answers by calling vLLM's own gate predicates inside the image, so it cannot drift from the code that will actually run. model-convert turns its recommendation into a runnable llm-compressor recipe. Both in docs/RUNNING.md.

run.sh is a convenience, not a requirement — the image dispatches for itself and carries its own ROCm settings, so a plain docker run <image> /path/to/model behaves identically under Kubernetes or Slurm. It prints what it is at every start (patch markers, arch check), because a half-patched image fails by being slow rather than by erroring.

compose.yaml is the other way in, for a deployment you run repeatedly. Both are covered in docs/RUNNING.md, along with the settings that are not optional on ROCm and the errors worth recognising.

If this host holds cards of more than one architecture, read Which GPUs the container sees before anything else. vLLM reads the GPU architecture once at import, from amdsmi, for physical device 0 — ignoring HIP_VISIBLE_DEVICES and ROCR_VISIBLE_DEVICES alike. Give a mixed host's container all of /dev/dri and vLLM can decide the box is gfx12, disable every gfx9 path, and run decode 2.4x slower (48.0 → 19.9 tok/s on Qwen3.8-27B-W8A8) with no error anywhere. run.sh selects the gfx90a render nodes for you; compose.yaml needs GPU_RENDER_NODE. A host whose cards are all gfx90a is unaffected.

On an HPC site without Docker, see docs/APPTAINER.md — Frontier's MI250X is gfx90a, the same architecture this image targets.

Everything is built from git sources

This repo stores no binaries and no build artifacts. Every compiled thing in the image is produced during the build from a git clone at an immutable tag:

This table mirrors VERSIONS, which is authoritative. If the two disagree, VERSIONS is what gets built and this table is the stale one.

component source how it is built
_rocm_C (attention.cu) davetha/vllm @ v0.27.2rc0+mi210.5 pip wheel in build/Dockerfile
AITER Python + C++ ROCm/aiter @ v0.1.19 pip install --no-build-isolation .
gfx90a ASM code objects davetha/aiter-cdna2 @ v1.3 repatch_gfx942_to_gfx90a.py, at build time

.github/workflows/sources-only.yml enforces this rather than asserting it: it rejects any committed binary or file over 256 KiB, requires BASE_IMAGE to be digest-pinned, and checks that all three refs still resolve as tags.

One caveat on that last check, because it is narrower than it sounds: it validates the refs in VERSIONS, not the ones written above. Both cells were wrong for a while and CI stayed green, because a stale value can still be a real tag. The aiter-cdna2 one mattered — VERSIONS records that v1.0 shipped the ATTENTION carve-out only, so images built from it served int8 checkpoints through the generic Triton kernel at roughly a third of the decode rate, with no warning. Treat a mismatch here as a bug in this table, and check it by eye against VERSIONS rather than trusting the job to catch it.

The one exception, and it is upstream's

AITER's ASM kernels are the exception, and no build anywhere compiles them from source. ROCm/aiter @ v0.1.19 contains 2,863 prebuilt .co code objects and zero .s/.S files — AMD ships these kernels only as binaries, for gfx942, gfx950 and gfx1250. There is no gfx90a build to run and no assembly to run it on.

So aiter-cdna2 transforms the code objects rather than compiling them, re-assembling every instruction to prove portability and reporting the kernels that do not translate instead of skipping them quietly. Those binaries are fetched from upstream's git during add-aiter.sh; none are stored here.

If AMD ever publishes the sources, this step becomes a compile and the repatcher can go.

Building it

./build/build.sh local/vllm-mi210:dev    # no GPU needed to build
./build/add-aiter.sh local/vllm-mi210:dev  # optional, requires gfx90a cards

build.sh runs tier 0 inside the build, then tiers 1-2 against the cards, and records the digest in VERSIONS only if everything passes. The qualification record it writes looks like this:

{ "result": "pass", "max_tier": 2, "arch": "gfx90a:sramecc+:xnack-",
  "hardware_validated": true,
  "unvalidated_claims": [
    "gfx12/RDNA4 head_size rule: derived from constexpr, not measured",
    "gfx942 block_size behaviour: inferred from upstream test cases, not measured"
  ] }

hardware_validated is false when the build host has no GPU, and the claims this hardware cannot check are listed rather than asserted.

The core image needs no GPU to build, which matters because most people who want it do not have a spare MI210 to build it on. Every patch this project carries works without AITER.

AITER is a separate step by choice. import aiter probes the GPU through rocminfo and plain docker build exposes no /dev/kfd — but BuildKit CDI can pass the cards into a build, verified working here and written up in docs/GPU-IN-BUILD.md. It is kept out so the core image needs no CDI setup, no labs Dockerfile frontend and no GPU, which is the difference between "anyone can rebuild this" and "anyone with an MI210 can rebuild this".

Measured against llama.cpp on the same cards

Same box, same model architecture (NemotronHForCausalLM, 512 experts, top-22), matched prompt lengths. vLLM here is dsa7-aiterint8 with --tensor-parallel-size 2; llama.cpp is davetha/llama.cpp-mi210 with its CDNA2 prefill patches.

prefill llama.cpp vLLM TP=2 gap
~4k 2102 3635 1.73×
~16k 2708 3563 1.32×

vLLM's figures include HTTP and tokenisation, so they are marginally pessimistic. Most of what is left is tensor parallelism: sampling gpu_busy_percent during a 16k prefill, vLLM holds both cards at 99.4% / 99.6% simultaneously for 99% of samples, while llama.cpp's -sm layer manages 85.6% aggregate at 16k and 62% at 4k. That is worth roughly 1.16×, and llama.cpp cannot currently claim it on Mamba-2 hybrids — ggml's split-state model cannot describe ggml_ssm_scan's fused output, which is why every SSM/hybrid arch is on its -sm tensor exclusion list. Written up in that repo.

Known limits

  • Model load needs GPU_PINNED_MIN_XFER_SIZE=67108864 (the image bakes it in). Above HIP's ~1 MiB pin threshold, .to(device) page-locks the caller's buffer and hsa_amd_memory_lock_to_pool costs ~1 s while the DMA is 14 ms. GLM-4.5-Air: 22 s with it, hours without. See docs/LOAD-TIME.md. Unreported upstream.

  • Quantizing needs the separate convert image, built by build/add-convert.sh. The serving image has no quantizer, and adding one naively breaks it: a plain pip install llmcompressor downgrades transformers 5.14.0 → 5.10.1 and moves compressed-tensors off vLLM's ==0.17.0 pin. --no-deps plus three upstream patches avoids that; the build asserts the stack is unmoved and that vLLM still imports before committing anything. Verified end to end on 2026-08-05: a self-quantized W4A16 model served and answered correctly. See docs/RUNNING.md.

  • GLM-5.2 (DSA sparse attention) runs, at 0.81 tok/s. Five blockers were cleared to get there; the recipe, the patches and the honest limits are in docs/DSA-GFX90A.md. It needs ~420 GB of host RAM and a --memory cgroup cap (without one, an overshoot takes the whole box down). Treat it as proof the architecture runs on CDNA2, not as a deployment.

  • The int4 interleave path is compressed-tensors only. AWQ and GPTQ MoE checkpoints do not reach it; the port to moe_wna16.py is not done, so those miss a measured 1.45–4.8x on the int4 MoE kernel. model-fastpath tells you which side of this a given checkpoint falls on.

Licence

Patches derived from vLLM carry vLLM's Apache-2.0 headers; those derived from AITER carry AITER's MIT header. See individual file headers.

About

Pinned, gate-verified vLLM deployment stack for AMD MI210 (gfx90a/CDNA2)

Resources

Contributing

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages