Deployment stack for running vLLM on 2x AMD Instinct MI210. It carries the patches this hardware needs — several upstream, several local — pinned, verified at build time, and documented with the measurements that justify them.
Status: builds and verifies. Built on 2x MI210 on 2026-08-04; all three verification tiers pass with
hardware_validated=true(23 numeric tests).DERIVED_IMAGEinVERSIONSis still empty because the image has not been pushed to a registry — runbuild/build.shto produce your own, or publish one and record its digest.
| patch | origin | what it does |
|---|---|---|
| multi-pass paged attention | vllm#39001 (Eugene Kuznetsov) | removes the 131,072-token context ceiling |
| gfx90a free-kernel guards | local | declines shapes that compute silently wrong attention on CDNA2 |
| single-pass 131k–262k | local | keeps the common long-context range off the multi-pass path |
| int4 interleave packing | vllm#43389 (amd-xavierwang), gate widened | 1.45–4.8x on the int4 MoE kernel, bit-identical output |
| wvSplitK stride guard | vllm#50618 (John Qin / Yanyuan Qin) | fixes an out-of-bounds read on strided activations |
| sharded_state TP guard | local | rejects checkpoints saved at a different TP size instead of half-loading them |
| benchmark_moe int8_w8a16 | local, grouping from vllm#31011 | the tuner could not run at all before this |
| W4A16 gfx90a tile table | local | gfx90a inherited the MI300 tiles, which assume 304 CUs against MI210's 104 |
| W4A16 magic-bias dequant | local | bit-trick dequant + scale hoist in the dense W4A16 GEMM inner loop; gfx90a has no bf16 VALU arithmetic |
| W4A16 narrow rung to M≤16 | local | 1.37–1.45x GB/s on M=9..16, by keeping the narrow tile active over that range instead of dropping to BLOCK_M=64 |
| W4A16 magic-bias gate on BLOCK_M | local | keys the gate on the tile rather than on M, where the cost crossover actually lives, so a future ladder change cannot silently mis-enable it |
| fp8 W8A16 Triton kernel | local | first ROCm entry in _POSSIBLE_WFP8A16_KERNELS (upstream ships that list empty), plus the CDNA2 dispatcher fix that was routing fp8 checkpoints into a torch._scaled_mm crash |
| NVFP4 W4A16 Triton kernel | local | packed e2m1 decode + per-16-group scales for gfx90a — not in the pinned tag: this one is post-VLLM_REF, so it is not in an image built from VERSIONS as it stands (see patches/registry.yaml) |
| fused GDN decode on CDNA | local | builds vLLM's CUDA-only fused gated-delta-net decode kernel for gfx90a, so GDN layers stop falling back to unfused Triton — not in the pinned tag either (post-VLLM_REF, see patches/registry.yaml) |
Every row except the last two is in the tag VERSIONS pins and therefore in
any image built from it. The NVFP4 and fused-GDN rows are not: both are marked
status: post-tag in patches/registry.yaml and land in the next pin.
Measurements behind each are in the commit messages on davetha/vllm, and the investigation history is in mi210-llm-stack.
The gfx90a work here builds on people who got there first:
- Eugene Kuznetsov (vllm#39001) — the multi-pass reduction. Developed on gfx942; this repo validated it on gfx90a and added the CDNA2 guards.
- amd-xavierwang (vllm#43389) — int4 interleave packing, gated to RDNA upstream. Measured here on CDNA2 and widened.
- rlrs (vllm#49888 and others) — upstreaming AITER attention for gfx90a on MI250X.
- The packaging pattern — digest-pinned base, overlays as source of truth, diffs as documentation — is taken from ryanzhou/deepseek-v4-flash-mi300x.
VERSIONS every pin; base image by DIGEST, never a tag
run.sh serve any model in one command -- docs/RUNNING.md
docs/RECIPE-QWEN38-27B-INT8.md
end-to-end: the fastest known config for this model on
one MI210, copy-pasteable
docs/SPEC-DECODE.md speculation: dflash N=12, the draft, and why depth is not
a fixed property of the model
docs/INT8-GFX90A.md INT8 W8A8: confirming the AITER kernel, and what a stock
W8A8 recipe leaves in BF16
gpu-nodes.sh picks the gfx90a cards on a mixed-GPU host; sourced by
run.sh and build/add-aiter.sh
compose.yaml the stack consumers run
build/ the ONE compiled layer + its gates
Dockerfile rebuilds _rocm_C for gfx90a. NEEDS NO GPU.
entrypoint.sh the image's own dispatch + arch preamble
build.sh build, verify on real cards, record the digest
add-aiter.sh OPTIONAL, needs the cards: AITER + gfx90a ASM kernels
verify.sh static markers -> runtime gates -> numeric tests
quant_nemotron_heretic_vast.sh W4A16 on a rented box; the ignore list is
NVIDIA's own, because quantising a Mamba-2 hybrid's SSM
projections yields fluent WRONG text rather than an error
watch_vast_quant.sh watchdog that kills the rental when it stops progressing
patches/ NO diffs here -- the patches are branches on the fork,
registry.yaml already merged into VLLM_REF. This is their index, with
README.md an obsolete_when predicate per patch. Read patches/README.md.
tuning/ tuned fused_moe configs, and why there are none
probe/ probe_image_patches.sh, also shipped inside the image
upgrade.sh triage every patch against a new upstream tag
A half-patched image fails by being slow, not by erroring. This project has already published a throughput ratio it had not earned, because a benchmark client image was patched and the server image was not — nothing errored, and the number looked reasonable.
So build/verify.sh runs during the build and refuses to produce an image
unless it passes, in three tiers that are not substitutes for one another:
- static markers — could this image possibly do X
- runtime gates — does the gate actually select it
- numeric acceptance — 11 + 4 + 8 tests against reference implementations
Two runtime gates encode findings specific to this hardware:
gate ACCEPTS head_size 256 Qwen3-Next needs it; stock vLLM refuses it
gate DECLINES block_size 544 Qwen3-Next safety; silently wrong otherwise
./run.sh /path/to/your/model # or an HF repo idServes an OpenAI-compatible API on :8000. Needs only docker run — no compose
plugin, no config file.
For Qwen3.8-27B specifically there is a complete recipe — image, checkpoint, draft model, device nodes, flags and the three greps that confirm it took the fast paths: docs/RECIPE-QWEN38-27B-INT8.md. It is worth 2.4x over a naive launch on a mixed-GPU host.
The rest of the CLI is there too, with the GPUs, mounts and ROCm environment already correct:
./run.sh bench latency --model /path/to/model
./run.sh complete --url http://localhost:8000/v1 --model /path/to/model --quick "hello"
./run.sh exec probe-image-patches # or shell, or any commandBefore downloading something large, ask what it will actually hit — a local
path, an HF repo id or a pasted Hub URL, reading config.json only:
./run.sh exec model-fastpath https://huggingface.co/Qwen/Qwen3-8B --tp 2
./run.sh exec model-convert /path/to/model --to W4A16 --out /models/outmodel-fastpath answers by calling vLLM's own gate predicates inside the image,
so it cannot drift from the code that will actually run. model-convert turns
its recommendation into a runnable llm-compressor recipe. Both in
docs/RUNNING.md.
run.sh is a convenience, not a requirement — the image dispatches for itself
and carries its own ROCm settings, so a plain docker run <image> /path/to/model
behaves identically under Kubernetes or Slurm. It prints what it is at every
start (patch markers, arch check), because a half-patched image fails by being
slow rather than by erroring.
compose.yaml is the other way in, for a deployment you run repeatedly. Both
are covered in docs/RUNNING.md, along with the settings that
are not optional on ROCm and the errors worth recognising.
If this host holds cards of more than one architecture, read
Which GPUs the container sees
before anything else. vLLM reads the GPU architecture once at import, from
amdsmi, for physical device 0 — ignoring HIP_VISIBLE_DEVICES and
ROCR_VISIBLE_DEVICES alike. Give a mixed host's container all of /dev/dri
and vLLM can decide the box is gfx12, disable every gfx9 path, and run decode
2.4x slower (48.0 → 19.9 tok/s on Qwen3.8-27B-W8A8) with no error anywhere.
run.sh selects the gfx90a render nodes for you; compose.yaml needs
GPU_RENDER_NODE. A host whose cards are all gfx90a is unaffected.
On an HPC site without Docker, see docs/APPTAINER.md — Frontier's MI250X is gfx90a, the same architecture this image targets.
This repo stores no binaries and no build artifacts. Every compiled thing in the image is produced during the build from a git clone at an immutable tag:
This table mirrors VERSIONS, which is authoritative. If the two disagree,
VERSIONS is what gets built and this table is the stale one.
| component | source | how it is built |
|---|---|---|
_rocm_C (attention.cu) |
davetha/vllm @ v0.27.2rc0+mi210.5 |
pip wheel in build/Dockerfile |
| AITER Python + C++ | ROCm/aiter @ v0.1.19 |
pip install --no-build-isolation . |
| gfx90a ASM code objects | davetha/aiter-cdna2 @ v1.3 |
repatch_gfx942_to_gfx90a.py, at build time |
.github/workflows/sources-only.yml enforces this rather than asserting it: it
rejects any committed binary or file over 256 KiB, requires BASE_IMAGE to be
digest-pinned, and checks that all three refs still resolve as tags.
One caveat on that last check, because it is narrower than it sounds: it
validates the refs in VERSIONS, not the ones written above. Both cells were
wrong for a while and CI stayed green, because a stale value can still be a
real tag. The aiter-cdna2 one mattered — VERSIONS records that v1.0 shipped
the ATTENTION carve-out only, so images built from it served int8 checkpoints
through the generic Triton kernel at roughly a third of the decode rate, with no
warning. Treat a mismatch here as a bug in this table, and check it by eye
against VERSIONS rather than trusting the job to catch it.
AITER's ASM kernels are the exception, and no build anywhere compiles them from
source. ROCm/aiter @ v0.1.19 contains 2,863 prebuilt .co code objects and
zero .s/.S files — AMD ships these kernels only as binaries, for gfx942,
gfx950 and gfx1250. There is no gfx90a build to run and no assembly to run it
on.
So aiter-cdna2 transforms the code objects rather than compiling them,
re-assembling every instruction to prove portability and reporting the kernels
that do not translate instead of skipping them quietly. Those binaries are
fetched from upstream's git during add-aiter.sh; none are stored here.
If AMD ever publishes the sources, this step becomes a compile and the repatcher can go.
./build/build.sh local/vllm-mi210:dev # no GPU needed to build
./build/add-aiter.sh local/vllm-mi210:dev # optional, requires gfx90a cardsbuild.sh runs tier 0 inside the build, then tiers 1-2 against the cards, and
records the digest in VERSIONS only if everything passes. The qualification
record it writes looks like this:
{ "result": "pass", "max_tier": 2, "arch": "gfx90a:sramecc+:xnack-",
"hardware_validated": true,
"unvalidated_claims": [
"gfx12/RDNA4 head_size rule: derived from constexpr, not measured",
"gfx942 block_size behaviour: inferred from upstream test cases, not measured"
] }hardware_validated is false when the build host has no GPU, and the claims
this hardware cannot check are listed rather than asserted.
The core image needs no GPU to build, which matters because most people who want it do not have a spare MI210 to build it on. Every patch this project carries works without AITER.
AITER is a separate step by choice. import aiter probes the GPU through
rocminfo and plain docker build exposes no /dev/kfd — but BuildKit CDI
can pass the cards into a build, verified working here and written up in
docs/GPU-IN-BUILD.md. It is kept out so the core image needs no CDI setup, no
labs Dockerfile frontend and no GPU, which is the difference between "anyone can
rebuild this" and "anyone with an MI210 can rebuild this".
Same box, same model architecture (NemotronHForCausalLM, 512 experts, top-22),
matched prompt lengths. vLLM here is dsa7-aiterint8 with --tensor-parallel-size 2;
llama.cpp is davetha/llama.cpp-mi210
with its CDNA2 prefill patches.
| prefill | llama.cpp | vLLM TP=2 | gap |
|---|---|---|---|
| ~4k | 2102 | 3635 | 1.73× |
| ~16k | 2708 | 3563 | 1.32× |
vLLM's figures include HTTP and tokenisation, so they are marginally pessimistic.
Most of what is left is tensor parallelism: sampling gpu_busy_percent
during a 16k prefill, vLLM holds both cards at 99.4% / 99.6% simultaneously for
99% of samples, while llama.cpp's -sm layer manages 85.6% aggregate at 16k and
62% at 4k. That is worth roughly 1.16×, and llama.cpp cannot currently claim it
on Mamba-2 hybrids — ggml's split-state model cannot describe ggml_ssm_scan's
fused output, which is why every SSM/hybrid arch is on its -sm tensor
exclusion list. Written up in that repo.
-
Model load needs
GPU_PINNED_MIN_XFER_SIZE=67108864(the image bakes it in). Above HIP's ~1 MiB pin threshold,.to(device)page-locks the caller's buffer andhsa_amd_memory_lock_to_poolcosts ~1 s while the DMA is 14 ms. GLM-4.5-Air: 22 s with it, hours without. Seedocs/LOAD-TIME.md. Unreported upstream. -
Quantizing needs the separate convert image, built by
build/add-convert.sh. The serving image has no quantizer, and adding one naively breaks it: a plainpip install llmcompressordowngrades transformers 5.14.0 → 5.10.1 and moves compressed-tensors off vLLM's==0.17.0pin.--no-depsplus three upstream patches avoids that; the build asserts the stack is unmoved and that vLLM still imports before committing anything. Verified end to end on 2026-08-05: a self-quantized W4A16 model served and answered correctly. Seedocs/RUNNING.md. -
GLM-5.2 (DSA sparse attention) runs, at 0.81 tok/s. Five blockers were cleared to get there; the recipe, the patches and the honest limits are in docs/DSA-GFX90A.md. It needs ~420 GB of host RAM and a
--memorycgroup cap (without one, an overshoot takes the whole box down). Treat it as proof the architecture runs on CDNA2, not as a deployment. -
The int4 interleave path is compressed-tensors only. AWQ and GPTQ MoE checkpoints do not reach it; the port to
moe_wna16.pyis not done, so those miss a measured 1.45–4.8x on the int4 MoE kernel.model-fastpathtells you which side of this a given checkpoint falls on.
Patches derived from vLLM carry vLLM's Apache-2.0 headers; those derived from AITER carry AITER's MIT header. See individual file headers.