Skip to content

perf(sm90): size grouped W4A8 tiles per expert - #129

Open
LopezCastroRoberto wants to merge 2 commits into
vllm-project:mgoin/sm90-w4a8-generalfrom
LopezCastroRoberto:w4a8-row-major-expert-tiles
Open

LopezCastroRoberto wants to merge 2 commits into
vllm-project:mgoin/sm90-w4a8-generalfrom
LopezCastroRoberto:w4a8-row-major-expert-tiles

Conversation

@LopezCastroRoberto

@LopezCastroRoberto LopezCastroRoberto commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

In #124, grouped-contiguous W4A8 still takes its M tile from the generic seed heuristic, which sizes tiles from total rows as for a dense GEMM. This PR sizes the tile per expert instead, in two commits.

1. Row-major tile model (17d138d). In #124 the packed N16 schedule is applied only where the seed tile reaches M128. This causes two problems:

  • Seed-schedule holes. For total rows in (4416, 6144], (8512, 9472], (12160, 12416] and (16000, 16192] (≈ M 552-768, 1064-1184, 1520-1552 and 2000-2024 tokens at top-k 8), the seed tile is M96-M120, so these sizes run the generic warp-N32 schedule. That is 1.3-1.5× slower than the N16 schedule (M=768, 1152 and 1536 below).
  • M128 splits unevenly routed experts. With random routing, experts that land slightly over 128 or 256 rows need an extra, mostly empty tile. At M=512 that is 45 tiles instead of 32 (+41%), and at M=1024 it is 81 instead of 64 (+27%).

On H200, row-major now uses the same expected-work tile model as the M-major path (_w4a8_tile_m_for_expert_rows): the model tile when it is ≤ M160, M176 only while one tile covers the expert (≤ 176 rows), M128 beyond that, and no raster_group_m.

2. Small-expert seed tile, both layouts (3be4376). At ≤ 96 rows per expert, both scale layouts keep the seed schedule, but its M tile now comes from rows per expert (rows + √rows, rounded up to 8), with a (256, 128) block N/K when N is a multiple of 256. Before this change:

  • The row-major seed tile was about twice the rows of one expert; on down-proj, the (32, 256, 256) tile for 320-512 rows was 1.35× slower than the tiles on either side.
  • M-major used [Perf] Optimize Hopper MXFP4 × FP8_BLOCK grouped MoE prefill #95's packed N16 schedule with an M64+ tile, which was 1.03-1.55× slower than the per-expert seed tile at these sizes.

The 96 cutoff comes from a dense M=8-512 sweep (step 8): the crossover with the N16 tile model sits at 96-100 rows per expert for both layouts, both shapes and both routings.

Other SM90 devices keep the existing selection, as for the M-major tile model. No CUDA changes. Accuracy against the KernelTestRunner FP32 reference is unchanged (both layouts, M=8-512). test_grouped_w4a8_ranges_match_direct_selection is updated, because small-expert M-major ranges on H200 now use the seed schedule.

Results (H200, median of 3; M-major includes #126's split cutoff):

python benchmarks/bench_humming.py --shape_n 4096 --shape_k 6144 \
  --a_dtype float8e4m3 --b_dtype float4e2m1 --c_dtype bfloat16 --bs_dtype float8e8m0 \
  --input_scale_group_size 128 --weight_scale_group_size 32 \
  --gemm_type grouped_contiguous --num_experts 32 --top_k 8 [--use_m_major_input_scale]

Gate/up (N=4096, K=6144), default (unbalanced) routing

M row-major #124 row-major #124 + #126 + #129 speedup M-major #124 + #126 + #129 row-major vs M-major
64 0.1722 0.1540 1.12× 0.1537 1.00×
256 0.2312 0.2180 1.06× 0.2214 1.02×
512 0.3888 0.3104 1.25× 0.3163 1.02×
640 0.4739 0.4212 1.13× 0.4044 0.96×
768 0.6615 0.5146 1.29× 0.5642 1.10×
1024 0.6812 0.5977 1.14× 0.5961 1.00×
1152 0.9765 0.6524 1.50× 0.6464 0.99×
1536 1.2784 0.8778 1.46× 0.8798 1.00×
2048 1.1710 1.1709 1.00× 1.1192 0.96×
4096 2.1756 2.1796 1.00× 2.1108 0.97×
16384 8.2034 8.2021 1.00× 7.9913 0.97×

Down (N=6144, K=2048), default (unbalanced) routing

M row-major #124 row-major #124 + #126 + #129 speedup M-major #124 + #126 + #129 row-major vs M-major
64 0.1277 0.0938 1.36× 0.0938 1.00×
256 0.1267 0.1277 0.99× 0.1308 1.02×
512 0.2197 0.1764 1.25× 0.1777 1.01×
640 0.2678 0.2276 1.18× 0.2218 0.97×
768 0.3621 0.2811 1.29× 0.2808 1.00×
1024 0.3803 0.3326 1.14× 0.3369 1.01×
1152 0.5323 0.3693 1.44× 0.3627 0.98×
1536 0.6902 0.4824 1.43× 0.4951 1.03×
2048 0.6505 0.6529 1.00× 0.5938 0.91×
4096 1.2156 1.2186 1.00× 1.1721 0.96×
16384 4.5628 4.5630 1.00× 4.4139 0.97×

Further experiments at lower Ms

Small M (seed path, ≤ 96 rows per expert), Gate/up (N=4096, K=6144), default (unbalanced) routing

M row-major #124 row-major #124 + #126 + #129 speedup M-major #124 + #126 M-major #124 + #126 + #129 speedup row-major vs M-major (#129)
8 0.1445 0.1329 1.09× 0.2091 0.1348 1.55× 1.01×
16 0.1584 0.1412 1.12× 0.2104 0.1431 1.47× 1.01×
32 0.1669 0.1497 1.11× 0.2165 0.1494 1.45× 1.00×
48 0.1719 0.1492 1.15× 0.2185 0.1489 1.47× 1.00×
64 0.1722 0.1543 1.12× 0.2202 0.1535 1.43× 0.99×
96 0.1862 0.1574 1.18× 0.2199 0.1592 1.38× 1.01×
128 0.2010 0.1649 1.22× 0.2231 0.1663 1.34× 1.01×
192 0.1950 0.1877 1.04× 0.2250 0.1908 1.18× 1.02×
256 0.2312 0.2184 1.06× 0.2619 0.2214 1.18× 1.01×
264 0.2308 0.2107 1.10× 0.2606 0.2056 1.27× 0.98×
288 0.2471 0.2243 1.10× 0.2613 0.2166 1.21× 0.97×
320 0.2854 0.2336 1.22× 0.2654 0.2377 1.12× 1.02×
352 0.3400 0.2511 1.35× 0.2931 0.2545 1.15× 1.01×
384 0.2855 0.2883 0.99× 0.2943 0.2853 1.03× 0.99×

Small M (seed path, ≤ 96 rows per expert), Down (N=6144, K=2048), default (unbalanced) routing

M row-major #124 row-major #124 + #126 + #129 speedup M-major #124 + #126 M-major #124 + #126 + #129 speedup row-major vs M-major (#129)
8 0.0791 0.0820 0.96× 0.1072 0.0835 1.28× 1.02×
16 0.0879 0.0859 1.02× 0.1138 0.0875 1.30× 1.02×
32 0.0944 0.0905 1.04× 0.1145 0.0900 1.27× 0.99×
48 0.1275 0.0909 1.40× 0.1154 0.0904 1.28× 0.99×
64 0.1277 0.0938 1.36× 0.1162 0.0942 1.23× 1.00×
96 0.1002 0.0962 1.04× 0.1176 0.0968 1.21× 1.01×
128 0.1075 0.0988 1.09× 0.1187 0.1001 1.19× 1.01×
192 0.1118 0.1128 0.99× 0.1205 0.1139 1.06× 1.01×
256 0.1267 0.1279 0.99× 0.1353 0.1310 1.03× 1.02×
264 0.1257 0.1167 1.08× 0.1344 0.1164 1.15× 1.00×
288 0.1376 0.1237 1.11× 0.1354 0.1215 1.11× 0.98×
320 0.1562 0.1302 1.20× 0.1370 0.1291 1.06× 0.99×
352 0.1884 0.1400 1.35× 0.1598 0.1431 1.12× 1.02×
384 0.1648 0.1591 1.04× 0.1610 0.1594 1.01× 1.00×

Row-major grouped-contiguous took its M tile from the generic seed
heuristic, which sizes tiles from total rows. The packed N16 schedule
was only applied where that seed tile reached M128, so seed-schedule
holes remained (e.g. M=552-768 and 1064-1184 tokens), and elsewhere a
fixed M128 tile split unevenly routed experts into extra, mostly empty
tiles (+41% tiles at M=512, +27% at M=1024 with random routing).

On H200, size the row-major tile from rows per expert with the same
expected-work model as the M-major path: keep the seed schedule at
<= 64 rows per expert, use the model tile up to M160, M176 only while
one tile covers the expert, and M128 beyond. Other SM90 devices keep
the existing selection.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com>
Up to 96 rows per expert, grouped-contiguous W4A8 runs the generic seed
schedule (row-major) or the packed N16 schedule with an M64+ tile
(M-major). The seed sizes its M tile from total rows, about twice the
rows of one expert, and on down-proj the (32, 256, 256) seed tile for
320-512 rows was 1.35x slower than its neighbours. The M-major N16
schedule is slower still at these sizes.

On H200, both scale layouts now keep the seed schedule there with the M
tile taken from rows per expert (rows + sqrt(rows), rounded up to 8) and
a (256, 128) block N/K when N is a multiple of 256. Accuracy against the
FP32 reference is unchanged. H200, 32 experts, top-k 8, M <= 256:
row-major 1.04-1.22x on gate/up and up to 1.40x on down; M-major
1.19-1.55x on gate/up and 1.03-1.30x on down. Keeping the seed path up to
96 rows per expert (M=384) adds 1.02-1.28x at M=264-376 for both layouts;
above that the N16 tile model wins. M >= 392 is unchanged.

Update test_grouped_w4a8_ranges_match_direct_selection: on H200 the
small-expert M-major ranges use the seed schedule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com>
@LopezCastroRoberto LopezCastroRoberto changed the title perf(sm90): size row-major grouped W4A8 tiles per expert perf(sm90): size grouped W4A8 tiles per expert Oct 9, 2026
@LopezCastroRoberto
LopezCastroRoberto force-pushed the w4a8-row-major-expert-tiles branch from 1f99919 to 3be4376 Compare October 9, 2026 15:38

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant