Repository navigation
perf(sm90): size grouped W4A8 tiles per expert - #129
Open
LopezCastroRoberto wants to merge 2 commits into
Open
LopezCastroRoberto wants to merge 2 commits into
LopezCastroRoberto wants to merge 2 commits into
Conversation
Row-major grouped-contiguous took its M tile from the generic seed heuristic, which sizes tiles from total rows. The packed N16 schedule was only applied where that seed tile reached M128, so seed-schedule holes remained (e.g. M=552-768 and 1064-1184 tokens), and elsewhere a fixed M128 tile split unevenly routed experts into extra, mostly empty tiles (+41% tiles at M=512, +27% at M=1024 with random routing). On H200, size the row-major tile from rows per expert with the same expected-work model as the M-major path: keep the seed schedule at <= 64 rows per expert, use the model tile up to M160, M176 only while one tile covers the expert, and M128 beyond. Other SM90 devices keep the existing selection. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: LopezCastroRoberto <rocastro@redhat.com>
LopezCastroRoberto
force-pushed
the
w4a8-row-major-expert-tiles
branch
from
October 9, 2026 12:07
ce08136 to
17d138d
Compare
Up to 96 rows per expert, grouped-contiguous W4A8 runs the generic seed schedule (row-major) or the packed N16 schedule with an M64+ tile (M-major). The seed sizes its M tile from total rows, about twice the rows of one expert, and on down-proj the (32, 256, 256) seed tile for 320-512 rows was 1.35x slower than its neighbours. The M-major N16 schedule is slower still at these sizes. On H200, both scale layouts now keep the seed schedule there with the M tile taken from rows per expert (rows + sqrt(rows), rounded up to 8) and a (256, 128) block N/K when N is a multiple of 256. Accuracy against the FP32 reference is unchanged. H200, 32 experts, top-k 8, M <= 256: row-major 1.04-1.22x on gate/up and up to 1.40x on down; M-major 1.19-1.55x on gate/up and 1.03-1.30x on down. Keeping the seed path up to 96 rows per expert (M=384) adds 1.02-1.28x at M=264-376 for both layouts; above that the N16 tile model wins. M >= 392 is unchanged. Update test_grouped_w4a8_ranges_match_direct_selection: on H200 the small-expert M-major ranges use the seed schedule. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: LopezCastroRoberto <rocastro@redhat.com>
LopezCastroRoberto
force-pushed
the
w4a8-row-major-expert-tiles
branch
from
October 9, 2026 15:38
1f99919 to
3be4376
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
In #124, grouped-contiguous W4A8 still takes its M tile from the generic seed heuristic, which sizes tiles from total rows as for a dense GEMM. This PR sizes the tile per expert instead, in two commits.
1. Row-major tile model (
17d138d). In #124 the packed N16 schedule is applied only where the seed tile reaches M128. This causes two problems:On H200, row-major now uses the same expected-work tile model as the M-major path (
_w4a8_tile_m_for_expert_rows): the model tile when it is ≤ M160, M176 only while one tile covers the expert (≤ 176 rows), M128 beyond that, and noraster_group_m.2. Small-expert seed tile, both layouts (
3be4376). At ≤ 96 rows per expert, both scale layouts keep the seed schedule, but its M tile now comes from rows per expert (rows + √rows, rounded up to 8), with a (256, 128) block N/K when N is a multiple of 256. Before this change:The 96 cutoff comes from a dense M=8-512 sweep (step 8): the crossover with the N16 tile model sits at 96-100 rows per expert for both layouts, both shapes and both routings.
Other SM90 devices keep the existing selection, as for the M-major tile model. No CUDA changes. Accuracy against the
KernelTestRunnerFP32 reference is unchanged (both layouts, M=8-512).test_grouped_w4a8_ranges_match_direct_selectionis updated, because small-expert M-major ranges on H200 now use the seed schedule.Results (H200, median of 3; M-major includes #126's split cutoff):
Gate/up (N=4096, K=6144), default (unbalanced) routing
Down (N=6144, K=2048), default (unbalanced) routing
Further experiments at lower Ms
Small M (seed path, ≤ 96 rows per expert), Gate/up (N=4096, K=6144), default (unbalanced) routing
Small M (seed path, ≤ 96 rows per expert), Down (N=6144, K=2048), default (unbalanced) routing