-
Notifications
You must be signed in to change notification settings - Fork 77
Pull requests: meta-pytorch/MSLK
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Add mxfp4_block_size parameter to f4f4bf16_grouped_mm for MXFP4_16 support (#496)
cla signed
meta-exported
#496
opened Aug 21, 2026 by
ghjeong12
Loading…
Add FP8×INT4 rowwise GEMM for ROCm (f8i4bf16_rowwise)
cla signed
module: rocm
#495
opened Aug 21, 2026 by
liligwu
Contributor
Loading…
Autotune async copy for rowwise FP8 GEMM (#494)
cla signed
meta-exported
#494
opened Aug 21, 2026 by
warrendeng
Loading…
[ROCm] Flydsl decode backend
cla signed
module: rocm
#483
opened Aug 13, 2026 by
avbokovoy
Collaborator
Loading…
Back the FP8 rowwise grouped GEMM ops with FlyDSL on ROCm
cla signed
module: rocm
#471
opened Aug 6, 2026 by
aryaman-gupta
Contributor
Loading…
Fix flash_attn_bench crash on the flash_attn.cute (FA4) backend: pass window_size=(-1, -1) instead of None
#397
opened Jun 19, 2026 by
Maurits-de-Groot
Loading…
[CUDA] [PERFORMANCE] Increase speed of bf16bf16bf16_grouped_wgrad via indicating that ElementC is void / nullptr
cla signed
#329
opened Apr 19, 2026 by
benediktjohannes
Loading…
ProTip!
no:milestone will show everything without a milestone.