Skip to content

Fix Qwen3.5/3.6 SSD expert streaming after the upstream sync (crashes + full weight load) - #69

Merged
solderzzc merged 3 commits into
mainfrom
fix/qwen35-moe-ssd-direct-reduction
Sep 25, 2026
Merged

solderzzc merged 3 commits into
mainfrom
fix/qwen35-moe-ssd-direct-reduction

Conversation

@solderzzc

@solderzzc solderzzc commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

Proposed changes

--stream-experts (SSD expert streaming) on Qwen3.5/3.6 MoE broke with the upstream sync (#62, which reached SwiftLM through SharpAI/SwiftLM#167 and ships in release b769). It crashes on the first request. Three fixes:

  1. Double unsort (prefill crash). Upstream Qwen direct reduction ml-explore/mlx-swift-lm#573 changed SwitchGLU.projectExperts to return the sorted output plus inverseOrder, and callers now unsort it. The fork-only SSD block still unsorted in place and returned inverseOrder, so callers unsorted a second time:
    Fatal error: [broadcast_shapes] Shapes (263,8,8,2048) and (263,8,1) cannot be broadcast.
    Fix: drop the in-place unsort, matching the other SSD early returns. Credit to the SwiftLM review session for spotting this.
  2. Eval inside a compiled trace (decode crash). Upstream Qwen3.5/3.6: run the decode step through compiled traces ml-explore/mlx-swift-lm#467/Generalise compiled decode segments to Qwen 3 Next ml-explore/mlx-swift-lm#569 run Qwen3.5 decode through compiled traces (MoE block, decoder layer, decodeStep segments). SSD streaming evals mid-layer to load experts:
    Fatal error: [eval] Attempting to eval an array during function transformations like compile or vmap is not allowed.
    Fix: skip those compiled paths while ExpertStreamingConfig.shared.isEnabled. GPU mode is unchanged.
  3. Every weight loaded eagerly (slow, swaps). The new concurrent loader (loadWeightArrays) evaluates every tensor, including the experts that streaming is meant to page in on demand. On a 32 GB Mac that kept ~20 GB resident (server log GPU_MEM=19.9GB, MEM_DEMAND=37GB), grew swap, and cut prefill 4×. Fix: keep the old per-file lazy load while streaming. GPU mode keeps the concurrent loader.

Qwen3Next gets the same guard on its compiled decodeStep segments, which it only uses for fp16 embeddings. It runs SwitchGLU there too, so fp16 Qwen3-Next with --stream-experts would hit the same eval-in-trace crash. This one comes from code reading by the SwiftLM review session and hasn't been reproduced: there's no fp16 Qwen3-Next checkpoint on the test machine.

Verification (Mac mini M6, 32 GB, unsloth/Qwen3.6-35B-A3B-UD-MLX-4bit, SwiftLM scripts/profiling/m6_bench.py, temperature 0, median of 3, needle check)

Bisect: SwiftLM 5ae50ec (mlx-swift-lm 50d35c1) works. b769 (7cc37a0) crashes, and so do b769 with the mlx-swift #17 compile change reverted and b769 with SwiftLM ml-explore#189 bypassed.

Build Mode Prompt tok Prefill tok/s Decode tok/s Peak GPU Swap Δ Needle
50d35c1 (before sync) SSD 548 315.5 13.32 5.64 GB 0 ok
main (7cc37a0) SSD 548 crash crash — — —
fixes 1+2 only SSD 548 77.2 11.20 5.35 GB +0.78 GB 3/3
this PR SSD 548 252.3 12.71 5.37 GB 0 3/3
this PR SSD 2,346 391.5 12.55 5.25 GB 0 3/3
50d35c1 (README) SSD 2,356 401.6 13.12 6.24 GB 0 ok
this PR GPU 548 / 2,346 714 / 970 47.16 / 45.76 19.9 / 20.1 GB 0 ok

On a fixed 611-token prompt at temperature 0, SSD output matches GPU output for the opening words, then diverges but stays coherent. SSD and GPU also diverge on 50d35c1 (different kernels), so bit-identity isn't the bar.

A small gap remains at short prompts (prefill −20% at 548 tokens, decode −4%). At 2.3K tokens it's within 3%. I haven't tracked the rest down.

No unit test is added: the SSD path needs real safetensors on disk plus preadInto, which the unit suite doesn't have. The check above is end to end.

Checklist

  • I have read the CONTRIBUTING document
  • I have run pre-commit run --all-files to format my code / installed pre-commit prior to committing changes
  • I have added tests that prove my fix is effective or that my feature works
  • I have updated the necessary documentation (if needed)

AI usage

  • I have read this PR description in full and approve it as my own, and it
    accurately describes the code changes.
  • AI usage disclosure: Written by Claude Code (Claude Opus 5.5) in the M6 benchmarking session, with the repo owner's approval to open this PR. Claude did the bisecting and wrote the code and this description. The double-unsort diagnosis came from the SwiftLM review session (also Claude). Benchmarks were run by Claude on the owner's Mac mini M6. pre-commit/swift-format is not installed on this machine, so formatting wasn't run.

🤖 Generated with Claude Code

simba and others added 2 commits September 25, 2026 13:00
Two regressions from the upstream sync (#62) broke --stream-experts on
Qwen3.5/3.6 MoE:

- The fork-only SSD block in SwitchGLU.projectExperts unsorted its output
  and still returned inverseOrder, so callers unsorted a second time.
  Prefill crashed with broadcast_shapes (N,8,8,D) vs (N,8,1).
- Decode now runs through compiled traces, but SSD streaming evals mid-layer
  to load experts, which a trace cannot contain. Skip the compiled MoE block,
  decoder layer and decode segments while streaming.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The concurrent weight loader from the upstream sync evaluates every tensor,
including the experts SSD streaming is meant to page in on demand. On a 32 GB
Mac that held ~20 GB resident, grew swap, and cut Qwen3.6-35B-A3B SSD prefill
from 315 to 77 tok/s. Keep the per-file lazy load while streaming.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@solderzzc solderzzc changed the title Fix Qwen3.5/3.6 SSD expert streaming crashes after the upstream sync Fix Qwen3.5/3.6 SSD expert streaming after the upstream sync (crashes + full weight load) Sep 25, 2026
Same failure as the Qwen3.5 decode path: fp16 Qwen3-Next runs SwitchGLU
inside compiled decode segments, and the SSD streaming path evals there.
Found by code reading (SwiftLM review session); not reproduced, since no fp16
Qwen3-Next checkpoint is on the test machine.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant