Skip to content

Commit 529cbe9

Browse files
authored
Merge pull request #182 from SharpAI/feat/gpu-layers-moe-warning
Warn when --gpu-layers puts MoE layers on the CPU; mark #176 fixed in README
2 parents 0a2b688 + 60c0bf4 commit 529cbe9

2 files changed

Lines changed: 11 additions & 1 deletion

File tree

‎README.md‎

Lines changed: 3 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -123,7 +123,9 @@ Every needle check passed. `--mtp` with the bf16 assistant (`gemma-4-26B-A4B-it-
123123

124124
> ⚠️ **`--turbo-kv` loses exact long-range recall on every chip:** once a prompt passes the 2,048-token compression threshold, Qwen3.8-27B-4bit with `--turbo-kv` gets exact lookups wrong. Asked how many numbered lines a prompt has, it answers "1,000", "1,314" or "14" instead of 315 / 500 / 700. Reproduced on both M5 and M6; without `--turbo-kv` it answers correctly. The cause is a cache-eviction bookkeeping regression: after compression, attention only sees the recent hot window, and positions restart. A fix is in progress; tracked in [#175](https://github.com/SharpAI/SwiftLM/issues/175). Until then, avoid `--turbo-kv` when exact recall matters. Also note that `--turbo-kv` currently has no effect when `--ctx-size` is set (the attention layers use `RotatingKVCache`).
125125
>
126-
> ⚠️ **Known issues:** `--gpu-layers N` (CPU/GPU layer partitioning) hits a Metal GPU timeout on the first request, on both M5 and M6 (repro: `--model mlx-community/gemma-4-26b-a4b-it-4bit --gpu-layers 23`). Tracked in [#176](https://github.com/SharpAI/SwiftLM/issues/176); the fix is [SharpAI/mlx-swift#17](https://github.com/SharpAI/mlx-swift/pull/17). Even once it's fixed, CPU-resident MoE layers are very slow (~0.4 tok/s prefill), so on a 32 GB Mac try `--stream-experts` first (Qwen3.6-35B-A3B: 13.2 tok/s decode, 7.7 GB GPU). QAT-quantized Gemma 4 MTP assistants (`…-qat-assistant-4bit`) fail with `unhandledKeys pre_projection/post_projection`; use `gemma-4-26B-A4B-it-assistant-bf16`.
126+
> ⚠️ **Known issues:** QAT-quantized Gemma 4 MTP assistants (`…-qat-assistant-4bit`) fail with `unhandledKeys pre_projection/post_projection`; use `gemma-4-26B-A4B-it-assistant-bf16`.
127+
>
128+
> ℹ️ **`--gpu-layers N` with MoE models:** the Metal GPU timeout on the first request ([#176](https://github.com/SharpAI/SwiftLM/issues/176)) is fixed in #177. CPU-resident MoE layers still run on a single core and are very slow (Gemma 4 26B-A4B with `--gpu-layers 23`: ~0.4 tok/s prefill, ~3 s per decoded token on both M5 and M6), and SwiftLM now warns about this at startup. Treat `--gpu-layers` as a way to avoid running out of memory, not as a speed trade-off. On a 32 GB Mac try `--stream-experts` first (Qwen3.6-35B-A3B: 13.2 tok/s decode, 7.7 GB GPU).
127129
128130
Reproduce:
129131

‎Sources/SwiftLM/Server.swift‎

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -781,9 +781,11 @@ struct MLXServer: AsyncParsableCommand {
781781
}
782782

783783
var partitionPlan: PartitionPlan?
784+
var modelIsMoE = false
784785
if let modelDir = modelDirectory {
785786
let profile = mainModelProfile ?? ModelProfiler.profile(modelDirectory: modelDir, modelId: modelId)
786787
if let profile = profile {
788+
modelIsMoE = profile.isMoE
787789
let system = ModelProfiler.systemProfile()
788790
let contextSize = self.ctxSize ?? 4096
789791
let plan = ModelProfiler.plan(model: profile, system: system, contextSize: contextSize, draftWeightBytes: draftFootprintBytes)
@@ -1132,6 +1134,12 @@ struct MLXServer: AsyncParsableCommand {
11321134
let total = partitionPlan?.totalLayers ?? actual
11331135
let cpuCount = total - actual
11341136
print("[SwiftLM] 🔀 Layer split active: \(actual) GPU / \(cpuCount) CPU")
1137+
if modelIsMoE && cpuCount > 0 {
1138+
// CPU-resident MoE layers run the quantized expert matmuls on a
1139+
// single core (~0.4 tok/s prefill on Gemma 4 26B-A4B, see #176).
1140+
print("[SwiftLM] ⚠️ \(cpuCount) MoE layers will run on the CPU, which is very slow (expect well under 1 tok/s).")
1141+
print("[SwiftLM] For MoE models that don't fit in GPU memory, prefer --stream-experts over --gpu-layers.")
1142+
}
11351143
// Update the partition plan to reflect actual split
11361144
partitionPlan?.gpuLayers = actual
11371145
} else {

0 commit comments

Comments
 (0)