Skip to content

Commit 5eecf22

Browse files
simbaclaude
authored andcommitted
docs: Mac mini M6 (base, 32 GB) benchmarks, recordings and demo client
- README: new "Mac mini M6 (base, 32 GB)" section covering four models (Gemma-4-26B-A4B 4-bit and 8-bit, Qwen3.6-35B-A3B, Qwen3.8-27B), the fixes 32 GB exposed, known issues, and reproduce commands. - docs/profiling/m6/: per-model tables (.md) and raw per-run results (.jsonl) from scripts/profiling/m6_bench.py, plus asciinema recordings (.cast) and GIFs under media/. Server logs are not included; they hold the full test prompts. - scripts/demo/stream_client.py: the streaming client used in the recordings. It prints live prefill progress with the server's memory and swap, then a stats line. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
1 parent 396c960 commit 5eecf22

18 files changed

Lines changed: 719 additions & 0 deletions

‎README.md‎

Lines changed: 82 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -51,6 +51,88 @@ Then start the server (models download automatically if not cached):
5151

5252
*(Add `--stream-experts` when running oversized MoE models to bypass macOS virtual memory swapping and stream expert layers directly from NVMe SSD.)*
5353

54+
## 📊 Performance: Mac mini M6 (base, 32 GB)
55+
56+
The first SwiftLM numbers from a **32 GB** Mac. Every other table in this README comes from 64 GB hardware. Running the same context lengths on this M6 exposed memory bugs that 64 GB machines had been hiding, and they are fixed in this release (see [What 32 GB exposed](#what-32-gb-exposed)).
57+
58+
![Gemma-4-26B-A4B streaming at 55 tok/s on a base Mac mini M6 32 GB](docs/profiling/m6/media/m6_gemma4_26b_a4b_stream.gif)
59+
60+
> *Hardware:* Mac mini (Mac18,5), Apple M6, 12-core GPU, 32 GB unified memory (170 GB/s), macOS 27.0. Metal working set 26.8 GB.
61+
> *Method:* [`scripts/profiling/m6_bench.py`](scripts/profiling/m6_bench.py). One warm-up, then the median of 3 runs (1 run at 32K and above), temperature 0. Every prompt starts with a unique nonce, so the prompt cache can't hit, and hides a code word that the answer must return. A memory guard aborts any case whose swap grows by more than 2 GB. Raw results: [`docs/profiling/m6/`](docs/profiling/m6/).
62+
63+
### What runs well on a 32 GB M6
64+
65+
| Model (4-bit unless noted) | Weights | Mode | Decode, short prompt | Longest prompt that passed | Peak GPU |
66+
|---|---|---|---|---|---|
67+
| **`gemma-4-26b-a4b-it-4bit`** (MoE, ~4B active) | 15.3 GB | GPU | **53.3 tok/s** | 80.7K tokens | 20.2 GB |
68+
| **`Qwen3.6-35B-A3B-UD-MLX-4bit`** (MoE, ~3B active) | 21.6 GB | GPU | **46.7 tok/s** | 40.8K tokens | 22.4 GB |
69+
| `Qwen3.6-35B-A3B-UD-MLX-4bit` | 21.6 GB | `--stream-experts` | 13.2 tok/s | 40.8K tokens | 7.7 GB |
70+
| `Qwen3.8-27B-4bit` (dense) | 11.3 GB | GPU | 9.1 tok/s | 40.8K tokens | 19.6 GB |
71+
| `gemma-4-26b-a4b-it-8bit` | ~26 GB | GPU | swaps (+3.1 GB on the first prompt) | — | — |
72+
| `gemma-4-26b-a4b-it-8bit` | ~26 GB | `--stream-experts` | 8.8 tok/s | 9.5K tokens (32K swapped) | 7.6 GB |
73+
74+
- **MoE models are the sweet spot at 32 GB.** Only the active experts are read for each token, so they decode 5–6× faster than a dense 27B. A 4-bit MoE with up to about 22 GB of weights runs entirely on the GPU.
75+
- **Qwen3.6-35B-A3B on a base M6 reaches 76%** of the M1 Ultra 64 GB decode speed below (46.7 vs 61.7 tok/s).
76+
- **Dense 27B decode is bandwidth-bound.** 9.1 tok/s × 11.3 GB is about 100 GB/s, roughly 60% of the M6's rated 170 GB/s.
77+
- **An 8-bit 26 GB model needs SSD streaming** and tops out at about 10K tokens of context.
78+
79+
### Gemma-4-26B-A4B 4-bit — by prompt length
80+
81+
| Prompt tokens | Prefill | TTFT | Decode | Peak GPU · swap growth |
82+
|---|---|---|---|---|
83+
| 534 | 859 tok/s | 0.7 s | **53.3 tok/s** | 15.2 GB · 0 |
84+
| 2,285 | **998 tok/s** | 2.4 s | 50.1 tok/s | 15.8 GB · 0 |
85+
| 9,544 | 914 tok/s | 10.7 s | 44.9 tok/s | 16.5 GB · 0 |
86+
| 39,773 | 750 tok/s | 53.9 s | 31.4 tok/s | 19.1 GB · 0 |
87+
| 80,711 | 625 tok/s | 130.6 s | 21.8 tok/s | 20.2 GB · 0 |
88+
89+
### Qwen3.6-35B-A3B 4-bit — GPU vs SSD streaming
90+
91+
| Prompt tokens | GPU prefill / decode (tok/s) | GPU peak | `--stream-experts` prefill / decode (tok/s) | SSD peak |
92+
|---|---|---|---|---|
93+
| ~550 | 808 / 46.7 | 20.4 GB | 321 / 13.2 | 6.0 GB |
94+
| ~2.3K | 969 / 45.5 | 21.0 GB | 402 / 13.1 | 6.2 GB |
95+
| ~9.8K | 849 / 43.4 | 21.1 GB | 403 / 12.9 | 6.6 GB |
96+
| 40.8K | 635 / 35.1 | 22.4 GB | 340 / 11.9 | 7.7 GB |
97+
98+
### Qwen3.8-27B-4bit (dense) — Vanilla vs TurboKV
99+
100+
| Prompt tokens | Vanilla prefill / decode (tok/s) | TurboKV prefill / decode (tok/s) | Peak GPU · swap growth |
101+
|---|---|---|---|
102+
| ~550 | 105 / 9.1 | 102 / 9.3 | 16.5 GB · 0 |
103+
| ~2.3K | 108 / 8.8 | 108 / 9.2 | 16.2 GB · 0 |
104+
| ~9.8K | 96 / 8.6 | 98 / 8.3 | 16.7 GB · 0 |
105+
| ~40.8K | 88 / 7.9 | 85 / 7.3 | 19.6 GB · +0.7 GB (TurboKV +1.5 GB) |
106+
107+
TurboKV doesn't help this model. Only 16 of its 64 layers use full attention (the other 48 are GatedDeltaNet), so the KV cache is already small.
108+
109+
### What 32 GB exposed
110+
111+
| 8.5K-token prompt, Qwen3.8-27B-4bit | Before | After |
112+
|---|---|---|
113+
| Prefill | 33.4 tok/s | **106.7 tok/s** (3.2×) |
114+
| Peak process memory | 38 GB | **19 GB** |
115+
| Swap growth | +15 GB | **0** |
116+
117+
1. **The MLX buffer cache was unbounded on full-GPU loads.** It could grow to the whole 26.8 GB working set. It is now sized from the RAM left after weights and KV.
118+
2. **The KV-cache estimate counted every layer as full attention.** Gemma 4 (25 of 30 layers use a 1,024-token sliding window) was overestimated 10×, and Qwen3.5/3.8 (48 of 64 layers are linear attention) 4×. On 32 GB that pushed Gemma into CPU/GPU layer partitioning, which crashed with a Metal GPU timeout.
119+
3. **An auto-detected VLM that failed to load exited the server.** `Qwen3.6-35B-A3B-UD-MLX-4bit` ships a `preprocessor_config.json` without `image_mean`. SwiftLM now falls back to text-only unless you pass `--vision`.
120+
4. **Vision-capable models skipped chunked prefill.** On the VLM path, a text-only prompt ran through the model in a single pass. This one is fixed on mlx-swift-lm `main` and lands here with the next mlx-swift-lm bump. Until then, long prompts on Qwen3.5/3.8 and Gemma 4 use far more memory than the tables above.
121+
122+
> ⚠️ **Known issues on 32 GB:** `--turbo-kv` with Gemma 4 crashes in Metal when a prompt-cache save runs during generation, and dropped the digits of the code word in 2 of 3 runs at 2K. The Gemma 4 MTP assistant (`gemma-4-26B-A4B-it-qat-assistant-4bit`) doesn't load on the current mlx-swift-lm pin. Both are expected to clear with the mlx-swift-lm bump.
123+
124+
Reproduce:
125+
126+
```bash
127+
./build.sh
128+
.build/release/SwiftLM --model mlx-community/gemma-4-26b-a4b-it-4bit --port 5431 &
129+
python3 scripts/demo/stream_client.py short
130+
python3 scripts/profiling/m6_bench.py --model mlx-community/gemma-4-26b-a4b-it-4bit \
131+
--config "Vanilla=" --contexts 512,2048,8192,32768,65536 --out docs/profiling/m6/gemma4_26b_a4b_4bit
132+
```
133+
134+
More recordings: [Gemma, 41K-token prompt (4×)](docs/profiling/m6/media/m6_gemma4_26b_a4b_41k_prompt_4x.gif) · [Qwen3.8-27B streaming](docs/profiling/m6/media/m6_qwen38_27b_stream.gif) · [Qwen3.8-27B, 8.6K-token prompt (4×)](docs/profiling/m6/media/m6_qwen38_27b_8k_prompt_4x.gif)
135+
54136
## 📊 Performance: MTP Speculative Decoding — Gemma 4-26B (MacBook Pro M5 Pro 64 GB)
55137

56138
Benchmarked with `gemma-4-26b-a4b-it-4bit` running three configurations across 512 / 40K / 100K token contexts.
Lines changed: 20 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,20 @@
1+
{"model": "mlx-community/gemma-4-26b-a4b-it-4bit", "config": "Vanilla", "context": 512, "run": -1, "warmup": true, "peak_gpu_gb": 14.59, "swap_delta_gb": 0.0, "min_free_pct": 42, "status": "OK", "prompt_tokens": 521, "prefill_tps": 471.2, "ttft_s": 1.14, "decode_tps": 53.37, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
2+
{"model": "mlx-community/gemma-4-26b-a4b-it-4bit", "config": "Vanilla", "context": 512, "run": 0, "warmup": false, "peak_gpu_gb": 14.92, "swap_delta_gb": 0.0, "min_free_pct": 44, "status": "OK", "prompt_tokens": 534, "prefill_tps": 723.9, "ttft_s": 0.78, "decode_tps": 53.46, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
3+
{"model": "mlx-community/gemma-4-26b-a4b-it-4bit", "config": "Vanilla", "context": 512, "run": 1, "warmup": false, "peak_gpu_gb": 14.96, "swap_delta_gb": 0.0, "min_free_pct": 44, "status": "OK", "prompt_tokens": 543, "prefill_tps": 858.5, "ttft_s": 0.68, "decode_tps": 53.27, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
4+
{"model": "mlx-community/gemma-4-26b-a4b-it-4bit", "config": "Vanilla", "context": 512, "run": 2, "warmup": false, "peak_gpu_gb": 15.15, "swap_delta_gb": 0.0, "min_free_pct": 43, "status": "OK", "prompt_tokens": 528, "prefill_tps": 910.4, "ttft_s": 0.62, "decode_tps": 52.56, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
5+
{"model": "mlx-community/gemma-4-26b-a4b-it-4bit", "config": "Vanilla", "context": 2048, "run": 0, "warmup": false, "peak_gpu_gb": 15.78, "swap_delta_gb": 0.0, "min_free_pct": 40, "status": "OK", "prompt_tokens": 2274, "prefill_tps": 969.8, "ttft_s": 2.42, "decode_tps": 50.58, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
6+
{"model": "mlx-community/gemma-4-26b-a4b-it-4bit", "config": "Vanilla", "context": 2048, "run": 1, "warmup": false, "peak_gpu_gb": 15.52, "swap_delta_gb": 0.0, "min_free_pct": 41, "status": "OK", "prompt_tokens": 2289, "prefill_tps": 1006.9, "ttft_s": 2.36, "decode_tps": 49.84, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
7+
{"model": "mlx-community/gemma-4-26b-a4b-it-4bit", "config": "Vanilla", "context": 2048, "run": 2, "warmup": false, "peak_gpu_gb": 15.61, "swap_delta_gb": 0.0, "min_free_pct": 41, "status": "OK", "prompt_tokens": 2285, "prefill_tps": 997.6, "ttft_s": 2.38, "decode_tps": 50.08, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
8+
{"model": "mlx-community/gemma-4-26b-a4b-it-4bit", "config": "Vanilla", "context": 8192, "run": 0, "warmup": false, "peak_gpu_gb": 16.2, "swap_delta_gb": 0.0, "min_free_pct": 37, "status": "OK", "prompt_tokens": 9529, "prefill_tps": 941.2, "ttft_s": 10.35, "decode_tps": 44.9, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
9+
{"model": "mlx-community/gemma-4-26b-a4b-it-4bit", "config": "Vanilla", "context": 8192, "run": 1, "warmup": false, "peak_gpu_gb": 16.54, "swap_delta_gb": 0.0, "min_free_pct": 38, "status": "OK", "prompt_tokens": 9544, "prefill_tps": 913.7, "ttft_s": 10.67, "decode_tps": 44.99, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
10+
{"model": "mlx-community/gemma-4-26b-a4b-it-4bit", "config": "Vanilla", "context": 8192, "run": 2, "warmup": false, "peak_gpu_gb": 16.3, "swap_delta_gb": 0.0, "min_free_pct": 39, "status": "OK", "prompt_tokens": 9558, "prefill_tps": 904.5, "ttft_s": 10.81, "decode_tps": 44.93, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
11+
{"model": "mlx-community/gemma-4-26b-a4b-it-4bit", "config": "Vanilla", "context": 32768, "run": 0, "warmup": false, "peak_gpu_gb": 19.13, "swap_delta_gb": 0.0, "min_free_pct": 33, "status": "OK", "prompt_tokens": 39773, "prefill_tps": 749.6, "ttft_s": 53.9, "decode_tps": 31.44, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
12+
{"model": "mlx-community/gemma-4-26b-a4b-it-4bit", "config": "Vanilla", "context": 65536, "run": 0, "warmup": false, "peak_gpu_gb": 20.19, "swap_delta_gb": 0.0, "min_free_pct": 18, "status": "OK", "prompt_tokens": 80711, "prefill_tps": 625.1, "ttft_s": 130.63, "decode_tps": 21.83, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
13+
{"model": "mlx-community/gemma-4-26b-a4b-it-4bit", "config": "TurboKV", "context": 512, "run": -1, "warmup": true, "peak_gpu_gb": 14.59, "swap_delta_gb": 0.0, "min_free_pct": 44, "status": "OK", "prompt_tokens": 527, "prefill_tps": 603.1, "ttft_s": 0.91, "decode_tps": 52.52, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
14+
{"model": "mlx-community/gemma-4-26b-a4b-it-4bit", "config": "TurboKV", "context": 512, "run": 0, "warmup": false, "peak_gpu_gb": 15.01, "swap_delta_gb": 0.0, "min_free_pct": 44, "status": "OK", "prompt_tokens": 536, "prefill_tps": 882.9, "ttft_s": 0.65, "decode_tps": 53.46, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
15+
{"model": "mlx-community/gemma-4-26b-a4b-it-4bit", "config": "TurboKV", "context": 512, "run": 1, "warmup": false, "peak_gpu_gb": 14.91, "swap_delta_gb": 0.0, "min_free_pct": 44, "status": "OK", "prompt_tokens": 532, "prefill_tps": 840.1, "ttft_s": 0.67, "decode_tps": 53.06, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
16+
{"model": "mlx-community/gemma-4-26b-a4b-it-4bit", "config": "TurboKV", "context": 512, "run": 2, "warmup": false, "peak_gpu_gb": 15.02, "swap_delta_gb": 0.0, "min_free_pct": 44, "status": "OK", "prompt_tokens": 532, "prefill_tps": 892.6, "ttft_s": 0.64, "decode_tps": 53.13, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
17+
{"model": "mlx-community/gemma-4-26b-a4b-it-4bit", "config": "TurboKV", "context": 2048, "run": 0, "warmup": false, "peak_gpu_gb": 15.77, "swap_delta_gb": 0.0, "min_free_pct": 42, "status": "OK", "prompt_tokens": 2280, "prefill_tps": 924.0, "ttft_s": 2.65, "decode_tps": 52.96, "gen_tokens": 128, "needle_ok": false, "degenerate": false}
18+
{"model": "mlx-community/gemma-4-26b-a4b-it-4bit", "config": "TurboKV", "context": 2048, "run": 1, "warmup": false, "peak_gpu_gb": 15.63, "swap_delta_gb": 0.0, "min_free_pct": 41, "status": "OK", "prompt_tokens": 2311, "prefill_tps": 886.8, "ttft_s": 2.74, "decode_tps": 53.22, "gen_tokens": 128, "needle_ok": false, "degenerate": false}
19+
{"model": "mlx-community/gemma-4-26b-a4b-it-4bit", "config": "TurboKV", "context": 2048, "run": 2, "warmup": false, "peak_gpu_gb": 15.44, "swap_delta_gb": 0.0, "min_free_pct": 41, "status": "OK", "prompt_tokens": 2299, "prefill_tps": 918.9, "ttft_s": 2.65, "decode_tps": 53.57, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
20+
{"model": "mlx-community/gemma-4-26b-a4b-it-4bit", "config": "TurboKV", "context": 8192, "run": 0, "warmup": false, "peak_gpu_gb": 14.86, "swap_delta_gb": 0.0, "min_free_pct": 42, "status": "OK", "prompt_tokens": 9572, "prefill_tps": 964.6, "ttft_s": null, "decode_tps": null, "gen_tokens": 1, "needle_ok": false, "degenerate": false}
Lines changed: 17 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,17 @@
1+
### `mlx-community/gemma-4-26b-a4b-it-4bit`
2+
3+
Apple M6 · 32 GB · runs=3 (long=1) · warmup=1 · gen=128 · temperature 0 · medians
4+
5+
| Config | Context (prompt tok) | Prefill tok/s | TTFT s | Decode tok/s | Peak GPU GB | Swap Δ GB | Min free % | Checks |
6+
|---|---|---|---|---|---|---|---|---|
7+
| Vanilla | 512 (534) | 858.5 | 0.68 | 53.27 | 15.15 | 0.0 | 43 | ok |
8+
| Vanilla | 2048 (2285) | 997.6 | 2.38 | 50.08 | 15.78 | 0.0 | 40 | ok |
9+
| Vanilla | 8192 (9544) | 913.7 | 10.67 | 44.93 | 16.54 | 0.0 | 37 | ok |
10+
| Vanilla | 32768 (39773) | 749.6 | 53.9 | 31.44 | 19.13 | 0.0 | 33 | ok |
11+
| Vanilla | 65536 (80711) | 625.1 | 130.63 | 21.83 | 20.19 | 0.0 | 18 | ok |
12+
| TurboKV | 512 (532) | 882.9 | 0.65 | 53.13 | 15.02 | 0.0 | 44 | ok |
13+
| TurboKV | 2048 (2299) | 918.9 | 2.65 | 53.22 | 15.77 | 0.0 | 41 | needle 1/3, degen 0 |
14+
| TurboKV | 8192 | — | — | — | 0.35 | 0.0 | 87 | **REQUEST_FAIL** <urlopen error [Errno 61] Connection refused> |
15+
| TurboKV | 32768 | — | — | — | — | — | — | **SKIPPED_AFTER_ABORT** |
16+
| TurboKV | 65536 | — | — | — | — | — | — | **SKIPPED_AFTER_ABORT** |
17+
| MTP | None | — | — | — | — | — | — | **START_FAIL** |
Lines changed: 10 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,10 @@
1+
{"model": "mlx-community/gemma-4-26b-a4b-it-8bit", "config": "SSD", "context": 512, "run": -1, "warmup": true, "peak_gpu_gb": 6.21, "swap_delta_gb": 0.0, "min_free_pct": 73, "status": "OK", "prompt_tokens": 534, "prefill_tps": 105.0, "ttft_s": 5.12, "decode_tps": 9.24, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
2+
{"model": "mlx-community/gemma-4-26b-a4b-it-8bit", "config": "SSD", "context": 512, "run": 0, "warmup": false, "peak_gpu_gb": 6.94, "swap_delta_gb": 0.0, "min_free_pct": 73, "status": "OK", "prompt_tokens": 543, "prefill_tps": 219.0, "ttft_s": 2.52, "decode_tps": 9.18, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
3+
{"model": "mlx-community/gemma-4-26b-a4b-it-8bit", "config": "SSD", "context": 512, "run": 1, "warmup": false, "peak_gpu_gb": 6.89, "swap_delta_gb": 0.0, "min_free_pct": 73, "status": "OK", "prompt_tokens": 528, "prefill_tps": 232.6, "ttft_s": 2.32, "decode_tps": 8.78, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
4+
{"model": "mlx-community/gemma-4-26b-a4b-it-8bit", "config": "SSD", "context": 512, "run": 2, "warmup": false, "peak_gpu_gb": 6.95, "swap_delta_gb": 0.0, "min_free_pct": 74, "status": "OK", "prompt_tokens": 526, "prefill_tps": 219.6, "ttft_s": 2.44, "decode_tps": 8.49, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
5+
{"model": "mlx-community/gemma-4-26b-a4b-it-8bit", "config": "SSD", "context": 2048, "run": 0, "warmup": false, "peak_gpu_gb": 7.26, "swap_delta_gb": 0.0, "min_free_pct": 73, "status": "OK", "prompt_tokens": 2269, "prefill_tps": 321.3, "ttft_s": 7.14, "decode_tps": 7.83, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
6+
{"model": "mlx-community/gemma-4-26b-a4b-it-8bit", "config": "SSD", "context": 2048, "run": 1, "warmup": false, "peak_gpu_gb": 7.32, "swap_delta_gb": 0.0, "min_free_pct": 70, "status": "OK", "prompt_tokens": 2293, "prefill_tps": 193.7, "ttft_s": 11.93, "decode_tps": 7.7, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
7+
{"model": "mlx-community/gemma-4-26b-a4b-it-8bit", "config": "SSD", "context": 2048, "run": 2, "warmup": false, "peak_gpu_gb": 7.06, "swap_delta_gb": 0.0, "min_free_pct": 69, "status": "OK", "prompt_tokens": 2291, "prefill_tps": 119.6, "ttft_s": 19.24, "decode_tps": 7.64, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
8+
{"model": "mlx-community/gemma-4-26b-a4b-it-8bit", "config": "SSD", "context": 8192, "run": 0, "warmup": false, "peak_gpu_gb": 7.58, "swap_delta_gb": 0.0, "min_free_pct": 66, "status": "OK", "prompt_tokens": 9543, "prefill_tps": 147.6, "ttft_s": 64.86, "decode_tps": 7.08, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
9+
{"model": "mlx-community/gemma-4-26b-a4b-it-8bit", "config": "SSD", "context": 8192, "run": 1, "warmup": false, "peak_gpu_gb": 7.19, "swap_delta_gb": 0.0, "min_free_pct": 60, "status": "OK", "prompt_tokens": 9518, "prefill_tps": 138.5, "ttft_s": 68.94, "decode_tps": 6.92, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
10+
{"model": "mlx-community/gemma-4-26b-a4b-it-8bit", "config": "SSD", "context": 8192, "run": 2, "warmup": false, "peak_gpu_gb": 7.26, "swap_delta_gb": 0.0, "min_free_pct": 53, "status": "OK", "prompt_tokens": 9577, "prefill_tps": 136.9, "ttft_s": 70.22, "decode_tps": 6.64, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
Lines changed: 14 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,14 @@
1+
### `mlx-community/gemma-4-26b-a4b-it-8bit`
2+
3+
Apple M6 · 32 GB · runs=3 (long=1) · warmup=1 · gen=128 · temperature 0 · medians
4+
5+
| Config | Context (prompt tok) | Prefill tok/s | TTFT s | Decode tok/s | Peak GPU GB | Swap Δ GB | Min free % | Checks |
6+
|---|---|---|---|---|---|---|---|---|
7+
| Vanilla | 512 | — | — | — | 3.33 | 3.06 | 28 | **MEM_ABORT** swap grew 3.1 GB |
8+
| Vanilla | 2048 | — | — | — | — | — | — | **SKIPPED_AFTER_ABORT** |
9+
| Vanilla | 8192 | — | — | — | — | — | — | **SKIPPED_AFTER_ABORT** |
10+
| Vanilla | 32768 | — | — | — | — | — | — | **SKIPPED_AFTER_ABORT** |
11+
| SSD | 512 (528) | 219.6 | 2.44 | 8.78 | 6.95 | 0.0 | 73 | ok |
12+
| SSD | 2048 (2291) | 193.7 | 11.93 | 7.7 | 7.32 | 0.0 | 69 | ok |
13+
| SSD | 8192 (9543) | 138.5 | 68.94 | 6.92 | 7.58 | 0.0 | 53 | ok |
14+
| SSD | 32768 | — | — | — | 5.8 | 2.09 | 32 | **MEM_ABORT** swap grew 2.1 GB |

0 commit comments

Comments
 (0)