Skip to content

Commit 1735cd7

Browse files
authored
Merge pull request #196 from SharpAI/fix/ssd-streaming-followups-2
fix(ssd): stream the directory we load, resume partial first-run downloads, keep exiting events on their own line
2 parents 60f05ec + 4e071f9 commit 1735cd7

6 files changed

Lines changed: 209 additions & 6 deletions

File tree

‎README.md‎

Lines changed: 5 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -69,7 +69,7 @@ The first SwiftLM numbers from a **32 GB** Mac. Every other table in this README
6969
| `Qwen3.6-35B-A3B-UD-MLX-4bit` | 21.6 GB | `--stream-experts` | 13.2 tok/s | 40.8K tokens | 5.8 GB |
7070
| `Qwen3.8-27B-4bit` (dense) | 11.3 GB | GPU | 9.3 tok/s | 40.8K tokens | 18.4 GB |
7171
| `gemma-4-26b-a4b-it-8bit` | ~26 GB | GPU | swaps (+3.1 GB on the first prompt) | — | — |
72-
| `gemma-4-26b-a4b-it-8bit` | ~26 GB | `--stream-experts` | 8.8 tok/s | 9.5K tokens (32K swapped) | 7.6 GB |
72+
| `gemma-4-26b-a4b-it-8bit` | ~26 GB | `--stream-experts` | 9.1 tok/s | 9.5K tokens (32K swapped) | 7.3 GB |
7373

7474
- **MoE models are the sweet spot at 32 GB.** Only the active experts are read for each token, so they decode 5–6× faster than a dense 27B. A 4-bit MoE with up to about 22 GB of weights runs entirely on the GPU.
7575
- **Qwen3.6-35B-A3B on a base M6 reaches 76%** of the M1 Ultra 64 GB decode speed below (47.0 vs 61.7 tok/s).
@@ -90,14 +90,16 @@ Every needle check passed. `--mtp` with the bf16 assistant (`gemma-4-26B-A4B-it-
9090

9191
### Qwen3.6-35B-A3B 4-bit — GPU vs SSD streaming
9292

93+
![Qwen3.6-35B-A3B with --stream-experts on release b782: 13.6 tok/s decode, 5.1 GB peak, no swap, on a base Mac mini M6 32 GB](docs/profiling/m6/media/m6_qwen36_35b_a3b_ssd_stream.gif)
94+
9395
| Prompt tokens | GPU prefill / decode (tok/s) | GPU peak | `--stream-experts` prefill / decode (tok/s) | SSD peak |
9496
|---|---|---|---|---|
9597
| ~550 | 714 / 47.0 | 19.8 GB | 256 / 13.2 | 5.6 GB |
9698
| ~2.3K | 968 / 45.7 | 20.1 GB | 403 / 13.0 | 5.6 GB |
9799
| ~9.8K | 858 / 43.4 | 20.4 GB | 401 / 12.7 | 5.6 GB |
98100
| 40.8K | 615 / 36.1 | 21.5 GB | 336 / 12.0 | 5.8 GB |
99101

100-
> ⚠️ **`--stream-experts` crashes on Qwen3.5/3.6 in releases b769 and b773** (`broadcast_shapes … (N,8,8,2048)` on the first request). The mlx-swift-lm upstream sync in #167 broke the SSD path. Earlier versions of this table were measured before that sync and were never re-checked afterwards. Fixed in SharpAI/mlx-swift-lm#69 and #71; the table above was re-measured with those fixes.
102+
> ⚠️ **`--stream-experts` crashes on quantized MoE models in releases b769 and b773** (`broadcast_shapes … (N,8,8,D)` on the first request). Reproduced on M6 with Qwen3.6-35B-A3B and Gemma 4 26B-A4B. The mlx-swift-lm upstream sync in #167 broke the SSD path. Earlier versions of this table were measured before that sync and were never re-checked afterwards. Fixed in SharpAI/mlx-swift-lm#69 and #71; the table above was re-measured with those fixes.
101103
102104
### Qwen3.8-27B-4bit (dense)
103105

@@ -121,7 +123,7 @@ Every needle check passed. `--mtp` with the bf16 assistant (`gemma-4-26B-A4B-it-
121123
1. **The MLX buffer cache was unbounded on full-GPU loads.** It could grow to the whole 26.8 GB working set. It is now sized from the RAM left after weights and KV.
122124
2. **The KV-cache estimate counted every layer as full attention.** Gemma 4 (25 of 30 layers use a 1,024-token sliding window) was overestimated 10×, and Qwen3.5/3.8 (48 of 64 layers are linear attention) 4×. On 32 GB that pushed Gemma into CPU/GPU layer partitioning, which crashed with a Metal GPU timeout.
123125
3. **An auto-detected VLM that failed to load exited the server.** `Qwen3.6-35B-A3B-UD-MLX-4bit` ships a `preprocessor_config.json` without `image_mean`. SwiftLM now falls back to text-only unless you pass `--vision`.
124-
4. **Vision-capable models skipped chunked prefill.** On the older mlx-swift-lm pin, a text-only prompt on the VLM path ran through the model in a single pass. It's fixed by the mlx-swift-lm bump in #167. Every number in this section was measured on `main` after that bump; the Qwen3.6 table was re-measured with SharpAI/mlx-swift-lm#69 and #71.
126+
4. **Vision-capable models skipped chunked prefill.** On the older mlx-swift-lm pin, a text-only prompt on the VLM path ran through the model in a single pass. It's fixed by the mlx-swift-lm bump in #167. Every number in this section was measured on `main` after that bump. The `--stream-experts` rows (Qwen3.6 and Gemma 4 8-bit) were re-measured with SharpAI/mlx-swift-lm#69 and #71, because b769/b773 crash in that mode.
125127

126128
> ℹ️ **`--turbo-kv` long-range recall is fixed** ([#175](https://github.com/SharpAI/SwiftLM/issues/175), SharpAI/mlx-swift-lm#65). Before the fix, once a prompt passed the 2,048-token compression threshold, attention only saw the recent hot window and positions restarted, so Qwen3.8-27B-4bit got exact lookups wrong. Attention now covers the compressed history too, which makes `--turbo-kv` slower than before (97 s vs 72 s on an 11.8K-token prompt on the M6). `--turbo-kv` still has no effect when `--ctx-size` is set (the attention layers use `RotatingKVCache`).
127129
>

‎Sources/SwiftLM/Server.swift‎

Lines changed: 42 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -281,6 +281,21 @@ final class ProgressTracker {
281281
init(modelId: String) {
282282
self.modelId = modelId
283283
}
284+
285+
/// True while a `\r` progress bar is on screen without its closing newline.
286+
nonisolated(unsafe) static var barOpen = false
287+
288+
/// Stops redrawing and ends the bar's line, so later output (and the `exiting`
289+
/// event, which the daemon parses per line) starts on a fresh line.
290+
func finish() {
291+
isDone = true
292+
trackingTask?.cancel()
293+
if Self.barOpen {
294+
print("")
295+
fflush(stdout)
296+
Self.barOpen = false
297+
}
298+
}
284299

285300
func getDownloadedBytes() -> Int64 {
286301
let home = FileManager.default.homeDirectoryForCurrentUser
@@ -358,11 +373,14 @@ final class ProgressTracker {
358373

359374
let msg = String(format: "\r[SwiftLM] Download: [%@] %@ %@ (%@ MB / %@ MB) %@", bars, pctStr, spinner, completedMB, totalMB, speedText)
360375

376+
if self.isDone { break }
361377
print(msg.padding(toLength: 100, withPad: " ", startingAt: 0), terminator: "")
362378
fflush(stdout)
363-
379+
Self.barOpen = true
380+
364381
if fraction >= 1.0 {
365382
print("")
383+
Self.barOpen = false
366384
self.isDone = true
367385
break
368386
}
@@ -398,6 +416,10 @@ func emitEvent(_ payload: [String: Any]) {
398416
Data("[SwiftLM] failed to encode event for stdout: \(payload)\n".utf8))
399417
return
400418
}
419+
if ProgressTracker.barOpen {
420+
print("")
421+
ProgressTracker.barOpen = false
422+
}
401423
print(json)
402424
fflush(stdout)
403425
}
@@ -644,6 +666,13 @@ struct MLXServer: AsyncParsableCommand {
644666
var modelDirectory =
645667
ModelStorage.validatedContentDirectory(for: modelId)
646668
?? resolveModelDirectory(modelId: modelId)
669+
// resolveModelDirectory doesn't check the weights are there; streaming must not be
670+
// activated for an empty or partial snapshot the loader won't read.
671+
if self.streamExperts, let dir = modelDirectory,
672+
!ModelStorage.validateLocalModelDirectory(dir)
673+
{
674+
modelDirectory = nil
675+
}
647676
if self.streamExperts, !self.info, modelDirectory == nil,
648677
!FileManager.default.fileExists(atPath: modelId)
649678
{
@@ -654,22 +683,30 @@ struct MLXServer: AsyncParsableCommand {
654683
.appendingPathComponent("MLX", isDirectory: true)
655684
.appendingPathComponent("HuggingFace", isDirectory: true))
656685
let localRepo = hub.localRepoLocation(Hub.Repo(id: modelId))
657-
if FileManager.default.fileExists(
658-
atPath: localRepo.appendingPathComponent("config.json").path)
686+
// Every shard must be present: an interrupted download (or the config.json
687+
// the architecture probe fetches) would otherwise plan with a partial size.
688+
if FileManager.default.fileExists(atPath: localRepo.path),
689+
ModelStorage.validateLocalModelDirectory(localRepo)
659690
{
660691
modelDirectory = localRepo
661692
} else {
662693
// First run. A failed download is a model problem, not a binary one.
663694
phase = .architectureProbe
664695
print("[SwiftLM] --stream-experts: downloading \(modelId) before loading...")
665696
let prefetchTracker = ProgressTracker(modelId: modelId)
697+
defer { prefetchTracker.finish() }
666698
modelDirectory = try await hub.snapshot(
667699
from: modelId, matching: ["*.safetensors", "*.json", "*.jinja"]
668700
) { progress in
669701
prefetchTracker.printProgress(progress)
670702
}
671703
}
672704
}
705+
// Streaming is activated for `modelDirectory`, and only a load of that exact
706+
// directory streams. Load from it, or a different lookup could pick another copy.
707+
if self.streamExperts, let dir = modelDirectory {
708+
modelConfig = ModelConfiguration(directory: dir)
709+
}
673710
var mainModelProfile: ModelProfile? = nil
674711
if self.streamExperts, let dir = modelDirectory {
675712
mainModelProfile = ModelProfiler.profile(modelDirectory: dir, modelId: modelId)
@@ -937,6 +974,7 @@ struct MLXServer: AsyncParsableCommand {
937974
return self.model
938975
}()
939976
let tracker = ProgressTracker(modelId: resolvedModelId)
977+
defer { tracker.finish() }
940978

941979
let isAudio = self.audio
942980
phase = .mainModelLoad
@@ -1003,6 +1041,7 @@ struct MLXServer: AsyncParsableCommand {
10031041
print("[SwiftLM] Note: the prompt cache is not used for VLM/Omni loads; each text request re-prefills its full prompt.")
10041042
}
10051043

1044+
tracker.finish()
10061045
print("[SwiftLM] Loaded model configuration. Inferred tool call format: \(String(describing: await container.configuration.toolCallFormat))")
10071046

10081047
// ── Check if target model supports DFlash ──

‎docs/profiling/m6/gemma4_26b_a4b_8bit.jsonl‎

Lines changed: 7 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -8,3 +8,10 @@
88
{"model": "mlx-community/gemma-4-26b-a4b-it-8bit", "config": "SSD", "context": 8192, "run": 0, "warmup": false, "peak_gpu_gb": 7.58, "swap_delta_gb": 0.0, "min_free_pct": 66, "status": "OK", "prompt_tokens": 9543, "prefill_tps": 147.6, "ttft_s": 64.86, "decode_tps": 7.08, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
99
{"model": "mlx-community/gemma-4-26b-a4b-it-8bit", "config": "SSD", "context": 8192, "run": 1, "warmup": false, "peak_gpu_gb": 7.19, "swap_delta_gb": 0.0, "min_free_pct": 60, "status": "OK", "prompt_tokens": 9518, "prefill_tps": 138.5, "ttft_s": 68.94, "decode_tps": 6.92, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
1010
{"model": "mlx-community/gemma-4-26b-a4b-it-8bit", "config": "SSD", "context": 8192, "run": 2, "warmup": false, "peak_gpu_gb": 7.26, "swap_delta_gb": 0.0, "min_free_pct": 53, "status": "OK", "prompt_tokens": 9577, "prefill_tps": 136.9, "ttft_s": 70.22, "decode_tps": 6.64, "gen_tokens": 128, "needle_ok": true, "degenerate": false}
11+
{"model": "mlx-community/gemma-4-26b-a4b-it-8bit", "config": "SSD", "context": 512, "run": -1, "warmup": true, "peak_gpu_gb": 5.87, "swap_delta_gb": 0.0, "min_free_pct": 71, "status": "OK", "prompt_tokens": 520, "prefill_tps": 70.6, "ttft_s": 7.39, "decode_tps": 9.19, "gen_tokens": 128, "needle_ok": true, "degenerate": false, "code": "MARLIN-60", "answer_head": "The secret code word is MARLIN-60.\n\n***\n\nElias Thorne had lived within the salt-crusted walls of the Blackwood Reach lighthouse for twenty-two years, a duration long enough for the rhythmic pulse of the Fresnel lens to become his own heartb"}
12+
{"model": "mlx-community/gemma-4-26b-a4b-it-8bit", "config": "SSD", "context": 512, "run": 0, "warmup": false, "peak_gpu_gb": 6.5, "swap_delta_gb": 0.0, "min_free_pct": 71, "status": "OK", "prompt_tokens": 533, "prefill_tps": 90.6, "ttft_s": 5.93, "decode_tps": 9.13, "gen_tokens": 128, "needle_ok": true, "degenerate": false, "code": "MARLIN-23", "answer_head": "The secret code word is MARLIN-23.\n\n***\n\nElias Thorne had lived within the salt-crusted walls of the Blackwood Reach lighthouse for twenty-two years, a span of time measured not by calendars, but by the rhythmic pulse of the great Fresnel l"}
13+
{"model": "mlx-community/gemma-4-26b-a4b-it-8bit", "config": "SSD", "context": 512, "run": 1, "warmup": false, "peak_gpu_gb": 6.31, "swap_delta_gb": 0.0, "min_free_pct": 71, "status": "OK", "prompt_tokens": 542, "prefill_tps": 91.8, "ttft_s": 5.95, "decode_tps": 9.09, "gen_tokens": 128, "needle_ok": true, "degenerate": false, "code": "PELICAN-81", "answer_head": "The secret code word is PELICAN-81.\n\n***\n\nElias Thorne had lived within the salt-crusted walls of the Blackwood Reach Lighthouse for twenty-two years, a span of time that had turned his skin into something resembling cured leather and his e"}
14+
{"model": "mlx-community/gemma-4-26b-a4b-it-8bit", "config": "SSD", "context": 512, "run": 2, "warmup": false, "peak_gpu_gb": 5.6, "swap_delta_gb": 0.0, "min_free_pct": 71, "status": "OK", "prompt_tokens": 527, "prefill_tps": 88.7, "ttft_s": 5.99, "decode_tps": 9.04, "gen_tokens": 128, "needle_ok": true, "degenerate": false, "code": "HERON-20", "answer_head": "The secret code word is HERON-20.\n\n***\n\nElias Thorne had lived within the salt-crusted walls of the Blackwood Reach Lighthouse for twenty-two years, a span of time that had turned his skin to leather and his eyes to the color of a storm-tos"}
15+
{"model": "mlx-community/gemma-4-26b-a4b-it-8bit", "config": "SSD", "context": 8192, "run": 0, "warmup": false, "peak_gpu_gb": 6.46, "swap_delta_gb": 0.0, "min_free_pct": 66, "status": "OK", "prompt_tokens": 9546, "prefill_tps": 148.4, "ttft_s": 64.52, "decode_tps": 8.22, "gen_tokens": 128, "needle_ok": true, "degenerate": false, "code": "PELICAN-68", "answer_head": "The secret code word is PELICAN-68.\n\nElias Thorne had lived within the salt-crusted walls of the Blackwood Reach Lighthouse for twenty-two years, a duration that felt less like a career and more like a slow merging of man and stone. His wor"}
16+
{"model": "mlx-community/gemma-4-26b-a4b-it-8bit", "config": "SSD", "context": 8192, "run": 1, "warmup": false, "peak_gpu_gb": 5.92, "swap_delta_gb": 0.0, "min_free_pct": 66, "status": "OK", "prompt_tokens": 9511, "prefill_tps": 145.4, "ttft_s": 65.61, "decode_tps": 8.18, "gen_tokens": 128, "needle_ok": true, "degenerate": false, "code": "HERON-85", "answer_head": "The secret code word is HERON-85.\n\nElias Thorne had lived within the salt-crusted walls of the Blackwood Reach lighthouse for twenty-two years, a duration that felt less like a career and more like a slow merging of man and stone. His world"}
17+
{"model": "mlx-community/gemma-4-26b-a4b-it-8bit", "config": "SSD", "context": 8192, "run": 2, "warmup": false, "peak_gpu_gb": 7.3, "swap_delta_gb": 0.0, "min_free_pct": 59, "status": "OK", "prompt_tokens": 9549, "prefill_tps": 145.7, "ttft_s": 65.73, "decode_tps": 8.16, "gen_tokens": 128, "needle_ok": true, "degenerate": false, "code": "MARLIN-30", "answer_head": "The secret code word is MARLIN-30.\n\nElias Thorne had lived within the salt-crusted walls of the Blackwood Rock Lighthouse for twenty-two years, a duration long enough for the rhythmic pulse of the Fresnel lens to become his own heartbeat. T"}

‎docs/profiling/m6/gemma4_26b_a4b_8bit.md‎

Lines changed: 10 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -12,3 +12,13 @@ Apple M6 · 32 GB · runs=3 (long=1) · warmup=1 · gen=128 · temperature 0 ·
1212
| SSD | 2048 (2291) | 193.7 | 11.93 | 7.7 | 7.32 | 0.0 | 69 | ok |
1313
| SSD | 8192 (9543) | 138.5 | 68.94 | 6.92 | 7.58 | 0.0 | 53 | ok |
1414
| SSD | 32768 | — | — | — | 5.8 | 2.09 | 32 | **MEM_ABORT** swap grew 2.1 GB |
15+
16+
#### `--stream-experts` re-measured with SharpAI/mlx-swift-lm#69 and #71 (b769/b773 crash in this mode)
17+
18+
Apple M6 · 32 GB · runs=3 (long=1) · warmup=1 · gen=128 · temperature 0 · medians
19+
20+
| Config | Context (prompt tok) | Prefill tok/s | TTFT s | Decode tok/s | Peak GPU GB | Swap Δ GB | Min free % | Checks |
21+
|---|---|---|---|---|---|---|---|---|
22+
| SSD | 512 (533) | 90.6 | 5.95 | 9.09 | 6.5 | 0.0 | 71 | ok |
23+
| SSD | 8192 (9546) | 145.7 | 65.61 | 8.18 | 7.3 | 0.0 | 59 | ok |
24+
| SSD | 32768 | — | — | — | 5.79 | 2.04 | 31 | **MEM_ABORT** swap grew 2.0 GB |
Lines changed: 145 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,145 @@
1+
{"version":3,"term":{"cols":112,"rows":40},"timestamp":1790449820,"idle_time_limit":2.0,"command":"/Users/simba/.cache/swiftlm-bench/demo/record_ssd.sh","title":"SwiftLM · Qwen3.6-35B-A3B --stream-experts · Mac mini M6 32GB","env":{"SHELL":"/bin/zsh"}}
2+
[0.329, "o", "\u001b[36mSwiftLM — github.com/SharpAI/SwiftLM\u001b[39m\r\n\u001b[3244mOpenAI-compatible LLM server for Apple Silicon, in Swift on MLX · release b782\u001b[39m\r\n\r\n"]
3+
[0.000, "o", "\u001b[3244m$\u001b[39m sysctl -n machdep.cpu.brand_string hw.memsize | paste -sd' ' -\r\n"]
4+
[0.641, "o", "Apple M6 · 32 GB · macOS 27.0\r\n\r\n"]
5+
[0.000, "o", "\u001b[3244m$\u001b[39m ./SwiftLM --model unsloth/Qwen3.6-35B-A3B-UD-MLX-4bit --stream-experts --port 5431 &\r\n"]
6+
[2.203, "o", "💾 SSD Expert Streaming enabled (lazy load + layer-sync)\r\n✅ Ready. Listening on http://127.0.0.1:5431\r\n"]
7+
[0.000, "o", "\r\n"]
8+
[0.000, "o", "\u001b[3244m$\u001b[39m python3 scripts/demo/stream_client.py warmup # first request after load (warm-up)\r\n"]
9+
[1.545, "o", "\r\u001b[K\u001b[32m1"]
10+
[0.070, "o", ","]
11+
[0.069, "o", " "]
12+
[0.073, "o", "2"]
13+
[0.075, "o", ","]
14+
[0.073, "o", " "]
15+
[0.072, "o", "3"]
16+
[0.070, "o", ","]
17+
[0.071, "o", " "]
18+
[0.072, "o", "4"]
19+
[0.071, "o", ","]
20+
[0.072, "o", " "]
21+
[0.072, "o", "5"]
22+
[0.071, "o", ","]
23+
[0.072, "o", " "]
24+
[0.072, "o", "6"]
25+
[0.069, "o", ","]
26+
[0.072, "o", " "]
27+
[0.073, "o", "7"]
28+
[0.075, "o", ","]
29+
[0.078, "o", " "]
30+
[0.071, "o", "8"]
31+
[0.072, "o", "\u001b[0m\r\n\r\n\u001b[36m\u001b[1m 22 tokens · decode 13.3 tok/s · TTFT 0.9s ·\u001b[0m\r\n\u001b[36m\u001b[1m peak SwiftLM 5.1 GB · swap +0.0 GB\u001b[0m\r\n"]
32+
[0.004, "o", "\u001b[3244m$\u001b[39m python3 scripts/demo/stream_client.py short # measured request\r\n"]
33+
[1.448, "o", "\r\u001b[K\u001b[32m//"]
34+
[0.068, "o", " Returns"]
35+
[0.070, "o", " the"]
36+
[0.070, "o", " n"]
37+
[0.070, "o", "-th"]
38+
[0.074, "o", " Fibonacci"]
39+
[0.079, "o", " number"]
40+
[0.074, "o", " using"]
41+
[0.068, "o", " an"]
42+
[0.069, "o", " iterative"]
43+
[0.071, "o", " approach"]
44+
[0.072, "o", "."]
45+
[0.068, "o", "\r\n"]
46+
[0.071, "o", "func"]
47+
[0.070, "o", " fib"]
48+
[0.072, "o", "(_"]
49+
[0.070, "o", " n"]
50+
[0.072, "o", ":"]
51+
[0.072, "o", " Int"]
52+
[0.070, "o", ")"]
53+
[0.076, "o", " ->"]
54+
[0.081, "o", " Int"]
55+
[0.073, "o", " {"]
56+
[0.072, "o", "\r\n"]
57+
[0.070, "o", " "]
58+
[0.070, "o", " guard"]
59+
[0.071, "o", " n"]
60+
[0.070, "o", " >"]
61+
[0.071, "o", " "]
62+
[0.071, "o", "0"]
63+
[0.069, "o", " else"]
64+
[0.071, "o", " {"]
65+
[0.072, "o", " return"]
66+
[0.072, "o", " "]
67+
[0.072, "o", "0"]
68+
[0.080, "o", " }"]
69+
[0.083, "o", "\r\n"]
70+
[0.071, "o", " "]
71+
[0.071, "o", " guard"]
72+
[0.071, "o", " n"]
73+
[0.069, "o", " >"]
74+
[0.073, "o", " "]
75+
[0.072, "o", "1"]
76+
[0.071, "o", " else"]
77+
[0.070, "o", " {"]
78+
[0.072, "o", " return"]
79+
[0.072, "o", " "]
80+
[0.073, "o", "1"]
81+
[0.070, "o", " }"]
82+
[0.075, "o", "\r\n"]
83+
[0.089, "o", " \r\n"]
84+
[0.082, "o", " "]
85+
[0.070, "o", " var"]
86+
[0.070, "o", " prev"]
87+
[0.068, "o", "2"]
88+
[0.070, "o", " ="]
89+
[0.070, "o", " "]
90+
[0.071, "o", "0"]
91+
[0.070, "o", "\r\n"]
92+
[0.070, "o", " "]
93+
[0.071, "o", " var"]
94+
[0.070, "o", " prev"]
95+
[0.070, "o", "1"]
96+
[0.073, "o", " ="]
97+
[0.086, "o", " "]
98+
[0.093, "o", "1"]
99+
[0.071, "o", "\r\n \r\n"]
100+
[0.070, "o", " "]
101+
[0.070, "o", " for"]
102+
[0.069, "o", " _"]
103+
[0.069, "o", " in"]
104+
[0.072, "o", " "]
105+
[0.071, "o", "2"]
106+
[0.070, "o", "..."]
107+
[0.070, "o", "n"]
108+
[0.071, "o", " {"]
109+
[0.071, "o", "\r\n"]
110+
[0.069, "o", " "]
111+
[0.074, "o", " let"]
112+
[0.091, "o", " current"]
113+
[0.076, "o", " ="]
114+
[0.076, "o", " prev"]
115+
[0.074, "o", "1"]
116+
[0.072, "o", " +"]
117+
[0.072, "o", " prev"]
118+
[0.072, "o", "2"]
119+
[0.070, "o", "\r\n"]
120+
[0.072, "o", " "]
121+
[0.069, "o", " prev"]
122+
[0.070, "o", "2"]
123+
[0.071, "o", " ="]
124+
[0.072, "o", " prev"]
125+
[0.074, "o", "1"]
126+
[0.087, "o", "\r\n"]
127+
[0.096, "o", " "]
128+
[0.072, "o", " prev"]
129+
[0.071, "o", "1"]
130+
[0.071, "o", " ="]
131+
[0.072, "o", " current"]
132+
[0.071, "o", "\r\n"]
133+
[0.073, "o", " "]
134+
[0.071, "o", " }"]
135+
[0.072, "o", "\r\n"]
136+
[0.072, "o", " \r\n"]
137+
[0.070, "o", " "]
138+
[0.071, "o", " return"]
139+
[0.072, "o", " prev"]
140+
[0.087, "o", "1"]
141+
[0.098, "o", "\r\n"]
142+
[0.070, "o", "}"]
143+
[0.072, "o", "\u001b[0m\r\n\r\n\u001b[36m\u001b[1m 110 tokens · decode 13.6 tok/s · TTFT 0.8s ·\u001b[0m\r\n\u001b[36m\u001b[1m peak SwiftLM 5.1 GB · swap +0.0 GB\u001b[0m\r\n"]
144+
[0.004, "o", "\r\n"]
145+
[0.032, "x", "0"]
127 KB
Loading

0 commit comments

Comments
 (0)