Skip to content

Commit fe9da73

Browse files
test: shrink Test 38 prompt to fit Gemma-4 prefill on CI runner
The 6000-number prompt tokenizes to ~42k tokens on Gemma-4. Its prefill materializes full attention in one Metal buffer (28 GB), above the CI runner's 3.5 GB max buffer size, so the server crashed and the stream never reached [DONE]. 1000 numbers (~7k tokens) stays well under the limit; the ': connected' check still catches the original regression. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
1 parent ac26d57 commit fe9da73

1 file changed

Lines changed: 4 additions & 1 deletion

File tree

‎tests/test-server.sh‎

Lines changed: 4 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1161,9 +1161,12 @@ rm -f /tmp/mlx_inflight_models.json /tmp/mlx_inflight_health.json /tmp/mlx_infli
11611161
# whole prefill inside container.perform before returning headers, so clients
11621162
# saw no bytes for the entire prefill (Bun fetch 10s idle → ECONNRESET/retry).
11631163
# A large prompt makes prefill multi-second; TTFB must stay well under that.
1164+
# ponytail: ~7k tokens. Gemma-4 prefill allocates O(n²) attention in one Metal
1165+
# buffer; the CI runner caps buffers at 3.5 GB (~15k tokens). 42k tokens
1166+
# crashed the server with a 28 GB malloc.
11641167
log "Test 38: streaming TTFB — ': connected' arrives before model prefill"
11651168

1166-
python3 -c "print(' '.join(f'{i:06d}' for i in range(6000)))" > /tmp/mlx_ttfb_prompt.txt
1169+
python3 -c "print(' '.join(f'{i:06d}' for i in range(1000)))" > /tmp/mlx_ttfb_prompt.txt
11671170
jq -nc --arg m "$MODEL" --rawfile p /tmp/mlx_ttfb_prompt.txt \
11681171
'{model:$m, stream:true, max_tokens:5, messages:[{role:"user",content:("Summarize this list in one word:\n"+$p)}]}' \
11691172
> /tmp/mlx_ttfb_body.json

0 commit comments

Comments
 (0)