> ⚠️ **`--turbo-kv` loses exact long-range recall on every chip:** once a prompt passes the 2,048-token compression threshold, Qwen3.8-27B-4bit with `--turbo-kv` gets exact lookups wrong. Asked how many numbered lines a prompt has, it answers "1,000", "1,314" or "14" instead of 315 / 500 / 700. Reproduced on both M5 and M6; without `--turbo-kv` it answers correctly. The cause is a cache-eviction bookkeeping regression: after compression, attention only sees the recent hot window, and positions restart. A fix is in progress; tracked in [#175](https://github.com/SharpAI/SwiftLM/issues/175). Until then, avoid `--turbo-kv` when exact recall matters. Also note that `--turbo-kv` currently has no effect when `--ctx-size` is set (the attention layers use `RotatingKVCache`).
0 commit comments