Everything here is opt-in; the defaults are chosen so a plain ./coli chat
is safe on any machine. See also SETTINGS.md and
ENVIRONMENT.md for the full variable inventory.
| knob | what it does |
|---|---|
--temp T |
token sampling temperature (default 0.7 + nucleus 0.90 — tuned for int4; 0 = greedy) |
--topp 0.7 |
adaptive expert top-p (30–40% less disk; lossy — prints a warning) |
--ngen N |
max tokens per answer (:more in chat continues a truncated one) |
--repin N |
adapt RAM/VRAM hot experts every N emitted tokens |
RAM_GB=<n> |
claim more RAM for the expert cache than the conservative auto-detect |
PIN=stats PIN_GB=g |
pin the hottest experts from a measured usage profile |
DRAFT=n |
MTP draft depth (0 disables speculation) |
GRAMMAR=g.gbnf |
grammar-forced drafts for constrained JSON/NDJSON output (docs) |
THINK=1 |
enable GLM-5.2's reasoning block |
PILOT=1 |
router-lookahead disk prefetch (see below) |
URING=1 |
Linux-only batched expert I/O (implies PIPE=1) |
PIPE=0 |
disable the async expert-load pool (default ON — overlaps pread with matmul, −18% disk service) |
DIRECT=1 |
O_DIRECT expert reads (measured +65% alone on a Strix Halo, #200) |
COLI_NUMA=1 |
interleave resident weights across NUMA nodes on multi-socket hosts (#82) |
CACHE_ROUTE=1 |
cache-aware max-rank routing (opt-in, #199) |
AUTOPIN=0 |
disable the learning cache's auto-pin |
CAP_RAISE=0 |
don't auto-grow the expert cache |
KVSAVE=0 |
disable KV-cache persistence |
TF=1 |
teacher-forcing validation |
Automatic history pinning and the adaptive LRU share the same expert RAM
budget. Colibri caps automatic pinning to preserve the no-pin LRU capacity;
explicit PIN and PIN_GB settings remain authoritative.
coli plan reports the planned hot (VRAM), warm (RAM), and cold backing (disk)
tiers, the reason for each placement, and the expected bottleneck. The default
--policy quality and --policy balanced modes preserve checkpoint quantization
and router decisions unless --topk or --topp is passed; those explicit lossy
overrides print a warning and proceed.
Auto-tier plans size OpenMP from physical cores and bind workers across cores.
Memory-bound quantized kernels can regress sharply when SMT siblings compete for
limited memory channels. The GLM, Kimi K3, and OLMoE engines also apply that
physical-core cap when launched directly; explicit OMP_NUM_THREADS and the
COLI_NO_OMP_TUNE kill switch always take precedence.
Note (#471): exporting
OMP_PROC_BIND/OMP_PLACESused to interact badly with the engine's one-time OpenMP tuning re-exec on Linux — the re-exec'd image inherited the first image's place-0 thread binding and the whole team landed on one core (~20× slowdown). The engine now resets its affinity to all online CPUs right before the re-exec, so explicitOMP_*pinning works as documented.COLI_OMP_TUNED=1remains the escape hatch that skips the re-exec entirely.
coli plan --model /models/glm52_i4 --policy quality
coli run --auto-tier --policy quality "Explain MoE offloading"
# Explicit research-only router reduction:
coli run --policy experimental-fast --topk 4 "Benchmark prompt"coli plan chooses a safe starting point from capacity and topology.
coli tune measures the remaining scheduling choices on the actual model and
machine, then saves a hardware/model/engine-specific profile:
coli tune --model /models/glm52_i4
coli run --model /models/glm52_i4 --auto-tier "Explain MoE offloading"The calibration prompt is generated once. Every candidate then teacher-forces
the same continuation, so answer length and sampling do not contaminate the
comparison. The bounded sweep only includes execution knobs such as OpenMP
thread count, NUMA placement, I/O overlap, direct I/O, and CUDA pipelining. It
never changes weights, quantization, router decisions, TOPK, TOPP, or
sampling.
A candidate is saved only when median throughput improves by at least 3% while
expert hit rate remains within 0.5 percentage points and p99 latency stays
within 20% of the baseline. The winner is then rerun before a final baseline;
this reverse-order gate gives the baseline any remaining warm-cache advantage
and rejects startup drift. Otherwise the baseline is recorded and no override
is applied. Saved profiles are loaded by --auto-tier; explicit environment
variables always win. Use --no-tune-profile to bypass a saved profile.
Profiles live under $XDG_CONFIG_HOME/colibri/tuning (normally
~/.config/colibri/tuning) or %LOCALAPPDATA%\colibri\tuning on Windows. A
change to the engine binary, model metadata, CPU topology, or GPU inventory
produces a new fingerprint instead of reusing stale measurements.
Disk is an immutable recovery source, not a normal decode target. If the plan leaves cold expert bytes on disk, speed depends on cache hit rate; output quality does not.
Cold expert reads can use a deferred pipeline: resident RAM/VRAM experts execute
while missing experts are loaded in a bounded background I/O pool, then the cold
results join before the layer completes. The pool engages only under PIPE=1;
PIPE_WORKERS=n sets its worker count (default 8). Profiling reports both disk
service time and the smaller foreground-visible wait time so overlap is explicit.
--policy balanced enables lossless live placement (REPIN=64). At safe request
boundaries, a per-layer LFRU score combines decaying session frequency with recent
access and replaces at most four sufficiently colder pinned experts. --policy quality leaves live replacement off by default; REPIN=0 always disables it.
The engine records which experts your usage actually routes to (.coli_usage
next to the model, updated every turn) and at startup automatically pins the
hottest ones in spare RAM — colibrì literally gets faster the more you use it.
PIN=auto seeds the pin directly from the live usage history
(#301).
The expert cache auto-sizes to your RAM (since 2026-07-10): the engine
raises the LRU cap to fill your --ram budget instead of only lowering it.
If you benchmarked colibrì before that date, rerun — your numbers were capped.
Live tier adaptation (--repin N, opt-in): at safe turn boundaries, a
decaying session heat map replaces cold pinned experts with hotter streamed
experts. A 25% hysteresis and a four-swap limit prevent tier thrashing.
Persistent .coli_usage remains the long-term signal and is not decayed.
The history's on-disk format, what happens when one engine is handed another
engine's history, and how PIN=<file> differs from PIN=auto in how much it
trusts a file are documented in
routing-telemetry.md.
GLM-5.2's expert routing is measurably predictable ahead of time — applying
layer L+1's router to layer L's post-attention state recalls 71.6% of the
true top-8 (vs 41.3% for "same experts as last token"). PILOT=1 issues
next-layer expert readahead from a dedicated I/O thread while the current layer
computes. PILOT_REAL=1 moves the prefetched loads off the critical path
(measured +11pp hit rate on a big-cache host), and PILOT_TWO=1 folds the
computed shared-expert into the prediction (+3% recall,
#200). On disk-saturated
hosts hint-only PILOT can be net negative — measure on yours.
Speculative decoding requires that the draft and verify paths compute the same
function — SPEC_PIN=1 (default since #294)
pins every forward issued while drafts are live to the platform's S=1 kernel
family. For byte-exact reproducibility across runs: DRAFT=0, plus IDOT=0 COLI_CUDA=0 if you also want kernel-family/GPU independence. Acceptance
percentages are not comparable across engine versions under --topp
(#163 has the full story).
coli chat persists the compressed MLA KV-cache to disk after every turn
(.coli_kv, ~182 KB/token, appended incrementally, crash-safe). Close the chat,
reopen it tomorrow — the model still remembers the whole conversation and zero
re-prefill happens: validated byte-identical to an uninterrupted session.
:reset clears it, KVSAVE=0 disables it.