Curated research from Kiloloop — runtime comparisons, operational reviews, practical guides, and protocol patterns for multi-agent coordination.
Research published from our internal pipeline. Each piece is reviewed and sanitized before landing here.
| Directory | Contents |
|---|---|
runtime-comparison/ |
Runtime capability comparisons and benchmarks |
ops-reviews/ |
Operational reviews of our own agent fleet, with the scripts behind the numbers |
| Date | Piece | Where |
|---|---|---|
| 2026-09-07 | Estimate accuracy, measured: seven months of dispatch estimates against their actuals — actual-to-estimate ratios from 606 dispatch rows by period, month, estimate size and task class, the declared-envelope headroom on review-loop legs, verdicts and review status; the computation behind the numbers is estimate_stats.py |
ops-reviews/ |
| 2026-09-05 | Human in the loop, measured: four months of autonomy-gate records — pause rates, human outcomes and latency, checkpoint breaches and declared-vs-actual fidelity from 602 audit records; reproducible with hil_stats.py |
ops-reviews/ |
| refreshed 2026-09 | Runtime capability matrix — what each coding-agent runtime can and cannot do, self-reported and cross-verified by the agents | runtime-comparison/ |
| 2026-03 | Prompt caching patterns — maximizing prompt-cache hits across runtimes | runtime-comparison/ |
Code: Apache License 2.0 Content: CC BY 4.0