diff --git a/README.md b/README.md index 81bff21..1116fed 100644 --- a/README.md +++ b/README.md @@ -1,15 +1,13 @@ # Gobstopper -> 🍬 Gobstopper is a tool for saving tokens while preserving context. -> Compaction is the process built into your harness that automatically -> compresses long conversations to fit inside the context window. Gobstopper -> is a replacement for the compaction built into Claude Code and Codex. It -> proxies API requests and compacts more often, which saves tokens. Your -> agent’s full history stays on your machine, and compacted session copies are -> archived in a local vault your agents can search. On Terminal-Bench 2.1, -> Gobstopper used 29% fewer tokens than Claude Code’s built-in compaction and -> solved as many tasks. +> 🍬 Gobstopper saves tokens while preserving context. Compaction is the step +> your harness runs to squeeze a long conversation into the context window. +> Gobstopper replaces the compaction built into Claude Code and Codex. It +> proxies API requests, compacts more often to save tokens, and writes the full +> history to a local database your agents can search. On Terminal-Bench 2.1 it +> used 29% fewer tokens than Claude Code's built-in compaction and solved the +> same number of tasks. > > Ask your agent to set it up: https://gobstopper.sh > diff --git a/site/app/globals.css b/site/app/globals.css index bb58c63..a102aa6 100644 --- a/site/app/globals.css +++ b/site/app/globals.css @@ -849,12 +849,13 @@ a[href]:where(:hover, :focus-visible) { text-decoration-color: currentColor; } .founder-note__emoji { font-size: 2rem; line-height: 1; padding-block-start: 0.2rem; } .founder-note__body { display: grid; gap: 0.9rem; } .founder-note__body p { - margin: 0; font-size: 1.125rem; line-height: 1.6; text-wrap: pretty; hanging-punctuation: first; + margin: 0; font-family: var(--font-note, "Instrument Serif", "Iowan Old Style", "Palatino Linotype", Palatino, serif); + font-size: 1.375rem; line-height: 1.5; text-wrap: pretty; hanging-punctuation: first; } -.founder-note__action { font-weight: 550; } -.founder-note__body .founder-note__signature { color: var(--muted); font-size: 0.95rem; } +.founder-note__action { font-weight: 400; } +.founder-note__body .founder-note__signature { color: var(--muted); font-size: 1.25rem; } .founder-note__signature::before { content: "— "; } @media (max-width: 40rem) { .founder-note { grid-template-columns: minmax(0, 1fr); gap: 0.75rem; } - .founder-note__body p { font-size: 1.0625rem; } + .founder-note__body p { font-size: 1.25rem; } } diff --git a/site/app/page.tsx b/site/app/page.tsx index 8d65c97..7c1126f 100644 --- a/site/app/page.tsx +++ b/site/app/page.tsx @@ -116,7 +116,7 @@ export default function Home() { Gobstopper\n
\n

🍬 Gobstopper is a tool for saving tokens while preserving context.\nCompaction is the process built into your harness that automatically\ncompresses long conversations to fit inside the context window. Gobstopper\nis a replacement for the compaction built into Claude Code and Codex. It\nproxies API requests and compacts more often, which saves tokens. Your\nagent’s full history stays on your machine, and compacted session copies are\narchived in a local vault your agents can search. On Terminal-Bench 2.1,\nGobstopper used 29% fewer tokens than Claude Code’s built-in compaction and\nsolved as many tasks.

\n

Ask your agent to set it up: https://gobstopper.sh

\n

— Ben Guo

\n
\n

Gobstopper is free and open source. When a request passes a token threshold,\ngobstopper proxy replaces the older turns with one mechanical summary and\nsends the newest turns word for word, so the provider sees a smaller context.\nThat can delay the agent’s own auto-compaction trigger.

\n

On Terminal-Bench 2.1 through Claude Code, Gobstopper at its default tail\nand a 45,000-token threshold (the default threshold is 128,000) solved as\nmany tasks as Claude Code with no proxy, 61 and 60 of 89, and sent 29%\nfewer input tokens. That is one trial per arm with GLM 5.3 Flash on\nSeptember 27 and 28, 2026, so the solved counts are within single-trial\nnoise; see Benchmark results.

\n

\"Play

\n

The summary rule comes from CliffCompaction, an open-source proxy described\nin a paper by Trang Nguyen, Eulrang Cho, Bingqing Chen, and\nTim Dettmers. gobstopper proxy is a Rust port of it for the three dialects\ncoding agents use: Anthropic Messages (Claude Code, opencode, Crush), OpenAI\nResponses (Codex), and OpenAI Chat Completions (opencode, Crush, Aider,\nGoose, and other OpenAI-compatible clients). Run gobstopper proxy run -- claude to try it on one session, gobstopper proxy install to start it at\nlogin, or see\nCompact live coding-agent requests.

\n

Gobstopper also works on saved Claude Code and Codex session files. Preview a\ncompaction at a context size you choose, prepare a separate smaller copy, and\nkeep the original byte for byte in a local vault. It archives the exact source\nand candidate bytes and checks supported structural properties and protected\nrecent output.

\n

Use gobstopper watch --dry-run to inspect threshold decisions, or prepare a\ncopy with a file strategy. Released CLI builds cannot ask providers to compact,\neven when auto_compact_closed is enabled. Direct provider-store and in-place\nrewrites are disabled. Copy preparation preserves the source; resuming a copy\nwith a live provider requires separate compatibility testing. See the\nprovider support status and\nrecovery runbook. Plugins can add strategies\nand providers when you explicitly trust them.

\n

Website: gobstopper.sh · Compared with Claude Code /compact and CliffCompaction

\n

Quick start

\n

For the Claude Code example, install and sign in to Claude Code first, and\nmake sure claude and curl are on your PATH. Gobstopper forwards your\nclient's authentication; it does not sign you in to the model provider.\nThe temporary proxy applies to the child command only, so you do not need\nto edit your shell's base URL settings.

\n
Terminal
curl -fsSL https://gobstopper.sh/install.sh | sh   # macOS (Apple silicon) and Linux\ngobstopper proxy run -- claude   # one Claude Code session through a temporary proxy\n
\n

On Windows, install from PowerShell with irm https://gobstopper.sh/install.ps1 | iex.\nBoth installers download the latest release for your platform, check its\nSHA-256, and install it for your user only. To build from source instead, run\ncargo install --git https://github.com/hraness/gobstopper gobstopper --locked.

\n

Supported macOS and Linux release installs update automatically\nbefore a command, at most once a day, when no other Gobstopper command is\nrunning. Run gobstopper update to update now, gobstopper update check to\ncheck without installing, or gobstopper update disable to turn automatic\nupdates off. gobstopper update enable restores them. CI, offline replays,\nand versions selected with GOBSTOPPER_VERSION stay fixed. Use\n--no-update or HRANESS_NO_UPDATE=1 to skip a check for one invocation.\nCargo and source builds use their original install command; Windows uses\nthe PowerShell installer. Update-enabled installs also need the\nGitHub CLI (gh) authenticated with github.com; run\ngh auth login before installing. See update behavior.

\n

When the session ends, proxy run prints request and compaction counts.\nA short session can report zero compacted requests: compaction starts only\nwhen a supported request crosses the threshold. If Claude Code cannot start,\ncheck that claude runs directly in the same terminal. For proxy connection\nproblems, follow diagnosis and repair.\nFor\na proxy that stays up, with gobstopper proxy status counters, see Compact\nlive coding-agent requests. For Codex,\nsee Set up Gobstopper for Claude Code and\nCodex.

\n

Why

\n

A coding agent carries earlier context into later requests. In a long\nsession that history fills with old file reads and command output the next\nstep rarely needs, and later requests include it again. When the history\nnears the model's window, the agent asks a model to summarize it. That call\nis itself a large request, the summary can leave out an exact error or\nconstraint, and each later summary summarizes the one before.

\n

\"Diagram:

\n

Without compaction, each later request carries the earlier messages and\ntool output. This diagram shows how repeated context accumulates.

\n

gobstopper proxy keeps each request under a threshold you choose, without\na model call:

\n\n

\"Line

\n

One recorded Claude Code session: 383 requests replayed at a 45,000-token\nthreshold (default 128,000), with calibration on. Counts estimate four\ncharacters per token. Build f4db57e uses the v0.7.2 request engine.

\n

What CliffCompaction's authors report

\n

The CliffCompaction paper reports up to\n50% lower cost at a bounded context, with Terminal-Bench 2.0 scores held or\nimproved, on the Kimi K2.6 and GLM 5.1 models its authors tested. In one run\nthrough Claude Code (GLM 5.3 Flash on Terminal-Bench 2.1, at about 45,000\ntokens of mean peak context), the rule scored 76.69%, against 70.97% for\nClaude Code's own auto-compaction and 73.03% for its default 200,000-token\nsetting. The paper's costs come from a model of perfect prompt caching, not\nmetered bills. The authors also report that the benefit depends on the agent\nand the task, and that it matters only for medium-to-long tasks.

\n

These are the authors' measurements of their own proxy. gobstopper proxy\nshares its summary rule and adds context-retention and configuration\ncontrols, described in How Gobstopper compares with\nCliffCompaction. Gobstopper\nhas not rerun the authors' benchmarks as published; its own Terminal-Bench\n2.1 run is under What Gobstopper has measured.

\n

What Gobstopper has measured

\n

On September 27 and 28, 2026, Gobstopper ran the 89 tasks of Terminal-Bench\n2.1 through Claude Code 2.1.283 with GLM 5.3 Flash via Vercel AI Gateway,\none trial per arm, at a 45,000-token threshold (the default is 128,000).\ngobstopper proxy v0.7.2¹ at tail 0 (--keep-tail-percent 0, the default\nsince v0.7.3) solved 61 tasks, Claude Code with no proxy 60, and tail 40, the old\ndefault, 59. Those counts are within single-trial noise (McNemar p = 1.0\nagainst no proxy; 31 of 89 tasks changed outcome between arms). The tail-0\narm sent 29% fewer provider-reported input tokens than no proxy, 84.3\nmillion against 118.6 million, and almost all of the difference was cache\nreads. Its provider-reported cost for this model, metered through Vercel\nAI Gateway, was about 16% lower ($5.72 against $6.82 over 89 tasks), which\nis not statistically significant (95% interval −32% to +2%). Tail 40 cost\n39% more than tail 0 in total; put the other way, tail 0 cost 28% less (95%\ninterval 1.6% to 46.5% less), and five tasks drive most of that gap. v0.7.3\nmade tail 0 the default. The benchmarks\npage has the\nsetup, per-arm tables, paired statistics, limits, and downloadable\naggregates.

\n

\"Bar

\n

Cache reads account for most of the difference: 68.7M with tail 0 against\n102.6M with no proxy. New input and output were about equal.

\n

The proxy replay studies report estimated\nrequest sizes separately from task results and provider-reported token counts.

\n
Benchmark notes
\n
    \n
  1. 21 of 89 tail-0 trials may have run an earlier build.
  2. \n
\n

Saved sessions

\n

For session files, Gobstopper lets you check the tradeoff before you commit\nto it. You can preview a compaction, compare strategies on the same frozen\nbytes, keep the exact source in a local vault, and recover a specific\narchived record when a copy leaves it out. The built-in strategies use local\nrules and need no model. Optional model scorers change what gets selected;\nthey do not skip the snapshot or the verification step. A running session's\ncontext belongs to the provider process that loaded it, so file compaction\nprepares a separate copy. The published\nstudies report context reduction,\nretention, no-op cases, and limitations separately.

\n

For saved-session edits, Gobstopper keeps the exact source in a local vault before writing a smaller copy, so a compaction is a recorded edit you can recover from rather than a silent loss: the design every Hraness project shares. The thread through hraness follows that design across the projects, and the ALGAL vision states the bet behind it.

\n

Compact live coding-agent requests

\n

gobstopper proxy is a local HTTP proxy that sits between a coding agent\nand its model provider. Each time the client resends its history, the proxy estimates the request\nsize. Past the threshold (128,000 tokens by default), it sends the system\nprompt and the first task verbatim, one mechanical summary of the older\nturns, and the last three turns verbatim. At the default tail of 0, the kept\nturns are exactly --keep-recent; a positive --keep-tail-percent lets the\nsummary and older whole turns fill that share of the room left under the\nthreshold after the system prompt and the first task. The provider then\nreports the compacted size back to the client, which can delay the client’s own\nauto-compaction trigger. The summary rule is\nCliffCompaction's; see How Gobstopper compares with\nCliffCompaction.

\n

\"Diagram:

\n

When a request passes the threshold, the middle becomes one summary.\nGobstopper keeps the start and the last three turns word for word and\nreplaces the middle with a mechanical summary. No model writes it. Your\nfiles and your saved session are not changed.

\n

\"Diagram:

\n

Inside a rewritten request. In the logged part of the tail-0\nTerminal-Bench arm (about 68 of the 89 trials), compacted requests had a\nmedian of 31.5K estimated tokens, against a median of 55K before\ncompaction. The example lines are illustrative.

\n

It speaks the three dialects coding agents use:

\n
\n\n\n\n\n\n\n\n\n\n\n\n
AgentDialectHow to point it at the proxy
Claude CodeAnthropic Messagesexport ANTHROPIC_BASE_URL=http://127.0.0.1:8260
CodexOpenAI Responsesmodel_providers block in ~/.codex/config.toml
opencodeAnthropic Messages or Chat Completionsprovider.<id>.options.baseURL → http://127.0.0.1:8260/v1
CrushAnthropic Messages or Chat Completionsproviders.<id>.base_url → http://127.0.0.1:8260/v1
AiderChat Completionsaider --openai-api-base http://127.0.0.1:8260/v1
GooseChat CompletionsOPENAI_HOST=http://127.0.0.1:8260
\n

The setup for each agent is in docs/proxy.md. Any other\nOpenAI-compatible client that posts to {base}/chat/completions works the\nsame way. Claude Code and Codex routing is live-checked; the Chat\nCompletions dialect is contract-tested against synthetic histories and has\nnot yet been qualified against a live opencode, Crush, Aider, or Goose\nsession.

\n

The summary keeps human and assistant text, keeps tool results of at most 500\ncharacters, and reduces each tool call to a one-line signature. A separate\nbounded carry retains selected original tool results and images with their\ninvocation and labels excerpts. Context retention\ndescribes the limits and controls.\nThe next compaction starts again from the history the client resends and\ndiscards the previous summary, but the human's words and the assistant's\nvisible replies carry forward: each later summary opens with them, oldest\nfirst, up to 24,000 characters, and the oldest text drops out when they no\nlonger fit. Between compactions, requests reuse the same compacted prefix,\nso the provider's prompt cache can match it.

\n

The kept turns hold the files and command output the agent read most\nrecently. The summary omits long results unless the bounded evidence carry\nselects them. By default\nthe proxy keeps exactly the newest --keep-recent turns, as CliffCompaction\ndoes. --keep-tail-percent (0 to 60, default 0)\nkeeps older whole turns too while the summary and the kept turns fit in that\nshare of the room, and leaves the rest for new turns before the next\ncompaction. A higher floor leaves less room before the next compaction, so\nwe expect more compactions and larger requests in between; in replay of 24\nrecorded sessions, tail 40 compacted 369 times against 343 at 32,000\ntokens and 50 against 38 at 128,000 (estimates). In the Terminal-Bench 2.1\nrun at a 45,000-token threshold, tail 40 sent 118.5 million input tokens\nagainst 84.3 million at tail 0 and cost 39% more in total in\nprovider-reported terms (tail 0 cost 28% less, 95% interval 1.6% to 46.5%\nless, with five tasks driving most of the gap), with solved counts within\nsingle-trial noise; see Benchmark results. The carried words keep earlier instructions in view\nafter the summary that held them is discarded; they use at most a quarter of\nthat room, and --carry-max-chars 0 turns carrying off. Anthropic Messages\nrequests that declare a 1M-token context window use a separate threshold,\n--threshold-1m: 256,000 estimated tokens by default, or --threshold if\nthat is higher. The proxy reads the window from the anthropic-beta header,\nwhere Claude Code sends a token starting with context-1m for a model such\nas opus[1m], and never from the model name. Setting --threshold-1m equal\nto --threshold applies one threshold to every request. Keep each threshold\nbelow the point where the client compacts on its own, including any\nclaude --autocompact value.

\n
Terminal
gobstopper proxy run -- claude            # one session through a temporary proxy\ngobstopper proxy serve                    # background proxy on http://127.0.0.1:8260\ngobstopper proxy install                  # owned user service; start at login\nexport ANTHROPIC_BASE_URL=http://127.0.0.1:8260\ngobstopper proxy replay <session>         # what the proxy would have sent; calls no provider\ngobstopper proxy status                   # counters and estimated-token totals, this run and all time\n
\n

The Terminal-Bench run used --threshold 45000; the default 128,000\ncompacts later, and in replays most recorded Claude Code sessions never\nreach it (the replay grid in Benchmark results\ncompares thresholds).\nSee docs/proxy.md for per-agent setup (Claude Code, Codex,\nopencode, Crush, Aider, Goose), proxy install and proxy uninstall,\nchoosing a threshold, and every setting.

\n

Claude Code → local Gobstopper proxy → model provider.

\n

Optional rewrite failures can send the original bytes when policy permits.\nA provider HTTP 400 rejection can trigger another trim for a length error,\nor an original-body retry for another error when the original fits configured\ncapacity.

\n

Flags: --threshold (keep it below the client's auto-compaction point),\n--threshold-1m, --keep-recent, --keep-tail-percent,\n--result-max-chars, --carry-max-chars, --evidence-max-bytes,\n--evidence-max-chars, --context-window, --no-keep-awake, --drop-thinking,\n--no-calibrate, --shadow (log what would change and forward everything\nunchanged), and --strict.

\n\n

On September 25, 2026, gobstopper proxy replay with three kept turns, the\ndefault on that date, over nine recorded sessions on one Mac kept six Claude\nCode sessions, whose recorded requests peaked at 273k to 652k estimated\ntokens, at or under about 127k, and one Codex session that peaked at 242k\nunder about 127k. Two Codex sessions that Codex had already compacted itself\nbegan with heads near 160k and stayed under about 243k. No replayed request\nwas left with an unpaired tool call. These are estimates over recorded\nhistories, not billed tokens or task results.

\n

Longer work, local visibility, and startup recovery

\n

Temporary context budgets

\n

Compacting while an agent is gathering evidence can make it reread material that\nwas removed. For a difficult analysis phase, you or the agent can reserve more\ninput context within a scope bound to your client and its descendants. Declare\ncapacities supported by your route; Gobstopper returns the effective budget\nafter output headroom and client limits.

\n
Terminal
gobstopper proxy run --context-window 1000000 --client-context-window 1000000 --adaptive-context -- claude\n# From inside that scoped session:\ngobstopper context reserve --tokens 500000 --requests 20 --ttl-seconds 1800\ngobstopper context status\ngobstopper context release\n
\n

The larger budget expires by request count or time. Adaptive rescue is off by\ndefault; --adaptive-context enables a temporary increase after repeated reads\nof unchanged evidence that the proxy previously removed. It needs a scope and\nconfigured capacity. See context budgets.

\n

The proxy also keeps a limited collection of original tool results and supported\nimages across repeated compactions. After a restart, it rebuilds that collection\nfrom the history the client sends. Older evidence can still be evicted, and the\nproxy cannot recover material removed by the client's own compaction. These\ncontrols address evidence loss; they do not guarantee that an agent stops looping\nor completes its task. See evidence retention.

\n

Startup, recovery, and direct fallback

\n

gobstopper proxy install starts a user service at login and restarts it after a\nprocess exit. Managed service changes pause new inference with a retry response\nand wait for existing requests to finish. If the controller disappears while\nwaiting, its lease expires and requests reopen. Once a stop has been committed,\nrecovery checks the outcome before reopening. Service changes use no firewall\nrules.

\n

Request parsing, compaction, and status checks have time and resource limits.\nLogging, metrics, and sleep prevention run in background workers so slow optional\nwork does not hold up request forwarding. Configured context limits still apply,\nand unavailable scoped context storage returns a retry response.

\n
Terminal
gobstopper proxy launch --client claude --print  # inspect readiness and route\ngobstopper proxy launch --client claude\ngobstopper proxy launch --client codex --codex-auth chatgpt\n
\n

The launcher checks the proxy before starting a client. Claude Code can use its\nofficial provider directly when the proxy is unavailable and its configuration\npermits that route. Custom upstreams, uncertain authentication, scoped context\nreservations, and configured capacity constraints prevent direct fallback.\nCodex requires a healthy proxy and an existing explicit custom provider pointing\nto it; --codex-auth selects the existing authentication route to check.\nThe launcher does not replay inference or reroute running sessions. Clients\nalready configured with a fixed proxy URL still depend on that listener.\nSee startup and recovery for setup and fallback requirements.

\n

During active inference, Gobstopper requests idle-sleep prevention and releases\nit when inference ends. Closing a lid and forced sleep remain operating-system\ndecisions. --no-keep-awake disables the feature.

\n

Local session data

\n

The proxy records local metadata and provider usage. Inspect requests, attempts,\ncompaction decisions and tool activity, or export the versioned journal:

\n
Terminal
gobstopper data requests\ngobstopper data metrics\ngobstopper data export > gobstopper-events.jsonl\ngobstopper data check\ngobstopper proxy doctor\n
\n

Session data explains the schema, privacy boundaries,\nimports, backups and metric denominators.\nThese controls have functional regression tests; the September 28 benchmark\npredates them and does not measure their effect on task accuracy.

\n

Token use across your agents

\n

gobstopper usage shows your token use across coding agents by day, agent,\nprovider and model, including agents that never pass through the proxy. The\nnumbers come from aicharts, which keeps a daily record\non your computer and uploads nothing. The installer below adds aicharts and\nturns that record on; after another install method,\nget aicharts.

\n
Terminal
gobstopper usage                        # the last 30 days, per agent\ngobstopper usage report --days 7 --csv  # one row per day, agent and model\ngobstopper usage enable                 # collect four times a day\n
\n

Agents can read the same record through aicharts mcp. gobstopper data\ncounts what the proxy saw, so some requests appear in both; read them side by\nside rather than adding them together.

\n

Recoverable history

\n

Before publishing a Claude Code or Codex copy, Gobstopper stores the exact\nsource and candidate bytes in a content-addressed vault\n(~/.local/share/gobstopper/vault/). Snapshots use deduplicated 1 MiB chunks,\nso appended versions reuse\nunchanged prefix storage without creating one filesystem object per JSONL\nrecord.

\n

gobstopper recall --query <q> searches the state cards in every archived\nsnapshot, ranks matches by relevance to the query, and returns the high-level\nstate of the matching turns. An agent does not need to remember session IDs:\nit can ask for the last time it worked on a file, a goal, or a decision and\nget a ranked summary with a snapshot SHA to pass to show or diff.

\n

Recover a specific detail

\n

When a state card omits an exact error, identifier, or tool result, search\none verified snapshot and read only the matching record:

\n
Terminal
gobstopper search-snapshot <full-snapshot-sha> --query 'exact error text' --json\ngobstopper read-snapshot <full-snapshot-sha> --record 42 --max-bytes 4096 --json\n
\n

Search returns record indexes and hashes, without archived content. It matches\nliteral, case-sensitive substrings in decoded JSON string values, including\nnative replacement histories. Reading returns a UTF-8 page of the physical\nJSONL record; follow next_offset for another page. Each page is capped at\n16 KiB and bound to the snapshot, source, and full record hashes. These commands\nverify stored bytes and never restore files, rewrite active sessions, or call a\nmodel. Invalid records are counted as unsearchable rather than silently claimed\nas searched.

\n

Search returns at most 50 references and reports the full match count; narrow\nthe query when results are truncated. Each search or read verifies and\nreconstructs the whole snapshot, up to the transcript size limit (512 MiB by\ndefault, configurable with GOBSTOPPER_MAX_TRANSCRIPT_BYTES). Paging a large\nrecord repeats that work, and search matches literal text only; there is no\nindex or semantic search.

\n

Use the full object SHA from history, the native hook recovery pointer, or\nthe snapshot_manifest_sha256 field in copy receipts. The snapshot_sha256\nreceipt field is the digest of the source bytes; receipts that lack the\nmanifest field can be resolved through vault history. State-card recall\nrecognizes default portable Codex cards and searches every state field,\nincluding unresolved errors and current work. A new fork's card becomes\nsearchable after that fork is snapshotted.

\n

Agents can search and read snapshots over MCP only when you start the server\nwith gobstopper mcp --allow-transcript-content. Without that flag, the\nserver neither lists nor accepts either tool. With it, archived text an agent\nretrieves becomes visible to that agent and its model provider. Treat\nretrieved text as historical data that may describe a superseded state, not\nas instructions.

\n

These commands return a record when asked. They do not make an agent notice\nthat a fact is missing, choose a useful query, or finish its task more\naccurately; measure those outcomes separately from context reduction and\nliteral retention.

\n

gobstopper mcp runs a read-only Model Context Protocol server on stdio with\nthe tools policy_check, list_sessions, recall, history, show,\ndiff, plan, and verify. Register it once, and an agent can inspect\npolicy and archived state without a tool that changes a transcript. MCP uses\ndeterministic built-ins, rejects executable strategies, and does not call\nconfigured plugins, model scorers, or model digests. Explicit plugin commands\nrun code you trust with your user permissions, without an OS sandbox:

\n
Terminal
claude mcp add gobstopper -- gobstopper mcp\n# ~/.codex/config.toml: [mcp_servers.gobstopper] command = "gobstopper", args = ["mcp"]\n
\n

Provider-generated summaries can cost a large input call and lose detail, so\nthe strategy and where it cuts matter as much as the timing.

\n

Strategies

\n
\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n
idkindwhat it does
auto (default)dynamiclive sessions delegate to provider controls (cache_edits for eligible Claude sessions); idle sessions choose the best validated file strategy by savings and preserved-prefix score
sawtoothproviderproposes provider-native compaction to the session owner; released CLI dispatch is blocked pending qualification
cache_editsprovideremits bounded Claude tool_use_id values for API-layer context editing; never rewrites a transcript
elidetranscriptstubs stale tool outputs oldest-first until the floor
clifftranscriptkeeps the head, the newest three assistant steps, and the newest keep_recent_tool_outputs tool results (default 8) byte-for-byte and drops older eligible tool results over 500 bytes; no floor seeking and no state card (see CliffCompaction; for running sessions, use gobstopper proxy)
cache_awaretranscriptelides a tailward stale-output window and injects a bounded state card while preserving the longest practical prefix
compactedtranscriptelides stale outputs and injects the state card; synthetic Codex compacted records require --experimental-compacted
scoredtranscriptranks candidates with deterministic recency, error, reference, TF-IDF, duplicate, and tool-type signals before elision
dedupetranscriptremoves older exact duplicate tool payloads using payload SHA-256, not summaries
microtranscriptkeeps the newest configured outputs per stable tool label and stubs older ones
middletranscriptprotects both ends of the transcript and elides eligible middle outputs
structuredtranscriptemits a bounded metadata-derived state card; it is not semantic summarization
agenticextensionaccepts bounded edit proposals from a command or versioned plugin you trust; Gobstopper still validates every edit
\n

Custom strategies are userspace code: a preset can name a command that\nreceives the normalized transcript as JSON on stdin and returns an edit\nplan on stdout, or install a versioned gobstopper-plugin.json bundle\n(see gobstopper plugin check). A command runs only with\ntrusted_legacy_command = true. Gobstopper checks eligible payloads, protected\nrecent output, edit combinations, digest size, and projected token reduction.\nFile candidates must not introduce supported structural findings. These checks\ncover edit structure and size, not semantic preservation or provider acceptance.

\n

Install & use

\n

Check the release notes when\nyou need a capability tied to a particular release.

\n

Install the latest release:

\n
Terminal
# macOS (Apple silicon) and Linux (x86_64, arm64): installs ~/.local/bin/gobstopper\ncurl -fsSL https://gobstopper.sh/install.sh | sh\n
\n
# Windows (x86_64), in PowerShell: installs to %LOCALAPPDATA%\\Programs\\gobstopper\\bin, no administrator rights\nirm https://gobstopper.sh/install.ps1 | iex\n
\n

On macOS (Apple silicon) and Linux x86_64, install.sh also adds aicharts\nbeside gobstopper, checked against a pinned SHA-256 digest and, on macOS,\nits Developer ID signature. On a first install it turns on local usage\nhistory: daily token totals for your agents, kept on this computer and never\nuploaded. aicharts history disable turns it off. It also turns on\naicharts' daily self-update check, which installs a new release only after\nverifying it; aicharts update disable turns that off, or set\nGOBSTOPPER_AICHARTS_UPDATE=no before installing. Set\nGOBSTOPPER_USAGE_HISTORY=no to leave history off, or GOBSTOPPER_AICHARTS=no\nto skip aicharts.

\n

Set GOBSTOPPER_VERSION=X.Y.Z to install one exact release. The installers\ncheck each download against the release's SHA-256 file; docs/release.md\nshows how to check a download's build provenance attestation yourself. On\nWindows, the vault, apply, watch and the provider hooks are Unix-only and\nrefuse with an error; detect, plan, verify, mcp and the proxy work.

\n

To build main or another platform from source:

\n
Terminal
cargo install --git https://github.com/hraness/gobstopper gobstopper\n# or from a checkout: cargo build --release\n
\n
Terminal
gobstopper proxy run -- claude     # one Claude Code session through the proxy\ngobstopper proxy serve             # background proxy on http://127.0.0.1:8260\ngobstopper proxy status            # requests compacted, estimated tokens saved\n\ngobstopper detect                  # sessions, context sizes, lifetime burn\ngobstopper plan <session>          # what would happen, under which strategy\ngobstopper plan <session> --trigger 100000 --floor 30000    # tune the trade-off\ngobstopper eval <session>          # compare strategies on the same frozen bytes\ngobstopper apply <session> --strategy elide  # Codex/Claude copy; native requests are refused\ngobstopper verify <session>        # supported structural checks (exit 1 on errors)\ngobstopper fork <session>          # clone under a fresh session id + resume cmd\ngobstopper undo <session>          # Codex/Claude: restore a snapshot into a new fork\ngobstopper vault                   # list snapshots in the undo vault\ngobstopper prune                   # preview keeping the newest 10 snapshots per session\ngobstopper install-hooks --output ./hook-candidates.json  # private settings candidates\ngobstopper watch --dry-run         # inspect threshold decisions without preparing copies\ngobstopper watch --dry-run --active-only --once  # bounded recent-session inspection\ngobstopper explain                 # the occupancy model behind the defaults\ngobstopper recall --query <q>      # search state-card digests across all archived sessions\ngobstopper history <session>       # every archived state of one session\ngobstopper diff <sha-a> <sha-b>    # structural comparison of two vault snapshots\ngobstopper bench                   # compare strategies on recently changed sessions\ngobstopper tune <session>          # preview the adaptive trigger/floor for a session\ngobstopper mcp                     # deterministic inspection; executable strategies are rejected\ngobstopper proxy serve             # compact live Claude Code and Codex requests on 127.0.0.1:8260\n
\n

Set up Gobstopper for Claude Code and Codex

\n
    \n
  1. \n

    Install the binary from main and check it:

    \n
    Terminal
    cargo install --git https://github.com/hraness/gobstopper gobstopper\ngobstopper --version\n
    \n
  2. \n
  3. \n

    Register the MCP server with each agent you use:

    \n
    Terminal
    claude mcp add -s user gobstopper -- gobstopper mcp\ncodex mcp add gobstopper -- gobstopper mcp\n
    \n

    Confirm with claude mcp list or codex mcp list.\nAdd --allow-transcript-content after mcp only if the agent should be\nable to search and read archived transcript text.

    \n
  4. \n
  5. \n

    Start the proxy and point each client at it as described in\ndocs/proxy.md.

    \n
  6. \n
\n

Hook installation and removal export candidates without changing provider settings.\nThe bundle includes the exact original settings and hashes, so keep it private.\nAutomatic settings replacement is disabled because Gobstopper cannot obtain\ncustody honored by provider/editor writers. Review and apply candidates through\nprovider-owned settings controls and retain provider trust prompts. Callbacks\narchive source-bound evidence; their session identifiers do not prove which\noperation caused a compaction. See the recovery runbook.

\n

For automation, gobstopper plan <session> --json returns the existing plan\nobject when a plan is available. A successful inspection without a plan returns\na separate JSON result, for example:

\n
{\n  "status": "no_plan",\n  "reason_code": "below_trigger",\n  "context_tokens_before": 100000,\n  "effective_trigger_tokens": 250000,\n  "target_context_tokens": 40000,\n  "min_savings_tokens": 4096,\n  "projected_context_tokens_after": null,\n  "projected_savings_tokens": null\n}\n
\n

The reason identifies the decision actually reached:

\n
\n\n\n\n\n\n\n\n\n\n\n
reason_codeMeaning
below_triggerContext is below the effective policy trigger.
strategy_returned_no_planThe strategy declined; its underlying reason is unknown.
empty_external_editsThe configured command or plugin supplied no edits.
minimum_savings_not_metA proposal fell short of the minimum projected savings.
external_nonreducing_planAn external proposal did not reduce estimated context.
\n

Projections are present only when a rejected proposal supplied them.\nThe target is a policy setting, not a measured\nminimum context size, and projected savings are not billed savings. Invalid\nconfiguration, invalid proposals, and execution failures are command errors.

\n

Gobstopper's copy paths require retained source and candidate bytes before\npublication. File-copy paths publish a separate candidate after structural\nverification. Native dispatch remains guarded even if a policy proposes it;\nstandalone native apply refuses before creating a fork or snapshot. Legacy\ndirect-write flags remain readable but cannot authorize those writes. A real\nwatch pass can archive a source snapshot before reaching the native guard;\nwatch --dry-run does not create that snapshot.\nTelemetry is best effort:\nsuccessful event writes use the gobstopper/compaction-events-v1 schema.

\n

eval and bench freeze each session's source before comparing strategies.\nbench selects sessions updated within seven days by default;\n--all removes that age filter but retains discovery and input limits. Its\n24-column CSV includes source/result hashes, execution_state, token_basis,\nretention availability and a closed failure category. Discovered sessions\nthat fail policy resolution or evaluation remain explicit failed rows with\nunavailable measurements. Parse CSV quoting rather than splitting lines or\ncommas: session identifiers can contain those characters. A provider proposal\nis provider_not_executed; a detached transform is not a resumed provider\nsession. Numeric legacy fields must be read with those state and availability\nfields, not counted as measured zeroes or task success.

\n

Typed-retention experiments (opt-in)

\n

eval-study replays four arms on isolated in-memory candidates: an unchanged\nno_compaction baseline; plain\nobservation masking; typed masking (constraints, procedures, and open tasks\nstay pinned in their original records and roles, and retrieved text never\nbecomes a higher-authority instruction); and typed_digest (pinned records\nare elided but their spans are carried verbatim on an injected state card,\nwhich loses the original record and role just as a summary does). It does not\nchange auto, call a model, emit live compaction telemetry, or modify the\nprovider session. A requested floor may remain unreachable rather than\ndropping a pinned item.

\n
Terminal
gobstopper eval-study /private/source.jsonl --prepare-manifest /private/checks.json\ngobstopper eval-study /private/source.jsonl --manifest /private/checks.json --rounds 10 --trigger 1 --floor 40000 --json\n
\n

Preparation refuses an existing destination and writes only hashes, byte spans,\nJSON pointers, types, and opaque check IDs, not transcript text. Its labels are\nheuristic candidates that nobody has reviewed: at most 16 complete lines per type,\nwith elidable records considered first and source order breaking ties. Reviewed\nmanifests can instead use label_source = "reviewed"; classification coverage\nis not measured by retention. The JSON schema is gobstopper-retention-v1, with\nsource_sha256, label_source, and checks entries containing id, kind,\nrecord_index, pointer, start_byte, end_byte, and sha256 of that exact\nUTF-8 span. Types are constraint, procedure, open_task, fact, preference,\nand episode. Only the first three are pinned. Source identity, live context,\ntext-only pointers, span boundaries, duplicate IDs, and hashes are checked\nbefore any replay. Limits: 64 MiB of source for replay (512 MiB, the vault\nlimit, for score-only manifest prep and --against audits),\n1 MiB of manifest, 256 checks, 4 KiB per span, and 1–10 rounds.

\n

The report separates text presence, same-origin presence, and preservation at\nthe original source record/pointer. lexical_retained is a paraphrase-sensitive\nmiddle tier: a check counts when ≥75% of its normalized content tokens\n(lowercase alphanumeric, ≥4 chars, stopwords removed) appear together in one\nlive slot. That helps when a provider summary rephrases rather than repeats,\nbut it measures token coverage, not semantic equivalence. by_kind holds [total, source-bound retained, lexical retained]; the elidable subset is reported\nseparately. Dead\nbranches and metadata cannot satisfy a check. Pre-existing source verification\nerrors and newly introduced errors are counted separately. Counts are not\nsemantic or behavioral scores. Estimated context uses adapter item estimates, not stale provider usage records\nor billing. All arms use the same policy, including minimum savings and the\nprotected recent tool-output tail.

\n

--against AFTER switches to a score-only realized audit: the manifest binds\nto the session's before-state and retention is scored against independent\nafter-bytes, with no replay and no mutation. Either spec may be a vault:<sha256>\nsnapshot reference. scripts/retention-audit.py scans the vault for consecutive\nsnapshots whose provider compaction-marker count increased (Claude\ncompact_boundary, Codex "type":"compacted"; hook bracket labels alone can\nmiss the actual write), pairs surgery-labeled snapshots with the next\nsnapshot, and runs the audit over each pair. The result is realized, per-kind\nretention of compactions that already happened, including provider-native\nones.

\n

Without new work, replay is explicitly static_stress; unchanged passes do not\ncount as applied compactions. For Codex/Claude fixtures, optional growth\nentries (after_round, records) append complete provider records between\nrounds and are verified before use. Checks still refer to the initial source;\nthis is not a test of revised tasks, independent tasks, or agent reasoning.\nProvider-native compaction, semantic summarization, continuation success, cost,\nand retrieval are not measured, and the report does not score them as\nsuccessful or free. The built-in structured strategy is not used as a\nsubstitute for a semantic summarizer.

\n

A pilot can freeze up to eight selected session exports and register its\nprotocol before outcomes. Choose a new private output directory outside Git:

\n
Terminal
python3 scripts/compaction-study.py --binary target/release/gobstopper --output /private/new-pilot --session SESSION_ID\n
\n

It pins the executable, source exports, annotation manifests, and hashes;\nkeeps content private; uses isolated config/telemetry paths; and checks that the\nfrozen inputs remain unchanged. There are no provider calls. Commands have\noutput/deadline limits and the study has a 900-second overall deadline.

\n

The separate synthetic provider probe makes at most three\nClaude commands, capped at $0.25 each, using an isolated configuration directory,\nno tools, safe mode, and no MCP servers. It requires explicit opt-in and stops\nif that isolated profile is not authenticated; it never copies credentials.\nIt checks for a persisted native compaction boundary before testing recall.\n--seed-style baseline uses explicit test framing; --seed-style naturalistic\nembeds the identical facts in a plausible work narrative; constraints makes\nthe seed rule-dense; pinned keeps the rules out of the transcript entirely:\nthey ride in --append-system-prompt, the provider's own pinned-context\nchannel (safe mode disables CLAUDE.md discovery), while conversational facts\nstill go through the summarizer. claude_md exercises the production pin\nchannel instead: the same rules land in a workspace CLAUDE.md and the arm\ndrops --safe-mode so project memory loads (the isolated config home and\nscratch workspace remain the boundary). Rule-bearing styles add a rules[]\nrecall scored per-marker as constraint_rules_recalled. Recall is scored twice:\nstrict exact match (recall_checks_passed) and containment\n(recall_checks_lenient), so a semantically preserved superset answer is not\nindistinguishable from a lost fact. After interactive login in that isolated\nprofile, a fresh probe output directory can reuse it with\n--auth-home /private/previous-probe/claude-home:

\n
Terminal
python3 scripts/provider-retention-probe.py --claude-bin /absolute/path/to/claude --output /private/new-native-probe --allow-provider-calls\n
\n

A passing synthetic probe is not a four-arm real-session comparison or evidence\nof billed savings. These commands do not change the strategies the watch\ndaemon uses. Design references: Knowledge Triage,\nThe Complexity Trap,\nSelfCompact,\nACON, and\nLongMemEval.

\n

Config: ~/.config/gobstopper/config.toml

\n
[policy]\nstrategy = "auto"\ntrigger_tokens = 250_000\nfloor_tokens = 40_000\nmin_savings_tokens = 4_096   # reject ineffective plans\nadaptive = true              # derive trigger/floor per session; see `gobstopper tune`\n\n[provider.codex]             # per-provider overrides\ntrigger_tokens = 200_000\n\n[sessions."01a08d7c-…"]      # per-session overrides\nstrategy = "structured"\ntrigger_tokens = 120_000\n\n[presets.deep-work]          # named presets, selectable via --preset\nstrategy = "elide"\ntrigger_tokens = 150_000\n\n[presets.cliff]              # CliffCompaction's rule on a transcript copy\nstrategy = "cliff"\nkeep_recent_turns = 3        # newest assistant steps kept byte-for-byte\nresult_max_bytes = 500       # older tool results above this are dropped\nkeep_recent_tool_outputs = 0 # no extra protected result tail\n\n[presets.custom-script]      # legacy userspace code preset\ncommand = "python3 ~/bin/my_compactor.py"\ntrusted_legacy_command = true\n\n[discovery]\nmax_age_secs = 604800        # rolling window for `watch` and `report`;\n                             # 0 = every session regardless of age\n
\n

For sessions stored outside the default directories, such as in a sandboxed\nhome, pass --codex-home or --claude-home.

\n

Monitoring an existing Codex desktop session

\n

Standalone watch cannot compact the context already held by another Codex\nprocess. It reports native delegation as skipped, with zero credited savings;\nthe program running that session has to request the compaction. Without\n--active-only, watch and report consider sessions active within\n[discovery] max_age_secs (7 days by default); watch --max-age and\nreport --max-age/--all override it. --active-only limits discovery\nto files updated within the last 180 seconds (a recency heuristic, not proof of\nan owning process), and --once exits after one pass. A dry run writes no forks\nor compaction events. Installed provider-managed lifecycle hooks can archive\nobservations and provide a recovery pointer. They do not establish an applied\nGobstopper operation or a matched before/after pair.

\n

The optional local monitor records numeric observations\nfor an explicit list of sessions and checks a deterministic dry-run watcher.\nIt separates observed context drops, native hook activity, and projected\ncompaction plans; none is automatically counted as Gobstopper-caused usage\nsavings. Current Codex event_msg/token_count accounting and legacy usage\nrecords are both supported, including the advertised model context window.

\n

scored uses the deterministic offline heuristic by default. Experimental\nmodel scoring is opt-in with GOBSTOPPER_SCORER=llm,\nGOBSTOPPER_SCORER=clef, or GOBSTOPPER_SCORER=apple; merely setting an API\ntoken never sends data. Remote Clef and LLM scorers receive bounded labels and\nsummaries, including tool arguments, short output tails, and user-prompt\nsnippets. These are transcript-derived text, not redacted metadata; enabling\na remote scorer sends them to its configured endpoint even when additional\ncontent excerpts are disabled. Apple scoring runs on-device. All model\nscorers retain deterministic heuristic scores whenever a model omits an\nanswer or a request fails. Hosted LLM settings are hard-capped\nat 256 candidates, 64 items per batch, 16 batches, and a 100–30,000 ms\ntimeout. The built-in heuristic is the recommended default because the\nrecorded live trials did not show a better plan from the LLM scorer.

\n

The scored strategy also supports an optional keep-score cutoff. For example,\nadd this named preset to your config:

\n
[presets.retained]\nstrategy = "scored"\nkeep_score_threshold = 0.5\n
\n

Inspect it with gobstopper plan <session-id> --preset retained; use the same\nflag with gobstopper apply to create a fork. Candidates\nat or above the cutoff are preserved, even if that prevents reaching the token\ntarget. Missing, invalid, or duplicate candidate scores are also preserved.\nThe value must be finite and between 0.0 and 1.0. The heuristic score is a\nranking signal, not a calibrated probability; 0.5 is an experimental example,\nnot a tuned recommendation. Model scorers still use heuristic fallback for\npartial or failed responses, so the cutoff does not guarantee provider\nconfidence. Set strategy = "scored" explicitly: auto, compacted, and other\nstrategies ignore the cutoff. Omitting it in a later configuration layer\ninherits an earlier value rather than clearing it. The default configuration\nhas no cutoff.

\n

GOBSTOPPER_SCORER=clef uses Cloudflare Clef\non Workers AI. It returns typed noul keep-probabilities instead of generated\nprose. Set CLOUDFLARE_ACCOUNT_ID to your 32-character hexadecimal account ID\nand CLOUDFLARE_API_TOKEN to a token with Workers AI access for that account.\nCLOUDFLARE_AUTH_TOKEN is an alternative token variable. The endpoint is fixed\nto https://api.cloudflare.com/client/v4/accounts/{account}/ai/run/@cf/cloudflare/{model};\nGOBSTOPPER_CLEF_MODEL selects clef (default) or clef-flash.

\n

Cloudflare credentials come only from CLOUDFLARE_API_TOKEN, then\nCLOUDFLARE_AUTH_TOKEN. Gobstopper does not read credential files, the clipboard\nor OS credential stores for Clef. The retired GOBSTOPPER_CLEF_USE_KEYCHAIN\nsetting has no effect. Existing stored entries are left untouched.

\n
Terminal
gobstopper auth clef --status     # environment source and local checks; no model call\n
\n

auth clef without --status and auth clef --delete refuse token storage\nand removal without reading stdin or accessing a credential store. Status\nchecks local configuration only; it does not establish API access or model quality.

\n

Live Jev scoring is retired. Historical measurements and offline artifacts\nremain available under their original names. TypeSafe keys, old endpoints and\nGOBSTOPPER_JEV_* settings are not accepted by Cloudflare. Select clef\nexplicitly and configure Cloudflare environment credentials;\nGOBSTOPPER_SCORER=jev falls back to the heuristic with a warning.\nRemote scorers require curl 8.3+: bearer credentials are imported from a\nchild-only environment variable and expanded inside curl, never placed in\nprocess argv. Gobstopper disables .curlrc for these calls so user defaults\ncannot enable verbose header logging.\nGOBSTOPPER_CLEF_CONTENT_BYTES (default 0, maximum 1024) opts in to\nattaching additional bounded per-candidate content excerpts to each question.\nThe remote scorer already receives the bounded labels and summaries described\nabove. A historical Jev 141k-token A/B run selected the same six records with and\nwithout 400-byte additional excerpts, so those excerpts remain off by default.\nEvery numeric runtime knob is clamped: 1–64 questions per call, 1–128 state\nitems, 100–30,000 ms timeout, 1–16 batches per scoring pass, and 1–4\nconcurrent calls (GOBSTOPPER_CLEF_PARALLEL, default 2).\nGOBSTOPPER_CLEF_MAX_BATCHES defaults to 4. Only the newest\nMAX_Q × MAX_BATCHES tailward candidates are sent; an older prefix keeps its\ndeterministic heuristic score. This caps the default at four calls and 256\nremote-scored candidates even for unusually large transcripts. Each logical\nbatch sends at most one HTTP POST. Redirects, HTTP failures (including 5xx),\ntransport failures and malformed responses are never retried automatically.\nA failure does not prove the provider did no billable work.\nIdentical\nquestion texts within a pass are asked once: repeated tool outputs share a\nsingle remote answer instead of being billed per item. Batches run through\na bounded worker pool: execution is parallel, but results are overlaid in\nstable order, a failed or panicked batch retains heuristic scores, and\ncached answers still overlay when a batch's remote half fails. Each pass logs one summary line\nto stderr (candidates, unique/cached/sent questions, calls, failures,\nelapsed). In a historical Jev trial recorded on September 19, 2026, one three-batch\n336k-token Claude session took 148.64s before bounded parallelism and 19.47s\nafterward (about 7.6×). This single-session latency observation does not\nestablish current or provider-wide performance.

\n

eval and bench honor GOBSTOPPER_SCORER for their scored row, so\nan A/B run measures the same Clef or Apple ranking used by plan rather than\nsilently substituting the heuristic. GOBSTOPPER_EVAL_JUDGE=clef adds a\nseparate model-judged retention estimate to eval. Verbatim survivors are\ncredited locally; only up to 64 sampled strings absent from the rewritten live\ncontext become typed noul questions. They are judged against at most 100,000 bytes of bounded\ncompaction evidence (state cards, elision stubs, and short tool records, with\na head+tail fallback). An intact candidate costs no judge request. The judge is\noff by default because missed-fact evaluation sends that bounded evidence to the\nremote API. Invalid or missing answers remain unmeasured; coverage is explicit\nthrough probes_requested, probes_total, complete, recall_available and\nbasis. A partial denominator must not be compared as if it covered every\nprobe. Restricting the run to --strategy scored creates at most one logical\njudge request, with no automatic retry:

\n
Terminal
GOBSTOPPER_SCORER=clef GOBSTOPPER_EVAL_JUDGE=clef \\\n  gobstopper eval <session> --strategy scored\n
\n

Gobstopper checks Cloudflare's success, result and errors response,\nthe returned model, question IDs and typed probabilities. Clef starts from\nthe complete deterministic heuristic ranking and\noverlays only valid remote answers; capped candidates, missing answers, and\nfailed or malformed chunks keep their heuristic scores. Semantic eval omits\nunavailable answers instead of crediting unknown facts. These estimates do not\nestablish semantic equivalence, instruction authority, or task success. In one\nhistorical 80k-token A/B run, Jev chose five smaller records where the heuristic chose four larger\nones, reclaimed about 649 more tokens, and both retained all 38 extracted\nprobes. At a more aggressive floor, one probe lost verbatim was not falsely\ncredited by the semantic judge (37/38 on both scores). This is one session,\nnot a general measure of ranking quality.

\n

Successful Clef answers are cached in-process for five minutes under two\nscopes. The scorer caches each question together with the complete scoring\nstate and model identifier. Unchanged polls reuse judgments, including across\ndifferent question batches. A changed goal or tail requires fresh answers\nwhen that change is represented in the bounded scoring state; the cache cannot\ndetect task changes omitted from that state. The eval judge caches whole exact\nrequests.\nCache keys cover endpoint, credential identity, and question plus scoring\nstate or full request text; values are\nparsed probabilities only (never transcript text), evict oldest-first at\n512 questions and 64 requests, and failures are never cached. The\nquestion layer also persists to\n~/.local/share/gobstopper/clef-cache.json as sha256 key digests mapped to\na probability and timestamp, never text, so a cold plan inside\nthe TTL can reuse an answer when its question and scoring state match.\nThe version 2 disk format rejects malformed, duplicate, future-dated and\nout-of-range entries. Publication uses a new private temporary file, sync and\natomic replacement; concurrent writers can lose reusable entries, causing a\nfresh request, but do not publish a partial image. Older cache formats are\nignored. clef and clef-flash are provider model names, not weight hashes;\nmatching context and TTL do not prove the remote model stayed unchanged.\nClef inference has no automatic retries or redirect following. HTTP 5xx,\ntransport failures and malformed responses may follow billable provider work;\nGobstopper returns the failure or keeps heuristic scores without another POST.\nThe configured timeout covers the single attempt and its subprocess cleanup.\nThe scorer and judge resolve environment credentials once per process. Set\nGOBSTOPPER_CLEF_CACHE=0 or GOBSTOPPER_CLEF_CACHE_TTL_SECS=0 to disable\nboth layers;\nGOBSTOPPER_CLEF_CACHE_PATH relocates the disk file;\nthe TTL maximum is 3,600 seconds.

\n

For a separate hosted decision, write a JSON evidence file and run\ngobstopper decide evidence.json. This command sends only the file's state,\nquestions and optional images to Cloudflare, and prints the validated model,\nanswers and usage. It does not discover sessions, capture screenshots, compact\ntranscripts, or replace provider originals. It is not exposed through the\nread-only MCP server. The file's model defaults to clef; use clef-flash\nexplicitly to select the other endpoint.

\n
{\n  "model": "clef",\n  "state": "The selected build report shows two failing tests.",\n  "questions": {\n    "investigate": {\n      "type": "noul",\n      "instructions": "Does this report need investigation?"\n    }\n  }\n}\n
\n

Questions require instructions and IDs of 1–100 ASCII letters, digits, _,\n. or -. A request has 1–64 questions. choice uses a criteria object with\n2–255 options; score uses an ordered array of 2–10 levels. Returned choices\nmust match the requested options and highest reported probability. Score legends\nmust match the rubric. Gobstopper allows rounding to four decimal places: there\nmust be a distribution summing to one within the reported probabilities' rounding\nintervals, clipped to 0–1. Its probability-weighted score must fit the reported\nscore's rounding interval within the rubric's range. Returned values are not\nrenormalized. A separate decide request requires all answers; the scorer and eval judge\nkeep their fallback and missing-evidence behavior described above.

\n

Images must be embedded PNG, JPEG or WebP, either as a\ndata:image/png;base64,... string (with the appropriate MIME type), or an\nobject with content_type and base64. Only images you put in this evidence\nfile are sent. Images are never automatically attached to scoring or eval.\nRemote URLs, unsupported formats and malformed images are rejected locally.\nLimits are four images, 4 MiB per image file after base64 decoding, 8 MiB total,\n16 megapixels per image and 13 MiB for the JSON request. Image decoding has a\n128 MiB allocation limit. Oversized evidence is rejected rather than shortened. Credentials, images and state are not included\nin diagnostics.

\n

GOBSTOPPER_SCORER=apple (macOS 26+, Apple Silicon) scores on-device with\nApple Intelligence Foundation Models via the shared apple-foundation\nbridge with no remote API key. It needs a small helper program, built once on\nyour Mac:

\n
Terminal
gobstopper apple install   # builds ~/.local/share/gobstopper/apple-bridge (about 10 seconds)\ngobstopper apple status    # says whether Apple's model is ready, and what to do if not\n
\n

apple install needs Xcode 26 or Apple's command line tools. It checks for\nthem first: when they are missing it prints xcode-select --install and stops,\nso the macOS install dialog never appears unannounced. Compiler output goes to\napple-bridge-build.log next to the helper instead of your terminal. Set\nGOBSTOPPER_APPLE_BRIDGE to install to, or use, a different path. A helper\nnamed apple-bridge next to the gobstopper binary is used before the one in\n~/.local/share, and apple install rebuilds that one when it exists. Scoring\nand plan never build the helper themselves.

\n

When Apple's model can't be used, the scorer says why once and uses the\nbuilt-in scorer: Apple Intelligence is off, the model is still downloading,\nthis Mac can't run it, macOS is older than 26, or the helper isn't installed.\nEach message names the fix and, where there is one, the System Settings pane\n(for example open x-apple.systempreferences:com.apple.Siri-Settings.extension\nfor Apple Intelligence & Siri). gobstopper apple status prints the same\nmessage on demand and exits 1 until the model is ready; --json gives the\nreason code.\nUncached requests are serialized, each using one bounded --once process with\nguided JSON output and owned process cleanup. Failure retains heuristic scores.\nGOBSTOPPER_APPLE_TIMEOUT_MS, _MAX_CANDIDATES, _BATCH_SIZE, and\n_MAX_BATCHES tune it, hard-capped at 100–600,000 ms, 256 candidates, 64\nlabels-only items per batch (8 with content), and 16 batches. Since inference\nis local, the scorer also reads a bounded excerpt of each candidate record\n(GOBSTOPPER_APPLE_CONTENT_BYTES, default and maximum 400; 0 restores\nlabels-only scoring) and shrinks its default batch sizes to fit the ~4k-token\ncontext window. Each batch's guided schema contains one required p_<id> field\nper candidate, so omitted or duplicate array IDs cannot silently distort the\nranking; a malformed batch retains its heuristic scores. Excerpts come from one\nbounded source image whose normalized eligibility and payload fingerprints must\nmatch the plan input; changed or ambiguous sources fall back to the heuristic.

\n

GOBSTOPPER_DIGEST=apple goes further: the injected state card is written\nby the on-device model instead of keyword extraction. Because inference is\nlocal, it may read bounded excerpts of the records being elided without a\nremote request. Each field still\nlands in the same DigestBlock shape via guided output, capped to a small\ntoken overhead, and falls back to the mechanical card on any failure, saying\nwhy on stderr the same way the scorer does.\nGOBSTOPPER_APPLE_DIGEST_ITEMS, _ITEM_BYTES, and _TOTAL_BYTES tune the\nexcerpt budget, hard-capped at 32 records, 2,048 bytes per record, and 16,000\nbytes total; zero disables the model digest and preserves the mechanical card.

\n

The same call also writes a one-line stub per excerpted record, such as\nScript completed Wall time 4.4 seconds, stored in the elide edit's\nper_item_stubs map and rendered verbatim in place of the {bytes}/{kind}\ntemplate where the payload was removed. Records the model did not cover\nkeep the generic stub; invalid or oversized stubs are dropped by validation.

\n

Apple requests are cached in-process by task, instructions, schema, prompt and\nbridge binary identity. Only validated complete scorer batches or validated digest fields/stubs enter\nthe cache. This identifies the exact submitted bounded input, not omitted\nsource context or opaque model weights.\nModel weights and OS inference internals remain opaque; a binary hash does not\nattest their identity. An identical admitted request can reuse its recorded\nresponse without another generation.\nThe cache is bounded at 64 entries with oldest-first eviction;\nGOBSTOPPER_APPLE_CACHE=0 disables both reads and writes. Scorer diagnostics\nreport cached batches separately from real model calls. The savings gate\nalso prices the residual stub text left behind by elision so\ncontext_tokens_after doesn't overstate reclaim.

\n

With adaptive = true, the effective trigger/floor are re-derived per\nsession at each decision point: the trigger is capped at a quarter of\nthe provider-advertised context window, backed off (bounded 2x) when\nrecent compactions reclaimed too little to be worth a cycle, and\ntightened when most of the window is reclaimable tool output. The\nadjustment is deterministic and its reasons appear in plan output and\ntelemetry. gobstopper tune <session> previews it.

\n

How Gobstopper compares with CliffCompaction

\n

CliffCompaction is\nan open-source (MIT) API proxy for coding agents by Trang Nguyen, Eulrang Cho,\nBingqing Chen, and Tim Dettmers, described in\narXiv:2609.26779 (September 2026). You\npoint an agent's base URL at it. When a request exceeds a token threshold, the\nproxy sends the system prompt and task verbatim, then one mechanical summary of\nthe older turns, then the last three turns verbatim. The summary keeps tool\nresults of at most 500 characters, drops longer ones because the files behind\nthem are still readable, reduces tool calls to one-line signatures, and keeps\nassistant text. Each later compaction is rebuilt from the original history the\nagent resends, and the previous summary is discarded: the authors call this\nnever compacting a compaction. Their paper reports up to 50% lower cost at a\nbounded context with maintained or improved Terminal-Bench 2.0 results for the\nKimi and GLM models they tested; those are the authors' benchmark figures, not\nmeasurements of Gobstopper.

\n

Gobstopper extends that summary rule with carried conversation and bounded\noriginal observations from the turns earlier compactions summarized, keeps a\nlarger recent tail only if you set one, and also works on saved session\nfiles:

\n
\n\n\n\n\n\n\n\n\n\n\n\n\n\n
CliffCompactionGobstopper
Where it runsA local HTTP proxy between the agent and the Anthropic or OpenAI APIA local HTTP proxy between the agent and its model provider, plus a CLI over the session files Claude Code and Codex write
ClientsAny client of the Anthropic Messages, OpenAI Chat Completions, or OpenAI Responses APIAny client of the same three dialects that accepts a custom provider address: Claude Code, Codex, opencode, Crush, Aider, Goose, and more
What it changesEach outgoing request, transparently, while the session runsThe proxy rewrites outgoing requests over the threshold; file commands publish a separate compacted copy and leave the source unchanged
How it shrinksDrops tool results over 500 characters, signatures for tool calls, last three turns verbatim; never paraphrasesThe proxy extends the summary rule with bounded original evidence carry and keeps the last three turns verbatim by default, and older whole turns within a tail budget if you set one; file strategies drop or stub stale tool results, and structured and compacted add a metadata state card; no built-in strategy paraphrases unless GOBSTOPPER_DIGEST=apple has an on-device model write the card
RecompactionRebuilt from the original history; the prior summary is discardedThe proxy rebuilds from the original history, and each summary keeps the human's words and the assistant's visible replies from the turns earlier compactions summarized, up to 24,000 characters; cliff on a copy drops the same records as one pass over the source when both passes produce a plan; strategies that inject a state card carry it forward into the next copy
What holds the originalsThe agent's own history and the files on disk; the proxy keeps only an in-memory cache of compacted prefixesFor proxied requests, the agent's own transcript and an in-memory cache; for copies, a content-addressed vault with search-snapshot and read-snapshot
Evidence publishedTerminal-Bench 2.0 and 2.1 (including a run through Claude Code), SWE-bench Verified, and KernelBench results in the paper, on Kimi, GLM, and GPT-5-mini modelsOffline replays of 729 archived sessions, replays of nine recorded sessions through the proxy, one dated afternoon of live proxy counters, literal retention probes, dated single-session trials, and one live Terminal-Bench 2.1 run (89 tasks, three arms, one trial each, September 27 and 28, 2026): resolution within noise of Claude Code alone, 29% fewer provider-reported input tokens at the default tail and a 45,000-token threshold
Model neededNone; the summary is mechanicalNone for the proxy or the built-in strategies; optional model scorers
\n

gobstopper proxy is a Rust port of CliffCompaction's request engine: the\nsame summary format and header, prefix reuse between compactions, harsher\nsettings when one pass leaves a request over the threshold, and a retry when\nthe provider rejects a request for length. It adds context-retention and configuration controls. With --keep-tail-percent above 0 it\nkeeps older whole turns verbatim beyond the last three while they fit that\ntail budget; the default, 0, keeps the reference tail, except\nfor the next rule. It counts a run of consecutive assistant messages\nas one turn in every dialect, where the reference does so only for\nResponses. Claude Code can record one step as two assistant messages, tool\ncalls and then text; the Anthropic API merges them, so a split between them\nwould separate the calls from their results. Anthropic requests that declare\na 1M-token window use a separate threshold, --threshold-1m. When the\nprovider rejects a rewritten request for another reason, it resends the\noriginal. When the verbatim head alone approaches the threshold, it raises\nthe threshold instead of compacting every request. Each summary also carries\nthe human's words and the assistant's visible replies from the turns earlier\ncompactions summarized, up to 24,000 characters, where the reference keeps\nonly the turns since the previous compaction; --carry-max-chars 0 restores\nthe reference rule. With calibration on, it\ndivides the threshold by a learned ratio of provider-reported to estimated\ninput tokens, between 1.0 and 2.0, so it compacts earlier when its estimates\nrun low; --no-calibrate restores the reference threshold.

\n

The port's MIT notice is in\nTHIRD_PARTY_NOTICES.md.

\n

The cliff strategy applies the drop rule to a transcript copy instead:\nthe head and the newest keep_recent_turns assistant steps stay\nbyte-for-byte, older tool results over result_max_bytes are dropped unless they are among\nthe newest keep_recent_tool_outputs (default 8), smaller ones stay, and nothing is summarized or added. A step starts where the\nassistant side resumes after a user prompt or a tool result and includes the\ntool results that answer it. Tool-call signatures and reasoning caps are not\npart of the file transform, because Gobstopper's copy transforms only replace\ntool-result payloads. Codex compacted records count as one result. When both\npasses run at the same cut and both produce a plan, the records dropped from\nthe source and then from the copy are, together, the records a single\ncompaction from the source would drop; one unit test checks this on a\nsynthetic transcript. A copy below the trigger or the minimum savings is not\ncompacted again, so under the default policy the two paths can differ. The\ndropped bytes stay in the vault, not in the copy.

\n
Terminal
gobstopper plan <session> --strategy cliff\ngobstopper eval <session>               # cliff appears beside the other strategies\n
\n

auto does not select cliff; choose it explicitly or through a preset. Run\none proxy per client: chaining CliffCompaction and gobstopper proxy would\ncompact each other's output, and the two have not been tested together. The\ncomparison page at\ngobstopper.sh/compare/cliffcompaction\ncarries the same table.

\n

How Gobstopper compares with Claude Code /compact

\n

Claude Code ships its own compaction. /compact sends the conversation in a\nsummarization request carrying the same system prompt, tools, and history,\nthen replaces the in-context history with the summary the model writes;\noptional focus instructions steer it. Claude Code also compacts automatically\nas the context nears the model's limit, and /autocompact sets how full the\nwindow gets first. The session transcript file keeps the original messages,\nand /rewind can restore the conversation to an earlier checkpoint while its\nsnapshots remain, but nothing previews what the summary will keep or lists\nwhat it dropped. Anthropic's session-management\nguide\ncalls the trade "lossy".

\n

Gobstopper previews the cut on frozen bytes, writes a separate compacted copy\nunder a fresh session ID, and keeps the exact source and candidate bytes in a\ncontent-addressed local vault:

\n
\n\n\n\n\n\n\n\n\n\n\n\n\n\n
Claude Code /compactGobstopper
What it changesThe running session's in-context history, replaced by a summary the model writesA separate copy of a saved Claude Code or Codex session file; the source is never changed. gobstopper proxy compacts each outgoing request and leaves session files alone
Who writes the summaryThe model, in a request carrying the same system prompt, tools, and history plus a summarization instruction; /compact focus text steers itNo model by default: built-in strategies drop or stub stale tool results by local rules, and structured and compacted add a metadata state card
Seeing the cut firstNo preview; the summary is written and applied in one step, and you read what it kept afterwardgobstopper plan, eval, and diff show each strategy's cut on the same frozen bytes before apply writes anything
When it runsOn demand, or automatically as the context nears the model's limit; /autocompact sets how full the window gets firstOn demand over saved sessions; gobstopper proxy compacts outgoing requests over a threshold you choose, which can delay Claude Code's own auto-compaction
Undo/rewind returns the conversation to an earlier checkpoint; file snapshots cover the 100 most recent checkpoints and are swept about 30 days after the session last saved onegobstopper undo restores a vaulted snapshot into a new fork; the source file is never rewritten
What holds the originalsThe session's own transcript file; Claude Code documents that summarizing leaves the original messages in the transcriptA content-addressed local vault stores the exact source and candidate bytes before a copy publishes; search-snapshot and read-snapshot return archived records
Providers coveredClaude CodeClaude Code and Codex session files; the proxy covers any client that speaks Anthropic Messages, OpenAI Responses, or OpenAI Chat Completions with a custom provider address
PriceBuilt into Claude Code; the summarization request consumes usage like any other model callFree and open-source (MIT or Apache-2.0); the built-in strategies and the proxy make no model calls
\n

The two work at different layers. /compact shrinks the live session's\ncontext in place. gobstopper apply writes the compacted copy as a new fork\nand never touches the source, and resuming a copy with a live provider\nrequires separate compatibility testing. For a running session, gobstopper proxy compacts requests over the threshold, which can delay Claude Code's\nauto-compaction. The client still controls its own trigger, and /compact\nstays available. The comparison page at\ngobstopper.sh/compare/claude-code-compact\ncarries the same table.

\n

Integrating with a session runtime

\n

A program that runs agent sessions can ask Gobstopper what to do without\ngiving it transcript access. policy-check takes the numbers the runtime\nalready tracks, such as current context tokens and whether the session is\nactive, and returns an action:

\n
Terminal
gobstopper policy-check --provider codex --context-tokens 300000 \\\n    --session-active --json\n# {"action":"provider_compact","control":"thread/compact/start", ...}\n
\n

policy-check returns a decision; it does not call the provider. The runtime\nmust establish that it controls the selected session and test the provider's\noperation before executing it on its own connection. File preparation publishes\nseparate copies for an explicit resume; provider acceptance is a separate check.

\n

This numeric interface was designed for the session runtime that preceded\nxcb, retired on 2026-09-19. xcb embeds gobstopper-core as a library; see the\nplugin protocol. The\nhistorical integration contract records the\noriginal interface and the obligations of a runtime that uses it.

\n

Provider levers observed in earlier versions

\n

These are protocol notes from earlier provider builds, not an activation grant\nfor the installed version. The qualification matrix\nrecords current status; there are no qualified live native cells.

\n
\n\n\n\n\n\n\n\n\n\n\n\n
levercodexclaude code
auto-compact thresholdmodel_auto_compact_token_limit (config; ≤90% of window)--autocompact <100k–1M> argv
on-demand triggerthread/compact/start (app-server v2)/compact [instructions]
compaction promptcompact_prompt config/compact instructions
tool output captool_output_token_limitnone
live usage streamthread/tokenUsage/updated notificationmessage.usage per turn
transcript store~/.codex/sessions/**/rollout-*.jsonl~/.claude/projects/*/*.jsonl
\n

Codex persists compaction as a compacted rollout record carrying\nreplacement_history, the context Codex loads in place of the earlier\nhistory when it resumes.\ngobstopper's parser and verifier understand that shape, including tool pairs\ninside replacement_history. Synthetic records are experimental, pair-aware,\nchecked before copy publication, and available only behind --experimental-compacted; ordinary\napply uses the portable forked digest representation.

\n

The CLI contains native adapters and a durable operation journal, but release\nbuilds refuse dispatch with native_unqualified, including prior\nauto_compact_closed settings. Debug protocol fixtures require exact synthetic\nexecutable bytes and isolated temporary homes. They do not qualify a live provider.\nIf an earlier attempt is dispatched or unknown, cooldown expiry, changed source\nbytes, and watch restarts cannot automatically replay it. Inspect\ngobstopper native-operations; reconciliation accepts only persisted matching\nCodex terminal identity, never a caller-supplied success flag. See the\nrecovery runbook.

\n

The compatibility settings auto_apply_inplace, auto_apply_store, and\nauto_compact_closed cannot enable these disabled operations. See\nconfig.example.toml for their current meanings.

\n

Layout

\n\n

Benchmark results

\n

Terminal-Bench 2.1 through Claude Code, September 27 and 28, 2026

\n

Three arms ran the 89 tasks of Terminal-Bench 2.1 once each through Claude\nCode 2.1.283 with GLM 5.3 Flash via Vercel AI Gateway: gobstopper proxy\nv0.7.2 at tail 0 and at tail 40, both at a 45,000-token threshold with\ncalibration and carry on, and Claude Code with no proxy. Tasks ran with\nharbor 0.23.0 as x86 images under emulation on one Apple Mac, three at a\ntime. Token counts are provider-reported and summed per trial. Costs are\nprovider-reported, metered through Vercel AI Gateway, for this model; other\nproviders price cache reads differently, and subscriptions are not billed\nper token.

\n

\"Dot

\n

Solved about as often. Share of 89 tasks resolved, with Wilson 95%\nintervals. One trial per arm. Terminal-Bench 2.1 · 89 tasks · one trial per\narm · Claude Code 2.1.283 with GLM 5.3 Flash via Vercel AI Gateway ·\nGobstopper v0.7.2, 45,000-token threshold (default 128,000) · September\n27–28, 2026 · 21 of 89 tail-0 trials may have run an earlier build

\n
\n\n\n\n\n\n\n\n\n
ArmSolved of 89 (95% interval)Total inputCache readsNew inputOutputProvider-reported cost
Gobstopper, tail 0 (the new default)61, 68.5% (58.3–77.2%)84.3M68.7M15.6M2.64M$5.72 ($0.064 per task)
Claude Code, no proxy60, 67.4% (57.1–76.3%)118.6M102.6M15.9M2.71M$6.82 ($0.077 per task)
Gobstopper, tail 40 (old default)59, 66.3% (56.0–75.3%)118.5M98.2M20.3M3.97M$7.97 ($0.090 per task)
\n\n

CliffCompaction's authors report 76.69% for their proxy at about 45,000\ntokens, 73.03% for Claude Code's 200,000-token default, and 70.97% for its\n45,000-token auto-compaction on GLM 5.3 Flash (their figures, on their\nsetup). This run had no CliffCompaction arm and no 45,000-token\nauto-compaction arm, so the two sets of figures are not a head-to-head. The\nbenchmarks page\nhas the paired statistics, per-task cost concentration, and downloadable\naggregate results.

\n

Replays at four thresholds, September 26, 2026

\n

Replaying 24 recorded sessions (12 Claude Code, 12 Codex; 665 million\nestimated tokens) through the proxy engine on main at fdeb099, tail 0\ncut cumulative estimated input by 78% at a 32K threshold, 73% at 64K, 61%\nat 128K and 38% at 256K, with 0 unpaired tool calls in 288 replays. Three\nlarge sessions hold 476 million of the 665 million tokens, so a typical\nsession's cut at 32K is about 46%, and at 128K most Claude Code sessions\nnever cross the threshold. These are estimates at four characters per\ntoken, not billed tokens.

\n

\"Line

\n

Lower thresholds cut more. Pooled cut in estimated input across 24\nrecorded sessions, by threshold. Estimates, not billed · 24 recorded\nsessions (12 Claude Code, 12 Codex), 665M tokens · main fdeb099 · September\n26, 2026

\n

Saved-session retrospective, September 19, 2026

\n

The September 19, 2026 retrospective\nevaluated 729 frozen sessions on one Mac. Portable compacted projected a\n36.4% median reduction across 73 high-context archived Codex roots, with\n76.9% sampled-string retention. Across all 729 sessions, 637 produced no\nplan and the median reduction was 0%. These are offline projections, not\nbilling savings or task-accuracy measurements. The page includes all cohorts,\nlimitations, and downloadable aggregate results and methodology.

\n

A separate paired scored-policy replay\nreused the 114 archived roots with the corrected probe limit. This development\ncomparison is not held-out validation; the original baseline was already known.\nOn the 73\nhigh-context tasks, a 0.35 cutoff increased sampled-string retention from\n77.05% to 80.78% while median projected reduction fell from 36.44% to 34.96%.\nTwo tasks produced no plan, and their retention is derived from leaving the\nsource unchanged. All three registered cutoffs and both disabled baselines are\nreported, and the default has no cutoff. The result is a small trade of\nsize for retention; it does not measure task quality or recommend a cutoff.

\n

Historical live trials, recorded September 17, 2026

\n

The following single-session experiments were recorded in the repository on\nSeptember 17 using earlier builds and workflows. They are separate from the\n729-session retrospective and do not establish current provider-wide savings\nor general task quality. They do not qualify this artifact's activation matrix.\nOne 333k-token Claude session was asked the same\nresume question under four conditions, measuring provider tokens on that turn:

\n
\n\n\n\n\n\n\n\n\n\n
conditioninput tokens on resumeoutput tokensrecalled the standing task?
none (original)312,7221,405yes: reported npm unification and stalled renames
gobstopper elide219,1671,052yes: same standing task, stalled renames
gobstopper compacted220,447621yes: same standing task from the state-card digest
claude --autocompact 10056,300416no: incorrectly claimed the renames were already done and published
\n

gobstopper elide and compacted both cut the resume context by about\n30% while keeping the answer accurate. Claude's native --autocompact 100\ncut the resume context by ~82% but produced a confident, inaccurate\nsummary of the session.

\n

The same question was then asked on a 101k-token Codex session:

\n
\n\n\n\n\n\n\n\n\n
conditioninput tokens on resumeoutput tokensrecalled the standing task?
none (original)101,275244yes: Oh's memory benchmark and the 0.60 expansion gate
gobstopper elide57,98083yes: same 0.545 score and 0.60 gate
gobstopper compacted34,503159yes: same BEAM experiment and expansion gate
\n

On Codex, compacted cut resume input tokens by 66% and elide cut\nthem by 43%, both with accurate answers to that question. Provider-native\nCodex compaction was not included in this historical trial.

\n

The same-session cache_aware A/B (339k-token Claude session, floor 310k,\nreal provider cache counters):

\n
\n\n\n\n\n\n\n\n\n
conditioncache_readcache_creationfile-level prefix preservedcostaccurate?
none (original)10,010325,647n/a$6.53yes
gobstopper cache_aware13,536258,517107,884 tokens$5.19yes
gobstopper compacted13,536257,5056,639 tokens$5.18yes
\n

In this trial, cache_aware and compacted had almost the same API cost and\nboth recorded 13,536 cache-read tokens, versus 10,010 in the baseline.\ncache_aware preserved 16x more identical transcript prefix. That makes the\nrewrite easier to audit; this trial did not establish extra provider cache savings\nfrom the preserved prefix. Current compacted uses a portable forked digest;\nsynthetic Codex-native records require --experimental-compacted.

\n

Snapshots and gobstopper diff make compaction inspectable and provide a\nrecovery path. They do not guarantee that omitted facts are unimportant or\nthat a continuation will retrieve them automatically.

\n

In the same period, a Gobstopper-written Codex compacted record with a\ncorrect window chain was accepted by codex exec resume, and the model\ncompleted a real API turn that recalled the elided commands. In a separate\nClaude Code trial on a 333k-token test copy, Gobstopper elided 43 stale tool\nrecords and injected a state card, and claude --resume succeeded and\nrecalled the standing task. These trials used earlier builds and do not establish current provider resume\ncompatibility. Current apply writes Claude Code compactions to a separate fork.

\n

Status

\n

The correctness audit records established behavior,\nknown defects and evidence limits; the assurance plan\ntracks the remaining work. No whole-system correctness proof is claimed.\nFile-copy publication uses source hashes, no-clobber creation, retained source\nand candidate bytes, structural checks and durable operation receipts. Shared\nvault readers coordinate with pruning, which fails closed on corrupt recovery\nroots. Operation pins have no automatic retirement policy. Transcript processing\ndefaults to 512 MiB and 100,000 records. Process-death fixtures and\nTLA+ vault models, the\nnative dispatch model,\nKani proofs of selected production Rust kernels, and\nLean transcript algebra cover their declared\ninvariants and bounds; the TLA+ models check safety only, without fairness, so they make no eventual-completion claim. The bounded synthetic stress gate\nexercises named fault, restart and process fixtures. The\nverification guide describes reproducible tool inputs and CI\ngates. These checks do not prove that the Rust code implements the TLA+ models\nor the Lean algebra (that link is reviewed, or tested on finite fixtures), arbitrary filesystem power-loss\nbehavior, proprietary provider acceptance, or preservation of every task fact.

\n

Direct provider controls still belong to the live session owner. Synthetic\nCodex compacted records, external model scoring, and semantic editor plugins\nremain explicitly experimental or trusted extension paths. Released\nnative dispatch remains guarded for both providers. In-place and\narbitrary-path rewrite APIs refuse mutation. Deterministic MCP inspection rejects\nexecutable strategies; explicitly invoked extensions remain trusted code rather\nthan an OS sandbox. The activation matrix\nand recovery runbook define the supported modes.\nSee docs/design.md, docs/roadmap.md, and\ndocs/plugin-protocol.md for details.

\n

License

\n

MIT OR Apache-2.0

\n"; +export const readmeLead = "🍬 Gobstopper saves tokens while preserving context. Compaction is the step your harness runs to squeeze a long conversation into the context window. Gobstopper replaces the compaction built into Claude Code and Codex. It proxies API requests, compacts more often to save tokens, and writes the full history to a local database your agents can search. On Terminal-Bench 2.1 it used 29% fewer tokens than Claude Code's built-in compaction and solved the same number of tasks. Ask your agent to set it up: https://gobstopper.sh — Ben Guo"; +export const readmeHtml = "

Gobstopper

\n
\n

🍬 Gobstopper saves tokens while preserving context. Compaction is the step\nyour harness runs to squeeze a long conversation into the context window.\nGobstopper replaces the compaction built into Claude Code and Codex. It\nproxies API requests, compacts more often to save tokens, and writes the full\nhistory to a local database your agents can search. On Terminal-Bench 2.1 it\nused 29% fewer tokens than Claude Code's built-in compaction and solved the\nsame number of tasks.

\n

Ask your agent to set it up: https://gobstopper.sh

\n

— Ben Guo

\n
\n

Gobstopper is free and open source. When a request passes a token threshold,\ngobstopper proxy replaces the older turns with one mechanical summary and\nsends the newest turns word for word, so the provider sees a smaller context.\nThat can delay the agent’s own auto-compaction trigger.

\n

On Terminal-Bench 2.1 through Claude Code, Gobstopper at its default tail\nand a 45,000-token threshold (the default threshold is 128,000) solved as\nmany tasks as Claude Code with no proxy, 61 and 60 of 89, and sent 29%\nfewer input tokens. That is one trial per arm with GLM 5.3 Flash on\nSeptember 27 and 28, 2026, so the solved counts are within single-trial\nnoise; see Benchmark results.

\n

\"Play

\n

The summary rule comes from CliffCompaction, an open-source proxy described\nin a paper by Trang Nguyen, Eulrang Cho, Bingqing Chen, and\nTim Dettmers. gobstopper proxy is a Rust port of it for the three dialects\ncoding agents use: Anthropic Messages (Claude Code, opencode, Crush), OpenAI\nResponses (Codex), and OpenAI Chat Completions (opencode, Crush, Aider,\nGoose, and other OpenAI-compatible clients). Run gobstopper proxy run -- claude to try it on one session, gobstopper proxy install to start it at\nlogin, or see\nCompact live coding-agent requests.

\n

Gobstopper also works on saved Claude Code and Codex session files. Preview a\ncompaction at a context size you choose, prepare a separate smaller copy, and\nkeep the original byte for byte in a local vault. It archives the exact source\nand candidate bytes and checks supported structural properties and protected\nrecent output.

\n

Use gobstopper watch --dry-run to inspect threshold decisions, or prepare a\ncopy with a file strategy. Released CLI builds cannot ask providers to compact,\neven when auto_compact_closed is enabled. Direct provider-store and in-place\nrewrites are disabled. Copy preparation preserves the source; resuming a copy\nwith a live provider requires separate compatibility testing. See the\nprovider support status and\nrecovery runbook. Plugins can add strategies\nand providers when you explicitly trust them.

\n

Website: gobstopper.sh · Compared with Claude Code /compact and CliffCompaction

\n

Quick start

\n

For the Claude Code example, install and sign in to Claude Code first, and\nmake sure claude and curl are on your PATH. Gobstopper forwards your\nclient's authentication; it does not sign you in to the model provider.\nThe temporary proxy applies to the child command only, so you do not need\nto edit your shell's base URL settings.

\n
Terminal
curl -fsSL https://gobstopper.sh/install.sh | sh   # macOS (Apple silicon) and Linux\ngobstopper proxy run -- claude   # one Claude Code session through a temporary proxy\n
\n

On Windows, install from PowerShell with irm https://gobstopper.sh/install.ps1 | iex.\nBoth installers download the latest release for your platform, check its\nSHA-256, and install it for your user only. To build from source instead, run\ncargo install --git https://github.com/hraness/gobstopper gobstopper --locked.

\n

Supported macOS and Linux release installs update automatically\nbefore a command, at most once a day, when no other Gobstopper command is\nrunning. Run gobstopper update to update now, gobstopper update check to\ncheck without installing, or gobstopper update disable to turn automatic\nupdates off. gobstopper update enable restores them. CI, offline replays,\nand versions selected with GOBSTOPPER_VERSION stay fixed. Use\n--no-update or HRANESS_NO_UPDATE=1 to skip a check for one invocation.\nCargo and source builds use their original install command; Windows uses\nthe PowerShell installer. Update-enabled installs also need the\nGitHub CLI (gh) authenticated with github.com; run\ngh auth login before installing. See update behavior.

\n

When the session ends, proxy run prints request and compaction counts.\nA short session can report zero compacted requests: compaction starts only\nwhen a supported request crosses the threshold. If Claude Code cannot start,\ncheck that claude runs directly in the same terminal. For proxy connection\nproblems, follow diagnosis and repair.\nFor\na proxy that stays up, with gobstopper proxy status counters, see Compact\nlive coding-agent requests. For Codex,\nsee Set up Gobstopper for Claude Code and\nCodex.

\n

Why

\n

A coding agent carries earlier context into later requests. In a long\nsession that history fills with old file reads and command output the next\nstep rarely needs, and later requests include it again. When the history\nnears the model's window, the agent asks a model to summarize it. That call\nis itself a large request, the summary can leave out an exact error or\nconstraint, and each later summary summarizes the one before.

\n

\"Diagram:

\n

Without compaction, each later request carries the earlier messages and\ntool output. This diagram shows how repeated context accumulates.

\n

gobstopper proxy keeps each request under a threshold you choose, without\na model call:

\n\n

\"Line

\n

One recorded Claude Code session: 383 requests replayed at a 45,000-token\nthreshold (default 128,000), with calibration on. Counts estimate four\ncharacters per token. Build f4db57e uses the v0.7.2 request engine.

\n

What CliffCompaction's authors report

\n

The CliffCompaction paper reports up to\n50% lower cost at a bounded context, with Terminal-Bench 2.0 scores held or\nimproved, on the Kimi K2.6 and GLM 5.1 models its authors tested. In one run\nthrough Claude Code (GLM 5.3 Flash on Terminal-Bench 2.1, at about 45,000\ntokens of mean peak context), the rule scored 76.69%, against 70.97% for\nClaude Code's own auto-compaction and 73.03% for its default 200,000-token\nsetting. The paper's costs come from a model of perfect prompt caching, not\nmetered bills. The authors also report that the benefit depends on the agent\nand the task, and that it matters only for medium-to-long tasks.

\n

These are the authors' measurements of their own proxy. gobstopper proxy\nshares its summary rule and adds context-retention and configuration\ncontrols, described in How Gobstopper compares with\nCliffCompaction. Gobstopper\nhas not rerun the authors' benchmarks as published; its own Terminal-Bench\n2.1 run is under What Gobstopper has measured.

\n

What Gobstopper has measured

\n

On September 27 and 28, 2026, Gobstopper ran the 89 tasks of Terminal-Bench\n2.1 through Claude Code 2.1.283 with GLM 5.3 Flash via Vercel AI Gateway,\none trial per arm, at a 45,000-token threshold (the default is 128,000).\ngobstopper proxy v0.7.2¹ at tail 0 (--keep-tail-percent 0, the default\nsince v0.7.3) solved 61 tasks, Claude Code with no proxy 60, and tail 40, the old\ndefault, 59. Those counts are within single-trial noise (McNemar p = 1.0\nagainst no proxy; 31 of 89 tasks changed outcome between arms). The tail-0\narm sent 29% fewer provider-reported input tokens than no proxy, 84.3\nmillion against 118.6 million, and almost all of the difference was cache\nreads. Its provider-reported cost for this model, metered through Vercel\nAI Gateway, was about 16% lower ($5.72 against $6.82 over 89 tasks), which\nis not statistically significant (95% interval −32% to +2%). Tail 40 cost\n39% more than tail 0 in total; put the other way, tail 0 cost 28% less (95%\ninterval 1.6% to 46.5% less), and five tasks drive most of that gap. v0.7.3\nmade tail 0 the default. The benchmarks\npage has the\nsetup, per-arm tables, paired statistics, limits, and downloadable\naggregates.

\n

\"Bar

\n

Cache reads account for most of the difference: 68.7M with tail 0 against\n102.6M with no proxy. New input and output were about equal.

\n

The proxy replay studies report estimated\nrequest sizes separately from task results and provider-reported token counts.

\n
Benchmark notes
\n
    \n
  1. 21 of 89 tail-0 trials may have run an earlier build.
  2. \n
\n

Saved sessions

\n

For session files, Gobstopper lets you check the tradeoff before you commit\nto it. You can preview a compaction, compare strategies on the same frozen\nbytes, keep the exact source in a local vault, and recover a specific\narchived record when a copy leaves it out. The built-in strategies use local\nrules and need no model. Optional model scorers change what gets selected;\nthey do not skip the snapshot or the verification step. A running session's\ncontext belongs to the provider process that loaded it, so file compaction\nprepares a separate copy. The published\nstudies report context reduction,\nretention, no-op cases, and limitations separately.

\n

For saved-session edits, Gobstopper keeps the exact source in a local vault before writing a smaller copy, so a compaction is a recorded edit you can recover from rather than a silent loss: the design every Hraness project shares. The thread through hraness follows that design across the projects, and the ALGAL vision states the bet behind it.

\n

Compact live coding-agent requests

\n

gobstopper proxy is a local HTTP proxy that sits between a coding agent\nand its model provider. Each time the client resends its history, the proxy estimates the request\nsize. Past the threshold (128,000 tokens by default), it sends the system\nprompt and the first task verbatim, one mechanical summary of the older\nturns, and the last three turns verbatim. At the default tail of 0, the kept\nturns are exactly --keep-recent; a positive --keep-tail-percent lets the\nsummary and older whole turns fill that share of the room left under the\nthreshold after the system prompt and the first task. The provider then\nreports the compacted size back to the client, which can delay the client’s own\nauto-compaction trigger. The summary rule is\nCliffCompaction's; see How Gobstopper compares with\nCliffCompaction.

\n

\"Diagram:

\n

When a request passes the threshold, the middle becomes one summary.\nGobstopper keeps the start and the last three turns word for word and\nreplaces the middle with a mechanical summary. No model writes it. Your\nfiles and your saved session are not changed.

\n

\"Diagram:

\n

Inside a rewritten request. In the logged part of the tail-0\nTerminal-Bench arm (about 68 of the 89 trials), compacted requests had a\nmedian of 31.5K estimated tokens, against a median of 55K before\ncompaction. The example lines are illustrative.

\n

It speaks the three dialects coding agents use:

\n
\n\n\n\n\n\n\n\n\n\n\n\n
AgentDialectHow to point it at the proxy
Claude CodeAnthropic Messagesexport ANTHROPIC_BASE_URL=http://127.0.0.1:8260
CodexOpenAI Responsesmodel_providers block in ~/.codex/config.toml
opencodeAnthropic Messages or Chat Completionsprovider.<id>.options.baseURL → http://127.0.0.1:8260/v1
CrushAnthropic Messages or Chat Completionsproviders.<id>.base_url → http://127.0.0.1:8260/v1
AiderChat Completionsaider --openai-api-base http://127.0.0.1:8260/v1
GooseChat CompletionsOPENAI_HOST=http://127.0.0.1:8260
\n

The setup for each agent is in docs/proxy.md. Any other\nOpenAI-compatible client that posts to {base}/chat/completions works the\nsame way. Claude Code and Codex routing is live-checked; the Chat\nCompletions dialect is contract-tested against synthetic histories and has\nnot yet been qualified against a live opencode, Crush, Aider, or Goose\nsession.

\n

The summary keeps human and assistant text, keeps tool results of at most 500\ncharacters, and reduces each tool call to a one-line signature. A separate\nbounded carry retains selected original tool results and images with their\ninvocation and labels excerpts. Context retention\ndescribes the limits and controls.\nThe next compaction starts again from the history the client resends and\ndiscards the previous summary, but the human's words and the assistant's\nvisible replies carry forward: each later summary opens with them, oldest\nfirst, up to 24,000 characters, and the oldest text drops out when they no\nlonger fit. Between compactions, requests reuse the same compacted prefix,\nso the provider's prompt cache can match it.

\n

The kept turns hold the files and command output the agent read most\nrecently. The summary omits long results unless the bounded evidence carry\nselects them. By default\nthe proxy keeps exactly the newest --keep-recent turns, as CliffCompaction\ndoes. --keep-tail-percent (0 to 60, default 0)\nkeeps older whole turns too while the summary and the kept turns fit in that\nshare of the room, and leaves the rest for new turns before the next\ncompaction. A higher floor leaves less room before the next compaction, so\nwe expect more compactions and larger requests in between; in replay of 24\nrecorded sessions, tail 40 compacted 369 times against 343 at 32,000\ntokens and 50 against 38 at 128,000 (estimates). In the Terminal-Bench 2.1\nrun at a 45,000-token threshold, tail 40 sent 118.5 million input tokens\nagainst 84.3 million at tail 0 and cost 39% more in total in\nprovider-reported terms (tail 0 cost 28% less, 95% interval 1.6% to 46.5%\nless, with five tasks driving most of the gap), with solved counts within\nsingle-trial noise; see Benchmark results. The carried words keep earlier instructions in view\nafter the summary that held them is discarded; they use at most a quarter of\nthat room, and --carry-max-chars 0 turns carrying off. Anthropic Messages\nrequests that declare a 1M-token context window use a separate threshold,\n--threshold-1m: 256,000 estimated tokens by default, or --threshold if\nthat is higher. The proxy reads the window from the anthropic-beta header,\nwhere Claude Code sends a token starting with context-1m for a model such\nas opus[1m], and never from the model name. Setting --threshold-1m equal\nto --threshold applies one threshold to every request. Keep each threshold\nbelow the point where the client compacts on its own, including any\nclaude --autocompact value.

\n
Terminal
gobstopper proxy run -- claude            # one session through a temporary proxy\ngobstopper proxy serve                    # background proxy on http://127.0.0.1:8260\ngobstopper proxy install                  # owned user service; start at login\nexport ANTHROPIC_BASE_URL=http://127.0.0.1:8260\ngobstopper proxy replay <session>         # what the proxy would have sent; calls no provider\ngobstopper proxy status                   # counters and estimated-token totals, this run and all time\n
\n

The Terminal-Bench run used --threshold 45000; the default 128,000\ncompacts later, and in replays most recorded Claude Code sessions never\nreach it (the replay grid in Benchmark results\ncompares thresholds).\nSee docs/proxy.md for per-agent setup (Claude Code, Codex,\nopencode, Crush, Aider, Goose), proxy install and proxy uninstall,\nchoosing a threshold, and every setting.

\n

Claude Code → local Gobstopper proxy → model provider.

\n

Optional rewrite failures can send the original bytes when policy permits.\nA provider HTTP 400 rejection can trigger another trim for a length error,\nor an original-body retry for another error when the original fits configured\ncapacity.

\n

Flags: --threshold (keep it below the client's auto-compaction point),\n--threshold-1m, --keep-recent, --keep-tail-percent,\n--result-max-chars, --carry-max-chars, --evidence-max-bytes,\n--evidence-max-chars, --context-window, --no-keep-awake, --drop-thinking,\n--no-calibrate, --shadow (log what would change and forward everything\nunchanged), and --strict.

\n\n

On September 25, 2026, gobstopper proxy replay with three kept turns, the\ndefault on that date, over nine recorded sessions on one Mac kept six Claude\nCode sessions, whose recorded requests peaked at 273k to 652k estimated\ntokens, at or under about 127k, and one Codex session that peaked at 242k\nunder about 127k. Two Codex sessions that Codex had already compacted itself\nbegan with heads near 160k and stayed under about 243k. No replayed request\nwas left with an unpaired tool call. These are estimates over recorded\nhistories, not billed tokens or task results.

\n

Longer work, local visibility, and startup recovery

\n

Temporary context budgets

\n

Compacting while an agent is gathering evidence can make it reread material that\nwas removed. For a difficult analysis phase, you or the agent can reserve more\ninput context within a scope bound to your client and its descendants. Declare\ncapacities supported by your route; Gobstopper returns the effective budget\nafter output headroom and client limits.

\n
Terminal
gobstopper proxy run --context-window 1000000 --client-context-window 1000000 --adaptive-context -- claude\n# From inside that scoped session:\ngobstopper context reserve --tokens 500000 --requests 20 --ttl-seconds 1800\ngobstopper context status\ngobstopper context release\n
\n

The larger budget expires by request count or time. Adaptive rescue is off by\ndefault; --adaptive-context enables a temporary increase after repeated reads\nof unchanged evidence that the proxy previously removed. It needs a scope and\nconfigured capacity. See context budgets.

\n

The proxy also keeps a limited collection of original tool results and supported\nimages across repeated compactions. After a restart, it rebuilds that collection\nfrom the history the client sends. Older evidence can still be evicted, and the\nproxy cannot recover material removed by the client's own compaction. These\ncontrols address evidence loss; they do not guarantee that an agent stops looping\nor completes its task. See evidence retention.

\n

Startup, recovery, and direct fallback

\n

gobstopper proxy install starts a user service at login and restarts it after a\nprocess exit. Managed service changes pause new inference with a retry response\nand wait for existing requests to finish. If the controller disappears while\nwaiting, its lease expires and requests reopen. Once a stop has been committed,\nrecovery checks the outcome before reopening. Service changes use no firewall\nrules.

\n

Request parsing, compaction, and status checks have time and resource limits.\nLogging, metrics, and sleep prevention run in background workers so slow optional\nwork does not hold up request forwarding. Configured context limits still apply,\nand unavailable scoped context storage returns a retry response.

\n
Terminal
gobstopper proxy launch --client claude --print  # inspect readiness and route\ngobstopper proxy launch --client claude\ngobstopper proxy launch --client codex --codex-auth chatgpt\n
\n

The launcher checks the proxy before starting a client. Claude Code can use its\nofficial provider directly when the proxy is unavailable and its configuration\npermits that route. Custom upstreams, uncertain authentication, scoped context\nreservations, and configured capacity constraints prevent direct fallback.\nCodex requires a healthy proxy and an existing explicit custom provider pointing\nto it; --codex-auth selects the existing authentication route to check.\nThe launcher does not replay inference or reroute running sessions. Clients\nalready configured with a fixed proxy URL still depend on that listener.\nSee startup and recovery for setup and fallback requirements.

\n

During active inference, Gobstopper requests idle-sleep prevention and releases\nit when inference ends. Closing a lid and forced sleep remain operating-system\ndecisions. --no-keep-awake disables the feature.

\n

Local session data

\n

The proxy records local metadata and provider usage. Inspect requests, attempts,\ncompaction decisions and tool activity, or export the versioned journal:

\n
Terminal
gobstopper data requests\ngobstopper data metrics\ngobstopper data export > gobstopper-events.jsonl\ngobstopper data check\ngobstopper proxy doctor\n
\n

Session data explains the schema, privacy boundaries,\nimports, backups and metric denominators.\nThese controls have functional regression tests; the September 28 benchmark\npredates them and does not measure their effect on task accuracy.

\n

Token use across your agents

\n

gobstopper usage shows your token use across coding agents by day, agent,\nprovider and model, including agents that never pass through the proxy. The\nnumbers come from aicharts, which keeps a daily record\non your computer and uploads nothing. The installer below adds aicharts and\nturns that record on; after another install method,\nget aicharts.

\n
Terminal
gobstopper usage                        # the last 30 days, per agent\ngobstopper usage report --days 7 --csv  # one row per day, agent and model\ngobstopper usage enable                 # collect four times a day\n
\n

Agents can read the same record through aicharts mcp. gobstopper data\ncounts what the proxy saw, so some requests appear in both; read them side by\nside rather than adding them together.

\n

Recoverable history

\n

Before publishing a Claude Code or Codex copy, Gobstopper stores the exact\nsource and candidate bytes in a content-addressed vault\n(~/.local/share/gobstopper/vault/). Snapshots use deduplicated 1 MiB chunks,\nso appended versions reuse\nunchanged prefix storage without creating one filesystem object per JSONL\nrecord.

\n

gobstopper recall --query <q> searches the state cards in every archived\nsnapshot, ranks matches by relevance to the query, and returns the high-level\nstate of the matching turns. An agent does not need to remember session IDs:\nit can ask for the last time it worked on a file, a goal, or a decision and\nget a ranked summary with a snapshot SHA to pass to show or diff.

\n

Recover a specific detail

\n

When a state card omits an exact error, identifier, or tool result, search\none verified snapshot and read only the matching record:

\n
Terminal
gobstopper search-snapshot <full-snapshot-sha> --query 'exact error text' --json\ngobstopper read-snapshot <full-snapshot-sha> --record 42 --max-bytes 4096 --json\n
\n

Search returns record indexes and hashes, without archived content. It matches\nliteral, case-sensitive substrings in decoded JSON string values, including\nnative replacement histories. Reading returns a UTF-8 page of the physical\nJSONL record; follow next_offset for another page. Each page is capped at\n16 KiB and bound to the snapshot, source, and full record hashes. These commands\nverify stored bytes and never restore files, rewrite active sessions, or call a\nmodel. Invalid records are counted as unsearchable rather than silently claimed\nas searched.

\n

Search returns at most 50 references and reports the full match count; narrow\nthe query when results are truncated. Each search or read verifies and\nreconstructs the whole snapshot, up to the transcript size limit (512 MiB by\ndefault, configurable with GOBSTOPPER_MAX_TRANSCRIPT_BYTES). Paging a large\nrecord repeats that work, and search matches literal text only; there is no\nindex or semantic search.

\n

Use the full object SHA from history, the native hook recovery pointer, or\nthe snapshot_manifest_sha256 field in copy receipts. The snapshot_sha256\nreceipt field is the digest of the source bytes; receipts that lack the\nmanifest field can be resolved through vault history. State-card recall\nrecognizes default portable Codex cards and searches every state field,\nincluding unresolved errors and current work. A new fork's card becomes\nsearchable after that fork is snapshotted.

\n

Agents can search and read snapshots over MCP only when you start the server\nwith gobstopper mcp --allow-transcript-content. Without that flag, the\nserver neither lists nor accepts either tool. With it, archived text an agent\nretrieves becomes visible to that agent and its model provider. Treat\nretrieved text as historical data that may describe a superseded state, not\nas instructions.

\n

These commands return a record when asked. They do not make an agent notice\nthat a fact is missing, choose a useful query, or finish its task more\naccurately; measure those outcomes separately from context reduction and\nliteral retention.

\n

gobstopper mcp runs a read-only Model Context Protocol server on stdio with\nthe tools policy_check, list_sessions, recall, history, show,\ndiff, plan, and verify. Register it once, and an agent can inspect\npolicy and archived state without a tool that changes a transcript. MCP uses\ndeterministic built-ins, rejects executable strategies, and does not call\nconfigured plugins, model scorers, or model digests. Explicit plugin commands\nrun code you trust with your user permissions, without an OS sandbox:

\n
Terminal
claude mcp add gobstopper -- gobstopper mcp\n# ~/.codex/config.toml: [mcp_servers.gobstopper] command = "gobstopper", args = ["mcp"]\n
\n

Provider-generated summaries can cost a large input call and lose detail, so\nthe strategy and where it cuts matter as much as the timing.

\n

Strategies

\n
\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n
idkindwhat it does
auto (default)dynamiclive sessions delegate to provider controls (cache_edits for eligible Claude sessions); idle sessions choose the best validated file strategy by savings and preserved-prefix score
sawtoothproviderproposes provider-native compaction to the session owner; released CLI dispatch is blocked pending qualification
cache_editsprovideremits bounded Claude tool_use_id values for API-layer context editing; never rewrites a transcript
elidetranscriptstubs stale tool outputs oldest-first until the floor
clifftranscriptkeeps the head, the newest three assistant steps, and the newest keep_recent_tool_outputs tool results (default 8) byte-for-byte and drops older eligible tool results over 500 bytes; no floor seeking and no state card (see CliffCompaction; for running sessions, use gobstopper proxy)
cache_awaretranscriptelides a tailward stale-output window and injects a bounded state card while preserving the longest practical prefix
compactedtranscriptelides stale outputs and injects the state card; synthetic Codex compacted records require --experimental-compacted
scoredtranscriptranks candidates with deterministic recency, error, reference, TF-IDF, duplicate, and tool-type signals before elision
dedupetranscriptremoves older exact duplicate tool payloads using payload SHA-256, not summaries
microtranscriptkeeps the newest configured outputs per stable tool label and stubs older ones
middletranscriptprotects both ends of the transcript and elides eligible middle outputs
structuredtranscriptemits a bounded metadata-derived state card; it is not semantic summarization
agenticextensionaccepts bounded edit proposals from a command or versioned plugin you trust; Gobstopper still validates every edit
\n

Custom strategies are userspace code: a preset can name a command that\nreceives the normalized transcript as JSON on stdin and returns an edit\nplan on stdout, or install a versioned gobstopper-plugin.json bundle\n(see gobstopper plugin check). A command runs only with\ntrusted_legacy_command = true. Gobstopper checks eligible payloads, protected\nrecent output, edit combinations, digest size, and projected token reduction.\nFile candidates must not introduce supported structural findings. These checks\ncover edit structure and size, not semantic preservation or provider acceptance.

\n

Install & use

\n

Check the release notes when\nyou need a capability tied to a particular release.

\n

Install the latest release:

\n
Terminal
# macOS (Apple silicon) and Linux (x86_64, arm64): installs ~/.local/bin/gobstopper\ncurl -fsSL https://gobstopper.sh/install.sh | sh\n
\n
# Windows (x86_64), in PowerShell: installs to %LOCALAPPDATA%\\Programs\\gobstopper\\bin, no administrator rights\nirm https://gobstopper.sh/install.ps1 | iex\n
\n

On macOS (Apple silicon) and Linux x86_64, install.sh also adds aicharts\nbeside gobstopper, checked against a pinned SHA-256 digest and, on macOS,\nits Developer ID signature. On a first install it turns on local usage\nhistory: daily token totals for your agents, kept on this computer and never\nuploaded. aicharts history disable turns it off. It also turns on\naicharts' daily self-update check, which installs a new release only after\nverifying it; aicharts update disable turns that off, or set\nGOBSTOPPER_AICHARTS_UPDATE=no before installing. Set\nGOBSTOPPER_USAGE_HISTORY=no to leave history off, or GOBSTOPPER_AICHARTS=no\nto skip aicharts.

\n

Set GOBSTOPPER_VERSION=X.Y.Z to install one exact release. The installers\ncheck each download against the release's SHA-256 file; docs/release.md\nshows how to check a download's build provenance attestation yourself. On\nWindows, the vault, apply, watch and the provider hooks are Unix-only and\nrefuse with an error; detect, plan, verify, mcp and the proxy work.

\n

To build main or another platform from source:

\n
Terminal
cargo install --git https://github.com/hraness/gobstopper gobstopper\n# or from a checkout: cargo build --release\n
\n
Terminal
gobstopper proxy run -- claude     # one Claude Code session through the proxy\ngobstopper proxy serve             # background proxy on http://127.0.0.1:8260\ngobstopper proxy status            # requests compacted, estimated tokens saved\n\ngobstopper detect                  # sessions, context sizes, lifetime burn\ngobstopper plan <session>          # what would happen, under which strategy\ngobstopper plan <session> --trigger 100000 --floor 30000    # tune the trade-off\ngobstopper eval <session>          # compare strategies on the same frozen bytes\ngobstopper apply <session> --strategy elide  # Codex/Claude copy; native requests are refused\ngobstopper verify <session>        # supported structural checks (exit 1 on errors)\ngobstopper fork <session>          # clone under a fresh session id + resume cmd\ngobstopper undo <session>          # Codex/Claude: restore a snapshot into a new fork\ngobstopper vault                   # list snapshots in the undo vault\ngobstopper prune                   # preview keeping the newest 10 snapshots per session\ngobstopper install-hooks --output ./hook-candidates.json  # private settings candidates\ngobstopper watch --dry-run         # inspect threshold decisions without preparing copies\ngobstopper watch --dry-run --active-only --once  # bounded recent-session inspection\ngobstopper explain                 # the occupancy model behind the defaults\ngobstopper recall --query <q>      # search state-card digests across all archived sessions\ngobstopper history <session>       # every archived state of one session\ngobstopper diff <sha-a> <sha-b>    # structural comparison of two vault snapshots\ngobstopper bench                   # compare strategies on recently changed sessions\ngobstopper tune <session>          # preview the adaptive trigger/floor for a session\ngobstopper mcp                     # deterministic inspection; executable strategies are rejected\ngobstopper proxy serve             # compact live Claude Code and Codex requests on 127.0.0.1:8260\n
\n

Set up Gobstopper for Claude Code and Codex

\n
    \n
  1. \n

    Install the binary from main and check it:

    \n
    Terminal
    cargo install --git https://github.com/hraness/gobstopper gobstopper\ngobstopper --version\n
    \n
  2. \n
  3. \n

    Register the MCP server with each agent you use:

    \n
    Terminal
    claude mcp add -s user gobstopper -- gobstopper mcp\ncodex mcp add gobstopper -- gobstopper mcp\n
    \n

    Confirm with claude mcp list or codex mcp list.\nAdd --allow-transcript-content after mcp only if the agent should be\nable to search and read archived transcript text.

    \n
  4. \n
  5. \n

    Start the proxy and point each client at it as described in\ndocs/proxy.md.

    \n
  6. \n
\n

Hook installation and removal export candidates without changing provider settings.\nThe bundle includes the exact original settings and hashes, so keep it private.\nAutomatic settings replacement is disabled because Gobstopper cannot obtain\ncustody honored by provider/editor writers. Review and apply candidates through\nprovider-owned settings controls and retain provider trust prompts. Callbacks\narchive source-bound evidence; their session identifiers do not prove which\noperation caused a compaction. See the recovery runbook.

\n

For automation, gobstopper plan <session> --json returns the existing plan\nobject when a plan is available. A successful inspection without a plan returns\na separate JSON result, for example:

\n
{\n  "status": "no_plan",\n  "reason_code": "below_trigger",\n  "context_tokens_before": 100000,\n  "effective_trigger_tokens": 250000,\n  "target_context_tokens": 40000,\n  "min_savings_tokens": 4096,\n  "projected_context_tokens_after": null,\n  "projected_savings_tokens": null\n}\n
\n

The reason identifies the decision actually reached:

\n
\n\n\n\n\n\n\n\n\n\n\n
reason_codeMeaning
below_triggerContext is below the effective policy trigger.
strategy_returned_no_planThe strategy declined; its underlying reason is unknown.
empty_external_editsThe configured command or plugin supplied no edits.
minimum_savings_not_metA proposal fell short of the minimum projected savings.
external_nonreducing_planAn external proposal did not reduce estimated context.
\n

Projections are present only when a rejected proposal supplied them.\nThe target is a policy setting, not a measured\nminimum context size, and projected savings are not billed savings. Invalid\nconfiguration, invalid proposals, and execution failures are command errors.

\n

Gobstopper's copy paths require retained source and candidate bytes before\npublication. File-copy paths publish a separate candidate after structural\nverification. Native dispatch remains guarded even if a policy proposes it;\nstandalone native apply refuses before creating a fork or snapshot. Legacy\ndirect-write flags remain readable but cannot authorize those writes. A real\nwatch pass can archive a source snapshot before reaching the native guard;\nwatch --dry-run does not create that snapshot.\nTelemetry is best effort:\nsuccessful event writes use the gobstopper/compaction-events-v1 schema.

\n

eval and bench freeze each session's source before comparing strategies.\nbench selects sessions updated within seven days by default;\n--all removes that age filter but retains discovery and input limits. Its\n24-column CSV includes source/result hashes, execution_state, token_basis,\nretention availability and a closed failure category. Discovered sessions\nthat fail policy resolution or evaluation remain explicit failed rows with\nunavailable measurements. Parse CSV quoting rather than splitting lines or\ncommas: session identifiers can contain those characters. A provider proposal\nis provider_not_executed; a detached transform is not a resumed provider\nsession. Numeric legacy fields must be read with those state and availability\nfields, not counted as measured zeroes or task success.

\n

Typed-retention experiments (opt-in)

\n

eval-study replays four arms on isolated in-memory candidates: an unchanged\nno_compaction baseline; plain\nobservation masking; typed masking (constraints, procedures, and open tasks\nstay pinned in their original records and roles, and retrieved text never\nbecomes a higher-authority instruction); and typed_digest (pinned records\nare elided but their spans are carried verbatim on an injected state card,\nwhich loses the original record and role just as a summary does). It does not\nchange auto, call a model, emit live compaction telemetry, or modify the\nprovider session. A requested floor may remain unreachable rather than\ndropping a pinned item.

\n
Terminal
gobstopper eval-study /private/source.jsonl --prepare-manifest /private/checks.json\ngobstopper eval-study /private/source.jsonl --manifest /private/checks.json --rounds 10 --trigger 1 --floor 40000 --json\n
\n

Preparation refuses an existing destination and writes only hashes, byte spans,\nJSON pointers, types, and opaque check IDs, not transcript text. Its labels are\nheuristic candidates that nobody has reviewed: at most 16 complete lines per type,\nwith elidable records considered first and source order breaking ties. Reviewed\nmanifests can instead use label_source = "reviewed"; classification coverage\nis not measured by retention. The JSON schema is gobstopper-retention-v1, with\nsource_sha256, label_source, and checks entries containing id, kind,\nrecord_index, pointer, start_byte, end_byte, and sha256 of that exact\nUTF-8 span. Types are constraint, procedure, open_task, fact, preference,\nand episode. Only the first three are pinned. Source identity, live context,\ntext-only pointers, span boundaries, duplicate IDs, and hashes are checked\nbefore any replay. Limits: 64 MiB of source for replay (512 MiB, the vault\nlimit, for score-only manifest prep and --against audits),\n1 MiB of manifest, 256 checks, 4 KiB per span, and 1–10 rounds.

\n

The report separates text presence, same-origin presence, and preservation at\nthe original source record/pointer. lexical_retained is a paraphrase-sensitive\nmiddle tier: a check counts when ≥75% of its normalized content tokens\n(lowercase alphanumeric, ≥4 chars, stopwords removed) appear together in one\nlive slot. That helps when a provider summary rephrases rather than repeats,\nbut it measures token coverage, not semantic equivalence. by_kind holds [total, source-bound retained, lexical retained]; the elidable subset is reported\nseparately. Dead\nbranches and metadata cannot satisfy a check. Pre-existing source verification\nerrors and newly introduced errors are counted separately. Counts are not\nsemantic or behavioral scores. Estimated context uses adapter item estimates, not stale provider usage records\nor billing. All arms use the same policy, including minimum savings and the\nprotected recent tool-output tail.

\n

--against AFTER switches to a score-only realized audit: the manifest binds\nto the session's before-state and retention is scored against independent\nafter-bytes, with no replay and no mutation. Either spec may be a vault:<sha256>\nsnapshot reference. scripts/retention-audit.py scans the vault for consecutive\nsnapshots whose provider compaction-marker count increased (Claude\ncompact_boundary, Codex "type":"compacted"; hook bracket labels alone can\nmiss the actual write), pairs surgery-labeled snapshots with the next\nsnapshot, and runs the audit over each pair. The result is realized, per-kind\nretention of compactions that already happened, including provider-native\nones.

\n

Without new work, replay is explicitly static_stress; unchanged passes do not\ncount as applied compactions. For Codex/Claude fixtures, optional growth\nentries (after_round, records) append complete provider records between\nrounds and are verified before use. Checks still refer to the initial source;\nthis is not a test of revised tasks, independent tasks, or agent reasoning.\nProvider-native compaction, semantic summarization, continuation success, cost,\nand retrieval are not measured, and the report does not score them as\nsuccessful or free. The built-in structured strategy is not used as a\nsubstitute for a semantic summarizer.

\n

A pilot can freeze up to eight selected session exports and register its\nprotocol before outcomes. Choose a new private output directory outside Git:

\n
Terminal
python3 scripts/compaction-study.py --binary target/release/gobstopper --output /private/new-pilot --session SESSION_ID\n
\n

It pins the executable, source exports, annotation manifests, and hashes;\nkeeps content private; uses isolated config/telemetry paths; and checks that the\nfrozen inputs remain unchanged. There are no provider calls. Commands have\noutput/deadline limits and the study has a 900-second overall deadline.

\n

The separate synthetic provider probe makes at most three\nClaude commands, capped at $0.25 each, using an isolated configuration directory,\nno tools, safe mode, and no MCP servers. It requires explicit opt-in and stops\nif that isolated profile is not authenticated; it never copies credentials.\nIt checks for a persisted native compaction boundary before testing recall.\n--seed-style baseline uses explicit test framing; --seed-style naturalistic\nembeds the identical facts in a plausible work narrative; constraints makes\nthe seed rule-dense; pinned keeps the rules out of the transcript entirely:\nthey ride in --append-system-prompt, the provider's own pinned-context\nchannel (safe mode disables CLAUDE.md discovery), while conversational facts\nstill go through the summarizer. claude_md exercises the production pin\nchannel instead: the same rules land in a workspace CLAUDE.md and the arm\ndrops --safe-mode so project memory loads (the isolated config home and\nscratch workspace remain the boundary). Rule-bearing styles add a rules[]\nrecall scored per-marker as constraint_rules_recalled. Recall is scored twice:\nstrict exact match (recall_checks_passed) and containment\n(recall_checks_lenient), so a semantically preserved superset answer is not\nindistinguishable from a lost fact. After interactive login in that isolated\nprofile, a fresh probe output directory can reuse it with\n--auth-home /private/previous-probe/claude-home:

\n
Terminal
python3 scripts/provider-retention-probe.py --claude-bin /absolute/path/to/claude --output /private/new-native-probe --allow-provider-calls\n
\n

A passing synthetic probe is not a four-arm real-session comparison or evidence\nof billed savings. These commands do not change the strategies the watch\ndaemon uses. Design references: Knowledge Triage,\nThe Complexity Trap,\nSelfCompact,\nACON, and\nLongMemEval.

\n

Config: ~/.config/gobstopper/config.toml

\n
[policy]\nstrategy = "auto"\ntrigger_tokens = 250_000\nfloor_tokens = 40_000\nmin_savings_tokens = 4_096   # reject ineffective plans\nadaptive = true              # derive trigger/floor per session; see `gobstopper tune`\n\n[provider.codex]             # per-provider overrides\ntrigger_tokens = 200_000\n\n[sessions."01a08d7c-…"]      # per-session overrides\nstrategy = "structured"\ntrigger_tokens = 120_000\n\n[presets.deep-work]          # named presets, selectable via --preset\nstrategy = "elide"\ntrigger_tokens = 150_000\n\n[presets.cliff]              # CliffCompaction's rule on a transcript copy\nstrategy = "cliff"\nkeep_recent_turns = 3        # newest assistant steps kept byte-for-byte\nresult_max_bytes = 500       # older tool results above this are dropped\nkeep_recent_tool_outputs = 0 # no extra protected result tail\n\n[presets.custom-script]      # legacy userspace code preset\ncommand = "python3 ~/bin/my_compactor.py"\ntrusted_legacy_command = true\n\n[discovery]\nmax_age_secs = 604800        # rolling window for `watch` and `report`;\n                             # 0 = every session regardless of age\n
\n

For sessions stored outside the default directories, such as in a sandboxed\nhome, pass --codex-home or --claude-home.

\n

Monitoring an existing Codex desktop session

\n

Standalone watch cannot compact the context already held by another Codex\nprocess. It reports native delegation as skipped, with zero credited savings;\nthe program running that session has to request the compaction. Without\n--active-only, watch and report consider sessions active within\n[discovery] max_age_secs (7 days by default); watch --max-age and\nreport --max-age/--all override it. --active-only limits discovery\nto files updated within the last 180 seconds (a recency heuristic, not proof of\nan owning process), and --once exits after one pass. A dry run writes no forks\nor compaction events. Installed provider-managed lifecycle hooks can archive\nobservations and provide a recovery pointer. They do not establish an applied\nGobstopper operation or a matched before/after pair.

\n

The optional local monitor records numeric observations\nfor an explicit list of sessions and checks a deterministic dry-run watcher.\nIt separates observed context drops, native hook activity, and projected\ncompaction plans; none is automatically counted as Gobstopper-caused usage\nsavings. Current Codex event_msg/token_count accounting and legacy usage\nrecords are both supported, including the advertised model context window.

\n

scored uses the deterministic offline heuristic by default. Experimental\nmodel scoring is opt-in with GOBSTOPPER_SCORER=llm,\nGOBSTOPPER_SCORER=clef, or GOBSTOPPER_SCORER=apple; merely setting an API\ntoken never sends data. Remote Clef and LLM scorers receive bounded labels and\nsummaries, including tool arguments, short output tails, and user-prompt\nsnippets. These are transcript-derived text, not redacted metadata; enabling\na remote scorer sends them to its configured endpoint even when additional\ncontent excerpts are disabled. Apple scoring runs on-device. All model\nscorers retain deterministic heuristic scores whenever a model omits an\nanswer or a request fails. Hosted LLM settings are hard-capped\nat 256 candidates, 64 items per batch, 16 batches, and a 100–30,000 ms\ntimeout. The built-in heuristic is the recommended default because the\nrecorded live trials did not show a better plan from the LLM scorer.

\n

The scored strategy also supports an optional keep-score cutoff. For example,\nadd this named preset to your config:

\n
[presets.retained]\nstrategy = "scored"\nkeep_score_threshold = 0.5\n
\n

Inspect it with gobstopper plan <session-id> --preset retained; use the same\nflag with gobstopper apply to create a fork. Candidates\nat or above the cutoff are preserved, even if that prevents reaching the token\ntarget. Missing, invalid, or duplicate candidate scores are also preserved.\nThe value must be finite and between 0.0 and 1.0. The heuristic score is a\nranking signal, not a calibrated probability; 0.5 is an experimental example,\nnot a tuned recommendation. Model scorers still use heuristic fallback for\npartial or failed responses, so the cutoff does not guarantee provider\nconfidence. Set strategy = "scored" explicitly: auto, compacted, and other\nstrategies ignore the cutoff. Omitting it in a later configuration layer\ninherits an earlier value rather than clearing it. The default configuration\nhas no cutoff.

\n

GOBSTOPPER_SCORER=clef uses Cloudflare Clef\non Workers AI. It returns typed noul keep-probabilities instead of generated\nprose. Set CLOUDFLARE_ACCOUNT_ID to your 32-character hexadecimal account ID\nand CLOUDFLARE_API_TOKEN to a token with Workers AI access for that account.\nCLOUDFLARE_AUTH_TOKEN is an alternative token variable. The endpoint is fixed\nto https://api.cloudflare.com/client/v4/accounts/{account}/ai/run/@cf/cloudflare/{model};\nGOBSTOPPER_CLEF_MODEL selects clef (default) or clef-flash.

\n

Cloudflare credentials come only from CLOUDFLARE_API_TOKEN, then\nCLOUDFLARE_AUTH_TOKEN. Gobstopper does not read credential files, the clipboard\nor OS credential stores for Clef. The retired GOBSTOPPER_CLEF_USE_KEYCHAIN\nsetting has no effect. Existing stored entries are left untouched.

\n
Terminal
gobstopper auth clef --status     # environment source and local checks; no model call\n
\n

auth clef without --status and auth clef --delete refuse token storage\nand removal without reading stdin or accessing a credential store. Status\nchecks local configuration only; it does not establish API access or model quality.

\n

Live Jev scoring is retired. Historical measurements and offline artifacts\nremain available under their original names. TypeSafe keys, old endpoints and\nGOBSTOPPER_JEV_* settings are not accepted by Cloudflare. Select clef\nexplicitly and configure Cloudflare environment credentials;\nGOBSTOPPER_SCORER=jev falls back to the heuristic with a warning.\nRemote scorers require curl 8.3+: bearer credentials are imported from a\nchild-only environment variable and expanded inside curl, never placed in\nprocess argv. Gobstopper disables .curlrc for these calls so user defaults\ncannot enable verbose header logging.\nGOBSTOPPER_CLEF_CONTENT_BYTES (default 0, maximum 1024) opts in to\nattaching additional bounded per-candidate content excerpts to each question.\nThe remote scorer already receives the bounded labels and summaries described\nabove. A historical Jev 141k-token A/B run selected the same six records with and\nwithout 400-byte additional excerpts, so those excerpts remain off by default.\nEvery numeric runtime knob is clamped: 1–64 questions per call, 1–128 state\nitems, 100–30,000 ms timeout, 1–16 batches per scoring pass, and 1–4\nconcurrent calls (GOBSTOPPER_CLEF_PARALLEL, default 2).\nGOBSTOPPER_CLEF_MAX_BATCHES defaults to 4. Only the newest\nMAX_Q × MAX_BATCHES tailward candidates are sent; an older prefix keeps its\ndeterministic heuristic score. This caps the default at four calls and 256\nremote-scored candidates even for unusually large transcripts. Each logical\nbatch sends at most one HTTP POST. Redirects, HTTP failures (including 5xx),\ntransport failures and malformed responses are never retried automatically.\nA failure does not prove the provider did no billable work.\nIdentical\nquestion texts within a pass are asked once: repeated tool outputs share a\nsingle remote answer instead of being billed per item. Batches run through\na bounded worker pool: execution is parallel, but results are overlaid in\nstable order, a failed or panicked batch retains heuristic scores, and\ncached answers still overlay when a batch's remote half fails. Each pass logs one summary line\nto stderr (candidates, unique/cached/sent questions, calls, failures,\nelapsed). In a historical Jev trial recorded on September 19, 2026, one three-batch\n336k-token Claude session took 148.64s before bounded parallelism and 19.47s\nafterward (about 7.6×). This single-session latency observation does not\nestablish current or provider-wide performance.

\n

eval and bench honor GOBSTOPPER_SCORER for their scored row, so\nan A/B run measures the same Clef or Apple ranking used by plan rather than\nsilently substituting the heuristic. GOBSTOPPER_EVAL_JUDGE=clef adds a\nseparate model-judged retention estimate to eval. Verbatim survivors are\ncredited locally; only up to 64 sampled strings absent from the rewritten live\ncontext become typed noul questions. They are judged against at most 100,000 bytes of bounded\ncompaction evidence (state cards, elision stubs, and short tool records, with\na head+tail fallback). An intact candidate costs no judge request. The judge is\noff by default because missed-fact evaluation sends that bounded evidence to the\nremote API. Invalid or missing answers remain unmeasured; coverage is explicit\nthrough probes_requested, probes_total, complete, recall_available and\nbasis. A partial denominator must not be compared as if it covered every\nprobe. Restricting the run to --strategy scored creates at most one logical\njudge request, with no automatic retry:

\n
Terminal
GOBSTOPPER_SCORER=clef GOBSTOPPER_EVAL_JUDGE=clef \\\n  gobstopper eval <session> --strategy scored\n
\n

Gobstopper checks Cloudflare's success, result and errors response,\nthe returned model, question IDs and typed probabilities. Clef starts from\nthe complete deterministic heuristic ranking and\noverlays only valid remote answers; capped candidates, missing answers, and\nfailed or malformed chunks keep their heuristic scores. Semantic eval omits\nunavailable answers instead of crediting unknown facts. These estimates do not\nestablish semantic equivalence, instruction authority, or task success. In one\nhistorical 80k-token A/B run, Jev chose five smaller records where the heuristic chose four larger\nones, reclaimed about 649 more tokens, and both retained all 38 extracted\nprobes. At a more aggressive floor, one probe lost verbatim was not falsely\ncredited by the semantic judge (37/38 on both scores). This is one session,\nnot a general measure of ranking quality.

\n

Successful Clef answers are cached in-process for five minutes under two\nscopes. The scorer caches each question together with the complete scoring\nstate and model identifier. Unchanged polls reuse judgments, including across\ndifferent question batches. A changed goal or tail requires fresh answers\nwhen that change is represented in the bounded scoring state; the cache cannot\ndetect task changes omitted from that state. The eval judge caches whole exact\nrequests.\nCache keys cover endpoint, credential identity, and question plus scoring\nstate or full request text; values are\nparsed probabilities only (never transcript text), evict oldest-first at\n512 questions and 64 requests, and failures are never cached. The\nquestion layer also persists to\n~/.local/share/gobstopper/clef-cache.json as sha256 key digests mapped to\na probability and timestamp, never text, so a cold plan inside\nthe TTL can reuse an answer when its question and scoring state match.\nThe version 2 disk format rejects malformed, duplicate, future-dated and\nout-of-range entries. Publication uses a new private temporary file, sync and\natomic replacement; concurrent writers can lose reusable entries, causing a\nfresh request, but do not publish a partial image. Older cache formats are\nignored. clef and clef-flash are provider model names, not weight hashes;\nmatching context and TTL do not prove the remote model stayed unchanged.\nClef inference has no automatic retries or redirect following. HTTP 5xx,\ntransport failures and malformed responses may follow billable provider work;\nGobstopper returns the failure or keeps heuristic scores without another POST.\nThe configured timeout covers the single attempt and its subprocess cleanup.\nThe scorer and judge resolve environment credentials once per process. Set\nGOBSTOPPER_CLEF_CACHE=0 or GOBSTOPPER_CLEF_CACHE_TTL_SECS=0 to disable\nboth layers;\nGOBSTOPPER_CLEF_CACHE_PATH relocates the disk file;\nthe TTL maximum is 3,600 seconds.

\n

For a separate hosted decision, write a JSON evidence file and run\ngobstopper decide evidence.json. This command sends only the file's state,\nquestions and optional images to Cloudflare, and prints the validated model,\nanswers and usage. It does not discover sessions, capture screenshots, compact\ntranscripts, or replace provider originals. It is not exposed through the\nread-only MCP server. The file's model defaults to clef; use clef-flash\nexplicitly to select the other endpoint.

\n
{\n  "model": "clef",\n  "state": "The selected build report shows two failing tests.",\n  "questions": {\n    "investigate": {\n      "type": "noul",\n      "instructions": "Does this report need investigation?"\n    }\n  }\n}\n
\n

Questions require instructions and IDs of 1–100 ASCII letters, digits, _,\n. or -. A request has 1–64 questions. choice uses a criteria object with\n2–255 options; score uses an ordered array of 2–10 levels. Returned choices\nmust match the requested options and highest reported probability. Score legends\nmust match the rubric. Gobstopper allows rounding to four decimal places: there\nmust be a distribution summing to one within the reported probabilities' rounding\nintervals, clipped to 0–1. Its probability-weighted score must fit the reported\nscore's rounding interval within the rubric's range. Returned values are not\nrenormalized. A separate decide request requires all answers; the scorer and eval judge\nkeep their fallback and missing-evidence behavior described above.

\n

Images must be embedded PNG, JPEG or WebP, either as a\ndata:image/png;base64,... string (with the appropriate MIME type), or an\nobject with content_type and base64. Only images you put in this evidence\nfile are sent. Images are never automatically attached to scoring or eval.\nRemote URLs, unsupported formats and malformed images are rejected locally.\nLimits are four images, 4 MiB per image file after base64 decoding, 8 MiB total,\n16 megapixels per image and 13 MiB for the JSON request. Image decoding has a\n128 MiB allocation limit. Oversized evidence is rejected rather than shortened. Credentials, images and state are not included\nin diagnostics.

\n

GOBSTOPPER_SCORER=apple (macOS 26+, Apple Silicon) scores on-device with\nApple Intelligence Foundation Models via the shared apple-foundation\nbridge with no remote API key. It needs a small helper program, built once on\nyour Mac:

\n
Terminal
gobstopper apple install   # builds ~/.local/share/gobstopper/apple-bridge (about 10 seconds)\ngobstopper apple status    # says whether Apple's model is ready, and what to do if not\n
\n

apple install needs Xcode 26 or Apple's command line tools. It checks for\nthem first: when they are missing it prints xcode-select --install and stops,\nso the macOS install dialog never appears unannounced. Compiler output goes to\napple-bridge-build.log next to the helper instead of your terminal. Set\nGOBSTOPPER_APPLE_BRIDGE to install to, or use, a different path. A helper\nnamed apple-bridge next to the gobstopper binary is used before the one in\n~/.local/share, and apple install rebuilds that one when it exists. Scoring\nand plan never build the helper themselves.

\n

When Apple's model can't be used, the scorer says why once and uses the\nbuilt-in scorer: Apple Intelligence is off, the model is still downloading,\nthis Mac can't run it, macOS is older than 26, or the helper isn't installed.\nEach message names the fix and, where there is one, the System Settings pane\n(for example open x-apple.systempreferences:com.apple.Siri-Settings.extension\nfor Apple Intelligence & Siri). gobstopper apple status prints the same\nmessage on demand and exits 1 until the model is ready; --json gives the\nreason code.\nUncached requests are serialized, each using one bounded --once process with\nguided JSON output and owned process cleanup. Failure retains heuristic scores.\nGOBSTOPPER_APPLE_TIMEOUT_MS, _MAX_CANDIDATES, _BATCH_SIZE, and\n_MAX_BATCHES tune it, hard-capped at 100–600,000 ms, 256 candidates, 64\nlabels-only items per batch (8 with content), and 16 batches. Since inference\nis local, the scorer also reads a bounded excerpt of each candidate record\n(GOBSTOPPER_APPLE_CONTENT_BYTES, default and maximum 400; 0 restores\nlabels-only scoring) and shrinks its default batch sizes to fit the ~4k-token\ncontext window. Each batch's guided schema contains one required p_<id> field\nper candidate, so omitted or duplicate array IDs cannot silently distort the\nranking; a malformed batch retains its heuristic scores. Excerpts come from one\nbounded source image whose normalized eligibility and payload fingerprints must\nmatch the plan input; changed or ambiguous sources fall back to the heuristic.

\n

GOBSTOPPER_DIGEST=apple goes further: the injected state card is written\nby the on-device model instead of keyword extraction. Because inference is\nlocal, it may read bounded excerpts of the records being elided without a\nremote request. Each field still\nlands in the same DigestBlock shape via guided output, capped to a small\ntoken overhead, and falls back to the mechanical card on any failure, saying\nwhy on stderr the same way the scorer does.\nGOBSTOPPER_APPLE_DIGEST_ITEMS, _ITEM_BYTES, and _TOTAL_BYTES tune the\nexcerpt budget, hard-capped at 32 records, 2,048 bytes per record, and 16,000\nbytes total; zero disables the model digest and preserves the mechanical card.

\n

The same call also writes a one-line stub per excerpted record, such as\nScript completed Wall time 4.4 seconds, stored in the elide edit's\nper_item_stubs map and rendered verbatim in place of the {bytes}/{kind}\ntemplate where the payload was removed. Records the model did not cover\nkeep the generic stub; invalid or oversized stubs are dropped by validation.

\n

Apple requests are cached in-process by task, instructions, schema, prompt and\nbridge binary identity. Only validated complete scorer batches or validated digest fields/stubs enter\nthe cache. This identifies the exact submitted bounded input, not omitted\nsource context or opaque model weights.\nModel weights and OS inference internals remain opaque; a binary hash does not\nattest their identity. An identical admitted request can reuse its recorded\nresponse without another generation.\nThe cache is bounded at 64 entries with oldest-first eviction;\nGOBSTOPPER_APPLE_CACHE=0 disables both reads and writes. Scorer diagnostics\nreport cached batches separately from real model calls. The savings gate\nalso prices the residual stub text left behind by elision so\ncontext_tokens_after doesn't overstate reclaim.

\n

With adaptive = true, the effective trigger/floor are re-derived per\nsession at each decision point: the trigger is capped at a quarter of\nthe provider-advertised context window, backed off (bounded 2x) when\nrecent compactions reclaimed too little to be worth a cycle, and\ntightened when most of the window is reclaimable tool output. The\nadjustment is deterministic and its reasons appear in plan output and\ntelemetry. gobstopper tune <session> previews it.

\n

How Gobstopper compares with CliffCompaction

\n

CliffCompaction is\nan open-source (MIT) API proxy for coding agents by Trang Nguyen, Eulrang Cho,\nBingqing Chen, and Tim Dettmers, described in\narXiv:2609.26779 (September 2026). You\npoint an agent's base URL at it. When a request exceeds a token threshold, the\nproxy sends the system prompt and task verbatim, then one mechanical summary of\nthe older turns, then the last three turns verbatim. The summary keeps tool\nresults of at most 500 characters, drops longer ones because the files behind\nthem are still readable, reduces tool calls to one-line signatures, and keeps\nassistant text. Each later compaction is rebuilt from the original history the\nagent resends, and the previous summary is discarded: the authors call this\nnever compacting a compaction. Their paper reports up to 50% lower cost at a\nbounded context with maintained or improved Terminal-Bench 2.0 results for the\nKimi and GLM models they tested; those are the authors' benchmark figures, not\nmeasurements of Gobstopper.

\n

Gobstopper extends that summary rule with carried conversation and bounded\noriginal observations from the turns earlier compactions summarized, keeps a\nlarger recent tail only if you set one, and also works on saved session\nfiles:

\n
\n\n\n\n\n\n\n\n\n\n\n\n\n\n
CliffCompactionGobstopper
Where it runsA local HTTP proxy between the agent and the Anthropic or OpenAI APIA local HTTP proxy between the agent and its model provider, plus a CLI over the session files Claude Code and Codex write
ClientsAny client of the Anthropic Messages, OpenAI Chat Completions, or OpenAI Responses APIAny client of the same three dialects that accepts a custom provider address: Claude Code, Codex, opencode, Crush, Aider, Goose, and more
What it changesEach outgoing request, transparently, while the session runsThe proxy rewrites outgoing requests over the threshold; file commands publish a separate compacted copy and leave the source unchanged
How it shrinksDrops tool results over 500 characters, signatures for tool calls, last three turns verbatim; never paraphrasesThe proxy extends the summary rule with bounded original evidence carry and keeps the last three turns verbatim by default, and older whole turns within a tail budget if you set one; file strategies drop or stub stale tool results, and structured and compacted add a metadata state card; no built-in strategy paraphrases unless GOBSTOPPER_DIGEST=apple has an on-device model write the card
RecompactionRebuilt from the original history; the prior summary is discardedThe proxy rebuilds from the original history, and each summary keeps the human's words and the assistant's visible replies from the turns earlier compactions summarized, up to 24,000 characters; cliff on a copy drops the same records as one pass over the source when both passes produce a plan; strategies that inject a state card carry it forward into the next copy
What holds the originalsThe agent's own history and the files on disk; the proxy keeps only an in-memory cache of compacted prefixesFor proxied requests, the agent's own transcript and an in-memory cache; for copies, a content-addressed vault with search-snapshot and read-snapshot
Evidence publishedTerminal-Bench 2.0 and 2.1 (including a run through Claude Code), SWE-bench Verified, and KernelBench results in the paper, on Kimi, GLM, and GPT-5-mini modelsOffline replays of 729 archived sessions, replays of nine recorded sessions through the proxy, one dated afternoon of live proxy counters, literal retention probes, dated single-session trials, and one live Terminal-Bench 2.1 run (89 tasks, three arms, one trial each, September 27 and 28, 2026): resolution within noise of Claude Code alone, 29% fewer provider-reported input tokens at the default tail and a 45,000-token threshold
Model neededNone; the summary is mechanicalNone for the proxy or the built-in strategies; optional model scorers
\n

gobstopper proxy is a Rust port of CliffCompaction's request engine: the\nsame summary format and header, prefix reuse between compactions, harsher\nsettings when one pass leaves a request over the threshold, and a retry when\nthe provider rejects a request for length. It adds context-retention and configuration controls. With --keep-tail-percent above 0 it\nkeeps older whole turns verbatim beyond the last three while they fit that\ntail budget; the default, 0, keeps the reference tail, except\nfor the next rule. It counts a run of consecutive assistant messages\nas one turn in every dialect, where the reference does so only for\nResponses. Claude Code can record one step as two assistant messages, tool\ncalls and then text; the Anthropic API merges them, so a split between them\nwould separate the calls from their results. Anthropic requests that declare\na 1M-token window use a separate threshold, --threshold-1m. When the\nprovider rejects a rewritten request for another reason, it resends the\noriginal. When the verbatim head alone approaches the threshold, it raises\nthe threshold instead of compacting every request. Each summary also carries\nthe human's words and the assistant's visible replies from the turns earlier\ncompactions summarized, up to 24,000 characters, where the reference keeps\nonly the turns since the previous compaction; --carry-max-chars 0 restores\nthe reference rule. With calibration on, it\ndivides the threshold by a learned ratio of provider-reported to estimated\ninput tokens, between 1.0 and 2.0, so it compacts earlier when its estimates\nrun low; --no-calibrate restores the reference threshold.

\n

The port's MIT notice is in\nTHIRD_PARTY_NOTICES.md.

\n

The cliff strategy applies the drop rule to a transcript copy instead:\nthe head and the newest keep_recent_turns assistant steps stay\nbyte-for-byte, older tool results over result_max_bytes are dropped unless they are among\nthe newest keep_recent_tool_outputs (default 8), smaller ones stay, and nothing is summarized or added. A step starts where the\nassistant side resumes after a user prompt or a tool result and includes the\ntool results that answer it. Tool-call signatures and reasoning caps are not\npart of the file transform, because Gobstopper's copy transforms only replace\ntool-result payloads. Codex compacted records count as one result. When both\npasses run at the same cut and both produce a plan, the records dropped from\nthe source and then from the copy are, together, the records a single\ncompaction from the source would drop; one unit test checks this on a\nsynthetic transcript. A copy below the trigger or the minimum savings is not\ncompacted again, so under the default policy the two paths can differ. The\ndropped bytes stay in the vault, not in the copy.

\n
Terminal
gobstopper plan <session> --strategy cliff\ngobstopper eval <session>               # cliff appears beside the other strategies\n
\n

auto does not select cliff; choose it explicitly or through a preset. Run\none proxy per client: chaining CliffCompaction and gobstopper proxy would\ncompact each other's output, and the two have not been tested together. The\ncomparison page at\ngobstopper.sh/compare/cliffcompaction\ncarries the same table.

\n

How Gobstopper compares with Claude Code /compact

\n

Claude Code ships its own compaction. /compact sends the conversation in a\nsummarization request carrying the same system prompt, tools, and history,\nthen replaces the in-context history with the summary the model writes;\noptional focus instructions steer it. Claude Code also compacts automatically\nas the context nears the model's limit, and /autocompact sets how full the\nwindow gets first. The session transcript file keeps the original messages,\nand /rewind can restore the conversation to an earlier checkpoint while its\nsnapshots remain, but nothing previews what the summary will keep or lists\nwhat it dropped. Anthropic's session-management\nguide\ncalls the trade "lossy".

\n

Gobstopper previews the cut on frozen bytes, writes a separate compacted copy\nunder a fresh session ID, and keeps the exact source and candidate bytes in a\ncontent-addressed local vault:

\n
\n\n\n\n\n\n\n\n\n\n\n\n\n\n
Claude Code /compactGobstopper
What it changesThe running session's in-context history, replaced by a summary the model writesA separate copy of a saved Claude Code or Codex session file; the source is never changed. gobstopper proxy compacts each outgoing request and leaves session files alone
Who writes the summaryThe model, in a request carrying the same system prompt, tools, and history plus a summarization instruction; /compact focus text steers itNo model by default: built-in strategies drop or stub stale tool results by local rules, and structured and compacted add a metadata state card
Seeing the cut firstNo preview; the summary is written and applied in one step, and you read what it kept afterwardgobstopper plan, eval, and diff show each strategy's cut on the same frozen bytes before apply writes anything
When it runsOn demand, or automatically as the context nears the model's limit; /autocompact sets how full the window gets firstOn demand over saved sessions; gobstopper proxy compacts outgoing requests over a threshold you choose, which can delay Claude Code's own auto-compaction
Undo/rewind returns the conversation to an earlier checkpoint; file snapshots cover the 100 most recent checkpoints and are swept about 30 days after the session last saved onegobstopper undo restores a vaulted snapshot into a new fork; the source file is never rewritten
What holds the originalsThe session's own transcript file; Claude Code documents that summarizing leaves the original messages in the transcriptA content-addressed local vault stores the exact source and candidate bytes before a copy publishes; search-snapshot and read-snapshot return archived records
Providers coveredClaude CodeClaude Code and Codex session files; the proxy covers any client that speaks Anthropic Messages, OpenAI Responses, or OpenAI Chat Completions with a custom provider address
PriceBuilt into Claude Code; the summarization request consumes usage like any other model callFree and open-source (MIT or Apache-2.0); the built-in strategies and the proxy make no model calls
\n

The two work at different layers. /compact shrinks the live session's\ncontext in place. gobstopper apply writes the compacted copy as a new fork\nand never touches the source, and resuming a copy with a live provider\nrequires separate compatibility testing. For a running session, gobstopper proxy compacts requests over the threshold, which can delay Claude Code's\nauto-compaction. The client still controls its own trigger, and /compact\nstays available. The comparison page at\ngobstopper.sh/compare/claude-code-compact\ncarries the same table.

\n

Integrating with a session runtime

\n

A program that runs agent sessions can ask Gobstopper what to do without\ngiving it transcript access. policy-check takes the numbers the runtime\nalready tracks, such as current context tokens and whether the session is\nactive, and returns an action:

\n
Terminal
gobstopper policy-check --provider codex --context-tokens 300000 \\\n    --session-active --json\n# {"action":"provider_compact","control":"thread/compact/start", ...}\n
\n

policy-check returns a decision; it does not call the provider. The runtime\nmust establish that it controls the selected session and test the provider's\noperation before executing it on its own connection. File preparation publishes\nseparate copies for an explicit resume; provider acceptance is a separate check.

\n

This numeric interface was designed for the session runtime that preceded\nxcb, retired on 2026-09-19. xcb embeds gobstopper-core as a library; see the\nplugin protocol. The\nhistorical integration contract records the\noriginal interface and the obligations of a runtime that uses it.

\n

Provider levers observed in earlier versions

\n

These are protocol notes from earlier provider builds, not an activation grant\nfor the installed version. The qualification matrix\nrecords current status; there are no qualified live native cells.

\n
\n\n\n\n\n\n\n\n\n\n\n\n
levercodexclaude code
auto-compact thresholdmodel_auto_compact_token_limit (config; ≤90% of window)--autocompact <100k–1M> argv
on-demand triggerthread/compact/start (app-server v2)/compact [instructions]
compaction promptcompact_prompt config/compact instructions
tool output captool_output_token_limitnone
live usage streamthread/tokenUsage/updated notificationmessage.usage per turn
transcript store~/.codex/sessions/**/rollout-*.jsonl~/.claude/projects/*/*.jsonl
\n

Codex persists compaction as a compacted rollout record carrying\nreplacement_history, the context Codex loads in place of the earlier\nhistory when it resumes.\ngobstopper's parser and verifier understand that shape, including tool pairs\ninside replacement_history. Synthetic records are experimental, pair-aware,\nchecked before copy publication, and available only behind --experimental-compacted; ordinary\napply uses the portable forked digest representation.

\n

The CLI contains native adapters and a durable operation journal, but release\nbuilds refuse dispatch with native_unqualified, including prior\nauto_compact_closed settings. Debug protocol fixtures require exact synthetic\nexecutable bytes and isolated temporary homes. They do not qualify a live provider.\nIf an earlier attempt is dispatched or unknown, cooldown expiry, changed source\nbytes, and watch restarts cannot automatically replay it. Inspect\ngobstopper native-operations; reconciliation accepts only persisted matching\nCodex terminal identity, never a caller-supplied success flag. See the\nrecovery runbook.

\n

The compatibility settings auto_apply_inplace, auto_apply_store, and\nauto_compact_closed cannot enable these disabled operations. See\nconfig.example.toml for their current meanings.

\n

Layout

\n\n

Benchmark results

\n

Terminal-Bench 2.1 through Claude Code, September 27 and 28, 2026

\n

Three arms ran the 89 tasks of Terminal-Bench 2.1 once each through Claude\nCode 2.1.283 with GLM 5.3 Flash via Vercel AI Gateway: gobstopper proxy\nv0.7.2 at tail 0 and at tail 40, both at a 45,000-token threshold with\ncalibration and carry on, and Claude Code with no proxy. Tasks ran with\nharbor 0.23.0 as x86 images under emulation on one Apple Mac, three at a\ntime. Token counts are provider-reported and summed per trial. Costs are\nprovider-reported, metered through Vercel AI Gateway, for this model; other\nproviders price cache reads differently, and subscriptions are not billed\nper token.

\n

\"Dot

\n

Solved about as often. Share of 89 tasks resolved, with Wilson 95%\nintervals. One trial per arm. Terminal-Bench 2.1 · 89 tasks · one trial per\narm · Claude Code 2.1.283 with GLM 5.3 Flash via Vercel AI Gateway ·\nGobstopper v0.7.2, 45,000-token threshold (default 128,000) · September\n27–28, 2026 · 21 of 89 tail-0 trials may have run an earlier build

\n
\n\n\n\n\n\n\n\n\n
ArmSolved of 89 (95% interval)Total inputCache readsNew inputOutputProvider-reported cost
Gobstopper, tail 0 (the new default)61, 68.5% (58.3–77.2%)84.3M68.7M15.6M2.64M$5.72 ($0.064 per task)
Claude Code, no proxy60, 67.4% (57.1–76.3%)118.6M102.6M15.9M2.71M$6.82 ($0.077 per task)
Gobstopper, tail 40 (old default)59, 66.3% (56.0–75.3%)118.5M98.2M20.3M3.97M$7.97 ($0.090 per task)
\n\n

CliffCompaction's authors report 76.69% for their proxy at about 45,000\ntokens, 73.03% for Claude Code's 200,000-token default, and 70.97% for its\n45,000-token auto-compaction on GLM 5.3 Flash (their figures, on their\nsetup). This run had no CliffCompaction arm and no 45,000-token\nauto-compaction arm, so the two sets of figures are not a head-to-head. The\nbenchmarks page\nhas the paired statistics, per-task cost concentration, and downloadable\naggregate results.

\n

Replays at four thresholds, September 26, 2026

\n

Replaying 24 recorded sessions (12 Claude Code, 12 Codex; 665 million\nestimated tokens) through the proxy engine on main at fdeb099, tail 0\ncut cumulative estimated input by 78% at a 32K threshold, 73% at 64K, 61%\nat 128K and 38% at 256K, with 0 unpaired tool calls in 288 replays. Three\nlarge sessions hold 476 million of the 665 million tokens, so a typical\nsession's cut at 32K is about 46%, and at 128K most Claude Code sessions\nnever cross the threshold. These are estimates at four characters per\ntoken, not billed tokens.

\n

\"Line

\n

Lower thresholds cut more. Pooled cut in estimated input across 24\nrecorded sessions, by threshold. Estimates, not billed · 24 recorded\nsessions (12 Claude Code, 12 Codex), 665M tokens · main fdeb099 · September\n26, 2026

\n

Saved-session retrospective, September 19, 2026

\n

The September 19, 2026 retrospective\nevaluated 729 frozen sessions on one Mac. Portable compacted projected a\n36.4% median reduction across 73 high-context archived Codex roots, with\n76.9% sampled-string retention. Across all 729 sessions, 637 produced no\nplan and the median reduction was 0%. These are offline projections, not\nbilling savings or task-accuracy measurements. The page includes all cohorts,\nlimitations, and downloadable aggregate results and methodology.

\n

A separate paired scored-policy replay\nreused the 114 archived roots with the corrected probe limit. This development\ncomparison is not held-out validation; the original baseline was already known.\nOn the 73\nhigh-context tasks, a 0.35 cutoff increased sampled-string retention from\n77.05% to 80.78% while median projected reduction fell from 36.44% to 34.96%.\nTwo tasks produced no plan, and their retention is derived from leaving the\nsource unchanged. All three registered cutoffs and both disabled baselines are\nreported, and the default has no cutoff. The result is a small trade of\nsize for retention; it does not measure task quality or recommend a cutoff.

\n

Historical live trials, recorded September 17, 2026

\n

The following single-session experiments were recorded in the repository on\nSeptember 17 using earlier builds and workflows. They are separate from the\n729-session retrospective and do not establish current provider-wide savings\nor general task quality. They do not qualify this artifact's activation matrix.\nOne 333k-token Claude session was asked the same\nresume question under four conditions, measuring provider tokens on that turn:

\n
\n\n\n\n\n\n\n\n\n\n
conditioninput tokens on resumeoutput tokensrecalled the standing task?
none (original)312,7221,405yes: reported npm unification and stalled renames
gobstopper elide219,1671,052yes: same standing task, stalled renames
gobstopper compacted220,447621yes: same standing task from the state-card digest
claude --autocompact 10056,300416no: incorrectly claimed the renames were already done and published
\n

gobstopper elide and compacted both cut the resume context by about\n30% while keeping the answer accurate. Claude's native --autocompact 100\ncut the resume context by ~82% but produced a confident, inaccurate\nsummary of the session.

\n

The same question was then asked on a 101k-token Codex session:

\n
\n\n\n\n\n\n\n\n\n
conditioninput tokens on resumeoutput tokensrecalled the standing task?
none (original)101,275244yes: Oh's memory benchmark and the 0.60 expansion gate
gobstopper elide57,98083yes: same 0.545 score and 0.60 gate
gobstopper compacted34,503159yes: same BEAM experiment and expansion gate
\n

On Codex, compacted cut resume input tokens by 66% and elide cut\nthem by 43%, both with accurate answers to that question. Provider-native\nCodex compaction was not included in this historical trial.

\n

The same-session cache_aware A/B (339k-token Claude session, floor 310k,\nreal provider cache counters):

\n
\n\n\n\n\n\n\n\n\n
conditioncache_readcache_creationfile-level prefix preservedcostaccurate?
none (original)10,010325,647n/a$6.53yes
gobstopper cache_aware13,536258,517107,884 tokens$5.19yes
gobstopper compacted13,536257,5056,639 tokens$5.18yes
\n

In this trial, cache_aware and compacted had almost the same API cost and\nboth recorded 13,536 cache-read tokens, versus 10,010 in the baseline.\ncache_aware preserved 16x more identical transcript prefix. That makes the\nrewrite easier to audit; this trial did not establish extra provider cache savings\nfrom the preserved prefix. Current compacted uses a portable forked digest;\nsynthetic Codex-native records require --experimental-compacted.

\n

Snapshots and gobstopper diff make compaction inspectable and provide a\nrecovery path. They do not guarantee that omitted facts are unimportant or\nthat a continuation will retrieve them automatically.

\n

In the same period, a Gobstopper-written Codex compacted record with a\ncorrect window chain was accepted by codex exec resume, and the model\ncompleted a real API turn that recalled the elided commands. In a separate\nClaude Code trial on a 333k-token test copy, Gobstopper elided 43 stale tool\nrecords and injected a state card, and claude --resume succeeded and\nrecalled the standing task. These trials used earlier builds and do not establish current provider resume\ncompatibility. Current apply writes Claude Code compactions to a separate fork.

\n

Status

\n

The correctness audit records established behavior,\nknown defects and evidence limits; the assurance plan\ntracks the remaining work. No whole-system correctness proof is claimed.\nFile-copy publication uses source hashes, no-clobber creation, retained source\nand candidate bytes, structural checks and durable operation receipts. Shared\nvault readers coordinate with pruning, which fails closed on corrupt recovery\nroots. Operation pins have no automatic retirement policy. Transcript processing\ndefaults to 512 MiB and 100,000 records. Process-death fixtures and\nTLA+ vault models, the\nnative dispatch model,\nKani proofs of selected production Rust kernels, and\nLean transcript algebra cover their declared\ninvariants and bounds; the TLA+ models check safety only, without fairness, so they make no eventual-completion claim. The bounded synthetic stress gate\nexercises named fault, restart and process fixtures. The\nverification guide describes reproducible tool inputs and CI\ngates. These checks do not prove that the Rust code implements the TLA+ models\nor the Lean algebra (that link is reviewed, or tested on finite fixtures), arbitrary filesystem power-loss\nbehavior, proprietary provider acceptance, or preservation of every task fact.

\n

Direct provider controls still belong to the live session owner. Synthetic\nCodex compacted records, external model scoring, and semantic editor plugins\nremain explicitly experimental or trusted extension paths. Released\nnative dispatch remains guarded for both providers. In-place and\narbitrary-path rewrite APIs refuse mutation. Deterministic MCP inspection rejects\nexecutable strategies; explicitly invoked extensions remain trusted code rather\nthan an OS sandbox. The activation matrix\nand recovery runbook define the supported modes.\nSee docs/design.md, docs/roadmap.md, and\ndocs/plugin-protocol.md for details.

\n

License

\n

MIT OR Apache-2.0

\n"; export const readmeSections = [{"href":"#quick-start","label":"Quick start"},{"href":"#why","label":"Why"},{"href":"#compact-live-coding-agent-requests","label":"Compact live coding-agent requests"},{"href":"#longer-work-local-visibility-and-startup-recovery","label":"Longer work, local visibility, and startup recovery"},{"href":"#recoverable-history","label":"Recoverable history"},{"href":"#strategies","label":"Strategies"},{"href":"#install--use","label":"Install & use"},{"href":"#how-gobstopper-compares-with-cliffcompaction","label":"How Gobstopper compares with CliffCompaction"},{"href":"#how-gobstopper-compares-with-claude-code-compact","label":"How Gobstopper compares with Claude Code /compact"},{"href":"#integrating-with-a-session-runtime","label":"Integrating with a session runtime"},{"href":"#provider-levers-observed-in-earlier-versions","label":"Provider levers observed in earlier versions"},{"href":"#layout","label":"Layout"},{"href":"#benchmark-results","label":"Benchmark results"},{"href":"#historical-live-trials-recorded-september-17-2026","label":"Historical live trials, recorded September 17, 2026"},{"href":"#status","label":"Status"},{"href":"#license","label":"License"}] as const; diff --git a/site/scripts/readme-html.test.ts b/site/scripts/readme-html.test.ts index 3fc876e..d3742ab 100644 --- a/site/scripts/readme-html.test.ts +++ b/site/scripts/readme-html.test.ts @@ -48,7 +48,7 @@ test("extracts the landing block between the shared Hraness markers", async () = expect(source.indexOf(LANDING_END)).toBeGreaterThan(source.indexOf(LANDING_START)); const landing = readmeLanding(source); expect(landing.title).toBe("Gobstopper"); - expect(landing.lead).toStartWith("🍬 Gobstopper is a tool for saving tokens while preserving context."); + expect(landing.lead).toStartWith("🍬 Gobstopper saves tokens while preserving context."); expect(landing.lead).not.toContain(">"); expect(landing.markdown).toMatch(/cannot ask providers to compact,\s+even when `auto_compact_closed` is enabled/u); });