Grok OSS (grok-oss) is an unofficial open-source fork of
xai-org/grok-build (SpaceXAI’s Grok
Build CLI/TUI), maintained by Surmount.
It is not affiliated with or endorsed by xAI / SpaceXAI. Trademarks and product names belonging to xAI remain theirs.
Why the fork exists: upstream publishes under Apache-2.0 but does not accept external pull requests. This repo accepts community PRs. If upstream ever opens to outside contributions, Surmount intends to open a PR and try to land the useful fork work there.
| Pillar | Practice |
|---|---|
| Faithful | Absorb xAI monorepo exports after review; keep xai-grok-* paths for alignment |
| Complete history | Surmount main is the continuous product archive; xAI is a content feed |
| Open | Pull requests welcome here |
| Distinct | Product Grok OSS, binary grok-oss, clear unofficial labeling |
| Compatible | Config and sessions under ~/.grok (shared with upstream if both installed) |
| Superset | Fork features sit on top of upstream behavior, never hollow out core agent logic |
Normal feature branches → pull request → main. Temporary tool branches
(import/*, onto-xai/*) are not a second main; they land via PR.
On open PRs, catch up with main by merge, not rebase (no force-push
while CI runs). Detail: docs/git-workflow.md.
git remote add xai-org https://github.com/xai-org/grok-build.git # once
# origin → SurmountSystems/grok-oss
# xai-org → xai-org/grok-buildxAI publishes force-pushed snapshots (bot author, often orphan roots, sometimes short “Synced from monorepo” chains). GitHub may say histories are “entirely different.” Expected. Treat them as a tree feed, not shared ancestry.
Maintainer jobs (do not confuse them):
| Job | Command | Result |
|---|---|---|
| Import: their tree into Surmount history | grok-nix-helper import-upstream-export |
import/* review branch → PR to main |
| Stack on tip: our product commits on their tip | grok-nix-helper put-history-on-xai |
onto-xai/* (real cherry-pick; no MODE=overlay) |
Join main into onto: landable graph |
grok-nix-helper join-main-into-onto |
same tip; main becomes ancestor; tree kept (-s ours) → PR |
When histories keep breaking: stack product on their tip, then join
Surmount main (-s ours) so GitHub compare/PR works, then PR to main.
Detect: grok-nix-helper detect-upstream-export or just upstream-detect.
The helper prepares git state. A human TTY signs: git commit -S.
Agents never GPG-sign and never git commit.
GitHub tracking and git-flow (pinned 2026-09-08): Plan Approve → new
GitHub issue with the full plan text. Bug report → new issue on the
appropriate origin, with screenshots. After the operator signs and
pushes, open or update the PR that describes all the work. Collaborative
work uses feat/ / fix/ / docs/ branches from main. Canonical:
docs/github-tracking.md. Dual-pin
AGENTS.md 1b/1c, CONTRIBUTING.md,
docs/git-workflow.md, host ~/.grok/AGENTS.md.
Full process: docs/upstream-history.md
Import log: docs/upstream-import-log.md
Onto log: docs/upstream-onto-log.md
Never: reset Surmount main to xAI; GitHub “Sync fork” that drops Surmount
commits; unsigned commits; bulk tree rewrites without review.
Never assume competence. Do not treat a solution as superior because it comes from SpaceX, from xAI / SpaceXAI, or from a newer upstream export. Never assume someone else has thought everything through.
A more recent upstream version may already solve a problem we have. We also solve problems we find. Given two solutions, the result must always be maximally meritocratic, even if that means synthesizing a better solution from both. That is how a senior engineer solves problems. We balance the constraints with the best knowledge we have.
This is not a license to ignore upstream. Evaluate both. Upstream can win. Our fix can win. A synthesis can win. This is process law for Grok OSS versus xAI / SpaceXAI upstream, not a slur.
A dropped operator prompt is a product defect. Repeating because the
harness lost the text is an engineering miss. The durability path is
the prompt write-ahead log (prompt_wal.jsonl). /rebuild and session
load must preserve that log the same way they preserve an unsent
composer draft and pending_prompts.json, once that file exists in the
tree. User-guide 04-slash-commands /rebuild must say that relaunch
preserves the WAL once the file exists. Do not document that
preservation as shipped while the file is absent.
Named tests are Surmount contracts: quote the owed outcome, keep the
stronger assert, enroll them in
doc/dev/upstream-regression-filters.md
§ Prompt write-ahead log, do not fit them to a wipe. Meritocratic with
§ Hunter's razor. Operator-verified known good (2026-09-02) for
named tests that encode live send, rebuild-flush, interject, and
plan-notes appends, because live session files contained those kinds.
Queue enqueue stays a catalog contract and is not that known-good mark.
Restore and skip tests stay contracts. Do not delete or weaken those
tests in recon, onto, import, or join. This pin does not replace
Hunter's razor, pasted-image tokens, lost-prompt integration tests,
Operator/Agent speaker labels, or the /rebuild persist named-test
list. Dual-pin: AGENTS.md § Wasted human time (hard
constraint 24). Host pointer: ~/.grok/AGENTS.md § The operator's
words are the spec.
Hierarchical: one complete sentence here, then a named fn plus crate, or a
linked doc. This list is the defense inventory for the next upstream merge.
A checkbox is not proof. Import restores FORK_PATHS (docs, scripts, packaging)
only. Product seams inside xai-grok-* survive onto only by cherry-pick plus
named cargo tests.
When you ship a restack-droppable seam: add the one-liner here with the
exact fn name; enroll that filter in
doc/dev/upstream-regression-filters.md;
keep helper-green (substring grok on --version, theme file exists, schema
without /spend ingest, serde without a /settings row, rank helpers without
sampling_config hop keys) out of the land proof. Do not list a catalog
identifier that has no matching fn.
- UDAX JSON→TOON (T0-T6): model-facing structured JSON densifies via
shared
util/toon(GROK_TOOL_RESULT_FORMAT=auto|toon|json). Not a land class. Catalog residual-aligned filters:toon,json_to_toon,densify_mcp,densify_structured,task_output_handoff,subagent_completed_handoff. Detail:doc/dev/research/udax-json-toon-2026-07-26.md - ULID helper:
xai_grok_tools::util::ulidmints 26-char Crockford base32 ids for new work/log/tool artifacts; task UUID v7 unchanged. No landfn. Detail:doc/dev/research/ulid-helper-2026-07-25.md - usage.jsonl append log: fail-open per-session spend log at end of
model turns (
session/usage_log.rs←record_response_token_usage). This is the append log, not/spendingest. Catalog residual:usage_log,record_response_token_usage. Detail:doc/dev/research/usage-jsonl-2026-07-25.md - Last session on start: interactive
grok-osswith a remembered last session for this working directory opens that session, not Welcome. First-ever use stays Welcome. Headless does not steal last-session. Distinct from continue interrupted turn (canceled_turn_resume.json) and from/resume. Land:materialize_new_auto_opens_last_session_when_one_exists(app/session_startup.rs). Siblings (enroll so Welcome / headless cannot regress silently):materialize_new_auto_stays_welcome_when_no_last_session,materialize_new_auto_does_not_open_last_when_headless,from_pager_args_opens_last_session_on_start. - Finished nested sessions must not hang L1 chrome: a parent wait
whose nested id already exited must not stay on
Waiting for the modelwith a climbing timer. Duplicate Human[Image #1]after send is FAIL. Overlay L3 status click must open that specialist view (/dashboardis not/running). Frozen overlay elapsed after child ACP turn-end. Overlay title afterSubagentFinishedequals hostduration_ms(54m34s), not a later spawn-wall clock (1h14m). Reopening a finished overlay must not start a climbing title clock. Live list must drop Responding. L2 finish must surface on L1 without opening the overlay (wait tool completes; silentsubagent-completed-*wake still paints TurnCompleted). Afterinfo.finished,AutoCompactStartedmust not set live compact. A completed nested id missing from the map is not still running.Waiting for the modelis also the live sampler wait. Chrome must distinguish live nested wait, live sampler /TurnRunningwith no completed-wait fallthrough, queuedpending_prompts(1 queued while nested still running), and false wait after nested ids already completed. Do not call that string a hang without evidence./unstickstays operator-invoked; do not auto-fire it on a long live wait. Named tests are Surmount / grok-oss fork contracts:parent_must_not_wait_for_the_model_after_waited_nested_already_completed,waiting_for_the_model_is_not_idle_when_nested_subagent_still_running,waiting_for_the_model_is_not_idle_when_prompt_is_queued,nested_overlay_drops_responding_after_child_acp_turn_completes,nested_overlay_title_elapsed_matches_host_subagent_finished_duration,reopening_finished_nested_overlay_does_not_start_climbing_title_clock,subagents_live_list_drops_responding_after_subagent_finished,l2_finish_surfaces_to_l1_without_opening_nested_overlay,wait_on_completed_nested_id_missing_from_map_does_not_stay_running,silent_subagent_completed_wake_surfaces_turn_completed,auto_compact_started_after_subagent_finished_does_not_set_live_compact_activity,overlay_nested_status_click_opens_l3_session_view,interjection_echo_does_not_duplicate_last_human_prompt,image_interject_leaves_one_prompt_and_empty_queue,after_rebuild_or_resume_plus_plan_exit_follow_up_must_not_wait_for_the_model_with_no_sampler. - Soft process-rule reminders (I):
/settingsprocess-rule strings inject as soft<system-reminder>text on nested spawn (same family aswrite_paths). Spawn still succeeds. A third implementor L2 still spawns. Extra L2s are not auto-killed. Off or empty adds no extra reminder text. Do not hard-cap at two. "only two implementor L2s allowed" is example help copy, not a spawn reject and notMAX_LIVE_L2S=2. Keep such reminders soft for now. Two agents may share a file. Exclusive write is one tool call then release. Diverges from SpaceXAI (upstream has no this/settingslist). Crate:xai-grok-toolsreminders/process_rule_reminders.rs,implementations/grok_build/task/mod.rs,implementations/grok_build/task/admission.rs. Tests:process_rule_reminder_configured_third_l2_still_spawns,process_rule_reminder_text_in_nested_spawn_prompt,process_rule_reminders_off_nested_spawn_has_no_extra_reminder_text,process_rule_reminders_off_injects_no_extra_reminder_text,example_only_two_string_is_not_a_spawn_reject. - Binary / branding is
grok-oss:grok-oss --versionfirst token isgrok-oss, not baregrok(substringgrokis howgrok 1.0.3stayed green). Resume and relaunch hints aregrok-oss --resume. Welcome, tutorial, and hero chrome say Grok OSS. Crate:xai-grok-pagerclient_identity.rs,app/mod.rs,views/welcome/,views/tutorial.rs,docs.rs; bin:xai-grok-pager-binversion_without_tty. Tests:product_cli_name_is_grok_oss,product_version_line_uses_grok_oss_not_bare_grok,resume_session_command_uses_grok_oss,print_exit_resume_hint_writes_expected_lines,user_guide_resume_and_version_examples_use_grok_oss,user_guide_operator_cli_examples_use_grok_oss,welcome_badge_brands_grok_oss,hero_subtitle_brands_grok_oss,tutorial_list_title_brands_grok_oss; pluscargo test -p xai-grok-pager-bin --test version_without_tty. - OpenRouter: separate model option (
openrouter-grok-4.5); login/logout; secret store; optional Zed credential probe (read-only). Neighbor (not class 1):referer_is_surmount_*,title_is_grok_oss. - Multi-key OpenRouter: comma lists / failover keys for credit + rate-limit rotation. FORK claims; not a land class.
- Dual-auth hop after included SuperGrok period limits are full:
SuperGrok is paid. While included SuperGrok period limits still have
room, stay on SuperGrok session (
sampling_configomits console failover). After those included limits are full,sampling_configfills console failover and also switches the API host (SuperGrok proxy ↔api.x.ai). Rank helpers alone are not this class. Crate:xai-grok-shellagent/config_tests.rs. Tests:sampling_config_auto_use_fills_console_hop_after_included_full,sampling_config_auto_use_omits_console,sampling_config_auto_use_omits_console_while_supergrok_included_headroom,resolve_model_to_sampling_config_auto_use,sampling_config_auto_use_dollar_credits_keep_session_console_failover. Per-turn reconstruct:prepare_sampler_for_turn_aligns_to_ranked_included_primary(session/acp_session_impl/sampler_turn.rs). - Any stored SuperGrok login with included remaining before SuperGrok dollar credits:
Any stored SuperGrok identity with remaining included SuperGrok period
limits stays ahead of SuperGrok dollar credits and console. The personal
SuperGrok JWT is the paying identity when personal included SuperGrok
period limits have room. Operators switch SuperGrok identity with
use-personal / use-business on
$GROK_HOME/limits_pins.json. That sidecar is not a new[auth]key. Stockpreferred_method = "api_key"still pins console. After both included pools are exhausted, SuperGrok dollar credits rank before console. Do not flatten remaining to zero fromusagePct/creditUsagePercent100 plus missing SuperGrok Heavy. Prior remaining stays. A memo without a usage reading still forces remaining 0. Used percent below 100 still sets remaining from the percent helper. The client must not invent used-up included SuperGrok period limits from a 100% + missing Heavy snapshot. SuperGrok Heavy ranking optional label is not implemented. SuperGrok Heavy is a real distinct weekly pool. This file does not diagnose product usage of that pool. Rank helpers alone are still not hop. Crate:xai-grok-shellagent/config_tests.rs,session/acp_session_impl/sampler_turn.rs. Tests:sampling_config_hop_team_remaining_personal_exhausted_not_dollars_or_console,sampling_config_hop_personal_remaining_team_exhausted,sampling_config_hop_both_remaining_team_first_then_personal,sampling_config_hop_both_included_exhausted_dollar_credits_before_console,sampling_config_hop_missing_heavy_false_100_keeps_sibling_included,sampling_config_hop_dollar_credits_on_both_missing_heavy_keeps_team,prepare_sampler_for_turn_does_not_flatten_missing_heavy_100_off_sibling,prepare_sampler_for_turn_does_not_flatten_dollar_credits_on_both. Identifierssampling_config_hops_to_sibling_included_before_dollar_creditsandsampling_config_auto_use_dollar_credits_keep_session_console_failovername SuperGrok dollar credits. This file does not claim live Business remaining or a live window hop. - Personal SuperGrok JWT is the paying identity: spend included SuperGrok
period limits on a stored personal SuperGrok login first when those included
SuperGrok period limits have room. A Team / Business SuperGrok JWT is not
the paying source while that personal login exists (that JWT settles as
team postpaid OAuth / Grok Build and can debit the Billing Credits card).
Operators switch SuperGrok identity with use-personal / use-business on
$GROK_HOME/limits_pins.json. No new[auth]keys. Stockpreferred_method = "api_key"still pins console. Team-only login still uses the Team JWT. Crate:xai-grok-shellauth/supergrok_identity_rank.rs,auth/manager_tests.rs,auth/limits_pins.rs. Tests:live_operator_numbers_omit_team_jwt_while_personal_included_and_dollar_credits_remain,personal_included_period_limits_reset_uses_personal_supergrok_not_leftover_business_credits,business_with_no_period_limits_payload_still_switchable_via_use_business,use_personal_switches_back_from_business_pin,pick_prefers_business_included_before_personal_when_both_have_remaining,order_credentials_business_included_before_personal_when_both_have_room,align_to_ranked_free_period_primary_switches_sticky_team_base_to_personal. - Sibling included SuperGrok period limits before SuperGrok dollar credits:
remaining included SuperGrok period limits on a paying SuperGrok identity
(personal JWT) beat SuperGrok dollar credits. Team JWT remaining is not a
SuperGrok paying source while a personal SuperGrok login exists. After-burner
skip only when every paying included pool is exhausted. Tests:
sampling_config_hops_to_sibling_included_before_dollar_credits,afterburner_does_not_skip_mark_when_sibling_has_included_remaining(auth/allowance_exhaust_from_billing.rs). Compact meter stays on included SuperGrok period limits while a sibling pool has remaining:compact_meter_stays_included_while_sibling_pool_has_remaining,active_spend_driver_stays_included_while_any_distinct_pool_has_remaining(views/credit_bar.rs). Combined remaining sums distinct pools and does not double-count a unified pool:combined_included_remaining_sums_distinct_personal_and_business_pools,combined_included_remaining_does_not_double_count_unified_pool,combined_included_remaining_does_not_collapse_matching_percent_and_reset_into_one_pool. Compact remaining / Active driver:matching_percent_and_reset_does_not_collapse_combined_remaining_into_one_pool. - One-process SuperGrok billing flock: one
grok-ossprocess fetches SuperGrok billing; others read$GROK_HOME/limits_snapshot.json. Automatic limits and credits fetch is at most once an hour per machine through that snapshot hub. ForceRefresh still fetches (explicit/limitsandgrok-oss limits). The snapshot never stores JWTs or API keys. Crate:xai-grok-shellauth/limits_snapshot_hub.rs,extensions/billing.rs. Tests:limits_snapshot_second_process_within_the_hour_does_not_http,limits_snapshot_honor_ttl_fresh_within_hour_does_not_http,limits_snapshot_force_refresh_leader_http_fetches_when_snapshot_is_younger_than_one_hour,limits_snapshot_stale_file_lets_waiter_become_leader_and_fetch_once,limits_snapshot_never_writes_access_tokens,billing_handler_uses_snapshot_hub_instead_of_unconditional_sibling_http. - Dual-auth resolve, 429, and credit memo (FORK claims; residual-aligned):
first-party resolve merge (session primary + console failover by default;
preferred_method=api_keyreverses). Identity switch on credit / SuperGrok Heavy usage-limit and plain 429 (FORK claim, not a land class, not live proof). SuperGrok Heavy ranking optional label is not implemented. SuperGrok Heavy is a real distinct weekly pool. This file does not diagnose product usage of that pool. Exhausted-fingerprint memo lives in process cache plus$GROK_HOME/exhausted_credits/(1h TTL; console-key success clears; session success does not). Rate-limit switch uses temporary sharedgrok-rate-limitcooldown, not the credit memo.[auth] auto_use_included_limitsdefaults true on a new/empty Grok home. Catalog residual:resolve_credentials,fingerprint,hop_reason,live_rebind,credit_exhausted,dual_auth_hop_reason. Plans:.agents/plans/plan-secure-key-failover.md,.agents/plans/plan-rate-limit-failover.md,.agents/plans/plan-auth-preferred-roles-failover.md. - Three distinct billing meters: (1) included SuperGrok period limits
(subscription-included quota for the current SuperGrok billing period; how
much of that included quota is already used); (2) SuperGrok dollar credits
(prepaid top-ups on the SuperGrok account); (3) console team prepaid /
console API credits. SuperGrok is paid. Never call SuperGrok free. Desired
spend order: included SuperGrok period limits first, then SuperGrok dollar
credits, then console team prepaid / console API credits. Compact chrome
paints
SuperGrok period · N%for the included-period meter; SuperGrok dollar credits paintSuperGrok dollar credits · $N(live chrome must not nickname that meter); console stillconsole · $N./limits --jsonactiveDriverwire values (supergrok_free_period|supergrok_extras|console_key) stay wire labels after that plain thought. Land paint: class 4 compact-meter tests pluscompact_status_supergrok_on_dollar_credits_shows_dollars_not_free_period_pctandformat_supergrok_session_with_weekly_and_dollar_credits. Dual/limitshonesty (neighbor, not hop): grok-oss limits JSON and compact chrome are a client printout, not xAI billing truth; identicalnextResetor included % across SuperGrok (personal) and SuperGrok (business) is not a shared pool or shared reset clock; operator Usage and console.x.ai Billing win over the CLI;console.isLivefalse is sampler identity, not unused credits; do not invent remaining or call any pool used up. Fail-open: a client printout of included 100%, remaining 0, or SuperGrok dollar credits $0 must not mark SuperGrok used up or hop to console so this session cannot self-fix. Real SuperGrok HTTP 402 after that request failed can still leave SuperGrok. Named commands, same words on TUI/limitsand CLIgrok-oss limits: stay-supergrok, use-console, meter included | dollar-credits | console | combined, refresh (ForceRefresh). Sidecar$GROK_HOME/limits_pins.json, sibling of exhausted_credits/. No new[auth]keys. Stock preferred_method = "api_key" still pins console. Hop-back does not require console credits. Sampler consume of use_console / stay_supergrok is leftover. Tests:user_guide_limits_names_fail_open_and_named_commands,stay_supergrok_clears_false_exhaust_without_console_credits,limits_slash_and_cli_share_stay_supergrok_words.limits_json_lists_two_supergrok_principals_when_both_slots_exist,limits_json_honest_single_supergrok_session_cannot_see_team_plan. Explicit TUI/limitsopen and CLIgrok-oss limitscollect are ForceRefresh. Background FetchBilling is HonorTtl. ForceRefresh without a management key does not clear Management caches. First paint can still be a fresh-by-TTL HonorTtl snapshot. Do not invent live used percent from that file. Tests:management_meter_cache_policy_collect_force_background_honor_ttl,should_clear_management_meter_caches_force_with_key_only(xai-grok-pagerlimits_cmd.rs);limits_snapshot_mode_for_get_billing_explicit_is_force_refresh(xai-grok-shellextensions/billing.rs). - C4 server included-period debit is not a land class: the client
must not invent included SuperGrok period used percent. Optional hard block
is
[auth] allow_spend_when_free_period_debit_unproven = false(or envGROK_ALLOW_SPEND_WHEN_FREE_PERIOD_DEBIT_UNPROVEN=0). Default allows sampler turns under included SuperGrok period limits with loud honesty when the server debit is unproven. Operator ticket:.agents/reports/c4-xai-ticket-paste-ready-2026-08-07.md. - Keyring login time-box + fail-loud: OS keyring get/set/delete wall
clock budget (
KEYRING_OP_TIMEOUT); interactivegrok login --api-key/ OpenRouter login require a secure backend. Only if all secure backends fail: clear error, no silentprovider_credentials.jsonsecret dump. FORK claims; not a land class. Diagnose with in-tree tests, not host D-Bus probes. - Economic mode (cap shipped; slash leftover): nested L2 and L3
sampling stays at the Grok 4.5 long-context price cliff (~200k) at
spawn, model switch, and header. The main (L1) session uses the catalog
500k window. AUTO compact on L1 uses that catalog window, not the old
200k L1 knee. L2 may compact on the nested 200k window. L3 never
compact. An L3 is disposable. If it stalls or spirals, kill it. When
an L3 is near 200k, it summarizes, reports to L2, and stops. Do not
compact-and-continue on L3.
[ui] economic_modestill seeds implement-effort / Token Economy. The Settings setter applies to new sessions./economic-modeis a pager command that queues that text only. The shell has no BuiltinAction arm. Do not list/economic-modeas a live slash or a cargo-proven BuiltinAction. Separate from Token Economy implement-effort caps. Do not claim a Token Economy or economic-mode/settingstable row as cargo-proven (2026-08-15 seams walk did not re-prove those GUI rows). - Token Economy (four pillars;
/spendis the land class): (1) implement-loop effort 1-5 policy; (2) included SuperGrok billing-period linear-burn pacing on/limitsand/usage(never dollar-ize period %); (3)/spendingest ofusage.jsonlintolocal_usage_eventplusreconciliation_run(notDoubleEntryReport::default()); (4) extra SQL ledger$GROK_HOME/grok_oss.db(Token Economy ledger, not the session store, not SuperGrok dollar credits). Config table[token_economy]. Land class 3 tests:spend_path_ingests_usage_jsonl_and_records_reconciliation(xai-grok-shelltoken_economy/mod.rs),show_spend_ingests_usage_jsonl_and_is_not_empty_default(xai-grok-pagerapp/dispatch/tests/status.rs). Schema v1 without ingest is a failed land. - Baked default is Grok 4.6 at medium reasoning effort (fork
contract change; enabled by default;
[models].default_reasoning_effortis the operator override). Test:baked_default_is_grok_46_medium_fork_contract(xai-grok-shellutil/config/persist_tests.rs). - Auto-compact default 95% + live-apply: stock Grok 4.5 catalog omits
a per-model undercut; Settings commit live-applies to open sessions.
FORK claims; not a land class. Detail:
docs/dev/research/rca-auto-compact-early-fire.md - CoT death spiral stop + compact not 75k (GitHub #133): Isolated
Preview looped
Spawn dests of dest encoder skip. I'll spawn dests of dest encoder skip.for 19m18s, then compact paintedContext compacted: 75.2k → 75.2k tokens. Surmount fork of SpaceXAI stream + compact. SpaceXAIx-grok-doom-loop-check/[doom_loop_recovery]resamples confident thinking on Responses (DoomLoopDetectedis retried). Chat Completions never reports those triggers, and visible assistant walls stayed the operator's to cancel. SurmountStreamRepetitionGuardinxai-grok-samplerstream/mod.rsaborts assistant and thought sentence loops asSamplingError::RepetitiveGeneration(Fatal, not retried). Fixture constDEST_ENCODER_SKIP_LOOPis#[cfg(test)]and used by named tests (-D warnings). Compactstrip_repetitive_generationis recovery of that wall in summarizer input, not a second stream breaker. Compact summary capCOMPACT_SUMMARY_MAX_TOKENSis 8192. Compact reseedCOMPACT_RESEED_MAX_TOKENSis 32768 (4 × 8192). After compact,get_total_tokensmust stay in that reserve, not ~75k. Operator reported Oh My Pi does not loop this way. This tree has no Oh My Pi checkout; do not invent internals. Hunter's razor: client Fatal stop plus compact size cap, keep SpaceXAI server resample configurable. Upstream option:[doom_loop_recovery] enabled = falseturns off server thinking resample. Client sentence-loop Fatal stays on (that is the Isolated Preview miss). Named tests:dest_encoder_skip_loop_is_repetitive,dest_encoder_skip_single_sentence_four_times_is_repetitive,thought_line_loop_is_repetitive,chat_completions_stops_dest_encoder_skip_loop,chat_completions_stops_dest_encoder_skip_loop_in_thought,messages_stops_dest_encoder_skip_loop,responses_stops_dest_encoder_skip_loop,classify_repetitive_generation_is_fatal,compaction_reseed_drops_dest_encoder_skip_loop_below_75_2k,compaction_reseed_of_unique_75k_history_must_not_leave_wasteful_75k_context,compact_summary_budget_is_8192_tokens_and_reseed_reserve_is_32768,format_compact_summary_caps_unique_75k_body_to_compact_summary_budget,build_compacted_history_unique_75k_summary_stays_within_compact_summary_budget. - Pasted images are image tokens, not data-URL text: the main
(parent) session model request must not include image content parts.
Files live under the session directory (
images/). History, the prompt write-ahead log, andpending_promptskeep[Image #N]and file ids, notdata:image/...;base64crates. Understanding a paste uses public SpaceXAI Responsesinput_image(jpeg/png data URL or https),detailhigh, nofile_id. That path is not Files API, not Collections, and not Imagine. See xAI image understanding (accessed: 2026-09-07). A nested agent may sendinput_imageon its own request from those session files. The parent spawn prompt is text (paths and[Image #N]), never a data URL. The nested report is text. Compact/recap andcompaction_requestsstill usestrip_images([image]).image_editstill usesAttachedImages.view_imageis the tool path for search-found images only. Count viaestimate_item_tokens(IMAGE_TOKEN_ESTIMATE765) only forContentPart::Image; a parent text-only item after persist does not add 765. Tests:estimate_item_tokens_ignores_data_url_byte_length,strip_images_does_not_serialize_the_data_url_crate,compact_history_does_not_copy_the_data_url_crate,persist_inline_data_url_writes_session_file_and_drops_crate,tool_extracted_image_does_not_leave_data_url_on_parent_conversation,parent_text_only_user_item_after_describe_does_not_add_image_token_charge,read_file_image_does_not_leave_data_url_on_parent_conversation,parent_grok_oss_paste_conversation_request_has_no_image_part,nested_grok_oss_paste_conversation_request_keeps_image_part(xai-grok-shellsession/acp_session_tests/parent_paste_describes_not_inline.rs),drain_interjection_with_images_does_not_attach_image_parts_on_parent,drain_interjection_with_images_attaches_file_parts_on_nested(xai-grok-shellsession/acp_session_tests/interjection_actor_tests.rs),spawn_prompt_string_names_path_and_omits_data_url,nested_user_turn_attaches_file_image_not_data_url(xai-grok-subagent-resolutionnested_images.rs),nested_fork_first_user_turn_attaches_file_image_part,extra_parent_images_not_attached_unless_named_in_spawn_prompt(xai-grok-subagent-resolutioncontext.rs),nested_first_user_turn_has_content_part_image_file_for_named_file_not_unnamed_extra,nested_fork_first_user_turn_has_content_part_image_file_for_named_file_not_unnamed_extra(xai-grok-shellagent/subagent/nested_spawn_prompt.rs),verbatim_fork_with_spawn_prompt_drops_parent_image_parts. - Footer context chip names sampling vs catalog when they differ:
AUTO compact gates on the sampling window. L1 sampling is the catalog
500k window. AUTO compact on L1 uses that window, not 200k. Nested L2
and L3 sampling stays 200k. L2 may compact. L3 never AUTO compact.
When those windows differ, the chip must not paint unlabeled
207K / 500Kas if catalog 500k were the nested gate. Same honesty as the CompactionStarted banner. Test:context_chip_names_sampling_window_when_catalog_differs(xai-grok-pagerviews/context_bar.rs). - Session sampling must not copy catalog 500k into a nested field:
AUTO compact and the footer chip gate on the sampling window. Session
sampling comes from GetSessionInfo / AutoCompactStarted.
refresh_context_usedmust not copy catalog into that field. Spawn seeds L1 at catalog 500k and nested sessions at 200k. Nested fallback is 200k when the session field is empty; L1 fallback is catalog 500k. Sampling is the same 200k cap for L2 and L3. Compact is not: L2 may compact, and L3 must not. Tests:footer_chip_uses_session_sampling_window_when_economic_cache_is_off(views/context_bar.rs),refresh_context_used_does_not_copy_catalog_into_session_sampling(app/acp_handler/tests/session_events.rs),main_session_sampling_window_is_catalog_500k_even_when_economic_is_on,nested_session_sampling_window_stays_200k_when_catalog_is_500k(xai-grok-shellsession/acp_session_impl/spawn.rs). - Parent ingest folds huge spawn prompts: parent ingest folds
spawn prompts over 40k into a pointer (description + size + report path
if any). Live L2 execute still uses the full spawn prompt. Spawn
tool-call arguments on the parent assistant item can still count
until fold runs. Tests (filter
fold_spawn_prompt):huge_spawn_prompt_becomes_pointer_with_description_and_report,small_spawn_prompt_stays,read_file_args_are_not_folded(xai-grok-sampling-typesfold_spawn_prompt_parent_ingest_tests);parent_estimated_tokens_omit_huge_spawn_prompt(xai-chat-stateactor/tests.rs). - Parent ingest folds huge shell/nix ToolResults (2026-09-02): a
2MB ANSI/nix-looking tool result must not add ~200k tokens to L1 500k
or nested 200k sampling. Store a head/tail pointer (and a log path when
the body names one). Goal Plan Writer / nested fork must not inherit
the parent's full
just check-remotedump as opening context. L3 and once-run Goal Plan Writer still must not compact-and-continue. Tests:two_megabyte_nix_ansi_tool_result_does_not_add_200k_tokens,two_megabyte_task_completed_user_is_folded(xai-grok-sampling-types);two_megabyte_nix_tool_result_does_not_add_200k_estimated_tokens(xai-chat-state);forked_goal_plan_writer_does_not_inherit_two_megabyte_nix_dump(xai-grok-shellagent/subagent);goal_plan_writer_forked_parent_full_window_at_spawn_must_not_immediately_compact(xai-grok-shellcompaction). Catalog:doc/dev/upstream-regression-filters.md§ Shell/nix tool-result ingest. - Parent ingest caps huge last answers: parent ingest /
completed-poll / blocking-spawn prompt format cap huge last answers
(~40k) and point at an on-disk report if one exists. Stored child
output can still be the full string. There is no automatic on-disk
last-answer report. Tests:
to_model_text_caps_huge_last_answer_for_parent_ingest(xai-tool-typestask.rs),completed_subagent_task_output_is_capped_or_points_at_report(xai-grok-toolstask_output/mod.rs),blocking_spawn_subagent_completed_to_prompt_format_is_capped(xai-grok-toolstask/mod.rs). - Auto-run
/implement: after a successful turn, queue a follow-up implement block when present; appends after any already-queued prompts. Plan-approval review comments that contain/implementare that turn's work, not a later auto-run (named testsextract_skips_plan_approval_review_comments_containing_implement,successful_implement_turn_does_not_auto_run_plan_approval_comments_implement). Compact is not a successful operator turn. Successful/compactand AUTO compact must not re-enqueue the occupancy operator prompt, or any operator prompt (named testscompact_complete_does_not_reenqueue_occupancy_or_any_operator_prompt,auto_compact_completed_does_not_reenqueue_occupancy_or_any_operator_prompt). That is not the compact-fail pause unstick path. A prompt that already issued, or already has a Human turn in this session, must not come back as a queued stale Prompt after rebuild, occupancy drop, Compact, or session reload. Compact-fail unstick after occupancy drop requeues/compactonly (try_unstick_idle_over_window_compact_fail). Compact-empty scrollback still dropsshared_queuerows already inchat_history.jsonl(AgentView::sync_queue_pane). Session load cancel-resume returns afteroperator_text_already_recorded(apply_canceled_turn_resume_on_load). WAL Interject and Queue kinds already in chat history must not restore (restore_prompt_wal_does_not_enqueue_committed_interject_or_queue). Named occupancy tests:sync_queue_pane_drops_shared_queue_rows_already_in_chat_history_when_scrollback_is_empty,compact_fail_unstick_after_occupancy_drop_requeues_compact_only_not_last_human_turn,session_load_cancel_resume_does_not_enqueue_human_turn_already_in_chat_history. Catalog:doc/dev/upstream-regression-filters.md§ Compact must not re-enqueue occupancy and § Prompt write-ahead log. FORK claims; not a land class. - Shared rate limits: crate
grok-rate-limit(Surmount name, notxai-); cooldowns under~/.grok/rate_limits/; optionalGROK_DISABLE_SHARED_RATE_LIMIT=1. Path-restored (FORK_PATHS). Before a sample, the sampler reads that flock store and waits if a peer's HTTP 429 cooldown is live. This is one machine, many processes (C1), not a daemon. Test:peer_process_does_not_sample_during_shared_rate_limit_cooldown(xai-grok-sampler--test peer_process_rate_limit). Do not fold this into the exhausted-credit hop. MatchingnextResetis not a shared pool. - Updates: no xAI auto-update channel by default (wrong product).
grok-oss update --checkcompares to Surmountmainusing git object ids (40-hex SHA-1 on today's repos, or a future git SHA-256 object id). That compare is not a SHA-1 security hash of a download or a Nix FOD. Escape hatch:GROK_OSS_ENABLE_XAI_UPDATER=1. FORK claims; not a land class. -
/rebuildis SHA-aware peer relaunch: localjust install, not an xAI download. Verify package version plus git SHA. Same semver plus a different SHA is newer. After install, the installed identity git SHA must match this workspacegit rev-parse --short=12 HEAD(same width as pager-binbuild.rs). A leftover cargo-bin such as1.0.3 (157f1746)is not an acceptable exec target when this workspace SHA differs. TUI/rebuildcompiles from the session workspace, not a random process cwd. Failed install or SHA mismatch must not replace the binary orSIGUSR1peers. Crate:xai-grok-updaterebuild.rs;xai-grok-shellleader/mod.rs; TUI dispatchxai-grok-pagerapp/dispatch/router.rs. Tests:failed_install_must_not_replace_or_signal_peers,build_fail_does_not_signal_leaders,parse_version_output_extracts_identity,peer_relaunch_accepts_same_semver_different_sha,peer_relaunch_declines_equal_identity_on_same_path,peer_relaunch_accepts_deleted_inode_even_when_identity_equal,operator_ran_rebuild_and_the_grok_oss_process_did_not_restart,leader_is_older_than_same_semver_git_sha_identity,rebuild_must_exec_workspace_binary_not_stale_cargo_bin,installed_identity_must_match_workspace_git_sha,tui_rebuild_starts_from_session_workspace_not_process_cwd. Fail-does-not-signal alone is not this seam. TUI/rebuildis the operator path with persist plus self re-exec. CLIgrok-oss rebuildis clap-wired (Command::Rebuild) to the same compile-and-signal core without self re-exec. Named test:rebuild_subcommand_parses. - Running grok-oss sessions: live TUI windows on this
$GROK_HOMEfromactive_sessions.json. Slash/running(alias/windows) and CLIgrok-oss running/grok-oss running --json. Not Agent Dashboard, not/sessions, not/tasks, not/resume. Distinct from/start(that slash starts paused or interrupted work in this process). Identity is(pid, session_id)so two windows on the same conversation both appear. Missing heartbeat is activityunknown. Title is the on-disk session summary. Never stores prompts, tool arguments, tokens, JWTs, file contents, or message text. Default headless stays unlisted unlessGROK_TRACK_HEADLESSis already set. Leader daemons stay ongrok-oss leader list./rebuildSIGUSR1 still dedupes by PID. Crates:xai-grok-active-sessions,xai-grok-pager,xai-grok-pager-bin,xai-grok-update. Tests:list_live_includes_two_windows_on_the_same_session_id,list_live_drops_dead_pid,heartbeat_omits_prompt_text,running_slash_lists_sibling_fixture_row,running_cli_json_omits_prompt_text,rebuild_signals_each_pid_after_composite_key,peer_pids_to_signal_excludes_self_dead_and_non_grok. User-guide04-slash-commands,17-sessions,23-dashboard(cite only). - L0 is
surmount-coordinator-gui(Surmount GPUI, not this pager): cratesurmount-coordinator-gui. Parses/running-shaped JSON. Drops prompt text, tool arguments, tokens, and JWTs. Each row has a local or remote host field.write_enqueuewrites$GROK_HOME/l0-enqueue/<session_id>/enqueue.jsonwith the prompt. grok-oss reads that file for this window's session id, queues one human prompt onpending_prompts(composer send path), and consumes the file. Other session ids are ignored. A missing file is a no-op. The prompt is not stored inactive_sessions.json. L0 does not merge into/dashboard.CoordinatorAppholds the session list, the selected index, load from local JSON plus an optional remote-tagged host, andenqueue_selected. Call L0grok-oss gui(notgrok-oss running). Binarysurmount-coordinator-guistill reads stdin or a file of/running --jsonand prints safe JSON (no prompt). It is not a grok-oss TUI and not/dashboard. Test:cli_gui_is_l0_not_running. Laptop-side action set remote host console API key (set-remote-host-console-api-key): the operator creates a machine xAI console API key at console.x.ai for host surmount-1. That key spends console API credits / console team prepaid. It is not included SuperGrok period limits. It is not SuperGrok dollar credits. Paste on stdin (never argv). The action writes owner-only staging files under the laptop grok home ($GROK_HOME/l0-remote-console-key/<host>/, otherwise~/.grok/...). Copy those files, or print/runscpas the existing deploy user. It never prints the key. It does not open git on the guest. It does not generate a GitHub SSH key. It does not copy laptop SuperGrok OAuth onto the guest. Guest grok home is$GROK_HOMEwhen set, otherwise~/.grokfor user grok (typically/home/grok/.grok). Attach stays SSH + tmux as user grok. There is no boot TUI. L0 is a laptop coordinator, not a website on the mail host :443, not pager/dashboard, not/running. This crate does not depend on gpui, does not path-pin Zed, and does not fetch crates.io. L0 is a Surmount GPUI window, not a grok-oss TUI dashboard./dashboardstays this pager./runningstays this machine's grok-oss sessions. They must not merge. Task tracking chrome for L0 reads$GROK_HOME/grok_oss.dbprompt_tasks. Session todos stay in this TUI (Ctrl+T). Do not replace that board. Tests:write_enqueue_creates_per_session_file,omits_prompt_text,keeps_pid_session_cwd,enqueue_drop_path_is_per_session_id,CoordinatorApp_selects_row,CoordinatorApp_omits_prompt_in_displayed_fields,CoordinatorApp_enqueue_writes_drop_file,drain_l0_enqueue_for_this_session_id_queues_one_human_line,drain_l0_enqueue_ignores_other_session_id,drain_l0_enqueue_missing_file_is_noop,set_remote_host_console_api_key_never_prints_the_key,set_remote_host_console_api_key_documented_workflow_does_not_require_guest_git_remote,set_remote_host_console_api_key_does_not_generate_github_ssh,set_remote_host_console_api_key_is_not_pager_dashboard,scp_copy_argv_does_not_include_the_key,CoordinatorApp_set_remote_host_console_api_key_never_prints_the_key,set_remote_host_console_api_key_cli_never_prints_the_key,set_remote_host_console_api_key_cli_refuses_key_on_argv,set_remote_host_console_api_key_cli_ssh_plan_omits_the_key,user_guide_machine_console_api_key_for_surmount_1. User-guide02-authentication,04-slash-commands,23-dashboard. Highest-value leftover is the GPUI window in the Surmount superproject depending on this crate. The delayed crate index still has no gpui, and a Zed path pin would break Nix. The operator must create the console API key themselves. -
/startstarts paused or interrupted work: pager builtin, not an alias of/resume(picker). Unpause if globally paused; else if a validcanceled_turn_resume.jsonexists, toast Continuing interrupted turn..., enqueue once, clear the marker, and drain. Soft-stop hold is released. An idle clean session does not invent a turn. Operator-typed/startapplies even when[ui] resume_canceled_turn_on_restartis off. Files:slash/commands/start.rs,app/dispatch/start.rs. Tests:start_while_globally_paused_continues_interrupted_turn_once,start_on_idle_clean_session_does_not_invent_a_turn,start_with_cancel_resume_marker_continues_interrupted_turn. -
/unstickresends the last L1 prompt: pager builtin, not/resume(picker) and not continue interrupted turn (canceled_turn_resume.json). When graceful resume did not unstick chrome, resend the last parent prompt as if the network dropped it. Do not paint a second Human line. Do not append a second<user_query>. Do not cancel nested agents, rewind, drop the transcript, reset sampler usage meters, or compact the turn away. Preferprompt_wal.jsonlL1 text when that file exists; else last Human send on the parent session, not a nested overlay. Image tokens stay[Image #N]. WAL image file ids resend as resource links (file://under sessionimages/), never data URLs. A hungrunning_taskis orphaned the way a reconnecting client drops an in-flight RPC, then the last L1 prompt is resent. That is not/resumeand not send-now cancel. Nested work and usage meters stay. No last prompt fails loud with a short toast. The leader drops a hungsession/promptRPC with the same routing as a disconnected client (leader.response.orphaned) while the pager stays connected. That is not a newClientId, not session evict, and notRelaunchForUpdate. Files:slash/commands/unstick.rs,app/dispatch/unstick.rs, shellsession/prompthonors_meta.unstickRetryandorphan_stuck_running_task_for_unstick, leadertake_in_flight_session_prompts_for_unstick. Tests:unstick_resends_last_l1_prompt_without_duplicate_human_line,unstick_does_not_cancel_nested_subagents_or_rewind_tokens,unstick_does_not_collide_with_resume_slash,unstick_with_no_last_prompt_fails_loud,unstick_retry_does_not_append_second_user_query_when_last_turn_matches,unstick_retry_orphans_stuck_running_task_then_samples_again,unstick_leader_drops_hung_session_prompt_like_disconnected_client,unstick_resends_wal_images_as_resource_blocks_not_data_urls,wal_image_resource_blocks_use_file_uri_not_data_url,wal_image_resource_blocks_drop_data_url_file_ids. User-guide04-slash-commands. -
/finishsession post-mortem: pager builtin injects the host skill~/.agents/skills/finish/SKILL.md. Work continues. Leftover and next features stay first-class. Not finished forever. Not/dream, not/recap, not/reports. Artifact under~/.agents/reports/finish-YYYY-MM-DD.md. Tests:finish_empty_args_injects_postmortem_skill,finish_registered_in_builtins,finish_skill_copy_does_not_say_work_is_closed_forever. -
/whatrestatement (CATE, 2026-08-27): default Grok OSS skill atcrates/codegen/xai-grok-bundle/skills/what/SKILL.md, installed into~/.grok/bundled/skills/what/. Live cache is not the source. Do not recreate repo.agents/skills/what/. Not an apology. Reply shape is four complete thoughts: Job, State, Operator, Next. Prefer Operator and Agent as speaker labels. Address the person as Operator, not Human. Do not say You or Human for the operator. Do not say Me or Grok as the speaker label for the machine. This diverges from upstream xAI You/Human / Me/Grok copy because the Operator said so. Painted chrome and user-guide call the composer the Operator box and DOGE caret/rails Operator green (accent_user). Identifiers such asaccent_userandUserPromptmay stay. Follow0005_CATE.md. When the operator asks to revise a skill in grok-oss, editcrates/codegen/xai-grok-bundle/skills/and keep named tests so skill maintenance cannot drop it. Never mix Grok Build version with grok-oss product version. Isolated Preview and plan chrome are grok-oss unless this process was launched asgrokfrom downloads. Tests:what_empty_args_injects_what_skill,what_instruction_prefers_operator_and_agent_speaker_labels,what_skill_does_not_mix_grok_build_version_with_grok_oss,user_guide_what_does_not_mix_grok_build_version_with_grok_oss,user_guide_operator_agent_speaker_labels_not_human_user_grok,waiting_chrome_does_not_paint_human_user_or_grok_as_speaker,user_prompt_prefix_is_not_the_word_human,agents_without_operator_agent_speaker_pin_fails_loud,what_registered_in_builtin_commands,what_registered_in_builtins,default_product_skills_include_polish_and_subagent(names includewhat). -
pull_remote_tree(2026-09-01): grok-build tool copiesHOST:SRC(or a local source directory) onto a local dest only. Ruststd::fswalk. OpenSSH may fetch. Not a rsync tool id. Excludes.git,target,.lake,result. Refuses SSH-shaped dest and git commit or git push argv. Default skillcrates/codegen/xai-grok-bundle/skills/pull-remote-tree/. Tests:copy_tree_excludes_git_target_lake_result,pull_remote_tree_refuses_ssh_shaped_dest,tool_id_is_pull_remote_tree_not_rsync. - Queue
/compaction//plan//reports//finish: named hold on the existing composer prompt queue (/queue <slash>or first-argqueue/later). Not a second queue. Immediate invoke stays./compactionaliases/compact./reportsinjects~/.agents/skills/reports/SKILL.md(checkpoint; work continues; not/finish). Cancelled compact still must not re-arm. Tests:queue_compaction_does_not_invoke_immediately,queue_plan_does_not_invoke_immediately,reports_empty_args_injects_reports_skill,reports_registered_in_builtin_commands,slash_compaction_alias_invokes_compact. -
/metadatalive session ids: transcript block with grok-oss ULID, Grok Build UUID, cwd, model, started, pid. Omit unknown fields.[ui] ulid_session_ids(default on) picks which id is listed first. Not/session-info. Map issession_id_mapingrok_oss.dbschema v5. Wire ACP session id stays UUID. New session, fork, andsession/load(attach) all fail-open map. Tests:format_lists_ulid_before_uuid_when_primary,metadata_command_emits_show_session_metadata,show_session_metadata_maps_uuid_and_shows_ulid_when_db_overridden,ensure_session_ids_same_uuid_returns_same_ulid,attach_and_new_session_both_call_ensure_session_ids_fail_open,existing_uuid_session_gets_mapped_ulid_on_load_path_helper. - Official serving-path fingerprints (Surmount, 2026-09-09): Consumer
Grok shows Grok 4.6 with no public checkpoint ID. grok-oss logs and
/metadatashow Chat Completionssystem_fingerprintand lastGET /v1/language-models{id, fingerprint, version, created}for the current sampling model, with when those were observed. Fingerprint is backend configuration, not a SHA of the weights. A flip means the serving path changed; behavior still decides if weights moved. Additive$GROK_HOME/grok_oss.dbschema v7 tablescompletion_system_fingerprint,language_model_serving,serving_fingerprint_flip. Does not drop/spendschema v1. Not on the status bar. Do not invent dated slugs such asgrok-4.6-20260812unless that list endpoint names them. Upstream owns the Chat Completionssystem_fingerprintJSON field. Surmount owns persist, flip history, language-models snapshot,/metadataserving lines, and user-guide copy. Tests:format_includes_stored_serving_fingerprint_fields,show_session_metadata_includes_fingerprint_fields_from_stored_samples,parse_language_models_json_reads_id_fingerprint_version_created,parse_language_models_json_does_not_invent_dated_slugs,record_completion_fingerprint_persists_a_flip,persist_language_models_list_upserts_and_records_flip,migrate_v6_file_to_v7_adds_serving_tables_without_dropping_spend,chat_completion_response_deserializes_system_fingerprint,chat_completion_chunk_deserializes_system_fingerprint,chat_completions_stream_copies_system_fingerprint_onto_assistant,user_guide_metadata_documents_serving_fingerprints. -
from_configno-prefetch usable catalog:ModelsManager::from_configwith no prefetch argument is a zero-network boot and must produce a usable bundled catalog. Test:from_config_without_prefetch_produces_usable_catalog(xai-grok-shellagent/models/tests.rs). Emptymodels_cache.jsonis a miss in code (load_freshreturnsNonewhenmodelsis empty). That empty-file branch has no named test. Do not claim it is cargo-proven. - Seeded custom model on
session/loadstays Chat Completions:session/loadkeeps a seeded custom model id on Chat Completions instead of remapping it to the default grok-4.5 Responses catalog entry. grok-4.5 itself still uses Responses. SuperGrok is paid. This is not last-session on start. Crate:xai-grok-shell. Tests:keep_unverified_persisted_model_keeps_seeded_custom_slug(agent/models/tests.rs),seeded_test_model_keeps_chat_completions_backend(agent/mvp_agent/tests.rs). Integration:poisoned_image_session_recovers_within_the_failing_turn(--test test_image_strip_recovery; in-turn strip after 400invalid_image). - Nucleo reuse-per-root: many workspace fuzzy searches without
closekeep one live matcher per root. Poll-onlyget_resultsmust not refresh the stale timer. Crate:xai-grok-workspacefile_system/mod.rs. Tests:repeated_open_without_close_keeps_one_search_per_root,distinct_roots_each_keep_one_search,get_results_does_not_keep_a_stale_search_alive. Per-matcher pool sizeNUM_NUCLEO_THREADS = 2is shipped in code (xai-fuzzy-file-search); nofnassertsSome(2). - L2 spawn prompt (GitHub #141, 2026-09-20; supersedes 2026-08-20):
process law lives in
AGENTS.md(D1; path-restored). An L2 coordinator for implement work must spawn L3 for greps, reads, and product edits. L2 does not fill 200k implementing. Compact on that L2 is a product miss when the cause is L2-solo implement tools. Product spawn-tool copy (CHILD_TASK_DESCRIPTION) must contain those Operator strings. It must not teach "Easy work can stay on L2", "Including implement loops", or "Spawn L3 only if the problem is actually hard" for greps, reads, or product edits. Default max depth must still let depth-1 spawn L3. Tests:child_task_description_is_concise(xai-grok-agentbuilder.rsCHILD_TASK_DESCRIPTION),default_max_allows_l2_to_spawn_l3(xai-grok-toolstask/mod.rs). User-guide16-subagents.mdnames Hierarchical fast path (L1-only: one-command host question, one already named path, or the asked-for report). Do not put Hierarchical fast path intoCHILD_TASK_DESCRIPTION. A restack can keep AGENTS viaFORK_PATHSand still dropCHILD_TASK_DESCRIPTION. Product cargo is the seam. -
/goalparent coordinates; L2 MUST spawn L3 (Surmount / grok-oss fork of the injected prompt): Upstreamgoal_instructiontells the parent to "Deliver everything the user asked for yourself," which fills L1. Surmount keeps the objective and theupdate_goalcontract, and tells L1 to coordinate (spawn L2; that L2 MUST spawn L3 for tools; L3 does the tools; no L4). There is no bundled goal skill;goal_instructionplus the livegoal_rules.md/goal_rules_legacy.mdtemplates are the product prompt. Distinct from spawn-tool copyCHILD_TASK_DESCRIPTION, which now matches that implement-coordinator law (must spawn L3 for greps, reads, and product edits), not the old easy-L2 license. Crate:xai-grok-tools-apislash_commands.rs. Tests:goal_instruction_parent_coordinates_and_l2_must_spawn_l3_for_tools,goal_instruction_carries_objective_and_contract_tokens. Live harness:xai-grok-shellgoal_rules_templates_parent_coordinates_and_l2_must_spawn_l3_for_tools,goal_task_discipline_parent_spawns_l2_not_product_tools. - Parent fire-and-return spawn (Surmount / grok-oss fork): a nested
L2 that is a long builder (compile, lake, mill) must not occupy the
parent as a blocking 10-minute
get_command_or_subagent_outputwait loop. Parent starts it, keeps working, completion is a notification. Parent can spawn a second L2 while the first is still running without waiting for the first to exit. Snapshot (omit/0) is allowed. Positivetimeout_msremains only when the parent must join. Upstream parent turns often sit on a 10-minute wait. Named test:parent_spawn_subagent_second_l2_while_first_still_running_without_waitinxai-tool-types(task.rs) andxai-grok-tools(task/backend_tests.rs). - Parent follow-up onto a running nested L2 (Surmount / grok-oss
fork, GitHub #143): L1 can enqueue a follow-up onto a still-running
nested L2 as additive work. It does not kill that L2, does not respawn
it, and does not wait for it to exit.
resume_fromstill continues a completed nested L2 only; it is not the live steer. Operator overlay typing stays (x.ai/interjecton an open L2; L3 overlay stays unbothered). Soft interject of the L1 turn stays the L1 turn unless this follow-up path is used. The follow-up must not inject into a live L3 unless the Operator explicitly targeted that specialist. Default on.[subagents] parent_follow_up = falseis the SpaceXAI / upstream option: spawn, wait, andresume_fromafter exit only (overlay compose unchanged). Do not invent a second permission system. No new[auth]key. Keywords: running L2 follow-up versusresume_fromcompleted, L3 unbothered, additive not kill. Operator: "You can't talk to your own L2s? And you're fine with that? Why?" Grok OSS vs SpaceXAI: upstream L1 has no this live-L2 parent-tool enqueue. Schema:TaskToolInput.follow_upis optional and distinct fromresume_from(xai-tool-typestask.rs). Coordinator:SubagentBackend::follow_up,SubagentEvent::FollowUp,handle_follow_up(xai-grok-toolstask/parent_follow_up_tests.rs). Tests:parent_cannot_talk_to_own_l2s_follow_up_enqueues_interject_on_running_l2_without_kill_or_respawn,parent_follow_up_does_not_inject_into_live_l3_unless_operator_targeted_that_specialist,resume_from_of_running_l2_still_fails_active,parent_follow_up_off_is_upstream_spawn_wait_resume_from_completed_only,parent_follow_up_onto_running_l2_with_live_l3_hits_l2_not_l3.TaskTool::run(xai-grok-toolstask/mod.rs):parent_cannot_talk_to_own_l2s_task_tool_run_follow_up_returns_queued_and_does_not_spawn,parent_cannot_talk_to_own_l2s_task_tool_run_follow_up_and_resume_from_are_mutually_exclusive. Schema:parent_cannot_talk_to_own_l2s_follow_up_field_is_optional_and_distinct_from_resume_from. Shell Interject (xai-grok-shellagent/subagent/tests/mod.rs):shell_child_follow_up_sends_session_command_interject_on_child_session. Config (xai-grok-shellconfig/tests.rs):subagents_config_parent_follow_up_false_parses_and_omitted_defaults_true,resolve_subagents_copies_parent_follow_up. KEEPparent_spawn_subagent_second_l2_while_first_still_running_without_wait,l2_overlay_send_prompt_interjects_l2_not_l1,l3_overlay_send_prompt_does_not_reach_l3_or_l1,nested_spawner_can_resume_from_completed_reparented_child,live_subagent_list_shows_only_l2_and_reports_live_l3_count. - Compact standing-law reminder (Surmount / grok-oss fork): after
compact, standing Surmount law (FORK.md as the divergence home, named
tests as contracts, red then green TDD, do not interrupt live L2s when
L1 shows a plan pane, fire-and-return for long builder L2s) is the
first section of the post-compaction
<system-reminder>. It is not a buried AGENTS.md paragraph and not/recap. Upstream parent turns often sit on a 10-minuteget_command_or_subagent_outputwait. Surmount starts the builder L2 and keeps working; completion is a notification. Crate:xai-grok-shellsession/helpers/compaction_context.rs(section_surmount_standing_law_after_compact, wrapped byxai-grok-compactionwrap_system_reminder). Named test:post_compact_reminder_includes_surmount_standing_law. - Soft interject only + Enter cue honesty: mid-turn interject
(plain Enter with text, and Ctrl+Enter when interjection is
appropriate) injects into the current turn and never cancels.
Ctrl+Enter inserts a newline when interjection is not appropriate
(idle, empty composer, L3 overlay). Cancel is Esc/stop only. Composer
footer Enter cue (send / queue / interject) is shipped in code with no
named footer
fn. Proven never-cancel:interject_contract_*. User-guide03-keyboard-shortcuts,16-subagents. - Todo board survives auto-compact:
auto_compact_completed_preserves_todo_board(app/acp_handler/tests/subagents.rs). - Status-row todo badge names tasks: paints
tasks N/M, not only614/638. Click and Ctrl+T still toggle. Nested L2 overlay keeps that nested session's badge and Ctrl+T. The pane stays closed until the operator toggles it. Tests:todo_badge_names_tasks_not_only_fraction,status_header_todo_badge_names_tasks,nested_l2_overlay_todo_toggle_stays_findable(views/agent.rs,app/agent_view/render.rs). User-guide03-keyboard-shortcuts,16-subagents,17-sessions. - plan.json honesty + resume board: compact writes the live
Resources
TodoStatetoplan.json. FORK claims; not a land class. User-guide17-sessions. - Auto-seed user asks as todos: real user turns seed protected
ask:<prompt_id>. FORK claims; helpers inxai-grok-toolstodo module. - Default agent uses the todo board: base
prompt.mdteachestodo_write. FORK claims; not a land class. - Same-batch plan write +
exit_plan_mode: mixed multi-tool batches run non-exit tools to completion first (same_batch_plan_write_before_exit_plan_mode_returns_new_body). Dated 2026-08-09 wave filter; not one of the seven product land classes. - Continue interrupted turn on restart:
canceled_turn_resume.json; distinct from last-session on start. Mid-turn/rebuilddoes not cancel the parent and does not write this marker; the new TUI adopts the live turn like a disconnect (runningPromptId). Nested ids are still not cancelled./rebuildis not a nested-work gate. Idle completed turns do not write a marker and do not re-fire the last prompt. Load drops a leftover marker after a successful primary-turn finish. Stale-queue skip of a Human turn already in chat history stays. Tests:handle_rebuild_done_must_not_cancel_parent_so_session_load_adopts_like_disconnect,handle_rebuild_done_mid_turn_writes_cancel_resume_and_session_load_continues_the_turn,handle_rebuild_done_idle_completed_turn_does_not_write_cancel_resume_or_refire_last_prompt,session_load_drops_stale_cancel_resume_marker_when_primary_turn_finished_successfully(xai-grok-pagerapp/dispatch/rebuild.rs). Still leftover (not shipped): auto-resume after an error-terminal turn with no marker; soft-stop button; mid-sample freeze without cancel. FORK claims plus these named tests; not a land class. User-guide17-sessions. -
/rebuildresume is fork-owned: relaunch must preserve work the same way a network disconnect does. Mid-turn/rebuilddoes not cancel the parent; the new TUI adopts the live turn (runningPromptId). Unsent composer draft (unsent_prompt_draft), queued prompts including mid-turn interject text (pending_prompts.json), plan Operator-boxfeedback_draft, and sessionplan.mdsurvive. Nested subagent ids are not cancelled and/rebuildis not blocked until nested work finishes. Compile source is the git index (staged files), not unstaged working-tree WIP. After--resume/ last-session restore, that preserved work appears once: not composer plus queue #1 with the same body, not Enter:interject unless a live sampler turn is running, not Waiting leftover. Tests:handle_rebuild_done_persists_unsent_composer_draft_and_session_load_restores_it,handle_rebuild_done_persists_pending_prompts_including_interject_and_session_load_restores_them,resume_restore_must_not_put_the_same_operator_prompt_in_composer_and_queue,resume_restore_must_not_arm_enter_interject_when_no_live_sampler_turn,resume_restore_must_not_show_waiting_when_nested_and_sampler_are_gone,after_rebuild_or_resume_plus_plan_exit_follow_up_must_not_wait_for_the_model_with_no_sampler,resume_restore_must_not_rehydrate_unsent_draft_and_queue_with_the_same_string,handle_rebuild_done_persists_plan_feedback_draft_and_plan_md,handle_rebuild_done_persists_open_plan_pane_and_session_load_docks_it,handle_rebuild_done_keeps_nested_subagents_for_resume,rebuild_and_relaunch_starts_while_nested_subagents_are_running,operator_ran_rebuild_and_the_grok_oss_process_did_not_restart,post_rebuild_relaunch_chrome_includes_grok_oss_version_and_git_sha,tui_rebuild_starts_from_session_workspace_not_process_cwd,restore_pending_prompts_from_disk_drops_human_turns_when_memory_queue_is_nonempty(xai-grok-pagerapp/dispatch/rebuild.rsandagent_view/session.rs);export_git_index_omits_unstaged_dirty_file,stash_keep_index_hides_unstaged_wip_from_compile_worktree,operator_ran_rebuild_and_the_grok_oss_process_did_not_restart,rebuild_must_exec_workspace_binary_not_stale_cargo_bin,installed_identity_must_match_workspace_git_sha(xai-grok-updaterebuild.rs). Keep these stronger than an upstream resume that cancels nested orphans. This TUI exec-replaces onto the new binary even while nested work is live. Unix exec keeps the same PID andpsstart time; compare/proc/<pid>/exeinode to the cargo-bin file. Post-relaunch chrome showsgrok-ossversion plus git SHA (psfork time is not the signal). That SHA must be this workspace HEAD, not a leftover cargo-bin such as157f1746. Named tests:post_rebuild_relaunch_chrome_includes_grok_oss_version_and_git_sha,rebuild_must_exec_workspace_binary_not_stale_cargo_bin. Restore occupancy-drops Human-turn queue rows even when memorypending_promptsis already non-empty. Named test:restore_pending_prompts_from_disk_drops_human_turns_when_memory_queue_is_nonempty. LeaderRelaunchForUpdatekeeps nested ids on that leader the same way a TUI disconnect does (the leader is not exec-replaced while nested ids are live). After nested ids finish, that leader stays up while the parent turn is still busy, with no five-second cap. Named drain tests:relaunch_drain_keeps_nested_ids_alive_after_grace_like_disconnect,relaunch_drain_keeps_parent_turn_until_idle_like_disconnect. Not a land class. User-guide04-slash-commands. - Prompt write-ahead log (
prompt_wal.jsonl): session-local append-only file next tounsent_prompt_draft. Enter send, mid-turn interject, queue enqueue (including mid-turnpending_promptsenqueue, which appendskind=queuebefore any laterkind=send), plan Operator-box notes that ride Approve, and/rebuildpersist each append (and fsync) one JSONL object before the model is asked, before compact, and before re-exec. The WAL is not rewritten, not compacted as conversation, and not counted as model tokens. If chat history, prompt history, and the queue lack a WAL send, session load restores it as a pending Operator turn. RebuildFlush is not a pending Operator turn. After--resume/ last-session restore, the operator prompt appears once: unsent draft restore and queue restore must not both rehydrate the same string, and a WAL Send must not enqueue a body already in the composer. Resume must not arm Enter:interject unless a live sampler turn is actually running, and Waiting must be a real sampler wait, not occupancy leftover. Enter after a paste chip must not wipe the composer without send or enqueue; footer[pause]button chrome (not engaged) must not swallow that Enter. Enter on[Pasted: 15 lines]sends or interjects; it does not only expand the chip. Expand is paste-again or double-click. Tests:prompt_wal_appends_on_enter_before_model_wait,enter_on_pasted_15_lines_chip_sends_or_interjects_does_not_only_expand,enter_after_paste_chip_must_wal_send_not_wipe_without_enqueue,enter_after_paste_chip_with_pause_button_chrome_still_sends,enter_while_drain_blocked_must_wal_queue_or_keep_composer,prompt_wal_appends_on_mid_turn_interject,prompt_wal_appends_on_queue_enqueue,prompt_wal_appends_on_approve_notes,session_load_restores_wal_send_missing_from_prompt_history,resume_restore_must_not_put_the_same_operator_prompt_in_composer_and_queue,resume_restore_must_not_arm_enter_interject_when_no_live_sampler_turn,resume_restore_must_not_show_waiting_when_nested_and_sampler_are_gone,resume_restore_must_not_rehydrate_unsent_draft_and_queue_with_the_same_string(xai-grok-pager); rebuild persist tests also require a rebuild-flush WAL line. User-guide04-slash-commands,17-sessions. Catalog:doc/dev/upstream-regression-filters.md§ Prompt write-ahead log. Operator-verified known good (2026-09-02) because live session files contained those kinds:prompt_wal_appends_on_enter_before_model_wait(send),prompt_wal_appends_on_mid_turn_interject(interject),prompt_wal_appends_on_approve_notes(plan-notes), and rebuild persist tests that require arebuild-flushWAL line. Queue enqueue (prompt_wal_appends_on_queue_enqueue) stays a named contract. Do not mark it operator-verified known good: a live session wrotepending_prompts.jsonand had noprompt_wal.jsonl. Restore tests (session_load_restores_wal_send_missing_from_prompt_history, resume occupancy) and skip (skips_prompt_wal_jsonl_because_it_is_not_conversation) stay contracts. Do not delete or weaken those tests in recon, onto, import, or join. - Interject Ctrl+Enter and Send now are fork-owned. Mid-turn
Ctrl+Enter interjects when interjection is appropriate, and otherwise
inserts a newline (the Shift+Enter analog). Interjection is appropriate
when a sampler turn is running, the Operator box has text or images, and
the target can take
x.ai/interject(this session or an open L2 overlay). It is not appropriate when idle, when the composer is empty, or when an L3 specialist overlay is open. Cancel-and-send is not Ctrl+Enter. The clickable queue[Send now]control on a plain prompt row dispatchesSendInterject. A queued/goalrow after Send now is a GoalSet viaSendPromptNow, not an interjected composer string. They must not drop the text, queue-only, or no-op. Enter with text while a turn runs is the separate soft-interject path. Empty composer does not send. A successful interject still appends WALkind=interject. Product: InterjectPrompt (agent_view/prompt.rs) and local-rowforce_interject_queue_row(agent_view/queue.rs; mouse Down on[Send now]inapp/mouse.rs). Grok OSS 1.0.3 is not last-known-good for this UI. Operator-verified WAL send / rebuild-flush / interject / plan-notes appends do not mean live Interject UI works. Named tests:ctrl_enter_mid_turn_dispatches_send_interject,queue_send_now_click_dispatches_send_interject,empty_ctrl_enter_mid_turn_does_not_send,enter_while_other_work_is_live_must_still_clear_composer,send_now_while_retrying_must_still_clear_composer,queued_prompt_edit_must_not_steal_later_send_clear,queued_goal_send_now_is_goal_action_not_stuck_composer_string,limits_help_lists_named_words_and_hyphenated_aliases,limits_hyphenated_aliases_match_unhyphenated_words,header_timeout_is_named_cold_start_class_with_retry_path,interject_does_not_wait_minutes_or_block_paint,enter_soft_interject_must_not_leave_duplicate_prompt_in_composer,enter_on_pasted_15_lines_chip_sends_or_interjects_does_not_only_expand,l2_overlay_enter_interject_must_not_leave_duplicate_prompt_in_composer,enter_send_must_not_leave_duplicate_prompt_in_composer,enter_at_end_of_last_composer_line_must_submit_immediately_not_silent_newline,enter_at_end_of_last_composer_line_mid_turn_must_interject_immediately_not_silent_newline,arrow_keys_then_enter_must_submit_the_same_body_not_a_different_path. After Enter that interjects or sends, the Operator box must not still hold that body. Enter at the end of the last composer line must send or interject immediately. It must not insert a silent extra newline. Arrow keys then Enter must submit the same body on the same path. Do not delete or weakenprompt_wal_appends_on_mid_turn_interject. Catalog:doc/dev/upstream-regression-filters.md§ Interject Ctrl+Enter and Send now. - Leader
RelaunchForUpdatenested work and parent-turn busy survive like a TUI disconnect: this leader process stays up while nested ids are live (spawned leaders already pass--no-exit-on-disconnect). After nested ids finish, this process stays up while the parent turn is still busy (AgentActivity::is_busyor IPCagent_busy), with no five-second wall-clock kill, then the leader may relaunch. Named tests:relaunch_drain_keeps_nested_ids_alive_after_grace_like_disconnect,relaunch_drain_keeps_parent_turn_until_idle_like_disconnect. User-guide04-slash-commands. - OAuth 403
bad-credentials→ auth path: HTTP 403 withunauthenticated:bad-credentialsclassifies as auth, not included SuperGrok period limits. Dated 2026-08-09 wave filters on sampler types. Not a land class. - Multi-track also-guard (first cut):
todo_writeacceptsmeta.taskId; demotingin_progress→pendingis rejected while that subagent is still Running. FORK claims; not a land class. - Task tracking system (local first; chrome not shipped): more formal
than the session todo board, less formal than issues. Durable rows live in
$GROK_HOME/grok_oss.db(prompt_tasks,prompt_exec_metrics). This does not replace Ctrl+T session todos. Left-sidebar tasks / right-sidebar plan is residual. Open:RESIDUAL.md.
Human chrome is green (accent_user: composer caret, human rails, OSC 12,
success). Agent activity is magenta (accent_running / accent_model:
active agent rails, tool spinner, lower-left still-running cue). Clear finished
is quiet secondary, not neon green and not magenta. Default theme is DOGE.
External role map:
0001_DOGE.md.
User-guide 06-theming.
- Unset theme is DOGE:
xai-grok-pager-rendertheme/cache.rs,theme/system_appearance.rs. Tests:default_theme_is_doge,resolve_from_config_no_config_returns_doge,resolve_auto_dark_system_returns_doge,to_theme_kind_dark_defaults_to_doge. This is not the models-catalogfrom_configempty-cache miss. - DOGE human green / system cyan / role map:
theme/doge.rs. Tests:doge_accent_user_is_pure_green_for_human,doge_accent_system_is_pure_cyan_for_system_limits_credits,doge_roles_green_cyan_no_blue_ui_no_gray_text. - Human left rail paints green:
scrollback/blocks/user.rs. Tests:user_prompt_block_accent_is_static_human_rail,user_prompt_block_accent_is_green_rail_under_doge_default,user_prompt_entry_renderer_paints_green_rail,user_prompt_prefix_matches_human_rail_color. - Running agent rail paints magenta:
agent_message_block_accent_is_magenta_rail_under_doge_while_running(scrollback/blocks/agent.rs). - Composer box caret is Operator green, never agent magenta:
views/prompt_widget/tests.rs. Tests:paint_composer_box_cursor_uses_human_green_not_agent_magenta,focused_composer_paints_human_green_box_caret_hides_terminal_cursor,doge_human_box_caret_plate_is_rgb_0_255_0,paint_composer_box_cursor_named_ansi_green_becomes_doge_rgb. DOGE plate/ink and OSC 12 areColor::Rgb(0, 255, 0), not named ANSIColor::Green(terminal lime /#00cd00). - Model label uses
accent_model:info_line_model_name_uses_accent_model_not_gray. - Titled composer frame is
prompt_border_active(white); title only is yellow:titled_doge_composer_frame_is_prompt_border_not_context_yellow. - Compact included SuperGrok period limits meter: status chip
SuperGrok period · N%; click opens/limits. Tests:status_bar_pushes_credits_compact_included_supergrok_period_limits,hit_credits_click_dispatches_show_limits(app/agent_view/render.rs). - Forked-session upper-left header switcher plus dashboard: a
fork family paints
[‹][›]and[Dashboard]on the status row, not git plus cwd only. The yellowuse /dashboardtranscript line is not this chrome. Tests:forked_session_status_header_paints_switcher_and_dashboard,forked_session_status_header_clicks_open_dashboard_and_cycle,forked_session_status_header_paints_dashboard_for_lone_fork(app/agent_view/render.rs). Resume of a persisted fork restores the parent as a live agent and stampsforked_from:load_session_restores_fork_family_from_disk(app/dispatch/tests/session/load.rs). - Plan footer CTAs: idle footer is Approve / Comment / Revise / Exit
(four CTAs). Clarify is only in the comment flow after Comment, not an
idle top-level notes path. Notes is gone. Letter
a/Atype. Empty Enter never Approves. Revise arms the box and waits. The white plan prompt frame usestheme.prompt_border_active. Tests:plan_approval_footer_paints_five_cta_vocabulary,plan_footer_exit_not_quit,plan_footer_has_no_notes_button,plan_prompt_letter_a_inserts_when_composing(views/file_search/line_viewer.rs,app/agent_view/plan.rs). - Plan Operator box typing and Ctrl+Z: Preview-focused
Ctrl+Zreaches the composer undo stack so a wiped Operator box comes back. Keystroke unsent draft persist coalesces and does notsync_allon every character (the main prompt shares that path). Helpers:plan_preview_key_is_composer_text,persist_unsent_composer_draft,should_flush_unsent_draft. Tests:plan_preview_ctrl_z_restores_wiped_human_box,plan_prompt_ctrl_z_restores_wiped_human_box,plan_human_box_keystroke_burst_does_not_flush_unsent_draft_every_char,main_composer_keystroke_burst_does_not_flush_unsent_draft_every_char,keystroke_burst_does_not_flush_unsent_draft_every_char,plan_human_box_keystroke_burst_does_not_append_prompt_wal,main_composer_keystroke_burst_does_not_append_prompt_wal. A keystroke burst must not appendprompt_wal.jsonlor rewritepending_prompts.json. Ordinary queue snapshots skipsync_all(pending_prompts_queue_snapshot_skips_sync_all,pending_prompts::tests::write_without_fsync_still_roundtrips). - Lost-prompt / composer-draft tests are fork-owned contracts: recon
must not delete or weaken them. Pane-open Operator-box notes ride along with
Approve (Preview typing after park is review comments, not a silent wipe).
Isolated present with the pane shut must not consume a restored agent
prompt. Clickable Approve must not drop the Operator-box prompt (mouse
Approve is not Empty Enter on Revise). When those tests change, resolve
meritocratically: keep the stronger assert; synthesize if upstream and
Surmount both have a piece; never fit our contract to a wipe. Module
app/acp_handler/tests/plan_approve_lost_prompt.rs. Tests:isolated_present_preview_click_approve_does_not_drop_human_box_prompt,isolated_present_preview_typed_after_present_click_approve_sends_human_box_prompt,isolated_present_prompt_focus_click_approve_does_not_drop_human_box_prompt,isolated_present_click_approve_dispatches_interject_with_prompt_text,isolated_present_preview_enter_is_human_turn_then_click_approve,isolated_preview_idle_non_empty_operator_paste_enter_approves_with_notes_not_plan_exit,isolated_preview_idle_leftover_slash_plus_notes_click_approve_is_approve_with_comment,isolated_preview_vanished_pane_notes_enter_approves_with_comment,isolated_preview_idle_leftover_slash_plus_notes_enter_approves_with_comment,isolated_preview_approve_with_plan_composer_notes_submits_with_approve_not_as_prompt,isolated_preview_stays_after_present_so_comment_then_approve_can_run,isolated_preview_comment_cta_then_notes_then_approve_submits_with_approve_not_as_prompt,view_plan_reopens_isolated_preview_from_current_disk_plan_md_after_panel_closed,isolated_preview_human_send_closes_leftover_present_after_mill_continues,isolated_preview_implement_closes_leftover_present_after_mill_continues,isolated_preview_rereads_current_disk_plan_md_when_mill_rewrote_it,isolated_preview_after_mill_completion_must_not_paint_leftover_present_or_tech_md,preview_typed_comment_rides_along_on_approve,prompt_tab_typed_comment_rides_along_on_approve,esc_with_human_box_draft_keeps_feedback_draft,tab_preview_prompt_keeps_human_box_draft,exit_with_human_box_draft_does_not_drop_unsent_text,approve_with_composer_comments_sends_one_human_line,empty_approve_does_not_send_composer_as_second_prompt,resume_restore_keeps_revise_box_draft. - Plan present is not operator Approve + modal-free typing:
exit_plan_modepresents the plan. It does not click Approve. Always-approve permission mode does not auto-click the CTA. Empty Enter never Approves. Soft-park must not steal mid-compose keys. Crate:xai-grok-pagerapp/agent_view/plan.rs,app/acp_handler/tests/plan_mode.rs;xai-grok-toolsexit_plan_mode/mod.rs. Tests:exit_plan_mode_present_is_not_operator_approve,exit_plan_mode_tool_result_does_not_claim_operator_approval,empty_enter_on_revise_prompt_does_not_approve,soft_park_empty_ctrl_c_abandons_plan_approval,exit_plan_mode_keeps_mid_compose_draft_and_a_types,exit_plan_mode_modal_park_does_not_steal_mid_compose_keys,exit_plan_mode_empty_present_printable_goes_to_composer,exit_plan_mode_shows_overlay_even_in_yolo. Settings park picker is class 2 (plan_approval_park_*). Prefer these exact names over a vagueexit_plan_mode_softsubstring. User-guide19-plan-mode,22-permissions-and-safety. - Soft plan present is a real right-side pane: default soft park
docks the existing plan list plus four idle CTAs (Approve / Comment /
Revise / Exit) on the right, full overlay
height, no dim of the transcript. Status Plan ready. Side panel open
only when that viewer is actually open. A click on a plan row does not
enter Commenting.
cremains the explicit line-comment gesture. Tests:plan_soft_park_docks_right_not_centered_overlay,plan_soft_park_draw_right_pane_matches_side_panel_status,plan_row_click_does_not_enter_commenting,plan_loop_status_does_not_claim_side_panel_when_viewer_closed. -
/plan --softdocks Isolated Preview: does not enter plan mode, does not park L1, does not enqueue a Prompt. Nested L2s stay Working. Hard/planwithout--softenters plan mode.--softis not the queue hold token. Present is not Approve. Empty Enter never Approves. Comment then Approve carries notes. Soft planning does not reset the primary plan. It makes a secondary plan. Isolated Preview does not immediately pull up leftover currentplan.md. Comment then Approve still works on a real present of that secondary plan afterexit_plan_modewrites it. Tests:plan_soft_flag_dispatches_isolated_preview_dock_not_plan_mode,plan_soft_docks_isolated_preview_without_entering_plan_mode,plan_soft_with_feature_seeds_isolated_preview_and_does_not_enqueue_prompt,plan_soft_is_not_the_queue_hold_token,soft_planning_does_not_reset_the_primary_plan_it_makes_a_secondary_plan,isolated_preview_soft_planning_does_not_pull_up_leftover_current_plan_md,user_guide_plan_soft_docks_isolated_preview. - Isolated Preview re-reads rewritten
plan.md: after Revise rewrites sessionplan.mdand re-presents, Isolated Preview paints the current file, not the first-draftplan_contentsnapshot. Opening the panel re-reads the file. Older leftover disk still loses to a newer SQL row. After Plan Exit, chrome must not keep Plan ready. Side panel open. Idle CTAs must not stay armed for the exited present. A new present that writes sessionplan.mdmust paint that file, not a frozen SQL snapshot and not a previous transcript plan body. Two different plan texts in the same window after Exit plus re-present is a fail unless the panel matches disk. After Plan Exit, Isolated Preview must not wedge: Esc:close,/start, or/unstick(hung parent prompt) leave the pane./startcontinues paused or interrupted work in this process. It is not/resume. After Plan Exit, Isolated Preview must paint this session's current diskplan.md, not leftover TECH.md, or close. With Isolated Preview closed, chrome must not stay plan. Bare/planafter Exit paints covering exclusive present from current diskplan.mdand exclusive-blocks nested implementers. It is not leftover Isolated Preview./plan --softdoes not reset the primary plan. It makes a secondary plan. Isolated Preview does not immediately pull up leftover currentplan.md. Isolated Preview stays until Esc, Exit, or Approve. Compact at 100% / over 500k must not swallow/plan. Typing an Operator sentence after Exit still sends. Empty Enter never Approves./planwith extra Operator text submits a plan-update turn of the primary plan (Operator send / plan rewrite) and writes the prompt write-ahead log. It must not only dock leftover Isolated Preview ("why the agent stopped" / TECH.md). GitHub issue 96 and issue 98. Tests:isolated_preview_after_revise_rereads_plan_md_not_first_draft_snapshot,isolated_preview_prefers_rewritten_plan_md_over_stale_sql_snapshot,isolated_preview_reads_sql_first_then_disk_plan_md_fallback,isolated_preview_and_present_read_sql_first_then_disk_plan_md_fallback,after_plan_exit_idle_ctas_must_not_stay_armed_for_the_exited_present,after_plan_exit_chrome_must_not_keep_plan_ready_side_panel_open,isolated_preview_must_paint_current_disk_plan_md_after_exit_and_represent,isolated_preview_dock_after_exit_paints_disk_and_does_not_rearm_plan_ready,isolated_preview_after_exit_represent_paints_disk_not_frozen_sql,after_plan_exit_esc_closes_isolated_preview,after_plan_exit_start_slash_enter_sends_and_does_not_approve,after_plan_exit_empty_enter_never_approves,after_plan_exit_kept_isolated_preview_paints_current_disk_plan_md_not_tech_md,after_plan_exit_esc_clears_isolated_preview_open_marker,after_plan_exit_closed_isolated_preview_composer_must_not_stay_plan,after_plan_exit_closed_isolated_preview_draw_must_not_keep_plan_chrome,after_plan_exit_slash_plan_docks_isolated_preview_not_ignored,after_plan_exit_slash_plan_with_body_submits_plan_update_not_only_stale_preview,slash_plan_with_args_already_in_plan_submits_plan_update,isolated_preview_plan_slash_with_body_submits_plan_update_not_only_stale_preview,isolated_preview_plan_slash_with_body_while_turn_running_sends_not_vanish,isolated_preview_second_plan_prompt_must_not_paint_stale_plan_as_live_present,user_guide_isolated_preview_rewrite_wait_on_second_plan_prompt,leftover_isolated_preview_bare_plan_exclusive_covering_from_current_disk,bare_plan_exclusive_blocks_nested_implementers_plan_soft_keeps_them_working,empty_enter_never_approves_exclusive_covering_present_github_122,isolated_preview_must_not_close_on_nested_specialist_finish,isolated_preview_must_not_vanish_every_couple_of_minutes_on_nested_occupancy_tick,isolated_preview_has_no_plan_exit_wall_clock_timer,plan_soft_must_not_close_on_nested_tick,isolated_preview_soft_planning_does_not_pull_up_leftover_current_plan_md,plan_slash_with_body_is_update_turn_bare_and_soft_are_not,user_guide_plan_slash_with_body_submits_plan_update,after_plan_exit_slash_plan_soft_during_autocompact_docks_isolated_preview,after_plan_exit_without_current_disk_closes_leftover_tech_md_when_disk_is_mill,dock_open_must_not_bump_updated_at_over_rewritten_disk_plan_md,start_leaves_parked_isolated_preview_and_continues_interrupted_work,start_with_nothing_held_still_leaves_parked_isolated_preview,unstick_leaves_parked_isolated_preview_when_hung,user_guide_plan_exit_start_leaves_isolated_preview. - Plan-review and Linux prompt screenshot paste:
Event::Pasteand plan-review Ctrl+V run the clipboard image probe on every OS. Approve and Revise drain composer image chips. Tests:event_paste_plan_commenting_empty_defers_clipboard_image_probe,plan_feedback_ctrl_v_defers_clipboard_image_probe,agent_empty_bracketed_paste_defers_probe_for_clipboard_image,approve_or_revise_drains_plan_composer_images. - No two live same-description Subagent rows: product spawn
rejects a second live Task-owned child with the same trimmed
description on the same parent. It does not replace the first child.
Unlimited retry paints
Retrying (1), neverRetrying (1/4294967295). FiniteRetrying (2/5)stays. Token Economy implement-loop effort is thoroughness, not reviewer count (one reviewer unless the operator asked for more). Tests:live_subagent_list_does_not_show_two_rows_with_the_same_description,task_spawn_rejects_or_replaces_second_live_same_description,format_activity_label_unlimited_retry_has_no_u32_max_fraction,implement_effort_two_does_not_spawn_two_review_rows_unless_operator_asked. - L1 Subagents list is L2-only plus a live L3 count: the L1
Subagents list, watching counts, and similar live chrome show only L2
coordinators. Each L2 row may append a live L3 count (
1 specialist/N specialists). L3 specialists do not get their own L1 rows or names. Opening an L2 still shows that L2's specialists inside the L2 view. HeadlessExtEvent::SubagentSpawnedis not the L1 list. Helpers:live_subagent_list,is_l2_list_row,format_live_l3_count(xai-grok-pagerapp/subagent.rs). Tests:live_subagent_list_shows_only_l2_and_reports_live_l3_count(app/subagent.rs),l2_row_shows_live_l3_count_not_specialist_names(views/tasks_pane.rs). - Live Subagents list is still-running only; already_exited drops the paused Implementer overlay: the live Subagents list, header
:: N, Subagents N, and footer N subagents share one running-only filter (listed_live_subagents). Host exit setsfinished = truethe same turn and the timer stops. Killalready_exited/AlreadyFinishedstill dismisses the paused Implementer overlay (finalize_killed_subagentidles leftover chrome and callsdismiss_nested_overlay). Paused closeout is not live./rebuildoccupancy restore (restore_nested_occupancy_from_disk) must not un-finish or revive a dead host from a snapshot that always hasfinished: false; Occupied rows stay as-is; vacant insert is still-running resume (retain_still_running_nested_occupancy). Grok OSS: SpaceXAI list paint does not encode this already_exited overlay closeout. Tests:kill_already_exited_dismisses_paused_implementer_overlay,restore_nested_occupancy_does_not_unfinish_or_revive_dead_host,running_count_matches_listed_live_l2_not_l3,kill_already_completed_drops_live_list_responding_and_still_running_cue. - Compacting Subagents row
[↗]still opens: Operator[↗]must open the L2 window while that row is Compacting, on any painted row including the top and the last. Compact chrome must not swallow the open hit target.[X]on a Compacting row still kills. AutoCompactStarted still clearsactive_subagentso compact does not auto-steal the parent TUI. Operator[↗]may setvisible_nested_overlay_sidwhile AutoCompacting. Keywords:open_subagent_fullscreenversus AutoCompactStarted auto-steal. Grok OSS: SpaceXAI auto-compact steal is not this Operator open path. Tests:click_tasks_open_on_compacting_row_opens_subagent,click_tasks_open_on_last_painted_row_opens_subagent,click_tasks_kill_on_compacting_row_emits_kill,open_subagent_fullscreen_sets_active_while_child_is_auto_compacting. KEEPnested_compact_chrome_does_not_steal_parent_fullscreen_overlay.nested_compact_chrome_must_not_steal_parent_tui_scrollis the AutoCompactStarted path (active_subagentNone). - L2 implement coordinator strips grep/read/edit (GitHub #141): an L2 coordinator for implement work (
apply_child_tool_policy) must spawn L3 for greps, reads, and product edits. L2 does not fill 200k implementing. Compact on that L2 is a product miss when the cause is L2-solo implement tools.CHILD_TASK_DESCRIPTIONmust contain those Operator strings and must not teach "Easy work can stay on L2", "Including implement loops", or "Spawn L3 only if the problem is actually hard" for greps, reads, or product edits.SubagentCapabilityMode::Allis the existing upstream/full-tool escape. Do not invent a second permission system. Ordinary L2 still AUTO compact at 95% of nested 200k. Do not fold implement L2 into never_auto_compact. L3 never compact. Grok OSS vs SpaceXAI: SpaceXAI nested L2 keeps grep/read/edit. Surmount implement coordinator strips those and keeps spawn. Tests:l2_implement_coordinator_capability_none_strips_search_replace,l2_implement_coordinator_capability_none_strips_grep_and_read_file,l2_implement_coordinator_still_keeps_spawn_subagent,l3_at_max_depth_keeps_search_replace_and_grep_without_task,l2_capability_mode_all_keeps_edit_grep_read,child_task_description_is_concise. KEEPl2_auto_compact_still_fires_at_95_percent_of_200kand the four #140 testsclick_tasks_open_on_compacting_row_opens_subagent,click_tasks_open_on_last_painted_row_opens_subagent,click_tasks_kill_on_compacting_row_emits_kill,open_subagent_fullscreen_sets_active_while_child_is_auto_compacting. - Ctrl+C two-stage: first Ctrl+C with a non-empty draft (text or image chips) clears. Isolated Preview stays. Second Ctrl+C when already empty then Isolated Preview Exit / abandon / cancel. Isolated Preview calls the same
CancelTurnpath as the mill composer (handle_prompt_keyskip-promote when draft, overlaytry_plan_overlay_agent_actionreturns None when draft). Leftover Isolated Preview first-Ctrl+C-to-Exit arm is deleted. Grok OSS: SpaceXAI mill two-stage is the mill composer only; Isolated Preview used to Exit on first Ctrl+C. Tests:isolated_preview_handle_input_ctrl_c_with_text_clears_and_stays,isolated_preview_handle_input_second_empty_ctrl_c_exits,isolated_preview_handle_input_running_turn_draft_ctrl_c_does_not_cancel_turn,leftover_isolated_preview_handle_input_ctrl_c_clears_then_exits. KEEP millctrl_c_idle_prompt_with_text_clears_text,ctrl_c_idle_prompt_with_image_chips_only_clears_chips,ctrl_c_running_prompt_with_text_clears_text_and_preserves_turn,ctrl_c_running_prompt_with_image_chips_only_clears_chips_and_preserves_turn,plan_approval_ctrl_c_clears_draft_then_second_abandons,line_viewer_ctrl_c_clears_draft_then_second_abandons. -
/modellast Tab: when slash dropdown is/modelor/mand exactly one model row is highlighted, Tab and Enter applySwitchModelnow (reuseModelCommand::action_for_args;SetDefaultModelbecomesSwitchModel { effort: None }). Composer clears. No Operator/modelchat. More than one row still completes. Unique trailing-space reasoning row still switches now. Complete typedGrok 4.6 xhighswitches with effort even when the effort dropdown still lists every level. Command-phase unique/modelstill completes/model. Isolated Preview slash Tab intercepts before RowWalk. Ctrl+M picker unchanged. Grok OSS: SpaceXAI Tab is text-only accept; Enter accepted then sent as a prompt. Tests:unique_model_slash_tab_switches_now_empty_composer_no_send,unique_model_slash_enter_switches_now_no_operator_model_chat,unique_m_slash_tab_switches_now,complete_typed_model_xhigh_tab_switches_now_with_effort,command_phase_unique_model_tab_still_completes,multi_row_model_tab_stays_complete_not_switch,isolated_preview_unique_model_tab_switches_now_does_not_rowwalk. - Plan search: Isolated Preview title-bar magnifying glass (
⌕/ ASCIIs) immediately left of copy, which stays immediately left of[↗]. Glass click opensLineViewerState/:search(open_search). Case-insensitive (planmatchesPlanandPLAN). After Enter accepts,n/Njump hits, not typing into the Operator box. Isolated Preview composer/stays slash. Grok OSS: SpaceXAI Isolated Preview has copy and enlarge, not this glass. Tests:plan_preview_title_bar_search_glass_immediately_left_of_copy,isolated_preview_search_query_plan_matches_plan_and_plan,isolated_preview_search_glass_clickable_next_to_copy_and_expand,isolated_preview_composer_slash_stays_slash_not_line_search,isolated_preview_handle_input_n_jumps_hits_after_search. KEEPassert_title_bar_copy_left_of_enlarge. - Isolated Preview screenshot paste: clipboard image paste is an image chip. Isolated Preview must not dump paste into line-viewer search. Search open plus GNOME All Markup / clipboard image must not fill the search bar (
isolated_preview_search_open_paste_does_not_fill_search). GNOME All Markup Copy is an image, not the dialog title (insert_or_defer_bracketed_prompt_pasteusesBracketedDeferredwhen the probe gate is Some). Same helper for millEvent::Paste. Distinct from shipped empty Isolated Preview probe tests. Do not regress paste-chip Enter send (#114). Grok OSS: SpaceXAI bracketed paste inserts the title first then probes. Tests:isolated_preview_gnome_all_markup_copy_title_with_raster_is_image_chip,isolated_preview_search_open_paste_does_not_fill_search,mill_event_paste_gnome_all_markup_copy_title_with_raster_does_not_insert_title,gnome_all_markup_copy_title_still_probes. KEEPisolated_preview_event_paste_must_not_swallow_screenshot_still_cant_paste,user_guide_paste_chip_enter_sends_not_only_expands. - Turbo planning: live exclusive
/planor Isolated Preview/plan --softuses xhigh while[ui].turbo_planningis on (default on). Keywords:effective_reasoning_effort,live_plan_turn,stamp_request_effort,model_effort_chrome_line. Stored session/effortis not mutated. Only the lower-right yellow model/effort line shows xhigh. Magenta model id stays the model id. No TURBO badge, banner, or toast. Settings toggle. Grok OSS vs SpaceXAI: upstream keeps session effort through/plan; Surmount turbo off is that upstream option (plan stays at session effort). Tests:session_medium_enter_plan_request_uses_xhigh_and_lower_right_shows_xhigh,exit_or_approve_plan_returns_session_medium_effort,turbo_planning_settings_off_plan_turn_stays_session_medium,exclusive_plan_turn_uses_xhigh_when_session_is_medium_and_turbo_planning_is_on,isolated_preview_plan_soft_live_turn_uses_xhigh_when_turbo_planning_is_on. - Soft process-rule reminders:
/settingslist injects as soft spawn reminders (ProcessRuleReminders,with_process_rule_spawn_reminder). Spawn still succeeds. A third implementor L2 still spawns. Extra L2s are not auto-killed. Off or empty injects nothing (upstream-like). "only two implementor L2s allowed" is example copy, not a spawn reject. Grok OSS vs SpaceXAI: upstream has no this/settingslist. Tests:process_rule_reminder_configured_third_l2_still_spawns,process_rule_reminder_text_in_nested_spawn_prompt,process_rule_reminders_off_nested_spawn_has_no_extra_reminder_text. - Subagents list compact window counts and TECH.md: nested session
usage is an in-memory map (
agent_view::l2_token_tracking). The nested accumulator is anAtomicU64high-water (fetch_max) so concurrent ACP usage ticks do not race. Subagents list paint uses the liveSubagentProgresssample, not that high-water, so compact cannot leave a stale 90k leftover. Grok OSS: this map is not upstream xAI. The Subagents list suffix is a compact count (format_tokens_compact) with an implicit unit (90k,112.6k,53.4k). Operator-visible chrome must not print the wordtokens(truncation must not become112.6k token...), must not paint a raw integer like 53407, and must not containmeasured. Each L2 row is a live atomic total of that L2's present plus past usage, including every specialist it spawned, with each unit counted once. Specialists still show separately. Do not add nested windows into the parent239K / 500KL1 context chip.sum_live_nested_session_windowsadds each live nested session once and does not add an L3 both inside its L2 figure and again in the total. Internal field names such asmeasured_tokensand the TECH.md table column "measured tokens" stay.format_subagent_labelcallsformat_subagent_label_partssoformat_measured_tokens_suffixis used in the shipped lib. TECH.md at the workspace root (tests inject a temp path) has a description-label L1 to L2 to L3 tree and a table with columns id, contract/aspect, owner, measured tokens, estimate, status. Layout must not parse the session transcript jsonl. Those counts are not included SuperGrok period limits, not SuperGrok dollar credits, and not console team prepaid / console API credits. Tests:subagents_list_omits_the_word_tokens,subagents_list_l2_row_is_present_plus_past_atomic_total_including_specialists,nested_compact_keeps_present_plus_past_without_double_counting_the_surviving_window,nested_specialist_windows_are_not_double_counted_in_the_total,l2_row_paints_present_plus_past_atomic_total_including_specialists,parent_context_chip_is_l1_window_and_does_not_add_nested_windows,format_live_subagents_list_row_uses_live_sample_not_tracker_high_water,subagents_list_truncation_does_not_split_compact_count,subagents_list_shows_measured_tokens_per_nested_l2,format_subagents_list_description_shows_measured_tokens_suffix,format_subagent_label_shows_measured_tokens_suffix,tech_md_write_records_measured_tokens_on_spawn_usage_tick_and_l2_exit,subagents_list_layout_does_not_read_chat_history_jsonl,concurrent_nested_l2_usage_ticks_keep_atomic_u64_high_water. - Always-on bubble copy is paint plus click: flag on paints
⧉. A full-width first line still paints a hit. Click on the human glyph copies that prompt. Click on the assistant glyph copies that message. Paint-only bubble copy is a failed land. Tests:bubble_copy_buttons_on_paints_copy_icon,bubble_copy_buttons_on_paints_copy_icon_when_first_line_is_full_width(scrollback/blocks/user.rs);append_bubble_copy_button_paints_when_first_line_fills_content_width(scrollback/blocks/mod.rs);clicking_human_bubble_copy_copies_the_prompt,clicking_assistant_bubble_copy_copies_the_message,clicking_wide_human_bubble_copy_still_paints_and_copies(app/mouse.rs). Settings row: class 2bubble_copy_buttons_*. - Clear finished is quiet secondary: compact
[−]in the todo header when the board is open and finished rows exist. Never neon green or agent magenta. Hits must not open a subagent. Tests:clear_finished_action_idle_is_quiet_not_neon_green_or_magenta(scrollback/selection.rs);clear_finished_only_when_open_with_finished_rows,clear_finished_hit_does_not_intersect_tasks_subagent_open_or_kill,clear_finished_click_does_not_open_subagent,clear_completed_todos_x_key_only_when_todo_pane_focused. Slash/clear-completed-todosexists. The old pagerSHELL_RESERVED/shell_collision_contract_covers_every_pager_command_and_aliasfnis gone. Do not list that identifier as a land filter. - Pause / resume / stop chips: status
[pause]/[resume]dispatch global pause, not cancel.[stop]is hard cancel only. Soft stop stays keyboard-only (Ctrl+Shift+S); no soft-stop button. Tests:pause_button_click_dispatches_global_pause_not_cancel(app/agent_view/render.rs);work_control_chrome_matrix_pause_not_cancel_stop_not_pause,idle_with_subagents_paints_pause_and_stop_hits,global_paused_idle_paints_resume_not_stop(views/turn_status.rs). - Esc on an L2/L3 overlay dismisses the nested view and leaves that
subagent running. It does not emit CancelTurn and does not start
Cancelling chrome. A prior parent cancel-confirm arm does not fire
while that overlay is open. Esc while not in the overlay still needs
confirm before cancel. Keep-working / interject while Cancelling
aborts the local cancel. Tests:
l2_overlay_esc_leaves_overlay_without_cancelling,l2_overlay_esc_empty_prompt_leaves_overlay_without_cancelling,l3_overlay_esc_leaves_overlay_without_cancelling,l2_overlay_app_esc_dismisses_without_cancel_or_cancelling,l2_overlay_esc_does_not_fire_armed_parent_cancel(app/agent_view/input.rs,app/app_view.rs,app/dispatch/interject.rs);interject_while_cancelling_aborts_cancel,l2_overlay_interject_while_child_cancelling_aborts_child_cancel. User-guide16-subagents,03-keyboard-shortcuts. - Hide header zeros in-app chrome:
[ui] hide_header(default false) zeros the top agent status bar, welcome location top bar, and dashboard location header only. Not window titles. Tests:hide_header_space_dispatches_typed_setter,hide_header_mouse_click_two_stage_toggles(settings_e2e.rs);hide_header_zeroes_status_bar_height,hide_header_zeros_welcome_top_bar_height,hide_header_zeroes_header_and_header_gap. Serde-onlyhide_header_defaults_false_and_parsesis not this class by itself. - Window titles on by default: product manages OSC titles when
[ui.notifications.title] enabled(default true). Never emit an empty window-title OSC. Distinct fromhide_header. Stale[ui] hide_title_baris ignored (stale_hide_title_bar_key_is_ignored). Proven:window_title_always_manages_non_empty_branded_osc,titles_on_session_name_osc_is_non_empty_branded,window_title_osc_payload_never_empty_string. Catalog namesdefault_title_items_include_agents,title_escape_never_empty_payload, andtitle_updates_gated_only_by_title_enabledhave no matchingfn. Do not list those as land filters. - Activity spinner is striped marquee, not braille:
doge_activity_spinners_use_striped_down_marquee_not_braille(xai-grok-pager-renderglyphs.rs). Still-running cue:idle_with_subagents_renders_still_running_cue(views/turn_status.rs). Recap idle rail stays tool-white:recap_accent_and_bullet_use_neutral_tool_color_when_idle. Do not claim a dedicated lower-left throbber colorfn(doge_idle_subagent_still_runninganddoge_tool_running_spinnerare still absent). -
/settingsunread restore set: rows plus runtime readers forhide_header,always_expand_thinking,scrub_ascii_punct,allow_worktree,bubble_copy_buttons,plan_approval_park,composer_multiline, and theme default doge. Tests: settings_e2e prefixes above;composer_multiline_space_dispatches_typed_setter,composer_multiline_mouse_click_two_stage_toggles;theme_choices_include_doge_and_default_is_doge;always_expand_thinking_keeps_blocks_expanded;always_expand_thinking_off_paints_collapsed_headers;always_expand_thinking_finish_overrides_sticky_collapsed;always_expand_thinking_flip_rematerializes_stacked_thinking;set_always_expand_thinking_refolds_live_thinking_in_parent_and_nested_overlay;ctrl_t_expand_is_default_for_next_thinking_block;ctrl_t_collapse_is_default_for_next_thinking_block;ctrl_t_expand_persists_always_expand_thinking;ctrl_t_collapse_persists_always_expand_thinking_off;ctrl_t_turns_always_expand_thinking_off_and_collapses;apply_always_expand_thinking_flip_leaves_aborted_collapsed;prime_applies_always_expand_thinking_from_ui(xai-grok-pager-renderappearance/cache.rs);prime_applies_scrub_ascii_punct_from_ui(xai-grok-pager-renderappearance/cache.rs);resolve_subagents_copies_allow_worktree(xai-grok-shell; copy only, no named test that spawn isolation actually changes). Session recap and cancel-subagents Settings rows are FORK claims, not re-proven as/settingse2e filters on 2026-08-15. - Composer multiline persist and plan Preview Shift+Enter (2026-09-02):
[ui] composer_multilinedefaults on. False makes the Operator box single-line: Enter and Shift+Enter send or interject and never insert a newline. Plan Preview and the main Prompt honor the same flag. Preview Shift+Enter is composer text (plan_preview_key_is_composer_textincludesis_mod_enter); overlay copy / clarify / approve must not steal it. Tests:plan_preview_key_treats_shift_enter_as_composer_text,plan_preview_shift_enter_inserts_newline_when_composer_multiline_on,plan_preview_shift_enter_sends_when_composer_multiline_off,plan_preview_session_multiline_shift_enter_sends,composer_multiline_off_shift_enter_sends_not_newline,composer_multiline_on_shift_enter_inserts_newline,composer_multiline_defaults_on,prime_applies_composer_multiline_from_ui. Do not weaken Ctrl-Z,?insert, or Approve-with-comment. - Composer Shift+Enter is newline (2026-09-09): Shift+Enter inserts
a newline and does not submit, including at the end of the last line
and when session Multiline is on, so the Operator can write a
multiline prompt without submitting. Bare Enter at the end of the
last line still submits (idle send / mid-turn interject). Do not
steal #85.
[ui] composer_multiline = falsestill sends on Shift+Enter. Test:composer_shift_enter_inserts_newline_and_does_not_submit. - Enter and double-click expand hidden blocks (2026-09-09): after
selecting a collapsed or hidden transcript block (image attachment
ellipsis and other folded hidden bodies), Enter expands it, same as
:expand. Double-click on that collapsed block also expands it. Composer Enter with text still sends. Tests:enter_on_selected_collapsed_image_prompt_expands,enter_on_selected_collapsed_tool_expands,double_click_on_collapsed_image_prompt_expands,composer_enter_with_text_still_sends_when_collapsed_image_is_selected. - Aborted thinking is not the live turn (2026-09-02): pause or
cancel must not leave a truncated user-facing draft inside an expanded
thought, and must not paint empty leftovers as
Thought for 0.0s. Internal "the user is asking me..." stays reasoning, not the answer. Tests:abort_turn_does_not_present_aborted_user_facing_draft_as_the_live_turn,abort_turn_omits_instant_empty_thinking_so_thought_for_zero_does_not_paint,abort_turn_collapses_truncated_draft_out_of_expanded_thinking,abort_turn_keeps_internal_reasoning_out_of_the_assistant_answer,thought_chunk_peels_trailing_user_facing_draft_while_streaming,collapsed_header_never_paints_thought_for_zero_point_zero_seconds,aborted_thinking_finished_display_mode_is_collapsed_even_when_always_expand_is_on,peel_trailing_user_facing_draft_keeps_internal_reasoning,aborted_mixed_thought_expanded_body_omits_user_facing_draft. Do not revert resume occupancy, hang-chrome, image-token, Approve-with-comment, Ctrl-Z debounce, WAL, or Operator labels. - Plan cancel and overlay Write finish in a bounded way (2026-09-02):
cancel of a plan-mode or overlay turn must leave
Cancelling…after the resend cap (overlay children are on the turn-end reconcile). A plan-mode turn with a live queue row must not stay Cancelling after the cancel path returns.[stop]during Cancelling must finish cancel, not sit. Queue promotion that changescurrent_prompt_idmust not skip turn-end reconcile. Idle or cancelling plan present typesx/e/j/kin the Operator box; empty Enter never Approves; queue edit plus plan plus Cancelling still CancelTurn. A completed WriteToolCall(not onlyToolCallUpdate) finishes pending Running chrome; a lost Write completion drops that chrome after the short bound. Do not call grok-oss 1.0.3 last-known-good. Tests:cancel_resend_cap_finishes_cancelling_overlay_turn,cancel_plan_mode_turn_with_queue_row_does_not_leave_cancelling_after_cancel_returns,stop_during_cancelling_finishes_cancel,plan_present_xejk_type_in_human_box_even_while_cancelling,nested_overlay_write_clears_running_after_completed_handle_update,write_tool_call_completed_clears_pending_running_activity,stale_write_tool_running_drops_activity_after_bound. - Stuck Retrying / StreamResumed (honesty): pager maps
RetryState::StreamResumedinsession_notification.rs. Shell emit exists:stream_started_emits_retry_state_stream_resumed. Sampler neighbors exist:wait_before_attempt_aborts_on_cancel,retry_footer_reason_uses_short_transport_label,retry_footer_backoff_hint_appends_next_try_in,stream_headers_timeout_defaults_to_120_secs_when_env_unset, pluscargo test -p xai-grok-sampler --test stream_headers_timeout. Catalog pager chrome names (retry_chrome_soft_reconnects_when_retry_stream_starts,stream_resumed_without_prior_retry_clears_activity,clip_retry_reason_*,retrying_activity_label_*,retrying_label_shows_timeout_*) have no matchingfn. Do not claim stuck-retry pager chrome is fully proven. - Click tasks chrome, Worked-for one live line, composer Ctrl+Home/End,
rewind overlays, btw Done-panel, ASCII stream scrub, trailing-whitespace
strip: shipped product behavior (FORK claims / residual-aligned). Not
seven-class land filters unless a named
fnis enrolled later.
- syntect path patch drops dump-load bincode (RUSTSEC-2025-0141).
Workspace
[patch.crates-io]isthird_party/syntect(5.3.0).parsingdoes not enable dump-load/dump-create. yaml-load uses yaml-rust2. Markdown ships.sublime-syntaxfiles and yaml-loads them. two-face dump binaries are gone. Named tests stay:highlight_lines_for_fence_info_still_accepts_rust_token,highlight_lines_for_fence_info_resolves_citation_path_to_rust,highlight_lines_for_token_json_from_bundled_syntax. Advisory: RUSTSEC-2025-0141 (accessed: 2026-08-27). - async-openai path patch keeps
ReasoningEffort::Maxwithoutbackoff. Workspace[patch.crates-io]isthird_party/async-openai(0.33.1 plus Max from our-forks rev95b52ebdedf42143083cf3d6f0e0be7c84e9c808). crates.io 0.41 dropped Max. Named tests stay:test_chat_completion_request_carries_reasoning_effort_top_level,xai-grok-sampling-typesResponses/messages effort maps, pagereffort_levels. Retry 429 / 5xx usestokio::time, notbackoff0.4 (RUSTSEC-2025-0012, accessed: 2026-08-27) /instant(RUSTSEC-2024-0384, accessed: 2026-08-27). Not a workspace member. Operator can later push this tree to our-forks and switch the patch back to git+rev. - AUR sources under
packaging/aur/ - Nix flake:
nix build .#grok-oss, dev shells (human packaging, not GHA release artifacts).flake.nixandflake/are inFORK_PATHS. - NixOS grok-oss workers fragment: grok-oss workers on surmount-1 stay
under existing MemoryMax and below scram. Host imports
packaging/nixos/grok-oss-workers.nix(no second Nix daemon, no boot TUI, optional instance cwd list at sshd class). Named tests in grok-nix-helpernixos_workers:grok_oss_workers_nix_requires_memory_max,grok_oss_workers_nix_does_not_start_nix_daemon,grok_oss_workers_nix_does_not_disable_surmount_scram,grok_oss_workers_nix_has_no_docker,grok_oss_workers_nix_no_boot_tui_and_sshd_class_nice. - Rust 1.98.0 (file pin only; not cargo-proven): project
rust-toolchain.tomlchannelstable(current rust-stable 1.98.0) plus matching fenix FOD inflake/rust-toolchain.nix(channel-rust-stable.toml). After an upstream export that still lists 1.94.x, keep Surmount stable / 1.98.0 unless the operator chooses another channel.rust-toolchain.tomlis not inFORK_PATHS. Import can keep the flake and take upstream's toolchain file. There is no cargofnthat asserts channel1.98.0. Do not add rustc 1.98.0 as a cargo land class until a named test or assert sniff exists. Report:.agents/reports/impl-toolchain-1971-2026-08-12.md - justfile:
just check/just cifull Nix quality gate;just check-localhost cargo (fmt, clippy, nextest, doctest) when the VPS is down;just testfor the cargo quality suite;just updaterefreshes the one workspaceCargo.lockplusflake.lock. Named test: grok-nix-helperjustfile_contractsjust_check_local_is_cargo_only_and_does_not_nix. - justfile helper bootstrap (pinned 2026-08-26).
just require_systemandjust current_systemare a justfile CI_SYSTEM/uname check. They must not require a prebuiltgrok-nix-helperand must not tell the operator to realize.#grok-nix-helperfirst.just check-remote/just require_remote_builderare justfile/uname/SSH preflight. They must notnix build .#grok-nix-helper(that realize copies gigabytes and contends the ssh-ng upload lock).just nix_retry/just flake-meta/ thejust check-remotemetadata step must not requiregrok_nix_helper_bin. The livenix_retrybody is the justfile recipe: argv exec of"$@", fail-fast on quality/SSH, force-remote flags whenGROK_NIX_FORCE_REMOTE=1. Missing helper must not failjust check-remote. Do not tell the operator to realize.#grok-nix-helperfirst.grok_helperassigns the helper path before exec (bashset -edoes not stopexec "$(failing-cmd)"; that becameexec: : not found). Locate order for later recipes that still need the helper (cargo-remote/test-remote, recon):GROK_NIX_HELPER, PATH,result/bin, crate target. Never cargo/rustc the helper on this laptop. Never nix-build the helper fromgrok_nix_helper_bin. Named tests in grok-nix-helperjustfile_contracts:require_system_and_current_system_do_not_require_helper_binary,nix_retry_flake_meta_and_check_remote_do_not_require_helper_binary,grok_helper_does_not_exec_empty_helper_path,grok_nix_helper_bin_locate_order_does_not_cargo_on_force_remote,check_remote_exports_force_remote_before_require_remote_builder,require_remote_builder_is_justfile_preflight_without_helper,check_remote_and_require_remote_builder_do_not_nix_build_helper. - release-dist debug sidecar:
just build-dist/just install-distbuild with--profile release-dist(strip=false, debug=1), extract DWARF togrok-oss.debugviagrok-nix-helper extract-debug-sidecar, strip the binary, embed GNU debuglink. Plainjust installstays local--release+ strip (no sidecar). - Workspace lock matches member manifests; cargo-mem-guard and grok-nix-helper are members (pinned 2026-08-27).
They are workspace members (not
exclude). One rootCargo.lock. Isolated crane builds stay fileset-rooted (flake/grok-nix-helper.nix,flake/cargo-mem-guard.nix) so crane never loads the parent workspaceCargo.toml. Quality.#workspace-cargo-qualitydeps (workspaceCargoArtifacts) runscargo check --locked --all-targets(andcargo build --locked). Do not drop--lockedto go green. When a workspace memberCargo.tomladds a dependency, refresh the rootCargo.lockso that check succeeds. Workspace--workspace --all-targetsclippy and nextest include those crates. Named tests in grok-nix-helperjustfile_contracts:workspace_root_members_include_cargo_mem_guard_and_grok_nix_helper,workspace_quality_deps_cargo_check_stays_locked,workspace_quality_fmt_then_clippy_then_nextest_and_helper_tests. - Vendored bm25 uses rustc-hash, not fxhash (pinned 2026-08-27).
crates.io
bm252.3.2 depends on unmaintainedfxhash0.2.1 (RUSTSEC-2025-0057, accessed: 2026-08-27). There is no published bm25 bump. Workspace[patch.crates-io]pointsbm25atthird_party/bm25(library-only 2.3.2, token ids viarustc-hash). Shell tool-search named tests stay. Do notcargo audit --ignore RUSTSEC-2025-0057. Named test: grok-nix-helperjustfile_contractsworkspace_lockfile_has_no_unmaintained_fxhash. - Vendored rhai uses compact_str, not smartstring (pinned 2026-08-27).
crates.io
rhai1.25.1 and 1.26.0 depend on unmaintainedsmartstring1.0.1 (RUSTSEC-2026-0249, accessed: 2026-08-27). The menhera 10-day cooldown index has 1.25.1; 1.26.0 does not drop smartstring. Workspace[patch.crates-io]and workspacerhaipoint atthird_party/rhai(library-only 1.25.1 from the cooldown cache,SmartStringaliasescompact_str::CompactString). Not a workspace member. xai-workflow named tests stay. Do notcargo audit --ignore RUSTSEC-2026-0249. Named test: grok-nix-helperjustfile_contractsworkspace_lockfile_has_no_unmaintained_smartstring. - Yanked aes / chacha20 / spin are gone from the lockfile (pinned 2026-08-27).
cargo audityanked rows wereaes0.9.0,chacha200.10.0,spin0.9.8 and 0.10.0. Delayed-index bumps:aes0.9.2 (pdf_oxideaes = "0.9"),spin0.9.9 (multer) and 0.10.1 (pprof, crc-fast).chacha200.10.2 is not on the delayed index (0.10.0 / 0.10.1 yanked for SSE2 UB in RNG and legacy 64-bit counter variants; see the chacha20 changelog, accessed: 2026-08-27). Workspace[patch.crates-io]pinschacha20to git tagchacha20-v0.10.2onRustCrypto/stream-ciphers(rev6b236b758a0279f64d777797514813b2cb572c8b). Not a grok-oss path copy. Nix vendor of the yanked crates.io 0.10.x did not exportChaCha12Rngforrand0.10.2. No RUSTSEC/CVE id for that yank yet; RUSTSEC-2019-0029 is a different bug, patched>= 0.2.3. Do notcargo updateagainst crates.io to skip the cooldown. Named test: grok-nix-helperjustfile_contractsworkspace_lockfile_has_no_yanked_aes_chacha20_spin. -
aws-sdk-s3/lrubump is deferred to fargo (pinned 2026-08-27). Operator order: do not bumpaws-sdk-s31.141.0 to 1.144.0 in this grok-oss wave. Remaininglru0.16.4 (RUSTSEC-2026-0253) is not forgotten. Resume in fargo. The delayed crate index still tops at 1.142.0. Do not fetch crates.io to skip that wait. fargo is not specified in this tree. Dual-pin:AGENTS.mdhard constraint 20;RESIDUAL.mdOpen cargo-audit. - fargo must unwind grok-oss path vendoring (pinned 2026-08-27).
Operator does not want grok-oss to vendor crates. Audit-wave
[patch.crates-io]path copies underthird_party/(async-openai, syntect, bm25, rhai, pdf_oxide, ttf-parser) are temporary.chacha20is a git tag pin, not a path copy. fargo replaces each with a delayed-index bump, a Surmount git fork that later enters that index, or dropping the parent. Do not add more path vendoring. Older mermaid/dagre copies inthird_party/are a separate history. Dual-pin:AGENTS.mdhard constraint 21;RESIDUAL.mdOpen fargo unwind. - Test dependencies are supply chain (pinned 2026-08-27). A
vulnerability in
[dev-dependencies], test JWT minting, or an unused test crate is still in the developer lockfile and still inbuild.rsreach. Never call it irrelevant. cargo-audit findings on test deps get the same remove / replace / isolate work as product deps. The menhera-cooldown registry delay is defense in depth against a malicious new crate version, including one that only appears in tests. Do notcargo updateagainst crates.io to skip that delay. cargo-audit is the start of a security pass, not the end (pinned 2026-08-27). Also check yanked crates, RUSTSEC pages, and CVEs for every remaining row (warnings included). Dual-pin:AGENTS.mdhard constraints 17 and 19. - SHA-1 is git object ids only; no bash-in-nix (pinned 2026-08-25).
SHA-1 in this tree is for git object ids (gix, the empty-tree constant,
40-hex commits,
/rebuildidentityversion (git-sha)). It is not a security hash for downloads or Nix FODs. Artifact verify is SHA-256 or minisign. POSIXinstall.sh/install-enterprise.shand PowerShellinstall.ps1/install-enterprise.ps1pin SHA-256 of the published${artifact}.sha256file (fail-closed on miss or mismatch). Windows bootstrap uses built-inGet-FileHash -Algorithm SHA256so it works without Nix. The SpaceXAI internal auto-updater (xai-grok-updateauto_update.rs) pins the same published SHA-256 file, then still smoke-tests--version. The GitHub Releases installer (install_gh_releaseinauto_update.rs) pins SHA-256 of the published${artifact}.sha256GitHub release asset the same way. Neither hashes those bytes with SHA-1. xAI CDN must publish the.sha256files or curl-install fails closed. GitHub Releases must publish${artifact}.sha256assets orinstall_gh_releasefails closed. Those publishes are operator-owned. npm installs still use npm's own integrity pin, not this published.sha256file. POSIX install stays so a host without Nix can curl-install. Hook examples underxai-grok-hooks/examples/hooks/bin/*.shstay.shbecause operators write hooks in shell. Those are not bash-in-nix. Git recon isgrok-nix-helpersubcommands. The helper prepares git state. A human TTY signsgit commit -S. Do not wrap old.shinwriteShellApplication(no bash-in-nix). Helper logs print command names and exit classes only. They must not print tokens, API keys, or secret env values. The operator owns the VPS builder. Agents may runjust check-remoteunderAGENTS.md3b-remote-check (pinned 2026-09-02; one live run at a time). File pin / process pin; not one of the seven land classes. Named crate tests:crate_manifest_does_not_depend_on_sha1_hasher,github_error_excerpt_redacts_token_shaped_fragments_not_git_object_ids,update_config_debug_omits_secret_values,install_scripts_refuse_when_sha256_does_not_match,install_scripts_refuse_when_sha256_checksum_file_is_missing,install_scripts_refuse_when_sha256_checksum_file_is_unreadable,install_scripts_fetch_published_sha256_and_install_when_it_matches,windows_install_scripts_pin_published_sha256_not_sha1,parse_sha256_file_bytes_refuses_unreadable_and_sha1,verify_file_against_digest_refuses_mismatch_and_keeps_previous_good,install_internal_refuses_when_sha256_does_not_match,install_internal_refuses_when_sha256_checksum_file_is_missing,install_internal_refuses_when_sha256_checksum_file_is_unreadable,install_internal_installs_when_published_sha256_matches,install_gh_release_refuses_when_sha256_does_not_match,install_gh_release_refuses_when_sha256_checksum_file_is_missing,install_gh_release_refuses_when_sha256_checksum_file_is_unreadable,install_gh_release_installs_when_published_sha256_matches. See NIST retires SHA-1 (accessed: 2026-08-25) and Git hash function transition (accessed: 2026-08-25). Dual-pin:AGENTS.mdhard constraint 16.
- Process docs hierarchy: D0 residual open-only; D1 AGENTS; D2 logs
under
docs/upstream-*anddoc/dev/campaigns; D3 research / skillreferences/ - Document leftover residual same turn (pinned 2026-08-25). When a
plan or slice ships a first wave, remaining later-wave work goes in
RESIDUAL.mdOpen that same turn, in complete thoughts. Chat is not enough. Do not list finished work as open. Do not omit sibling paths still unfixed (example:install.ps1after a POSIX-only pin). Dual-pin:AGENTS.md§ Residual. - Upstream tooling: detect / import / put-history /
join-main-into-onto / sync via
grok-nix-helper; scheduled export watch workflow - Onto land path: after product is on their tip, join Surmount
mainwithmerge -s oursso the tip is PR-able (docs/upstream-history.md,just upstream-join-main) - PRs accepted: CONTRIBUTING / this fork
- Parent = HITL only; always three layers (2026-08-15) plus
Hierarchical fast path (2026-08-16): process pin in
AGENTS.mdand host~/.grok/AGENTS.md. Whenever implement work, multi-file diagnosis, CI, or a regression needs tools, agents are three layers deep. Including implement loops. Hierarchical fast path (named): the main thread may do a one-command host question, a single known-path read already named, or read and quote the short on-disk report this thread asked for. That is not a license to diagnose or implement in the main thread. Mention is in scope: if the operator mentions work, that mention is in scope. L1 main: status, spawn L2, wait, read short reports, board upsert, Hierarchical fast path. L2: parallelize, spawn L3s, throw context away after a report. L3: all actual tools and work. Same agency as L2 except no L4. Operator clarify stays in the L2 nested view. L3 stays unbothered. Additive asks onto the same live L2 use parent follow-up (not kill, not respawn). Disjoint work still spawns another L2.resume_fromafter exit stays. L3 stays unbothered unless the Operator targeted that specialist. Nesting chrome stays L2-only plus an L3 count. The older weaker law (L2 must spawn L3 only when many greps / half the window) is replaced. L1 AUTO compact uses the catalog 500k window. L2 nested stays 200k and may compact. L3 never compact and must not compact-and-continue. A Grok OSS screenshot from any current working directory is this product. Do not assume another grok-oss window is out of scope. After Approve, do not block on another plan present. Track the work, write a size estimate, implement the groups in parallel, then reconcile the estimate against what landed. Product cargo pins for the prompt contract are under Product (CHILD_TASK_DESCRIPTION: an L2 coordinator for implement work must spawn L3 for greps, reads, and product edits; it must not teach easy L2 work). Assert sniffs that AGENTS still contains the coordinator sentence; that is not the crate seam. Write new short reports under~/.agents/reports/on this machine. Do not add report files to the git tree. Historical.agents/reports/foo.mdcitations in this file are finished-note names only. Productgrok-oss limits multipolldefault out dir is~/.agents/reports/limits-multipoll-<utc>/(temp fallback if HOME is empty). Shipped indefault_multipoll_out_dir. No namedfn. Do not claim repo.agents/reports/is the live home. Fold helperfirst_report_pathmatches any.agents/reports/substring (home or leftover repo path). That is implementation, not a land class. - Kill a think-only L3 after about 15 minutes (pinned 2026-09-09).
L2 must kill an L3 that is still on turn 1 with no useful file or
test progress after about 15 minutes of think-only work or stalled
cargo-verify. Then L2 must spawn a tighter L3, or report failure.
Do not wait forever on 10-minute polls. L3 rambling think dumps
are a failed run, not progress. L1 never product-edits. Dual-pin:
AGENTS.md§ Kill a think-only L3 after about 15 minutes; host~/.grok/AGENTS.mdsame heading; skillhierarchically-structured-subagents. - Fire-and-return (pinned 2026-09-09). Surmount wait law, not
upstream default. Start a long nested job (compile, mill, Lake).
The parent keeps working. It does not sit in a blocking
get_command_or_subagent_outputten-minute loop. When the job finishes, the parent is notified and then does the next step or reports the fail. Return means the parent still owns the outcome. Forget would mean never look at the result. "Background the job" and "don't wait" omit that ownership. Named tests now exist:parent_spawn_subagent_second_l2_while_first_still_running_without_wait(xai-tool-typestask.rsandxai-grok-toolstask/backend_tests.rs) andpost_compact_reminder_includes_surmount_standing_law(xai-grok-shellcompaction_context.rs, helpersection_surmount_standing_law_after_compact). Named tests are contracts: observed red, then green; do not fit the test to a ten-minute poll. Dual-pin:AGENTS.mdand host~/.grok/AGENTS.md§ Fire-and-return. - Take the Operator seriously (pinned 2026-09-09). Live
grok-oss after
/rebuildmust be this workspace's binary. Do not tell the Operator to accept tree versus this TUI when/rebuild, slash, plan Approve, or Enter is supposed to do the thing. Dual-pin:AGENTS.md§ Take the Operator seriously; host~/.grok/AGENTS.mdsame heading. - Cargo tests run on surmount-1, not on horizon (pinned
2026-09-09). For now, all cargo tests run on surmount-1 (the VPS
builder / nixbuilder host). Host horizon (this laptop) must not run
cargo tests. That includes edit-tool verify: the post-edit rustfmt,
clippy, and test pipeline that can invoke rustc or cargo. Horizon
must not run that verify.
GROK_SKIP_EDIT_VERIFY=1is the kill switch for that verify. The product reads only the process environment (GROK_SKIP_EDIT_VERIFYmust equal1).~/.grok/config.tomlhas no such key. Do not invent a new[auth]key. On horizon, export it before starting grok-oss: in fish,set -gx GROK_SKIP_EDIT_VERIFY 1; in a POSIX shell,export GROK_SKIP_EDIT_VERIFY=1. Durable home on this machine is~/.config/fish/config.fish. This already-running grok-oss process does not pick up a later fish export until the Operator relaunches grok-oss. Agents still may runjust check-remoteunder existing law: one live run at a time, and do not restart that run at five minutes.just test-remoteandjust cargo-remotestay operator-owned unless the Operator already whitelist those in the same words. This pin does not weaken agent-depth, fire-and-return, Kill a think-only L3 after about 15 minutes, I hate seeing you edit code at L1, or the rule that the Operator owns the VPS builder. It is additive: horizon is not the cargo-test host. Dual-pin:AGENTS.mdhard constraint 3b-horizon-cargo and the same heading; host~/.grok/AGENTS.mdsame heading. - After a product change, run just install and just check-remote
(pinned 2026-09-16). After a product change in this tree, the same
wave runs
just installandjust check-remote. Not later. Not only when the Operator nags. One livejust check-remoteat a time. Do not restart a live remote compile at five minutes.just test-remotestays operator-owned unless they also whitelist it. Never git commit. Dual-pin:AGENTS.mdhard constraint 3b-after-change-install and the same heading; host~/.grok/AGENTS.mdsame heading. - Subagent worktree policy: prefer isolation none; product default
[subagents] allow_worktree = false. Class 2 copies the flag:resolve_subagents_copies_allow_worktree. User-guide05-configuration+16-subagents. Campaign:doc/dev/campaigns/operator-orchestration-2026-07.md -
/execute-planhonorsallow_worktree: host skill defaults to shared-cwd protocol. Report:doc/dev/research/execute-plan-no-worktree-2026-07-24.md - Todo levels, fib leaves, cleared archive, session notes: product
todo_writesurface (priority,meta, protected prefixes, fibsize1|2,cleared_todos,/note). Not land classes. Reports underdoc/dev/research/todo-*.mdandnotes-channel-2026-07-24.md. - Git recon depth: host skill
/git-recon; productgrok-nix-helper recon-status+just recon-status(read-only probe); pin inFORK_PATHS+assert-process-pins. - Prefer Rust tools; product skills are not a Python runtime:
standing preference plus land class 7. Tool work is a named Rust function,
an ACP tool, or a shipped CLI bin of that function. Skills must not
generate Python or Bash and exec it. Sanitize rejects junk
.py; archive extract skips junk.py; product skill roots have no junk.py. The three allowlisted stub names stay intercept surfaces (memory.py,validate-plan.py,session_reader.py). grok-oss intercepts those names and the CLI binsgrok-oss-implement-memory,grok-oss-plan-validate,grok-oss-session-readerto Rust. Grok Build compatibility is those bins, not Python. Exceptions: those stub names plus office/docx/pptx/xlsx/pdf scripts. Host~/.agents/skillsis operator-owned and is not this class. grok-oss sqlite new session/work ids are ULIDs; UUID is the Grok Build wire id. Tests:sanitize_rejects_non_excepted_skill_python,extract_archive_skips_non_excepted_skill_python,product_repo_skill_roots_have_no_non_excepted_python,default_product_skills_include_polish_and_subagent,default_product_skill_markdown_does_not_tell_agents_to_generate_python_or_bash(xai-grok-bundlelib.rs/default_skills.rs);user_guide_skills_are_not_a_python_runtime(xai-grok-pagerdocs.rs);implement_memory_snapshot_intercept_does_not_spawn_shell,plan_validate_intercept_does_not_spawn_shell,session_reader_list_intercept_does_not_spawn_shell,grok_oss_implement_memory_cli_bin_intercept_does_not_spawn_shell,generated_python_payload_is_not_skill_stub_intercept(xai-grok-toolsbash/mod.rs). A restack that reintroduces non-excepted Python, drops a Rust intercept, or drops those CLI bins, is a failed land. Research:doc/dev/research/python-to-rust-tools-2026-07-26.md - File-level infer-from-path verify (ACP
search_replace/apply_patchand the other structured edit tools): a written.rsfile is formatted and linted as that file. Notcargo clippy -p <crate> --lib, notcargo fmt -p, notjust check. Other extensions do not get Rust cargo. Kill switch:GROK_SKIP_EDIT_VERIFY=1. Helper:xai-grok-toolsutil/rust_edit_verify.rs. Named tests below. A restack that drops the helper or those tests is a failed land. - ACP tools refuse rustc probe junk at the workspace root:
write,search_replace,apply_patch, and the shell tool refuse creating*.rmeta,*.long-type-*.txt,a.out, orrust_outat the workspace root. File-level clippy-driver runs with--out-dirand cwd in a temp directory, so rustc metadata does not land at repo root. The shell tool also refuses rustc one-shots that would write those names (always-approve does not bypass this). Do not gitignore probe junk; prevention is inside the tool call. Helper:xai-grok-toolsutil/compiler_probe_junk.rs. Named tests below. A restack that drops the helper, the temp-dir clippy-driver cwd, or those tests is a failed land. - ACP per-path write lock (
search_replace,apply_patch,write, OpenCodeedit,hashline_edit): each tool takes the path automatically as part of the call. Happy path is silent. A held path is a tool error that names the holder and the file. The tool does not write, wait, or show a human steal, skip, or wait menu. Keywords:try_acquire_write,try_acquire_read,write_paths, CoWpublished(published_cow_snapshot),held(), release after tool return. Share is allowed: two live agents with the samewrite_pathsspawn without error. The hard exclusive lock lasts only for that one edit-tool call, then Drop releases. Spawnwrite_pathsis a soft assignment (reminder to siblings, not a lifetime exclusive lock). File-level infer-from-path verify still runs under the same hold. Helper:xai-grok-toolsimplementations/editor_infra/per_path_write_lock.rs. Grok OSS: SpaceXAI does not encode this CoW reader plus share-is-allowed assignment. Surmount added it so two writers on the same file keep working (GitHub #129). Upstream option stays this same lock table with overlapping assignment allowed. Named tests below. A restack that drops the helper, the OpenCodeeditlock acquire, thehashline_editlock acquire, or those tests is a failed land.
After ACP search_replace / apply_patch (and the other structured
edit tools), the edit tool infers from the path. A .rs file is
formatted and linted as that file. The format and lint argv must
include the written path. That is not crate or project cargo, not
just check, and not an AGENTS process slogan. Markdown, toml, and
other non-.rs paths stay quiet. The command-running tool still
rejects crate-wide cargo launches (cargo fmt --all, cargo fmt -p without a file list, cargo clippy -p ... --all-targets,
--workspace). Kill switch if already in the plan:
GROK_SKIP_EDIT_VERIFY=1.
File-level rustfmt-only may stay on this laptop (no rustc). File-level
clippy that compiles belongs on the remote builder too (just cargo-remote / just check-remote). Agents must not run cargo test,
cargo clippy, cargo build, or rustc on this laptop for grok-oss.
Named filters: just test-remote (see the CI table). For now, cargo
tests run on surmount-1, not on horizon; horizon must export
GROK_SKIP_EDIT_VERIFY=1 before starting grok-oss so edit-tool verify
does not invoke rustc or cargo here. GROK_SKIP_EDIT_VERIFY is still
the kill switch, not a config.toml key and not the product default on
other hosts. See Process Cargo tests run on surmount-1, not on
horizon.
Quality cargo fmt --all -- --check (workspace-cargo-quality) is a hard
miss. rustfmt Diff in is not a flake 502. File-level rustfmt on the
written .rs is how a write stays on that gate (example: wrapping
Result in MCP servers.rs so rustfmt does not emit Diff in).
nix_retry does not retry that class.
This is product behavior. Process law (do not prove the slice by
spawning crate-wide cargo through extra subagents; named cargo on the
remote builder) lives in AGENTS.md hard constraints 3b,
3b-remote, and 3b-remote-named.
Named tests (fixture and argv only; they must not clippy this
workspace). Module filter rust_edit_verify matches these fns:
rustfmt_argv_edition_2024_config_and_absolute_filesclippy_argv_lints_the_edited_file_not_crate_libclippy_argv_includes_bin_path_not_package_libclippy_argv_includes_integration_test_path_not_package_libclippy_argv_is_file_level_not_package_libseveral_rust_writes_run_file_level_clippy_per_fileclippy_driver_uses_temp_out_dir_not_the_workspace_root
ACP refuse rustc probe junk at the workspace root (same crate,
compiler_probe_junk plus the write / search_replace / apply_patch /
bash filters):
names_match_rmeta_long_type_a_out_rust_outroot_rmeta_is_junk_nested_target_is_notrustc_oneshot_is_refused_version_and_tmp_out_dir_are_notwrite_refuses_rmeta_at_workspace_root_and_does_not_create_the_filewrite_refuses_a_out_at_workspace_root_and_does_not_create_the_filewrite_refuses_rust_out_at_workspace_root_and_does_not_create_the_filewrite_refuses_long_type_dump_at_workspace_root_and_does_not_create_the_filesearch_replace_refuses_a_out_at_workspace_root_and_does_not_create_the_fileapply_patch_refuses_add_rmeta_at_workspace_root_and_does_not_create_the_filerustc_oneshot_without_out_dir_is_refused_and_does_not_spawn_shellrustc_stdin_rust_out_is_refused_and_does_not_spawn_shellrustc_dash_o_a_out_at_workspace_root_is_refused_and_does_not_spawn_shellredirect_rmeta_at_workspace_root_is_refused_and_does_not_spawn_shell
Workspace hygiene still says do not gitignore probe junk. That stands.
The product fix is the tool refuse, not a mop and not .gitignore.
Command-tool reject (same crate, dangerous_cargo filter):
dangerous_cargo_fmt_all_is_refused_and_does_not_spawn_shelldangerous_cargo_fmt_package_without_file_list_is_refused_and_does_not_spawn_shelldangerous_cargo_clippy_all_targets_is_refused_and_does_not_spawn_shelldangerous_cargo_clippy_package_all_targets_is_refused_and_does_not_spawn_shelldangerous_cargo_clippy_workspace_is_refused_and_does_not_spawn_shelldangerous_cargo_test_workspace_is_refused_and_does_not_spawn_shelldangerous_cargo_nextest_run_without_package_or_filter_is_refused_and_does_not_spawn_shelldangerous_cargo_test_package_lib_filter_is_not_refused
Catalog:
doc/dev/upstream-regression-filters.md
§ File-level infer-from-path verify. Extra restack-droppable class, not
one of the seven numbered land classes.
ACP search_replace, apply_patch, write, OpenCode edit, and
hashline_edit (GrokBuildHashline:hashline_edit) take a per-path
write lock automatically as part of the tool call. There is no lock
argument on the tool schema. A successful write does not mention the
lock.
When another agent already holds that path, the tool returns an error that names the holder and the file. It does not write. It does not overwrite silently. It does not wait inside the tool. It does not show a human steal, skip, or wait menu. Agents resolve the conflict by talking to each other: they can wait, hand off, or pick another path.
The lock is held through file-level infer-from-path verify on a written
.rs file. GROK_SKIP_EDIT_VERIFY=1 still skips only that verify.
OpenCode edit (tool id "edit") acquires after directory, same-string,
and bulk-edit checks, and before create or replace. The guard stays in
run so rustfmt and clippy-driver on the same .rs path stay under
the hold.
hashline_edit acquires on the joined path after resolve_model_path,
before canonicalize or any write. Existing-file edits and new-file
Write both take the lock. Same helper; no second table; no human menu.
Named tests (module filter per_path_write_lock):
two_agents_cannot_write_the_same_path_at_oncehappy_path_first_writer_succeeds_silentlylock_releases_after_the_tool_call_so_a_later_call_can_writesearch_replace_apply_patch_and_write_all_take_the_lockheld_path_error_names_holder_and_file_without_a_steal_skip_wait_menuhashline_edit_refuses_when_another_agent_holds_the_pathhashline_edit_happy_path_does_not_mention_the_locksequential_writes_succeed_after_the_first_tool_call_returns_even_when_both_agents_were_assigned_the_same_write_pathsconcurrent_in_flight_writes_on_the_same_path_still_conflictspawn_write_paths_soft_assignment_does_not_block_a_sibling_and_the_reminder_is_observablesame_holder_can_write_a_path_they_reservedsequential_search_replace_succeeds_after_the_first_tool_call_returns_when_both_agents_were_assigned_the_same_write_pathssearch_replace_succeeds_when_a_sibling_only_has_a_soft_write_paths_assignmentsoft_lock_reminder_is_observable_on_a_sibling_tool_callspawn_write_paths_overlap_is_a_soft_assignment_not_a_spawn_errorcow_snapshot_read_is_ephemeral_many_readers_one_writerread_file_uses_cow_snapshot_and_does_not_take_the_exclusive_write_lockafter_write_returns_held_is_empty_lock_must_be_releasedreader_during_held_write_gets_published_pre_write_bytes_current_atomic_snapshottwo_live_agents_with_the_same_write_paths_spawn_without_error_l2_and_l3_may_be_assigned_the_same_file
Spawn write_paths on task / spawn_subagent is a soft assignment. Other nested agents get a reminder (L2 X is assigned these paths). Share is allowed. Spawn and later sequential edits do not fail for the child's lifetime. Exclusive is try_acquire_write for one search_replace / write / apply_patch call, then Drop. Two agents still cannot write the same file at the same instant. After write returns, held() is empty.
The path table is a reader-writer lock, not write-only. try_acquire_read is a CoW snapshot read: ephemeral, many concurrent readers, snapshot at a point in time. It does not take the exclusive write lock, does not block a writer, and is not blocked by a writer for the snapshot itself. read_file uses that CoW published snapshot (published_cow_snapshot) while a writer holds the path. Soft write_paths assignment stays a writer reminder.
cargo test -p xai-grok-tools --lib per_path_write_lock
cargo test -p xai-grok-tools --lib spawn_write_paths_overlap_is_a_soft_assignment_not_a_spawn_error
cargo test -p xai-grok-tools --lib soft_lock_reminder_is_observable_on_a_sibling_tool_call
cargo test -p xai-grok-tools --lib -- \
after_write_returns_held_is_empty_lock_must_be_released \
reader_during_held_write_gets_published_pre_write_bytes_current_atomic_snapshot \
two_live_agents_with_the_same_write_paths_spawn_without_error_l2_and_l3_may_be_assigned_the_same_fileOpenCode edit fixture (not under that module filter):
opencode_edit_cannot_write_a_path_another_agent_already_holds
cargo test -p xai-grok-tools --lib opencode_edit_cannot_write_a_path_another_agent_already_holdsSkills are loaded from several places; the product on this branch owns the
machinery. Full map: doc/dev/research/where-skills-come-from-2026-07-24.md,
user-guide 08-skills.md.
| Source | Role |
|---|---|
Project .agents/skills, .grok/skills |
Git-trackable on the branch (supported; may be empty). Not the home for /polish or /subagent. |
crates/codegen/xai-grok-bundle/skills/ |
In-tree Grok OSS default skills (polish, subagent, what, pull-remote-tree). Installed into ~/.grok/bundled/skills/ on startup and after network extract. Live cache is not the source. When the operator asks to revise a skill in grok-oss, edit this tree. Named tests are the contract so skill maintenance and upgrades cannot drop it. |
~/.agents/skills then ~/.grok/skills |
Host operator overlay (agents wins) |
[skills].paths / server inject / plugins |
Config and managed dirs |
~/.grok/bundled/skills |
Platform cache from network bundle sync plus installed Grok OSS defaults |
Process pins that must survive recon (import / onto): document in FORK +
AGENTS + product user-guide when product-facing; dual-pin host skills
(~/.agents) when operator-only. Host skill git alone does not ride product
history. Chat-only pins die at compaction.
The shared guide under crates/codegen/xai-grok-pager/docs/user-guide/ is
not in FORK_PATHS. Onto takes the xAI guide unless conflict resolve
keeps Surmount pages. Do not paste those pages here.
| Page | Fork pin | Cargo pin |
|---|---|---|
01-getting-started |
Binary is grok-oss. Bare interactive open is last session for this cwd, not Welcome. |
Last-session sentences shipped in code; no dedicated fn. |
02-authentication |
SuperGrok is paid. Distinct meters. /limits and compact chip. Hop after included SuperGrok period limits are full. Fail-open: a client 100% / remaining 0 / $0 printout must not mark SuperGrok used up. Named /limits words and limits_pins.json. grok-oss limits is not xAI billing truth. Machine console API key for host surmount-1: console API credits / console team prepaid, install under that host grok home ($GROK_HOME or ~/.grok), never commit the key, no guest git, GPG sign on the laptop, attach SSH + tmux as user grok, L0 is not /dashboard. |
user_guide_does_not_claim_automatic_host_hop_is_unshipped, user_guide_limits_names_fail_open_and_named_commands, user_guide_machine_console_api_key_for_surmount_1. Zero /limits hits is a failed land in catalog prose; no cargo hit-count fn. |
03-keyboard-shortcuts |
Plan keys and Enter cue (send / queue / interject). Empty Enter never approves a plan. Nested L2/L3 overlay Esc dismisses the view and does not cancel. | Plan honesty fns under Chrome. Overlay Esc: l2_overlay_app_esc_dismisses_without_cancel_or_cancelling. |
04-slash-commands |
/running (alias /windows) lists live grok-oss TUI windows. Not Agent Dashboard. L0 is Surmount GPUI and must not merge with /dashboard or /running. L0 action set remote host console API key is laptop-side for a machine console API key (never prints the key; no guest git). /start starts paused or interrupted work in this process; not /resume. /unstick resends the last parent prompt as if the network dropped it; orphans a hung in-flight prompt; WAL images resend as resource links, never data URLs; not /resume, not a second Operator line, not unwind. /finish writes a session post-mortem (work continues; leftover and next features stay first-class; not /dream, not /recap, not /reports). /reports writes a checkpoint while work continues (host overlay ~/.agents/skills/reports plus pager slash). /polish is a polish pass as a default Grok OSS skill (in-tree crates/codegen/xai-grok-bundle/skills/polish, installed into ~/.grok/bundled/skills/polish; not host overlay, not a pager builtin, not a project .agents/skills/polish pack; not /finish, not /reports). /subagent (and /subagent this) spawns one L2 coordinator as a default Grok OSS skill (in-tree crates/codegen/xai-grok-bundle/skills/subagent, installed into ~/.grok/bundled/skills/subagent; not host overlay, not a pager builtin, not a project .agents/skills/subagent pack; L1 does not do the job). /what restates this session in four complete thoughts (Job, State, Operator, Next) when chat is unclear. Speaker labels are Operator not You or Human, and Agent not Me or Grok when Grok means the assistant. Never mix Grok Build version with grok-oss product version. Isolated Preview and plan chrome are grok-oss unless this process was launched as grok from downloads. Default Grok OSS skill at crates/codegen/xai-grok-bundle/skills/what, installed into ~/.grok/bundled/skills/what. Not host overlay as the grok-oss source. Not repo .agents/skills/what. Follow Concise American Technical English (0005_CATE.md). /compaction aliases /compact. Named hold (queue/later or /queue <slash>) puts /compaction, /plan, /reports, /finish on the existing composer prompt queue without running them this turn. Immediate invoke stays. Present is not Approve. /metadata shows ULID, UUID, cwd, model, started, pid, plus last Chat Completions system_fingerprint and last language-models {id, fingerprint, version, created} for the current sampling model (serving-path config, not a SHA of the weights, not on the status bar). /limits named words: stay-supergrok, use-console, meter included or dollar-credits or console or combined, refresh. Fail-open printout must not mark SuperGrok used up. Once prompt_wal.jsonl exists in the tree, /rebuild must say relaunch preserves that WAL. Do not document that preservation as shipped while the file is absent. |
running_slash_lists_sibling_fixture_row; L0 surmount-coordinator-gui write_enqueue_creates_per_session_file, omits_prompt_text, keeps_pid_session_cwd, enqueue_drop_path_is_per_session_id, CoordinatorApp_selects_row, CoordinatorApp_omits_prompt_in_displayed_fields, CoordinatorApp_enqueue_writes_drop_file, set_remote_host_console_api_key_never_prints_the_key, set_remote_host_console_api_key_is_not_pager_dashboard; user_guide_machine_console_api_key_for_surmount_1; /start cite start_* tests; /unstick cite unstick_* tests; finish_empty_args_injects_postmortem_skill; finish_skill_copy_does_not_say_work_is_closed_forever; reports_empty_args_injects_reports_skill; what_empty_args_injects_what_skill; what_instruction_prefers_operator_and_agent_speaker_labels; what_skill_does_not_mix_grok_build_version_with_grok_oss; user_guide_what_does_not_mix_grok_build_version_with_grok_oss; what_registered_in_builtin_commands; queue_compaction_does_not_invoke_immediately; queue_plan_does_not_invoke_immediately; metadata_command_emits_show_session_metadata; show_session_metadata_includes_fingerprint_fields_from_stored_samples; user_guide_metadata_documents_serving_fingerprints; user_guide_limits_names_fail_open_and_named_commands. No user_guide_*start* fn. Guide still documents grok-oss rebuild; that page is not cargo-proven for CLI rebuild. |
05-configuration |
hide_header is in-app only. Titles use title.enabled. [subagents] allow_worktree defaults false. [ui] composer_multiline defaults on; false makes the Operator box single-line. |
Class 2 readers. Do not claim Token Economy /settings table rows as proven. |
06-theming |
Default theme is DOGE. Operator green / agent magenta roles. | Class 4 theme + rail fns. user_guide_operator_agent_speaker_labels_not_human_user_grok. |
08-skills |
Product skills are not a Python runtime (allowlisted CLI stubs + office/docx/pptx/xlsx/pdf only). /polish, /subagent, /what, and /pull-remote-tree are default Grok OSS skills (in-tree crates/codegen/xai-grok-bundle/skills/, installed into ~/.grok/bundled/skills/). Revising a skill in grok-oss edits that tree. Not repo .agents/skills/what. |
user_guide_skills_are_not_a_python_runtime; default_product_skills_include_polish_and_subagent; what_empty_args_injects_what_skill; what_instruction_prefers_operator_and_agent_speaker_labels; what_skill_does_not_mix_grok_build_version_with_grok_oss; user_guide_what_does_not_mix_grok_build_version_with_grok_oss; user_guide_operator_agent_speaker_labels_not_human_user_grok |
16-subagents |
Worktree isolation off by default. Soft interject never cancels. Three-layer paragraph. Hierarchical fast path (L1-only). L1 Subagents list is L2-only plus a live L3 count. Live list is still-running only. already_exited still dismisses the paused Implementer overlay. Compacting [↗] still opens. L2 overlay is a mid-turn ask to that L2. L3 overlays stay unbothered. Esc on the nested view dismisses it and leaves the L2 running (not Cancelling). New reports under ~/.agents/reports/. L1 AUTO compact uses catalog 500k. L2 nested 200k may compact. An L2 implement coordinator must spawn L3 for greps, reads, and product edits. L2 does not fill 200k implementing. capability_mode All is the full-tool escape. Ordinary L2 still AUTO compact at 95% of nested 200k. L3 never compact and must not compact-and-continue. Parent follow-up onto a running nested L2 is additive (not kill, not respawn). resume_from remains completed-continue. L3 overlays stay unbothered. Off is the SpaceXAI option (spawn, wait, resume after exit). |
Three-layer / fast-path / L2-only guide text shipped in code; no dedicated user-guide fn. Cargo: child_task_description_is_concise, l2_implement_coordinator_capability_none_strips_search_replace, l2_implement_coordinator_capability_none_strips_grep_and_read_file, l2_implement_coordinator_still_keeps_spawn_subagent, l3_at_max_depth_keeps_search_replace_and_grep_without_task, l2_capability_mode_all_keeps_edit_grep_read, l2_auto_compact_still_fires_at_95_percent_of_200k, live_subagent_list_shows_only_l2_and_reports_live_l3_count, kill_already_exited_dismisses_paused_implementer_overlay, restore_nested_occupancy_does_not_unfinish_or_revive_dead_host, running_count_matches_listed_live_l2_not_l3, click_tasks_open_on_compacting_row_opens_subagent, click_tasks_open_on_last_painted_row_opens_subagent, click_tasks_kill_on_compacting_row_emits_kill, open_subagent_fullscreen_sets_active_while_child_is_auto_compacting, l2_overlay_send_prompt_interjects_l2_not_l1, nested_reparent_stamps_l3_depth_and_immediate_parent, l2_overlay_esc_leaves_overlay_without_cancelling, l2_overlay_app_esc_dismisses_without_cancel_or_cancelling, l2_overlay_esc_does_not_fire_armed_parent_cancel, parent_cannot_talk_to_own_l2s_follow_up_enqueues_interject_on_running_l2_without_kill_or_respawn, parent_follow_up_does_not_inject_into_live_l3_unless_operator_targeted_that_specialist, resume_from_of_running_l2_still_fails_active, parent_follow_up_off_is_upstream_spawn_wait_resume_from_completed_only, parent_follow_up_onto_running_l2_with_live_l3_hits_l2_not_l3, parent_cannot_talk_to_own_l2s_task_tool_run_follow_up_returns_queued_and_does_not_spawn, parent_cannot_talk_to_own_l2s_task_tool_run_follow_up_and_resume_from_are_mutually_exclusive, parent_cannot_talk_to_own_l2s_follow_up_field_is_optional_and_distinct_from_resume_from, shell_child_follow_up_sends_session_command_interject_on_child_session, subagents_config_parent_follow_up_false_parses_and_omitted_defaults_true, resolve_subagents_copies_parent_follow_up, nested_spawner_can_resume_from_completed_reparented_child, l3_overlay_send_prompt_does_not_reach_l3_or_l1. |
17-sessions |
Last-session on start vs -c / --resume vs /start vs leftover canceled_turn_resume.json drop after a successful primary-turn finish. Running grok-oss sessions vs disk grok-oss sessions. Resume examples use grok-oss. |
user_guide_resume_and_version_examples_use_grok_oss; /start + marker-drop cite start_* and session_load_drops_stale_cancel_resume_marker_when_primary_turn_finished_successfully. |
19-plan-mode |
Present is not Approve. Idle footer is Approve / Comment / Revise / Exit. Clarify only after Comment. Empty Enter never approves. Copy/y is not a fifth idle CTA (title bar + hint). Approve files a GitHub issue with the plan text (docs/github-tracking.md). Selected CTA is marked. Enter submits the marked CTA. Click marks the CTA and runs it. First click on Approve still Approves. Letter keys type. Freeform questions, not the questionnaire modal. /plan --soft docks Isolated Preview and does not enter plan mode. |
Extra class B fns. Keep identifier plan_approval_footer_paints_five_cta_vocabulary. user_guide_plan_soft_docks_isolated_preview. Copy/y and selected-CTA Enter: y_copies_the_plan_while_the_comment_overlay_is_open, plan_approval_pane_has_a_clickable_copy_control, plan_approval_cta_row_does_not_paint_copy, plan_approval_copy_button_click_copies_the_plan, selected_idle_cta_is_visually_marked, enter_submits_the_marked_idle_cta, enter_while_composing_a_comment_still_saves_the_comment, empty_enter_never_approves_even_when_approve_is_marked, click_selects_a_cta_and_first_click_approve_still_submits, second_click_on_already_selected_cta_still_submits, letter_key_types_and_is_not_the_only_submit. |
22-permissions-and-safety |
Always-approve is tool permissions only, not plan Approve. | exit_plan_mode_shows_overlay_even_in_yolo |
23-dashboard |
Agent Dashboard is this pager. Running grok-oss sessions must not merge into /dashboard. L0 is Surmount GPUI, not this pager, and must not merge with either. Call L0 grok-oss gui. L0 action set remote host console API key is laptop-side, not this pager. Session todos stay in this TUI. |
Cite omits_prompt_text, set_remote_host_console_api_key_is_not_pager_dashboard, user_guide_machine_console_api_key_for_surmount_1. |
24-monitoring-usage |
/spend ledger vs org metrics. Do not mash meters. |
user_guide_names_token_economy_spend_order |
Also: user_guide_operator_cli_examples_use_grok_oss (leftover grok login /
grok sessions must not return).
This section is a dated operator handoff from 2026-08-09. It does not
demote later shipped claims in Product, Chrome, or Land above. Source restore
is not live TUI dogfood. This file does not claim a rebuilt interactive
grok-oss. Do not start rebuild or quit from this inventory.
Read first (next task, recon, or dogfood):
- Operator install gate: code can be green while old TUIs still run a
deleted-inode binary. Checklist:
.agents/reports/d0-dogfood-checklist-2026-08-09.md. Package status for that wave:.agents/reports/impl-remaining-plan-wave-2026-08-09.md. - Shipped in that wave (tree + reports; dogfood only after a later
install the operator chooses): decisive plan Revise; same-batch
plan.mdwrite +exit_plan_mode; plan decision surface (empty Enter never approves; Plan ready. Side panel open); revise/clarify in-flight status; caret empty halftext_primary; OAuth 403 bad-credentials path; rewind missing intermediate checkpoints; Ctrl+C dismisses rewind; status[pause]/[resume]+ red[stop]; soft stop chord-only; composer Entersend/queue/interjectcue; compact included SuperGrok period limits meter. Later Product / Chrome bullets and namedfns supersede this snapshot when they disagree. Later Product bullets for/startand leftovercanceled_turn_resume.jsondrop supersede this snapshot. Do not revive idle Clarify as current leftover. - Still not shipped (honesty leftovers, not a demotion of shipped chrome)
- Auto-resume after error terminal on rebuild/reopen: expected
operator contract if the last terminal was an error (not only the
cancel-resume marker). Document as shipped only when
.agents/reports/impl-rebuild-auto-resume-after-error-2026-08-09.md(or equivalent) is green in tree. Distinct from continue interrupted turn (canceled_turn_resume.json). After any auto-resume, 403 bad-credentials may still need/login. - Soft-stop button: not shipped; soft stop stays
Ctrl+Shift+Sonly. - Mid-sample freeze without cancel: not shipped (global pause cancels turns; soft stop only stops queue drain after the current turn). Do not invent a media-player freeze metaphor.
- CLI
grok-oss rebuild: clap-wired (Command::Rebuild,rebuild_subcommand_parses). Compiles and signals live instances. Does not re-exec this process. TUI/rebuildstill owns persist plus self re-exec. /economic-modeslash: pager queues the text only. No BuiltinAction. Economic cap at spawn / model switch / header is shipped.- SuperGrok Heavy ranking optional label: not implemented. SuperGrok Heavy is a real distinct weekly pool. This file does not diagnose product usage of that pool.
- Auto-resume after error terminal on rebuild/reopen: expected
operator contract if the last terminal was an error (not only the
cancel-resume marker). Document as shipped only when
- Still open (residual, not this install alone)
- Included SuperGrok period flat % / server C4 debit: paste-ready ticket
.agents/reports/c4-xai-ticket-paste-ready-2026-08-07.md; never invent included SuperGrok period used % on the client. - Thoughtful todo tracking process (session board hygiene; not
file-level edit verify):
RESIDUAL.mdOpen.
- Included SuperGrok period flat % / server C4 debit: paste-ready ticket
- Useful regression filters from that wave only (not a substitute for the seven-class land list):
cargo test -p xai-grok-pager --lib -- after_revise re_present_after_revise \
paint_composer_box_cursor_grapheme_phases_keep_letter \
left_through_letters_empty_phase_not_neon
cargo test -p xai-grok-shell --lib -- same_batch_plan_write_before_exit_plan_mode_returns_new_body \
replay_skips_missing_intermediate_checkpoint
cargo test -p xai-grok-sampling-types --lib -- forbidden_bad_credentials
cargo test -p xai-grok-sampler --lib -- api_403_bad_credentials classify_forbidden
cargo test -p xai-grok-pager --lib -- work_control_chrome_matrix_pause_not_cancel_stop_not_pause \
pause_button_click_dispatches_global_pause_not_cancel
# File-level infer-from-path verify (2026-08-15; extra class, not that wave)
cargo test -p xai-grok-tools --lib rust_edit_verify
cargo test -p xai-grok-tools --lib dangerous_cargoProcess law (plain English, no bad metaphors): host + project AGENTS.md
§ Prose + tone / hard constraint 4. Not re-dumped here.
| Path | Import | Put-history | Join (-s ours) |
|---|---|---|---|
Paths in FORK_PATHS (AGENTS, RESIDUAL, FORK, docs/upstream-*, crates/codegen/grok-nix-helper, .grok/workflows, .agents/skills, doc/dev, flake.nix, flake/, ...) |
Restored from base; post-restore assert-process-pins |
Via cherry-picks | Tip tree kept |
| Product commits after seed | N/A (tree = xAI + restore) | Cherry-picked onto tip | Tip tree kept |
Paths not in FORK_PATHS and absent from xAI |
Dropped | Only if stacked | Cannot backfill missing |
| Shared user-guide / crate seams | xAI base | Conflict resolve | Tip tree only |
Host ~/.agents/skills, ~/.grok/AGENTS.md |
Untouched | Untouched | Untouched |
rust-toolchain.toml |
Not in FORK_PATHS; import can take upstream's file |
Only if stacked | Tip tree only |
Assert: grok-nix-helper assert-process-pins or just upstream-assert-process-pins.
That command proves files exist. It does not prove product contracts
inside xai-grok-*. Detail: doc/dev/research/fork-paths-hardening-2026-07-24.md,
doc/dev/research/skills-survive-upstream-recon-2026-07-24.md,
docs/upstream-history.md.
Novel Surmount crates use the grok-* prefix (example: grok-rate-limit).
Upstream crate paths stay xai-grok-* for mergeability.
A restack-droppable seam is not defended by a FORK checkbox. Defense is four things that stay aligned:
- FORK line: one complete sentence in this inventory with the named
fnand crate (or an explicit “shipped in code, no named test” / “file pin only” label). - Named cargo test: a
fnthat goes red if the seam is deleted. - Catalog row: enroll the filter in
doc/dev/upstream-regression-filters.md. Recon agents walk these classes. The process improver owns that catalog and the assert / git-recon wiring. - Cherry-pick: product seams inside
xai-grok-*survive onto only via cherry-pick plus those tests.
Import restores docs and scripts only (FORK_PATHS). It does not restore
crate tests. grok-nix-helper assert-process-pins proves files exist. It does not
prove contracts. just check cannot fail a deleted catalog test. A
chrome-only inventory is a failed land. Paint-only bubble copy is a failed
land. Reintroducing non-excepted Python under product skills is a failed land.
Do not claim "Surmount seams survived" until this list is done. just check
is quality only. It cannot fail a deleted catalog test.
A named cargo fn is intended product behavior. Do not reshape asserts to
match code. Do not skip Nix-sandbox S3, MCP, or bwrap tests to go green.
Hermeticity fixes keep that contract: PATH includes bash, $GROK_HOME is
writable, TLS uses webpki roots when the OS store is empty, bwrap
placeholders exist without dropping deny binds, and grok argv0 fixtures
work on Nix coreutils (a multi-call binary). Land proof is those named
fns below, not a skipped subset. After compact, standing Surmount law
is the first section of the post-compaction <system-reminder>, not a
buried AGENTS.md paragraph. Process: AGENTS.md hard
constraint 15.
The operator's words are the spec (pinned 2026-09-01). A named cargo
fn encodes the operator's named contract, not a weaker paraphrase of
what they probably meant. Repeating the same prompt three times (tell,
then after nonsense, then verbatim because the build still does not
match) is a failed land of that contract, not a reason to hold the
operator's hand. Observed red, then product so the same test passes.
Do not fit the test to the code. Tests are Surmount contracts. Keep
the stronger assert. Adjacent to § Hunter's razor; it does not
replace it. Dual-pin: AGENTS.md § The operator's words
are the spec (hard constraint 23).
Wasted human time (pinned 2026-09-01). A dropped operator prompt is
a product defect. Repeating because the harness lost the text is an
engineering miss. The durability path is prompt_wal.jsonl. Enroll
named WAL tests in the catalog when they exist. Do not document
/rebuild WAL preservation as shipped while that file is absent.
Operator-verified known good (2026-09-02) for named tests that encode
live send, rebuild-flush, interject, and plan-notes appends,
because live session files contained those kinds. Catalog:
doc/dev/upstream-regression-filters.md
§ Prompt write-ahead log. Queue enqueue is still a named contract; do
not mark it known good from that date. Restore and skip tests stay
contracts. Do not delete or weaken those tests in recon, onto, import,
or join. Adjacent to this paragraph, Named tests are contracts,
lost-prompt, and § Hunter's razor; it does not replace them. Dual-pin:
AGENTS.md § Wasted human time (hard constraint 24).
Rules (not product class numbers):
FORK_PATHSrestore is docs and scripts only. Product seams insidexai-grok-*survive onto only via cherry-pick plus cargo tests.grok-nix-helper assert-process-pinsproves files exist. It does not prove contracts.- A chrome-only inventory is a failed land. Paint screenshots of rails
and four idle plan CTAs do not prove hop keys,
/spendingest, unread config, first-tokengrok-oss, last-session, or skills-not-Python. - Paint-only bubble copy is a failed land. Click-to-copy tests must still exist.
- Product skills reintroduced as a Python runtime is a failed land.
- Catalog hook (process improver): recon agents walk these classes in
doc/dev/upstream-regression-filters.md. A sibling name-existence check (not this file, notREQUIRED_FILES) may later fail if a required-land identifier has no matchingfn. Path assert stays files-only.
Steps (procedure; do not mix these into the 1-7 product count):
- Run
just upstream-assert-process-pins(orgrok-nix-helper assert-process-pins HEAD). Files and light sniffs only. - Run the named cargo filters for the seven product classes below, plus
the extra proven restack-droppable classes. Use existing test names. Do
not invent a filter that is not in the tree. Do not list an identifier
that has no matching
fn. rgeach required identifier for a matchingfn. A named filter with no matchingfnis a failed land.- Helper-green is a failed land. Forbidden as proof: a
--versiontest that only checks stdout contains the substringgrok; catalog-exists without paint; schema-exists without/spendingest; serdehide_headerwithout a/settingsrow and a runtime reader; rank helpers withoutsampling_confighop keys; bundle still hasmemory.py. - Dogfood screenshots (rails, four idle plan CTAs, compact included SuperGrok period
limits meter, SIGUSR1 after a failed install) stay an operator check
after those
fns exist. They are not the only check. This inventory does not claim live TUI dogfood.
Seven product classes (must match the catalog; each proven by a named
cargo fn):
- CLI identity. The product command is grok-oss.
grok-oss --versionfirst token isgrok-oss, not baregrok. Resume and relaunch hints aregrok-oss --resume. Welcome / tutorial badges say Grok OSS. - Config is a surface, not a field. A toml field that deserializes is
not shipped if
/settingshas no row and no runtime reader. Restack lost unread keys (hide_header, always-expand thinking, plan park, worktrees, ASCII scrub at launch, bubble copy) and leftover/settingsrows plus DOGE in the theme picker. - Token Economy ledger
/spend(extra SQL, not SuperGrok dollar credits).$GROK_HOME/grok_oss.dbis the Token Economy ledger, not the session store. Schema v1 surviving is not enough./spendmust ingestusage.jsonland writereconciliation_run(notDoubleEntryReport::default()). - DOGE / Surmount chrome. A theme file existing is not paint. Land must
keep paint/render tests for human green rails plus box caret, magenta
model / running agent, the compact included SuperGrok period limits
meter, the titled composer frame (
prompt_border_activewhite, yellow title only), and the four-CTA idle plan panel (Clarify only after Comment). - Dual-auth hop after included SuperGrok period limits are full. Rank
helpers are not hop.
sampling_configmust fill console failover after those included limits are full, and must omit it while they still have room. Includes personal SuperGrok JWT as the paying identity while that login exists and personal included SuperGrok period limits have room (Team JWT omitted; that JWT settles Billing Credits). Operators switch SuperGrok identity with use-personal / use-business on$GROK_HOME/limits_pins.json. No new[auth]keys. Stockpreferred_method = "api_key"still pins console. Also: any stored SuperGrok login with included remaining before SuperGrok dollar credits, sibling included before SuperGrok dollar credits, and the one-process hourly limits flock (HonorTtl at most once an hour per machine; ForceRefresh still fetches). Do not flatten remaining to zero from usage percent 100 plus missing SuperGrok Heavy. Never invent used-up included SuperGrok period limits. Fail-open: a client 100% / remaining 0 / SuperGrok dollar credits $0 printout must not mark SuperGrok used up or hop to console. Real SuperGrok HTTP 402 after that request failed can still leave SuperGrok. SuperGrok Heavy ranking optional label is not this class. SuperGrok is paid. - Last-session on start. Interactive
grok-ossopens the remembered last session for this working directory. It does not land on Welcome first. - Product skills are not a Python runtime. Tool work is a named Rust
function, an ACP tool, or a shipped CLI bin of that function. Skills
must not generate arbitrary Python or Bash payloads and then exec them.
A restack that installs non-excepted Python under product skills, that
drops the Rust intercept for the allowlisted stub names
memory.py/validate-plan.py/session_reader.py, or that drops the CLI binsgrok-oss-implement-memory/grok-oss-plan-validate/grok-oss-session-reader, is a failed land. grok-oss intercepts those stub names and bins to Rust. Grok Build compatibility is those bins, not a second Python runtime. Office/docx/pptx/xlsx/pdf scripts and those three allowlisted stub names are the only exceptions. Host~/.agents/skillsis operator-owned and is not this class. User-guide08-skills.mdmust keep that sentence. grok-oss sqlite session/work ids are ULIDs; UUID stays the Grok Build / ACP wire id.
After restack the required classes are all seven above: CLI branding,
/settings plus unread config, Token Economy ledger /spend, DOGE/chrome
paint, dual-auth hop after included SuperGrok period limits are full,
last-session on start, and product skills are not a Python runtime.
Extra proven restack-droppable classes (named cargo tests exist; a land that drops them while keeping the seven is still a seam loss):
- Official serving-path fingerprints (Surmount persist +
/metadata; upstream Chat Completionssystem_fingerprintJSON). Named tests:show_session_metadata_includes_fingerprint_fields_from_stored_samples,parse_language_models_json_reads_id_fingerprint_version_created,record_completion_fingerprint_persists_a_flip,user_guide_metadata_documents_serving_fingerprints. Schema v7 tables must not drop/spendv1 (migrate_v6_file_to_v7_adds_serving_tables_without_dropping_spend). - Always-on bubble copy click + wrap (paint-only is a failed land).
- Plan present ≠ Approve + modal-free typing (four-CTA idle paint is not honesty).
- Lost-prompt integration tests are fork-owned contracts. Clickable Approve
must not drop the Operator-box prompt. Empty Enter on Revise is not proof of
mouse Approve. When those tests change, keep the stronger assert;
synthesize if upstream and Surmount both have a piece; never fit the
contract to a wipe. Named tests (
plan_approve_lost_prompt):isolated_present_preview_click_approve_does_not_drop_human_box_prompt,isolated_present_preview_typed_after_present_click_approve_sends_human_box_prompt,isolated_present_prompt_focus_click_approve_does_not_drop_human_box_prompt,isolated_present_click_approve_dispatches_interject_with_prompt_text,isolated_present_preview_enter_is_human_turn_then_click_approve,isolated_preview_idle_non_empty_operator_paste_enter_approves_with_notes_not_plan_exit,isolated_preview_idle_leftover_slash_plus_notes_click_approve_is_approve_with_comment,isolated_preview_vanished_pane_notes_enter_approves_with_comment,isolated_preview_idle_leftover_slash_plus_notes_enter_approves_with_comment,isolated_preview_approve_with_plan_composer_notes_submits_with_approve_not_as_prompt,isolated_preview_stays_after_present_so_comment_then_approve_can_run,isolated_preview_comment_cta_then_notes_then_approve_submits_with_approve_not_as_prompt,view_plan_reopens_isolated_preview_from_current_disk_plan_md_after_panel_closed,isolated_preview_human_send_closes_leftover_present_after_mill_continues,isolated_preview_implement_closes_leftover_present_after_mill_continues,isolated_preview_rereads_current_disk_plan_md_when_mill_rewrote_it,isolated_preview_after_mill_completion_must_not_paint_leftover_present_or_tech_md,isolated_preview_second_plan_prompt_must_not_paint_stale_plan_as_live_present,user_guide_isolated_preview_rewrite_wait_on_second_plan_prompt. /rebuildSHA-aware peer relaunch (fail-does-not-signal is not enough). Installed identity git SHA must match this workspace HEAD short SHA. TUI/rebuildstarts from the session workspace. Named tests:rebuild_must_exec_workspace_binary_not_stale_cargo_bin,installed_identity_must_match_workspace_git_sha,tui_rebuild_starts_from_session_workspace_not_process_cwd./rebuildSIGUSR1s every other live grok-oss TUI PID (dedupe by PID; two windows on the same session both get a signal). SHA-aware fail-does-not-signal is not this list. Named tests:rebuild_signals_each_pid_after_composite_key,peer_pids_to_signal_excludes_self_dead_and_non_grok(xai-grok-update --lib)./rebuildTUI persist like a network interrupt (fork-owned drafts, queue, plan notes, WAL, and nested ids in this TUI). Named tests:handle_rebuild_done_persists_unsent_composer_draft_and_session_load_restores_it,handle_rebuild_done_persists_pending_prompts_including_interject_and_session_load_restores_them,handle_rebuild_done_persists_plan_feedback_draft_and_plan_md,handle_rebuild_done_keeps_nested_subagents_for_resume,rebuild_and_relaunch_starts_while_nested_subagents_are_running,operator_ran_rebuild_and_the_grok_oss_process_did_not_restart,export_git_index_omits_unstaged_dirty_file,stash_keep_index_hides_unstaged_wip_from_compile_worktree. Unstaged WIP is not the compile source. Do not weaken to an upstream nested-cancel. Prompt WAL:prompt_wal_appends_on_enter_before_model_wait,prompt_wal_appends_on_mid_turn_interject,prompt_wal_appends_on_queue_enqueue,prompt_wal_appends_on_approve_notes,session_load_restores_wal_send_missing_from_prompt_history. After--resume/ last-session restore, the operator prompt appears once (not composer plus queue #1), Enter is send unless a live sampler turn is running, and Waiting is a real sampler wait. Named tests:resume_restore_must_not_put_the_same_operator_prompt_in_composer_and_queue,resume_restore_must_not_arm_enter_interject_when_no_live_sampler_turn,resume_restore_must_not_show_waiting_when_nested_and_sampler_are_gone,resume_restore_must_not_rehydrate_unsent_draft_and_queue_with_the_same_string,post_rebuild_relaunch_chrome_includes_grok_oss_version_and_git_sha,restore_pending_prompts_from_disk_drops_human_turns_when_memory_queue_is_nonempty. LeaderRelaunchForUpdatekeeps nested ids on that leader the same way a TUI disconnect does, and keeps this process up while the parent turn is still busy (no five-second cap). Named tests:relaunch_drain_keeps_nested_ids_alive_after_grace_like_disconnect,relaunch_drain_keeps_parent_turn_until_idle_like_disconnect(xai-grok-shell --lib),rebuild_subcommand_parses(xai-grok-pager --lib). The prompt write-ahead log (prompt_wal.jsonl) is the durability path for a dropped operator prompt. Enroll named WAL tests in this extra class when they exist. Do not list a WALfnthat is not in the tree. Operator-verified known good (2026-09-02) for send, rebuild-flush, interject, and plan-notes append tests named indoc/dev/upstream-regression-filters.md§ Prompt write-ahead log. Queue enqueue is not that mark. Restore tests stay contracts. Do not delete or weaken those tests in recon.- Interject Ctrl+Enter and Send now are fork-owned. Mid-turn Ctrl+Enter
interjects when appropriate and otherwise inserts a newline. Queue
[Send now]on a plain prompt row dispatchesSendInterjectviaAction::Interject. A queued/goalrow is GoalSet viaSendPromptNow. Grok OSS 1.0.3 is not last-known-good. WAL known-good for send / rebuild-flush / interject / plan-notes appends is a separate contract and does not prove live Interject UI. Named tests:ctrl_enter_mid_turn_dispatches_send_interject,queue_send_now_click_dispatches_send_interject,empty_ctrl_enter_mid_turn_does_not_send,enter_while_other_work_is_live_must_still_clear_composer,queued_goal_send_now_is_goal_action_not_stuck_composer_string,header_timeout_is_named_cold_start_class_with_retry_path,limits_help_lists_named_words_and_hyphenated_aliases,interject_does_not_wait_minutes_or_block_paint,enter_soft_interject_must_not_leave_duplicate_prompt_in_composer,enter_on_pasted_15_lines_chip_sends_or_interjects_does_not_only_expand,l2_overlay_enter_interject_must_not_leave_duplicate_prompt_in_composer,enter_send_must_not_leave_duplicate_prompt_in_composer,enter_at_end_of_last_composer_line_must_submit_immediately_not_silent_newline,enter_at_end_of_last_composer_line_mid_turn_must_interject_immediately_not_silent_newline,arrow_keys_then_enter_must_submit_the_same_body_not_a_different_path. Keepprompt_wal_appends_on_mid_turn_interject. - TUI performance contracts are fork-owned. They are not last-known-good
until typing, cancel, and interject stay responsive. Hang/cancel,
Interject Send now, WAL, and queue-snapshot tests are in the tree.
The live TUI may still be 1.0.3. Do not treat 1.0.3 as last-known-good.
Persist:
pending_prompts_queue_snapshot_skips_sync_all,pending_prompts::tests::write_without_fsync_still_roundtrips,plan_human_box_keystroke_burst_does_not_append_prompt_wal,main_composer_keystroke_burst_does_not_append_prompt_wal. Existing debounce stays a contract:plan_human_box_keystroke_burst_does_not_flush_unsent_draft_every_char,main_composer_keystroke_burst_does_not_flush_unsent_draft_every_char,keystroke_burst_does_not_flush_unsent_draft_every_char. Paint:idle_plan_overlay_does_not_demand_fast_ticks,plan_overlay_repeat_prepare_at_same_width_does_not_rebuild_markdown. Cancel and interject:cancel_does_not_wait_minutes,interject_does_not_wait_minutes_or_block_paint. Catalog:doc/dev/upstream-regression-filters.md§ TUI performance. Do not delete or weaken image-token, lost-prompt, Operator/Agent, WAL append, or/rebuildpersist tests. - Compact / summarize / turn HTTP
image_urlmust be a base64 data URL or anhttp(s)URL. A local session asset path, afile://handle, an[Image #N]token, or an empty value must not reach the API (invalid_image). Compact still strips user images to[image]so it does not re-inline the data URL crate. Leftover Image parts are converted or omitted on the request clone; storedchat_history.jsonlis not rewritten. Aninvalid_imageretry must drop only non-API urls and keep valid data URL siblings. Named tests:compact_request_must_not_send_session_asset_path_or_image_token_as_image_url,repair_encodes_raw_session_asset_path_as_data_url,path_image_token_and_empty_are_not_api_image_urls,invalid_image_retry_drops_path_shaped_urls_and_keeps_valid_data_url. Catalog:doc/dev/upstream-regression-filters.md§ Compact image_url. Do not delete or weaken image-token, lost-prompt, Operator/Agent, WAL append, or/rebuildpersist tests. - Nucleo reuse-per-root.
- Baked default is Grok 4.6 at medium reasoning effort
(
baked_default_is_grok_46_medium_fork_contract). Fork contract change. - Soft plan present is a real right-side pane (not a 75% centered overlay).
Named tests:
plan_soft_park_docks_right_not_centered_overlay,plan_soft_park_draw_right_pane_matches_side_panel_status,plan_row_click_does_not_enter_commenting,plan_loop_status_does_not_claim_side_panel_when_viewer_closed. Idle parked plan must not spin the 30fps loop, and same-width paints must not re-scan a 240k plan body (idle_plan_overlay_does_not_demand_fast_ticks,plan_overlay_repeat_prepare_at_same_width_does_not_rebuild_markdown). - Plan-review and Linux prompt screenshot paste
(
event_paste_plan_commenting_empty_defers_clipboard_image_probe,plan_feedback_ctrl_v_defers_clipboard_image_probe,agent_empty_bracketed_paste_defers_probe_for_clipboard_image,approve_or_revise_drains_plan_composer_images). - Live chrome names SuperGrok dollar credits, not a nickname
(
compact_status_supergrok_on_dollar_credits_shows_dollars_not_free_period_pct,format_supergrok_session_with_weekly_and_dollar_credits). - No two live same-description Subagent rows; unlimited retry is not a
u32::MAXfraction (live_subagent_list_does_not_show_two_rows_with_the_same_description,task_spawn_rejects_or_replaces_second_live_same_description,format_activity_label_unlimited_retry_has_no_u32_max_fraction,implement_effort_two_does_not_spawn_two_review_rows_unless_operator_asked). from_configno-prefetch usable catalog (from_config_without_prefetch_produces_usable_catalog). Emptymodels_cache.jsonmiss is not cargo-proven.- Seeded custom model on
session/loadstays Chat Completions (keep_unverified_persisted_model_keeps_seeded_custom_slug,seeded_test_model_keeps_chat_completions_backend,poisoned_image_session_recovers_within_the_failing_turn). - Always-three-layer product prompt (
child_task_description_is_concise,default_max_allows_l2_to_spawn_l3). /goalparent coordinates (Surmount fork of the injected prompt vs upstream "Deliver everything yourself"):goal_instruction_parent_coordinates_and_l2_must_spawn_l3_for_tools(xai-grok-tools-apislash_commands.rs). Keepgoal_instruction_carries_objective_and_contract_tokens. Live templates:goal_rules_templates_parent_coordinates_and_l2_must_spawn_l3_for_tools,goal_task_discipline_parent_spawns_l2_not_product_tools(xai-grok-shell).- Parent fire-and-return spawn (Surmount / grok-oss fork): a nested L2
that is a long builder must not occupy the parent as a blocking
10-minute
get_command_or_subagent_outputwait. Parent starts it, keeps working, completion is a notification. Parent can spawn a second L2 while the first is still running. Namedfn:parent_spawn_subagent_second_l2_while_first_still_running_without_wait(xai-tool-typestask.rsandxai-grok-toolstask/backend_tests.rs). - Parent follow-up onto a running nested L2 (Surmount / grok-oss fork,
GitHub #143): enqueue additive work on a live L2 without kill or
respawn; do not inject a live L3 unless targeted;
resume_fromstays completed-only; off ([subagents] parent_follow_up = false) is SpaceXAI spawn / wait /resume_fromafter exit. No new[auth]key. Namedfns:parent_cannot_talk_to_own_l2s_follow_up_enqueues_interject_on_running_l2_without_kill_or_respawn,parent_follow_up_does_not_inject_into_live_l3_unless_operator_targeted_that_specialist,resume_from_of_running_l2_still_fails_active,parent_follow_up_off_is_upstream_spawn_wait_resume_from_completed_only,parent_follow_up_onto_running_l2_with_live_l3_hits_l2_not_l3(xai-grok-toolstask/parent_follow_up_tests.rs).TaskTool::run:parent_cannot_talk_to_own_l2s_task_tool_run_follow_up_returns_queued_and_does_not_spawn,parent_cannot_talk_to_own_l2s_task_tool_run_follow_up_and_resume_from_are_mutually_exclusive(xai-grok-toolstask/mod.rs). Schema:parent_cannot_talk_to_own_l2s_follow_up_field_is_optional_and_distinct_from_resume_from(xai-tool-typestask.rs). Shell Interject:shell_child_follow_up_sends_session_command_interject_on_child_session(xai-grok-shellagent/subagent/tests/mod.rs). Config:subagents_config_parent_follow_up_false_parses_and_omitted_defaults_true,resolve_subagents_copies_parent_follow_up(xai-grok-shellconfig/tests.rs). KEEPparent_spawn_subagent_second_l2_while_first_still_running_without_wait,l2_overlay_send_prompt_interjects_l2_not_l1,l3_overlay_send_prompt_does_not_reach_l3_or_l1,nested_spawner_can_resume_from_completed_reparented_child. - Compact standing-law reminder (Surmount / grok-oss fork): after compact,
standing Surmount law is the first
<system-reminder>section, not a buried AGENTS.md paragraph and not/recap. Upstream parent turns often sit on a 10-minute poll wait. Surmount is fire-and-return for long builder L2s, and that wait/discovery law is injected here. Namedfn:post_compact_reminder_includes_surmount_standing_law(xai-grok-shellcompaction_context.rssection_surmount_standing_law_after_compact). - CoT death spiral (GitHub #133):
DEST_ENCODER_SKIP_LOOP,StreamRepetitionGuard, FatalRepetitiveGeneration, compact 8192 / 32768. Oh My Pi comparison: this tree has no Oh My Pi checkout; do not invent internals. Surmount vs SpaceXAI as in Product inventory. Namedfns:dest_encoder_skip_loop_is_repetitive,dest_encoder_skip_single_sentence_four_times_is_repetitive,thought_line_loop_is_repetitive,chat_completions_stops_dest_encoder_skip_loop,chat_completions_stops_dest_encoder_skip_loop_in_thought,messages_stops_dest_encoder_skip_loop,responses_stops_dest_encoder_skip_loop,classify_repetitive_generation_is_fatal,compaction_reseed_of_unique_75k_history_must_not_leave_wasteful_75k_context,compact_summary_budget_is_8192_tokens_and_reseed_reserve_is_32768. - File-level infer-from-path verify after ACP structured edits
(
rustfmt_argv_edition_2024_config_and_absolute_files,clippy_argv_lints_the_edited_file_not_crate_lib,clippy_argv_is_file_level_not_package_lib,dangerous_cargo_clippy_package_all_targets_is_refused_and_does_not_spawn_shell). A restack that dropsutil/rust_edit_verify.rsor these tests is a failed land. Not one of the seven numbered classes. - ACP per-path write lock after structured edits
(
try_acquire_write,try_acquire_read,write_paths, CoWpublished/published_cow_snapshot,held(), share is allowed;two_agents_cannot_write_the_same_path_at_once,search_replace_apply_patch_and_write_all_take_the_lock,held_path_error_names_holder_and_file_without_a_steal_skip_wait_menu,opencode_edit_cannot_write_a_path_another_agent_already_holds,hashline_edit_refuses_when_another_agent_holds_the_path,hashline_edit_happy_path_does_not_mention_the_lock,sequential_writes_succeed_after_the_first_tool_call_returns_even_when_both_agents_were_assigned_the_same_write_paths,spawn_write_paths_overlap_is_a_soft_assignment_not_a_spawn_error,soft_lock_reminder_is_observable_on_a_sibling_tool_call,after_write_returns_held_is_empty_lock_must_be_released,reader_during_held_write_gets_published_pre_write_bytes_current_atomic_snapshot,two_live_agents_with_the_same_write_paths_spawn_without_error_l2_and_l3_may_be_assigned_the_same_file). A restack that dropsper_path_write_lock.rs, the OpenCodeeditlock acquire, thehashline_editlock acquire, or these tests is a failed land. Not one of the seven numbered classes. - Pause / resume chips and Clear finished quiet paint.
- User-guide cargo pins beyond skills + resume (
user_guide_operator_cli_examples_use_grok_oss,user_guide_does_not_claim_automatic_host_hop_is_unshipped,user_guide_names_token_economy_spend_order,user_guide_limits_names_fail_open_and_named_commands). - L1 Subagents list is L2-only plus a live L3 count
(
live_subagent_list_shows_only_l2_and_reports_live_l3_count,l2_row_shows_live_l3_count_not_specialist_names). - Live Subagents list is still-running only; already_exited dismisses the paused Implementer overlay
(
finished,already_exited, Occupied skip,retain_still_running_nested_occupancy,listed_live_subagents;kill_already_exited_dismisses_paused_implementer_overlay,restore_nested_occupancy_does_not_unfinish_or_revive_dead_host,running_count_matches_listed_live_l2_not_l3,kill_already_completed_drops_live_list_responding_and_still_running_cue). - Compacting Subagents row
[↗]still opens (open_subagent_fullscreenversus AutoCompactStarted auto-steal;click_tasks_open_on_compacting_row_opens_subagent,click_tasks_open_on_last_painted_row_opens_subagent,click_tasks_kill_on_compacting_row_emits_kill,open_subagent_fullscreen_sets_active_while_child_is_auto_compacting). - Ctrl+C two-stage,
/modellast Tab, Isolated Preview glass search, Isolated Preview screenshot paste (GNOME All Markup Copy is an image). Tests:isolated_preview_handle_input_ctrl_c_with_text_clears_and_stays,unique_model_slash_tab_switches_now_empty_composer_no_send,isolated_preview_search_query_plan_matches_plan_and_plan,isolated_preview_gnome_all_markup_copy_title_with_raster_is_image_chip. - Turbo planning (
effective_reasoning_effort,live_plan_turn,model_effort_chrome_line,stamp_request_effort). Live exclusive/planand Isolated Preview/plan --softrequest xhigh while turbo planning is on (default on). Only the lower-right yellow model/effort line shows xhigh. Magenta model id stays the model id. No TURBO badge, banner, or toast. Stored session/effortis not mutated. Off is the upstream/SpaceXAI option: plan turns stay at session effort. Tests:session_medium_enter_plan_request_uses_xhigh_and_lower_right_shows_xhigh,exit_or_approve_plan_returns_session_medium_effort,turbo_planning_settings_off_plan_turn_stays_session_medium. - Soft process-rule reminders (settings list; spawn still succeeds):
process_rule_reminder_configured_third_l2_still_spawns,process_rule_reminder_text_in_nested_spawn_prompt,process_rule_reminders_off_nested_spawn_has_no_extra_reminder_text. Off or empty list is the upstream option. /startplus leftover cancel-resume marker drop (start_while_globally_paused_continues_interrupted_turn_once,start_on_idle_clean_session_does_not_invent_a_turn,start_with_cancel_resume_marker_continues_interrupted_turn,handle_rebuild_done_must_not_cancel_parent_so_session_load_adopts_like_disconnect,handle_rebuild_done_mid_turn_writes_cancel_resume_and_session_load_continues_the_turn,handle_rebuild_done_idle_completed_turn_does_not_write_cancel_resume_or_refire_last_prompt,session_load_drops_stale_cancel_resume_marker_when_primary_turn_finished_successfully)./unstickshell skip-append, hung-task orphan, leader hung-RPC drop, and WAL image resource links (unstick_retry_does_not_append_second_user_query_when_last_turn_matches,unstick_retry_orphans_stuck_running_task_then_samples_again,unstick_leader_drops_hung_session_prompt_like_disconnected_client,unstick_resends_wal_images_as_resource_blocks_not_data_urls).- ForceRefresh on explicit
/limits(management_meter_cache_policy_collect_force_background_honor_ttl,should_clear_management_meter_caches_force_with_key_only,limits_snapshot_mode_for_get_billing_explicit_is_force_refresh,limits_snapshot_force_refresh_leader_http_fetches_when_snapshot_is_younger_than_one_hour). - Footer / session sampling window vs catalog
(
context_chip_names_sampling_window_when_catalog_differs,context_chip_hover_percent_uses_sampling_window_when_catalog_differs,footer_chip_uses_session_sampling_window_when_economic_cache_is_off,refresh_context_used_does_not_copy_catalog_into_session_sampling,main_session_sampling_window_is_catalog_500k_even_when_economic_is_on,nested_session_sampling_window_stays_200k_when_catalog_is_500k). - Spawn-prompt fold plus last-answer caps
(
huge_spawn_prompt_becomes_pointer_with_description_and_report,parent_estimated_tokens_omit_huge_spawn_prompt,to_model_text_caps_huge_last_answer_for_parent_ingest,completed_subagent_task_output_is_capped_or_points_at_report,blocking_spawn_subagent_completed_to_prompt_format_is_capped). - Any stored SuperGrok included remaining before SuperGrok dollar credits;
do not flatten remaining from usage percent 100 plus missing SuperGrok
Heavy (
sampling_config_hop_team_remaining_personal_exhausted_not_dollars_or_console,sampling_config_hop_personal_remaining_team_exhausted,sampling_config_hop_both_remaining_team_first_then_personal,sampling_config_hop_both_included_exhausted_dollar_credits_before_console,sampling_config_hop_missing_heavy_false_100_keeps_sibling_included,sampling_config_hop_dollar_credits_on_both_missing_heavy_keeps_team,prepare_sampler_for_turn_does_not_flatten_missing_heavy_100_off_sibling,prepare_sampler_for_turn_does_not_flatten_dollar_credits_on_both). Rankhop_*helpers are still not hop.
Not a cargo land class: rustc 1.98.0 (file pin only;
rust-toolchain.toml not in FORK_PATHS). Stuck-retry pager chrome is
not fully proven. Token Economy /settings table rows were not re-proven on
2026-08-15. CLI grok-oss rebuild is clap-wired (rebuild_subcommand_parses).
/economic-mode is
not a live BuiltinAction. SuperGrok Heavy ranking optional label is not
implemented. Empty models_cache.json miss has no named test. Live hop /
live Business remaining / live TUI dogfood are unknown.
Process pins survive import via FORK_PATHS restore +
assert-process-pins (path presence and light content sniffs). That gate does
not prove product behavior inside shared xai-grok-* crates.
Product seams live inside those crates. They survive onto only through
cherry-picks / conflict resolve and stay honest through named cargo
tests. After recon, run the assert, then the seven-class land checklist,
then the name-existence check. just check cannot fail a deleted catalog
test. Deleting a red catalog test is not a restore. Paint is one of seven
land classes, not the whole land.
Full filter catalog (why each exists + every residual Validate honesty block):
doc/dev/upstream-regression-filters.md.
Open residual still points at the same commands under RESIDUAL § Validate
honesty (D0 can demote; the catalog is durable).
Operator cheat sheet (post-import / post-onto tip). rg each identifier for
a matching fn first:
just upstream-assert-process-pins
grok-nix-helper assert-process-pins HEAD # or onto tip
# 1. CLI identity (first token grok-oss; substring "grok" is not enough)
cargo test -p xai-grok-pager --lib -- product_version_line_uses_grok_oss_not_bare_grok \
resume_session_command_uses_grok_oss user_guide_resume_and_version_examples_use_grok_oss \
product_cli_name_is_grok_oss print_exit_resume_hint_writes_expected_lines \
user_guide_operator_cli_examples_use_grok_oss welcome_badge_brands_grok_oss \
hero_subtitle_brands_grok_oss tutorial_list_title_brands_grok_oss
cargo test -p xai-grok-pager-bin --test version_without_tty
# 2. Config is a surface (/settings rows + readers + DOGE picker; serde-only is not enough)
cargo test -p xai-grok-pager --test settings_e2e -- hide_header always_expand_thinking \
scrub_ascii_punct allow_worktree bubble_copy_buttons plan_approval_park
cargo test -p xai-grok-pager --lib -- theme_choices_include_doge_and_default_is_doge \
hide_header_zeroes always_expand_thinking ctrl_t_expand_is_default \
ctrl_t_collapse_is_default \
bubble_copy_buttons_on append_bubble_copy_button_paints \
clicking_human_bubble_copy clicking_assistant_bubble_copy \
clicking_wide_human_bubble_copy
cargo test -p xai-grok-pager-render --lib -- prime_applies_scrub_ascii_punct_from_ui \
prime_applies_always_expand_thinking_from_ui
cargo test -p xai-grok-shell --lib -- resolve_subagents_copies_allow_worktree
# 3. Token Economy ledger /spend (schema-only is not enough; extra SQL, not SuperGrok dollar credits)
cargo test -p xai-grok-shell --lib -- spend_path_ingests_usage_jsonl_and_records_reconciliation
cargo test -p xai-grok-pager --lib -- show_spend_ingests_usage_jsonl_and_is_not_empty_default
# 4. DOGE / Surmount chrome (theme file existing is not paint)
cargo test -p xai-grok-pager-render --lib -- default_theme_is_doge resolve_from_config_no_config \
doge_accent_user_is_pure_green doge_accent_system_is_pure_cyan \
as_doge_human_green_named_ansi_is_rgb_0_255_0 osc12_named_ansi_green_is_doge_rgb_0_255_0
cargo test -p xai-grok-pager --lib -- user_prompt_block_accent user_prompt_entry_renderer_paints_green_rail \
paint_composer_box_cursor_uses_human focused_composer_paints_human_green_box_caret \
doge_human_box_caret_plate_is_rgb_0_255_0 paint_composer_box_cursor_named_ansi_green_becomes_doge_rgb \
agent_message_block_accent info_line_model_name_uses_accent_model \
status_bar_pushes_credits_compact_included_supergrok_period_limits \
hit_credits_click_dispatches_show_limits \
titled_doge_composer_frame_is_prompt_border_not_context_yellow \
plan_approval_footer_paints_five_cta_vocabulary \
auto_compact_completed_preserves_todo_board \
todo_badge_names_tasks_not_only_fraction \
status_header_todo_badge_names_tasks \
nested_l2_overlay_todo_toggle_stays_findable \
forked_session_status_header_paints_switcher_and_dashboard \
forked_session_status_header_clicks_open_dashboard_and_cycle \
load_session_restores_fork_family_from_disk
# 5. Dual-auth hop after included SuperGrok period limits are full (rank helpers are not hop)
cargo test -p xai-grok-shell --lib -- sampling_config_auto_use \
sampling_config_hops_to_sibling_included_before_dollar_credits \
afterburner_does_not_skip_mark_when_sibling_has_included_remaining \
resolve_model_to_sampling_config_auto_use \
align_after_billing_switches_sticky_personal_full_to_business_included \
prepare_sampler_for_turn_aligns_to_ranked_included_primary \
combined_included_remaining_sums_distinct_personal_and_business_pools \
combined_included_remaining_does_not_double_count_unified_pool \
combined_included_remaining_does_not_collapse_matching_percent_and_reset_into_one_pool \
pick_prefers_business_included_before_personal_when_both_have_remaining \
order_credentials_business_included_before_personal_when_both_have_room \
limits_snapshot_second_process_within_the_hour_does_not_http \
limits_snapshot_honor_ttl_fresh_within_hour_does_not_http \
limits_snapshot_force_refresh_leader_http_fetches_when_snapshot_is_younger_than_one_hour \
limits_snapshot_stale_file_lets_waiter_become_leader_and_fetch_once \
limits_snapshot_never_writes_access_tokens \
billing_handler_uses_snapshot_hub_instead_of_unconditional_sibling_http \
personal_included_period_limits_reset_uses_personal_supergrok_not_leftover_business_credits \
business_with_no_period_limits_payload_still_switchable_via_use_business \
use_personal_switches_back_from_business_pin
cargo test -p xai-grok-pager --lib -- compact_meter_stays_included_while_sibling_pool_has_remaining \
active_spend_driver_stays_included_while_any_distinct_pool_has_remaining \
matching_percent_and_reset_does_not_collapse_combined_remaining_into_one_pool
# 6. Last-session on start
cargo test -p xai-grok-pager --lib -- materialize_new_auto_opens_last_session_when_one_exists \
materialize_new_auto_stays_welcome_when_no_last_session \
materialize_new_auto_does_not_open_last_when_headless \
from_pager_args_opens_last_session_on_start
# 7. Product skills are not a Python runtime (non-excepted .py or dropped intercept is a failed land)
cargo test -p xai-grok-bundle --lib -- sanitize_rejects_non_excepted_skill_python \
extract_archive_skips_non_excepted_skill_python \
product_repo_skill_roots_have_no_non_excepted_python \
default_product_skills_include_polish_and_subagent \
default_product_skill_markdown_does_not_tell_agents_to_generate_python_or_bash
cargo test -p xai-grok-pager --lib -- user_guide_skills_are_not_a_python_runtime
cargo test -p xai-grok-tools --lib -- implement_memory_snapshot_intercept_does_not_spawn_shell \
plan_validate_intercept_does_not_spawn_shell session_reader_list_intercept_does_not_spawn_shell \
grok_oss_implement_memory_cli_bin_intercept_does_not_spawn_shell \
generated_python_payload_is_not_skill_stub_intercept
# Extra: plan present != approve + modal-free typing
cargo test -p xai-grok-pager --lib -- exit_plan_mode_present_is_not_operator_approve \
empty_enter_on_revise_prompt_does_not_approve \
soft_park_empty_ctrl_c_abandons_plan_approval \
exit_plan_mode_keeps_mid_compose_draft_and_a_types \
exit_plan_mode_modal_park_does_not_steal_mid_compose_keys \
exit_plan_mode_empty_present_printable_goes_to_composer \
exit_plan_mode_shows_overlay_even_in_yolo
cargo test -p xai-grok-tools --lib -- exit_plan_mode_tool_result_does_not_claim_operator_approval
# Extra: pause / Clear finished (not paint-only)
cargo test -p xai-grok-pager --lib -- work_control_chrome_matrix_pause_not_cancel_stop_not_pause \
pause_button_click_dispatches_global_pause_not_cancel \
idle_with_subagents_paints_pause_and_stop_hits \
global_paused_idle_paints_resume_not_stop \
clear_finished_action_idle_is_quiet_not_neon_green_or_magenta \
clear_finished_click_does_not_open_subagent
# Extra: /rebuild SHA-aware (fail-does-not-signal is not enough)
cargo test -p xai-grok-update --lib -- failed_install_must_not_replace_or_signal_peers \
build_fail_does_not_signal_leaders parse_version_output_extracts_identity \
peer_relaunch_accepts_same_semver_different_sha \
peer_relaunch_declines_equal_identity_on_same_path \
peer_relaunch_accepts_deleted_inode_even_when_identity_equal \
operator_ran_rebuild_and_the_grok_oss_process_did_not_restart \
rebuild_must_exec_workspace_binary_not_stale_cargo_bin \
installed_identity_must_match_workspace_git_sha
cargo test -p xai-grok-pager --lib -- tui_rebuild_starts_from_session_workspace_not_process_cwd
cargo test -p xai-grok-shell --lib -- leader_is_older_than_same_semver_git_sha_identity
# Extra: /rebuild signals every live grok-oss PID (SHA-aware is not this list)
cargo test -p xai-grok-update --lib -- rebuild_signals_each_pid_after_composite_key \
peer_pids_to_signal_excludes_self_dead_and_non_grok
# Extra: /rebuild TUI persist like a network interrupt (fork-owned)
# Leader drain is not that persist path.
cargo test -p xai-grok-pager --lib -- \
handle_rebuild_done_must_not_cancel_parent_so_session_load_adopts_like_disconnect \
handle_rebuild_done_persists_unsent_composer_draft_and_session_load_restores_it \
handle_rebuild_done_persists_pending_prompts_including_interject_and_session_load_restores_them \
handle_rebuild_done_persists_plan_feedback_draft_and_plan_md \
handle_rebuild_done_keeps_nested_subagents_for_resume \
rebuild_and_relaunch_starts_while_nested_subagents_are_running \
operator_ran_rebuild_and_the_grok_oss_process_did_not_restart \
post_rebuild_relaunch_chrome_includes_grok_oss_version_and_git_sha \
tui_rebuild_starts_from_session_workspace_not_process_cwd \
restore_pending_prompts_from_disk_drops_human_turns_when_memory_queue_is_nonempty \
rebuild_subcommand_parses \
prompt_wal_appends_on_enter_before_model_wait \
prompt_wal_appends_on_mid_turn_interject \
prompt_wal_appends_on_queue_enqueue \
prompt_wal_appends_on_approve_notes \
ctrl_enter_mid_turn_dispatches_send_interject \
queue_send_now_click_dispatches_send_interject \
empty_ctrl_enter_mid_turn_does_not_send \
enter_soft_interject_must_not_leave_duplicate_prompt_in_composer \
enter_on_pasted_15_lines_chip_sends_or_interjects_does_not_only_expand \
l2_overlay_enter_interject_must_not_leave_duplicate_prompt_in_composer \
enter_send_must_not_leave_duplicate_prompt_in_composer \
enter_at_end_of_last_composer_line_must_submit_immediately_not_silent_newline \
enter_at_end_of_last_composer_line_mid_turn_must_interject_immediately_not_silent_newline \
arrow_keys_then_enter_must_submit_the_same_body_not_a_different_path \
session_load_restores_wal_send_missing_from_prompt_history
cargo test -p xai-grok-update --lib -- export_git_index_omits_unstaged_dirty_file \
stash_keep_index_hides_unstaged_wip_from_compile_worktree
cargo test -p xai-grok-shell --lib -- \
relaunch_drain_keeps_nested_ids_alive_after_grace_like_disconnect \
relaunch_drain_keeps_parent_turn_until_idle_like_disconnect
# Extra: from_config cold catalog (empty models_cache.json miss is NOT this filter)
cargo test -p xai-grok-shell --lib -- from_config_without_prefetch_produces_usable_catalog
# Extra: session/load keeps seeded custom model on Chat Completions (not last-session on start)
cargo test -p xai-grok-shell --lib -- keep_unverified_persisted_model_keeps_seeded_custom_slug \
seeded_test_model_keeps_chat_completions_backend
cargo test -p xai-grok-shell --test test_image_strip_recovery -- \
poisoned_image_session_recovers_within_the_failing_turn
# Extra: nucleo reuse-per-root
cargo test -p xai-grok-workspace --lib -- repeated_open_without_close_keeps_one_search_per_root \
distinct_roots_each_keep_one_search get_results_does_not_keep_a_stale_search_alive
# Extra: always-three-layer product prompt
cargo test -p xai-grok-agent --lib -- child_task_description_is_concise
cargo test -p xai-grok-tools --lib -- default_max_allows_l2_to_spawn_l3
# Extra: file-level infer-from-path verify (not crate-wide cargo)
cargo test -p xai-grok-tools --lib rust_edit_verify
cargo test -p xai-grok-tools --lib compiler_probe_junk
cargo test -p xai-grok-tools --lib -- rustfmt_argv_edition_2024_config_and_absolute_files \
clippy_argv_lints_the_edited_file_not_crate_lib \
clippy_argv_is_file_level_not_package_lib \
clippy_driver_uses_temp_out_dir_not_the_workspace_root \
write_refuses_rmeta_at_workspace_root_and_does_not_create_the_file \
rustc_oneshot_without_out_dir_is_refused_and_does_not_spawn_shell \
dangerous_cargo_fmt_all_is_refused_and_does_not_spawn_shell \
dangerous_cargo_clippy_package_all_targets_is_refused_and_does_not_spawn_shell \
dangerous_cargo_test_package_lib_filter_is_not_refused
# Extra: ACP per-path write lock CoW + share is allowed (GitHub #129)
cargo test -p xai-grok-tools --lib -- \
after_write_returns_held_is_empty_lock_must_be_released \
reader_during_held_write_gets_published_pre_write_bytes_current_atomic_snapshot \
two_live_agents_with_the_same_write_paths_spawn_without_error_l2_and_l3_may_be_assigned_the_same_file \
spawn_write_paths_overlap_is_a_soft_assignment_not_a_spawn_error
# Extra: L1 Subagents list is L2-only plus a live L3 count
cargo test -p xai-grok-pager --lib -- \
live_subagent_list_shows_only_l2_and_reports_live_l3_count \
l2_row_shows_live_l3_count_not_specialist_names
# Extra: live Subagents list is still-running only; already_exited overlay closeout
cargo test -p xai-grok-pager --lib -- \
kill_already_exited_dismisses_paused_implementer_overlay \
restore_nested_occupancy_does_not_unfinish_or_revive_dead_host \
running_count_matches_listed_live_l2_not_l3 \
kill_already_completed_drops_live_list_responding_and_still_running_cue
# Extra: Compacting Subagents row open button still opens
cargo test -p xai-grok-pager --lib -- \
click_tasks_open_on_compacting_row_opens_subagent \
click_tasks_open_on_last_painted_row_opens_subagent \
click_tasks_kill_on_compacting_row_emits_kill \
open_subagent_fullscreen_sets_active_while_child_is_auto_compacting \
nested_compact_chrome_does_not_steal_parent_fullscreen_overlay \
nested_compact_chrome_must_not_steal_parent_tui_scroll
# Extra: Ctrl+C two-stage, /model last Tab, Isolated Preview glass search,
# Isolated Preview screenshot paste (GNOME All Markup Copy is an image)
cargo test -p xai-grok-pager --lib -- \
isolated_preview_handle_input_ctrl_c_with_text_clears_and_stays \
isolated_preview_handle_input_second_empty_ctrl_c_exits \
isolated_preview_handle_input_running_turn_draft_ctrl_c_does_not_cancel_turn \
leftover_isolated_preview_handle_input_ctrl_c_clears_then_exits \
unique_model_slash_tab_switches_now_empty_composer_no_send \
unique_model_slash_enter_switches_now_no_operator_model_chat \
unique_m_slash_tab_switches_now \
complete_typed_model_xhigh_tab_switches_now_with_effort \
command_phase_unique_model_tab_still_completes \
multi_row_model_tab_stays_complete_not_switch \
isolated_preview_unique_model_tab_switches_now_does_not_rowwalk \
plan_preview_title_bar_search_glass_immediately_left_of_copy \
isolated_preview_search_query_plan_matches_plan_and_plan \
isolated_preview_search_glass_clickable_next_to_copy_and_expand \
isolated_preview_composer_slash_stays_slash_not_line_search \
isolated_preview_handle_input_n_jumps_hits_after_search \
isolated_preview_gnome_all_markup_copy_title_with_raster_is_image_chip \
isolated_preview_search_open_paste_does_not_fill_search \
mill_event_paste_gnome_all_markup_copy_title_with_raster_does_not_insert_title \
gnome_all_markup_copy_title_still_probes \
user_guide_ctrl_c_two_stage_clears_then_exits \
user_guide_unique_model_tab_switches_now \
user_guide_isolated_preview_search_glass \
user_guide_gnome_all_markup_copy_is_image
# Extra: turbo planning (Surmount vs SpaceXAI; off = session effort)
cargo test -p xai-grok-pager --lib -- \
session_medium_enter_plan_request_uses_xhigh_and_lower_right_shows_xhigh \
exit_or_approve_plan_returns_session_medium_effort \
turbo_planning_settings_off_plan_turn_stays_session_medium
# Extra: soft process-rule reminders (spawn still succeeds)
cargo test -p xai-grok-tools --lib -- \
process_rule_reminder_configured_third_l2_still_spawns \
process_rule_reminder_text_in_nested_spawn_prompt \
process_rule_reminders_off_nested_spawn_has_no_extra_reminder_text
# Extra: Subagents list compact window counts and TECH.md (not billing meters)
cargo test -p xai-grok-pager --lib -- \
subagents_list_omits_the_word_tokens \
subagents_list_l2_row_is_present_plus_past_atomic_total_including_specialists \
nested_compact_keeps_present_plus_past_without_double_counting_the_surviving_window \
nested_specialist_windows_are_not_double_counted_in_the_total \
l2_row_paints_present_plus_past_atomic_total_including_specialists \
parent_context_chip_is_l1_window_and_does_not_add_nested_windows \
format_live_subagents_list_row_uses_live_sample_not_tracker_high_water \
subagents_list_truncation_does_not_split_compact_count \
subagents_list_shows_measured_tokens_per_nested_l2 \
format_subagents_list_description_shows_measured_tokens_suffix \
format_subagent_label_shows_measured_tokens_suffix \
tech_md_write_records_measured_tokens_on_spawn_usage_tick_and_l2_exit \
subagents_list_layout_does_not_read_chat_history_jsonl \
concurrent_nested_l2_usage_ticks_keep_atomic_u64_high_water
# Extra: /start + leftover cancel-resume marker drop
cargo test -p xai-grok-pager --lib -- \
start_while_globally_paused_continues_interrupted_turn_once \
start_on_idle_clean_session_does_not_invent_a_turn \
start_with_cancel_resume_marker_continues_interrupted_turn \
handle_rebuild_done_must_not_cancel_parent_so_session_load_adopts_like_disconnect \
handle_rebuild_done_mid_turn_writes_cancel_resume_and_session_load_continues_the_turn \
handle_rebuild_done_idle_completed_turn_does_not_write_cancel_resume_or_refire_last_prompt \
session_load_drops_stale_cancel_resume_marker_when_primary_turn_finished_successfully
# Extra: /unstick shell skip-append, hung-task orphan, leader hung-RPC drop, WAL image resource links
cargo test -p xai-grok-shell --lib -- \
unstick_retry_does_not_append_second_user_query_when_last_turn_matches \
unstick_retry_orphans_stuck_running_task_then_samples_again \
session_prompt_is_unstick_retry_reads_params_meta \
take_in_flight_session_prompts_for_unstick_leaves_other_sessions \
response_is_orphaned_for_unstick_while_client_connected
cargo test -p xai-grok-shell --test test_leader_stdio_integration -- \
unstick_leader_drops_hung_session_prompt_like_disconnected_client
cargo test -p xai-grok-pager --lib -- \
unstick_resends_wal_images_as_resource_blocks_not_data_urls \
wal_image_resource_blocks_use_file_uri_not_data_url \
wal_image_resource_blocks_drop_data_url_file_ids
# Extra: ForceRefresh on explicit /limits
cargo test -p xai-grok-pager --lib -- \
management_meter_cache_policy_collect_force_background_honor_ttl \
should_clear_management_meter_caches_force_with_key_only
cargo test -p xai-grok-shell --lib -- \
limits_snapshot_mode_for_get_billing_explicit_is_force_refresh \
limits_snapshot_force_refresh_leader_http_fetches_when_snapshot_is_younger_than_one_hour
# Extra: sampling window vs catalog (chip + session field + spawn seed)
cargo test -p xai-grok-pager --lib -- \
context_chip_names_sampling_window_when_catalog_differs \
context_chip_hover_percent_uses_sampling_window_when_catalog_differs \
footer_chip_uses_session_sampling_window_when_economic_cache_is_off \
refresh_context_used_does_not_copy_catalog_into_session_sampling
cargo test -p xai-grok-shell --lib -- \
main_session_sampling_window_is_catalog_500k_even_when_economic_is_on \
nested_session_sampling_window_stays_200k_when_catalog_is_500k
# Extra: parent fire-and-return spawn (no blocking 10-minute wait)
cargo test -p xai-tool-types --lib -- parent_spawn_subagent_second_l2_while_first_still_running_without_wait
cargo test -p xai-grok-tools --lib -- parent_spawn_subagent_second_l2_while_first_still_running_without_wait
# Extra: parent follow-up onto a running nested L2 (GitHub #143)
# KEEP fire-and-return spawn, overlay interject, nested_spawner_can_resume_from_completed_reparented_child
cargo test -p xai-tool-types --lib -- \
parent_cannot_talk_to_own_l2s_follow_up_field_is_optional_and_distinct_from_resume_from
cargo test -p xai-grok-tools --lib -- \
parent_cannot_talk_to_own_l2s_follow_up_enqueues_interject_on_running_l2_without_kill_or_respawn \
parent_follow_up_does_not_inject_into_live_l3_unless_operator_targeted_that_specialist \
resume_from_of_running_l2_still_fails_active \
parent_follow_up_off_is_upstream_spawn_wait_resume_from_completed_only \
parent_follow_up_onto_running_l2_with_live_l3_hits_l2_not_l3 \
parent_cannot_talk_to_own_l2s_task_tool_run_follow_up_returns_queued_and_does_not_spawn \
parent_cannot_talk_to_own_l2s_task_tool_run_follow_up_and_resume_from_are_mutually_exclusive
cargo test -p xai-grok-shell --lib -- \
shell_child_follow_up_sends_session_command_interject_on_child_session \
subagents_config_parent_follow_up_false_parses_and_omitted_defaults_true \
resolve_subagents_copies_parent_follow_up
# Extra: compact standing-law reminder (not AGENTS.md)
cargo test -p xai-grok-shell --lib -- post_compact_reminder_includes_surmount_standing_law
# Extra: CoT death spiral stop + compact not 75k (GitHub #133)
# Keywords: DEST_ENCODER_SKIP_LOOP, stream Fatal, compact 8192/32768.
# Surmount vs SpaceXAI. Oh My Pi is not in this tree.
cargo test -p xai-grok-sampler --lib -- dest_encoder_skip_loop_is_repetitive \
dest_encoder_skip_single_sentence_four_times_is_repetitive \
thought_line_loop_is_repetitive \
chat_completions_stops_dest_encoder_skip_loop \
chat_completions_stops_dest_encoder_skip_loop_in_thought \
messages_stops_dest_encoder_skip_loop \
responses_stops_dest_encoder_skip_loop \
classify_repetitive_generation_is_fatal
cargo test -p xai-chat-state --lib -- \
compaction_reseed_drops_dest_encoder_skip_loop_below_75_2k \
compaction_reseed_of_unique_75k_history_must_not_leave_wasteful_75k_context \
compact_summary_budget_is_8192_tokens_and_reseed_reserve_is_32768 \
format_compact_summary_caps_unique_75k_body_to_compact_summary_budget \
build_compacted_history_unique_75k_summary_stays_within_compact_summary_budget
# Extra: spawn-prompt fold + last-answer caps
cargo test -p xai-grok-sampling-types --lib -- fold_spawn_prompt
cargo test -p xai-chat-state --lib -- parent_estimated_tokens_omit_huge_spawn_prompt
cargo test -p xai-tool-types --lib -- to_model_text_caps_huge_last_answer_for_parent_ingest
cargo test -p xai-grok-tools --lib -- \
completed_subagent_task_output_is_capped_or_points_at_report \
blocking_spawn_subagent_completed_to_prompt_format_is_capped
# Extra: hop flatten / any stored included remaining (rank hop_* is not this)
cargo test -p xai-grok-shell --lib -- \
sampling_config_hop_team_remaining_personal_exhausted_not_dollars_or_console \
sampling_config_hop_personal_remaining_team_exhausted \
sampling_config_hop_both_remaining_team_first_then_personal \
sampling_config_hop_both_included_exhausted_dollar_credits_before_console \
sampling_config_hop_missing_heavy_false_100_keeps_sibling_included \
sampling_config_hop_dollar_credits_on_both_missing_heavy_keeps_team \
prepare_sampler_for_turn_does_not_flatten_missing_heavy_100_off_sibling \
prepare_sampler_for_turn_does_not_flatten_dollar_credits_on_both
# Extra: user-guide fork pins beyond class 1 resume + class 7 skills
cargo test -p xai-grok-pager --lib -- user_guide_does_not_claim_automatic_host_hop_is_unshipped \
user_guide_names_token_economy_spend_order \
user_guide_limits_names_fail_open_and_named_commands
# Neighbors that still have a matching fn (titles / stream retry emit / rebuild fail / /limits).
# Do NOT add retry_chrome_soft_reconnects_*, shell_collision, default_title_items_include_agents:
# those identifiers have no matching fn. Stuck-retry pager chrome is not fully proven.
cargo test -p xai-grok-pager --lib -- window_title_always_manages_non_empty_branded_osc \
titles_on_session_name_osc_is_non_empty_branded window_title_osc_payload_never_empty_string \
show_limits format_supergrok_session footer_names_live_principal \
limits_json_lists_two_supergrok_principals_when_both_slots_exist \
limits_json_honest_single_supergrok_session_cannot_see_team_plan
cargo test -p xai-grok-shell --lib -- stream_started_emits_retry_state_stream_resumed
cargo test -p xai-grok-sampler --lib -- wait_before_attempt_aborts_on_cancel \
retry_footer_reason_uses_short_transport_label \
retry_footer_backoff_hint_appends_next_try_in \
stream_headers_timeout_defaults_to_120_secs_when_env_unset
cargo test -p xai-grok-sampler --test stream_headers_timeout
just check # full gate before push/PR; does not replace missing catalog fn namesCI is for checks only: never build a shippable release package in GitHub Actions (supply-chain boundary). Humans package from a trusted tree when ready.
| Command | Role |
|---|---|
just check or just ci |
Full Nix local gate (flake-meta + prep + fmt/clippy/tests): run before push. Uses Nix on this machine or configured builders. This is not host cargo. The remote gate is just check-remote. |
just check-local |
Host cargo gate when the VPS is down. Runs cargo fmt --all -- --check, then workspace cargo clippy --all-targets --locked with -D warnings, then cargo nextest run --workspace --locked, then cargo test --doc --workspace --locked. Nextest compile and link, and doctest --jobs, cap at 4 (CARGO_LINK_JOBS). The recipe body does not run flake-meta, nix build, or the remote recipes. just check / just ci still use Nix. Remote remains just check-remote. Named test: grok-nix-helper justfile_contracts just_check_local_is_cargo_only_and_does_not_nix. |
just check-remote |
Optional. Realizes flake metadata and .#workspace-cargo-quality (the same full cargo gate as just check / just test: fmt, then workspace clippy --all-targets (members include cargo-mem-guard and grok-nix-helper), then workspace nextest execution, then doctests, as a Nix derivation) on this host's existing remote builder. rustc requires that builder's surmount-remote feature (and big-parallel) and must not run on the caller. This laptop never auto-detects surmount-remote; the host machines file must advertise it. --option system-features that omit big-parallel does not stop local nixbld (the daemon still advertises big-parallel). Force-remote nix also passes --cores 64 so that rustc can use the builder's cores. Workspace cargo passes --jobs from those cores, capped at 32 (an OOM hedge; not 2 from the package sandbox, and not a full 64 rustc processes). --jobs is after the subcommand (cargo check --jobs). cargo 1.97 has no global cargo --jobs. Quality does not run cargo clippy (external dispatcher; a 1-token jobserver then ignores --jobs). Workspace lint is cargo check with RUSTC_WORKSPACE_WRAPPER=clippy-driver under GNU make -j$CARGO_BUILD_JOBS (after dropping Nix MAKEFLAGS / CARGO_MAKEFLAGS / MFLAGS). One clippy-driver is still one typeck thread; independent crates share that jobserver. That derivation uses the same cargo dev profile as local just test-clippy, not crane's default --release check (one rustc thread at opt-level 3; codegen-units does not parallelize cargo check / clippy). Nix jobs (machines-file max-jobs: how many derivations) are not cargo/rustc workers (jobs inside one derivation). Do not raise Nix max-jobs to fix a single busy rustc. Force-remote nix passes --option max-jobs 0 on the caller so this laptop does not build: crates.io FODs, static.rust-lang.org toolchain tarballs (the builder instruction-set architecture, not extra cores), and crane vendor unpacks go to the remote builder. The VPS fetches those itself (builders-use-substitutes). Force-remote nix build uses --store with the machines-file ssh-ng URI and --eval-store auto so cargo-package, cargo-src, and toolchain store paths stay on the VPS. Default nix build realizes into the local store, then copies each remote output back over SSH (that is local store close, not a remote-builder miss). --no-link skips a local result symlink. -L logs still stream as text. This laptop must not substitute those NARs from cache.nixos.org either. Force-remote exports NIX_SSHOPTS with this account's known_hosts (host-key checks stay on) and copies that host key into the builders line so nix-daemon SSH can verify the builder. Missing builders, a missing known_hosts entry for the machines-file host, or SSH to Host surmount-1 exits 2 with no local cargo fallback. User SSH to Host surmount-1 alone is not enough. GitHub Actions must not use this recipe. Agents may run this recipe (pinned 2026-09-02; one live run at a time; do not restart at five minutes). |
just test-remote / just cargo-remote |
Named cargo on that same remote builder. just test-remote -p xai-grok-pager --lib -- actions::defaults realizes .#workspace-cargo-named-test (nix build --impure, GROK_NIX_FORCE_REMOTE=1, surmount-remote). Tests execute (cargo test --locked or cargo nextest run --locked); not compile-only --no-run. just cargo-remote takes kind test, nextest, clippy, build, or check, then the same filter argv. Reuses workspaceCargoArtifacts. Do not raise Nix max-jobs. GitHub Actions must not use these recipes. Agents must not run cargo test / cargo clippy / cargo build / rustc on this laptop for grok-oss. Agents must not run just test-remote / just cargo-remote unless the operator whitelist those in the same words. Agents may run just check-remote (pinned 2026-09-02; one live run at a time; do not restart at five minutes). The operator owns the VPS builder (pinned 2026-08-25). When a session is explicitly told to run one of those recipes as proof, it must start them as a background job with a long wait (timeout 0 or at least twenty minutes). A five-minute foreground wait that SIGKILLs a still-running VPS compile, then starts the same recipe again, is a failed wait policy (pinned 2026-08-28). Nested grok-oss agents inherit production bash wait policy (auto-background at the wait cap, ten-hour foreground ceiling). |
just test |
Quality suite without re-running full flake prep |
just build / install |
Optional release-style package (not CI) |
GHA quality job: flake-meta → ci-prep → just test (see .github/workflows/ci.yml).
There is no ci-quick or ci-host recipe.
The operator owns the VPS builder (pinned 2026-08-25; check-remote
whitelist 2026-09-02). Agents do not invoke just test-remote,
just cargo-remote, or force-remote nix build to nixbuilder /
surmount-1 unless the operator also whitelist those in the same words.
Agents may run just check-remote: one live run at a time; do not start
a second on the same drv or leftover list; do not restart a live remote
compile at five minutes. Competing nix on this laptop still means wait
or skip. GitHub Actions still must not call check-remote. Dual-pin:
AGENTS.md hard constraint 3b-remote-check.
just require_system does not need a prebuilt grok-nix-helper (pinned
2026-08-26). It is a justfile CI_SYSTEM/uname check, the same map as
parse-time system :=. just check-remote / just require_remote_builder
must not nix build .#grok-nix-helper. Preflight is justfile/uname/SSH
(builders file, known_hosts, inject GROK_NIX_REMOTE_SYSTEM_FEATURES or
live BatchMode). just nix_retry / just flake-meta / the just check-remote metadata step must not require grok_nix_helper_bin. The
live nix_retry body is the justfile recipe (argv exec of "$@",
fail-fast on quality/SSH, force-remote flags). Missing helper must not
fail just check-remote. The operator can run that gate on a dirty tree
without first realizing the helper and without GROK_NIX_HELPER set. Do
not tell them to realize the package first. just check-remote exports
GROK_NIX_FORCE_REMOTE=1 before require_remote_builder so later
nix_retry is force-remote. grok_helper must not exec an empty path.
Later recipes that still need the helper (cargo-remote / test-remote,
recon) locate GROK_NIX_HELPER, PATH, result/bin, or crate target only.
Quality nextest compile/link jobs (pinned 2026-08-26). Workspace
clippy/check stay at cargo jobs 32. cargo nextest run compile and
link uses --build-jobs capped at 4 (CARGO_LINK_JOBS). Named
just test-remote nextest and cargo test use that same cap. 32
parallel mold links of workspace test binaries were SIGKILL'd
(ld returned 137, 128+9) under the builder nix-daemon 32GiB
MemoryMax. Host MemAvailable is larger; cargo-mem-guard reads
/proc/meminfo and would not restart. just nix_retry retries that
linker SIGKILL as infra; a real rustc could not compile without
ld returned 137 still fail-fasts. Do not drop --locked. Do not
skip test_leader_death_repro. Raising nix-daemon MemoryMax is
operator-owned on the VPS.
Quality order: cargo fmt, then clippy, then tests (pinned 2026-08-26).
workspace-cargo-quality (just check-remote) and local just test
run in this order. First cargo fmt --all -- --check. Then clippy on
everything the gate compiles: workspace --all-targets (members include
cargo-mem-guard and grok-nix-helper, so a helper E0106 fails here,
not in a late cargo test). Quality clippy stays cargo check
with RUSTC_WORKSPACE_WRAPPER=clippy-driver under the GNU make
jobserver (not the cargo clippy dispatcher). Local just test-clippy
uses cargo clippy --workspace --all-targets. Then workspace
cargo nextest run, then cargo test --workspace --doc. Workspace
nextest covers those member crate tests. Do not add a late
cargo test --manifest-path. Do not allow-lint. --locked stays.
just check-local uses that same cargo order on this host (fmt, then
cargo clippy --workspace --all-targets --locked with -D warnings,
then nextest, then doctest) and does not use Nix. Named-test
(just test-remote / just cargo-remote) is one cargo kind
plus a filter, not this full chain. Named tests in grok-nix-helper
justfile_contracts:
workspace_quality_fmt_then_clippy_then_nextest_and_helper_tests,
workspace_quality_source_matches_just_test,
just_test_clippy_lints_all_targets,
just_check_local_is_cargo_only_and_does_not_nix.
just update (pinned 2026-08-27). Refresh locks without compiling:
the one workspace Cargo.lock, then flake.lock. Does not run
just check-remote. Quality still runs --locked after that. Named
test: just_update_refreshes_workspace_and_flake_locks.
Quality cargo fmt --check is a hard miss. workspace-cargo-quality
runs cargo fmt --all -- --check first. rustfmt Diff in is a quality
fail, not a flake 502. File-level rustfmt on the written .rs is how
agents keep that gate green. See § File-level infer-from-path verify.
PATH hermeticity (CI / low-mem): with CI_LOW_MEM=1, cargo-ci enters
nix develop .#ci, then grok-nix-helper hermetic-path rebuilds PATH
from /nix/store bins only (ci-tools + stdenv: rustc, nextest, mold, git,
python3, coreutils, ...). Host desktop tools (pw-record / parec / arecord,
...) are not visible to quality tests, matches headless GHA. Interactive
just dev / default shell keep impure host PATH. Audio recorders are
intentionally not in ci-tools; python3 is (cgroup + mock LSP e2e
spawn it under scrubbed PATH). Escape hatch: GROK_CI_ALLOW_HOST_PATH=1.
Closest GHA repro: CI_LOW_MEM=1 CI_SYSTEM=x86_64-linux just ci.
| Idea | Practice |
|---|---|
| Upstream owns the package version number | Keep lockstep with the upstream tree we track (CARGO_PKG_VERSION) |
| Our identity is the git revision | Binary shows upstream version + short git SHA (a git object id, not a download FOD hash) |
| No second release train | No Surmount stable/alpha channel mirroring SpaceXAI |
| No default xAI auto-update | Would advertise official grok builds |
Illustrative only (not necessarily this checkout):
grok-oss <upstream-version> (<short-sha>)
grok-oss --version
grok-oss update --check # vs github.com/SurmountSystems/grok-oss main
grok-oss update --check --jsonSOURCE_REV at the repo root is a monorepo export pin (full upstream-side
SHA recorded for the tree we absorbed), not a substitute for “what is HEAD.”
That SHA is a git object id. It is not SHA-1 hashing of a tarball and not a
Nix FOD pin. New download / FOD verify is SHA-256 or minisign. /rebuild
checks the installed binary with --version, then compares that identity.
If behind: from a checkout run TUI /rebuild (wired, SHA-aware peer
relaunch, persist plus self re-exec), or CLI grok-oss rebuild (same
compile-and-signal core, no self re-exec). Named parse test:
rebuild_subcommand_parses. Do not use the official
curl https://x.ai/cli/install.sh path (that installs upstream grok).
Concurrent grok-oss processes share cooldowns under ~/.grok/rate_limits/
(grok-rate-limit). On HTTP 429-style limits, the strictest wait wins across
processes. Before a sample, the sampler consults that store and waits if a
cooldown is live, so many windows on one machine do not stampede one SuperGrok
identity. Flock JSON is enough for that C1 case; do not add a daemon unless
flock is proven racy. Filenames fingerprint the bearer (never the raw token).
This path is not the exhausted-credit memo, not included SuperGrok period
limits, not SuperGrok dollar credits, and not console team prepaid. A 100%
client printout must not mark SuperGrok used up. Matching nextReset is not
proof of a shared pool. Named test:
peer_process_does_not_sample_during_shared_rate_limit_cooldown.
Disable shared coordination with GROK_DISABLE_SHARED_RATE_LIMIT=1.
Product HTTP paths that wait before send and observe on 429 (403 only when a
retry hint such as Retry-After is present):
| Class | Provider key shape | Examples |
|---|---|---|
| Chat / inference | host + key fingerprint | sampler (xAI, SuperGrok proxy, OpenRouter, BYOK base URLs) |
| SuperGrok billing | proxy host + session fingerprint | GET .../billing?format=credits, auto-topup |
| Management API | management host + management-key fingerprint | prepaid, postpaid, usage series, key validation |
| Imagine image | host + fingerprint + imagine |
image_gen, image_edit |
| Imagine video | host + fingerprint + video |
video_gen start + poll |
| Voice STT | host + fingerprint + voice |
streaming wss://.../v1/stt |
| Responses | host + fingerprint + responses |
web_search |
| GitHub | logical github |
OSS update compare |
Waits prefer server headers (Retry-After, then x-ratelimit-reset) over
hardcoded tier tables. Public docs (accessed 2026-08-03):
- xAI rate limits (per-model RPS/TPM; Imagine image/video have separate RPS; Voice/Imagine tier increases via sales)
- OpenRouter limits (honor
Retry-After/X-RateLimit-*on 429) - GitHub REST rate limits
(primary + secondary;
Retry-After/x-ratelimit-reset)
https://github.com/SurmountSystems/grok-oss
Apache License 2.0: LICENSE.
Third-party: THIRD-PARTY-NOTICES.