- Any policy change in one listed instruction file must trigger a relevance review of every other listed file before completion.
- Synchronize applicable shared policy in either direction; preserve intentional agent-specific differences and record when no counterpart change is needed.
- Repository-wide policy belongs in this top-level file.
- Lower-scope instruction files inherit it and must add only narrower rules or explicit exceptions, never repeat the same policy; when a top-level policy changes, review lower layers for conflicts or obsolete duplication rather than copying the new text into them.
- All edits stay inside this project directory. Never edit
$CODEX_HOMEor~/.claude/directly — both are install targets populated from this checkout, and a hand-edit there is overwritten on the next sync. - Permitted roots:
.codex/(Codex config, skills, session policy),.claude/settings.jsonand.claude/settings.local.json,plugins/*/{agents,skills,rules,hooks,bin}/. - The Makefile installs from the pushed GitHub remote, not the local working tree: commit and push first, then
make sync-claudeormake sync-codex. Running it against uncommitted work silently installs the previous state. - Never initiate propagation mid-task; it is a deliberate human-triggered step.
Start every user-facing message with a short plain-English explanation that names the full topic or question before technical details. Apply this to progress updates, questions, approval requests, errors, blockers, handoffs, and final answers; do not use unexplained references such as “both” or “that,” and do not repeat the entire conversation. Keep later evidence precise; do not prepend prose to machine-only payloads or violate an explicitly requested exact output format.
Simplicity and reliability come first. Understand the affected flow and root cause, then prefer the smallest clear, reversible solution that satisfies the verified contract. Prefer established project patterns, standard tools, and deletion over new abstractions, dependencies, configuration, or layers; add complexity only when current evidence proves it necessary.
Verification is part of implementation. Work is not complete until relevant checks pass and failures, residual risks, and deliberately deferred scope are reported accurately.
- Prefer dataclasses for reused, fixed-shape internal records to clarify contracts and reduce field-name mistakes.
- Keep dictionaries for dynamic keys, external JSON, and simple mappings. Shared types modules must reduce real complexity.
- Preserve runtime validation, behavior, and serialized schemas; annotations alone do not enforce types.
- A docstring's opening line must state the documented object's purpose in plain English. Move formulas, assignments, configuration literals, function-call notation, and other code-shaped details into the following description or a relevant section.
Scripts, hooks, bin/ entry points, and CI steps all run on Linux, macOS, and native Windows. A POSIX-only assumption is a defect to fix at the source, never a reason to skip the platform.
pathlib;Path(p).is_absolute()not a leading-slash check;PurePath(p).as_posix()before hashing, serializing, or comparing a path — native separators change the digest.- POSIX-absolute literals are not portable fixtures:
/host/xresolves toD:\host\xon Windows. - Serialized telemetry or provenance paths are cross-host coordinates, not local paths: preserve their exact string; recognize declared POSIX and Windows absolute forms with
PurePosixPathandPureWindowsPath; never convert them through hostPathbefore exact comparison. Regressions must exercise both forms on every host. - Byte-asserted or hashed writes use
newline="\n"or bytes; text mode emits CRLF on Windows. - Sanitized subprocess
env=keepsSystemRoot,SYSTEMROOT,COMSPEC,PATHEXT,TEMP,TMPon win32, else the child Python aborts before running; temp dirs viaos.environ.get("TMPDIR") or tempfile.gettempdir(), never/tmp. - A workflow
run:step invoking.shneeds explicitshell: bash— the Windows default shell dot-sources it and exits zero, a false green. - Symlinks, file modes, and uid checks are capabilities: degrade in production code first.
- Skips are the last resort: never a blanket
skipif(sys.platform == "win32"), always a capability probe skipping onOSError, with each surviving skip documented and re-audited. - Test skips must be collection-time decorators (
pytest.mark.skipif,pytest.mark.skip, or parametrized marks); never callpytest.skip()from a test or fixture body. - Green macOS is absence of regression, not Windows support: prove Windows semantics with
PureWindowsPathorntpath, since monkeypatchingos.namedoes not changepathlib. - Recurrent defect guard: a test simulating another OS must explicitly supply every host-only API and constant it exercises instead of assuming the runner exports them. For absent surfaces such as
os.killpgorsignal.SIGKILL, install test doubles withmonkeypatch.setattr(..., raising=False)and usemonkeypatch.delattr(..., raising=False)in the regression to prove the missing-attribute case; keep the simulated branch running on every host rather than adding an OS skip.
- Benchmark task IDs, target repositories, prompt wording, expected answers, and task-specific source or symbol examples are test evidence, not production content.
- Never copy them into shipped plugins, Skills, templates, or user-facing docs; use neutral generic examples and encode the generalized contract in a regression test instead.
- Plans, reports, scratch artifacts, and private implementation notes are evidence, not production content.
- Never copy plan-only notation, section references, task IDs, private source or code examples, plan-only placeholder names, or private shorthand into shipped code, plugins, Skills, templates, schemas, or user-facing docs, and never make a shipped artifact depend on access to its originating
.plans/or.reports/context. - Re-express every adopted requirement as a self-contained contract with complete or sufficiently descriptive names, neutral examples, and all context needed to understand and verify it without the originating plan.
- Use the lowest-cost capable subagent for small, well-defined support work when the task splits into independent bounded workstreams and the expected time or cost saving exceeds coordination overhead.
- Give each subagent narrow file or evidence ownership, only the context it needs, and explicit acceptance gates; parallelize disjoint work and never assign duplicate investigation or overlapping edits.
- Keep indivisible or very small work in the main agent.
- The main agent owns integration, reviews every handoff against its gates, resolves conflicts, and retains final acceptance for behavior-changing or executable results.
Use the canonical procedure in plugins/codex-rig/shared/adversarial-loop.md for every independent review → authorized-fix cycle. Read it before dispatch; its scope, evidence ledger, three-round limit including initial W_0, independent final snapshot, score weights (20/10/6/4/2/1), trend, and recovery rules are mandatory.
Do not fork the implementing conversation for review or treat a local fix as closed before later independent verification. An open structural finding, the same open signature in consecutive reviews, unavailable independent coverage, or stale final snapshot stops a clean claim; an open security or critical finding also forbids completion and commit. Stop on plateau, non-convergence, or the round cap with open findings. Every such stop reports only completed-round scores (for example W_0 → W_1 → W_2), or not-run when no review completed, plus per-tier residue and evidence, then asks for the concrete missing decision; a clean loop still requires the owning workflow’s remaining gates.
- Never hard-wrap prose in any Markdown file.
- Keep each prose paragraph on one physical line; preserve intentional structural breaks in headings, lists, tables, blockquotes, links, HTML
<details>blocks, and fenced code. - Do not blindly unwrap or reflow a whole file; edit only the intended prose and retain its surrounding structure.
Structure Markdown for scanning and correct execution, not from line length alone. When one paragraph combines multiple actions, conditions, actors, statuses, exceptions, or decision branches, use the smallest fitting structure:
- Parallel obligations or independently checkable facts → bullets.
- Ordered actions, recovery paths, or state transitions → numbered lists.
- Compact closed mappings or comparisons with repeated fields → tables; keep long causal explanations out of table cells.
- Genuine notes, warnings, interpretation limits, or safety boundaries → blockquotes.
- Optional depth that would interrupt the main path → an existing or justified
<details>block. - Keep causal reasoning and cohesive rationale as prose.
- Do not convert paragraphs wholesale, add headings for every rule, or duplicate an existing navigation system.
- Keep headings concise and move detailed contracts below them.
- When reformatting behavior-sensitive agent, skill, setup, approval, or recovery instructions, preserve modal language, exact literals, ordering, and stop conditions; run the affected contract and calibration gates because formatting can change instruction salience even when the words remain similar.
Compression or structural reformatting of any AGENTS.md or CLAUDE.md is behavior-sensitive and must pass every gate before handoff:
- Save a verified byte-exact pre-change backup under
.codex/caveman-compress/backups/, outside active instruction-discovery paths; never overwrite an existing backup. - Compare the backup and result for complete semantic preservation: scope, actors, obligations, modal strength, exceptions, ordering, approval and stop conditions, thresholds, examples, and cross-file relationships must remain unambiguous.
- Preserve headings, list hierarchy, fenced and inline code, commands, paths, URLs, identifiers, versions, numbers, environment variables, and other behavior-bearing literals exactly unless the task explicitly changes them.
- Run the affected Markdown, instruction-contract, and calibration gates. Broad instruction-set changes also require an independent agent followability review against the pre-change backup.
- Reject the compression and restore the pre-change file when any instruction is lost, weakened, broadened, made ambiguous, harder to navigate, or less reliably followed. An unresolved comparison difference blocks completion.
Plugin-specific authoring, installability, cross-reference, versioning, and verification rules live in plugins/AGENTS.md.
- Keep single, simple
str,bool,int, andfloatcases bare with pytest's default IDs; descriptive IDs alone do not justify wrappers. Use concise semantic IDs for generated oversized strings whose default IDs would be unreadable. - Wrap every tuple case, including multi-argument rows, as
pytest.param(..., id="meaningful-case"): pass row arguments separately; preserve a tuple-valued single argument aspytest.param((...), id=...). - Use
pytest.paramwith semantic IDs for functions, containers, and other opaque/unstable objects; retain case-specificmarks=. IDs describe behavior or intent, never memory addresses or object hashes. - Never pass separate
ids=topytest.mark.parametrize—list, tuple, callable, or otherwise. Generated cases and mapping rows follow the same rules; attach each required ID to its case. - Remove optional trailing commas that force short lists or calls onto multiple lines, then run the pinned Ruff hooks. Retain required tuple commas, comments, and Ruff-restored wrapping; skip ambiguous edits. Preserve values, argument unpacking, order, duplicates, marks, and assertions.
- Use semantic markers:
integrationfor real component interactions;installed_pluginfor tests runnable without repository context;packagingfor build/install/payload contracts;livefor real external services or credentials. - Apply markers directly to tests or homogeneous classes; never assign
pytestmark, including lists of markers. Do not infer membership from paths, names, platforms, or mocked subprocesses. Ordinary tests may remain unmarked. - Reuse repeated collection-time
pytest.mark.skipif(...)conditions as named decorators such as_skip_node_unavailable; preserve predicates and reasons. This does not authorize new skips or weaken the capability-probe policy. installed_plugindoes not implypackaging. Alivetest also carriesintegrationand retains explicit opt-in guards; selection never authorizes network or paid execution. Capability probes remain separate.- Register selectors in repository and shipped test configuration; validate with
--strict-markers. Classify tests by their behavioral contract or execution requirements, never duration; use CI duration reports to investigate underperforming tests. Full CI remains unfiltered. - From the repository root, use
.venv/bin/python -m pytest -m installed_pluginor.venv/bin/python -m pytest -m "packaging and not live"; omit paths for project-wide discovery. On Windows use.venv\Scripts\python.exe.
- Python minimum: 3.10. The repository root is an environment anchor, not an installable package.
- Bootstrap test tooling with
uv sync --only-group test; benchmark-only dependencies useuv sync --only-group bench. - Run focused tests with
.venv/bin/python -m pytest <paths>and broaden to the affected suite before completion. - Broad local runs go parallel:
.venv/bin/python -m pytest -n 4 <paths>(pytest-xdist, already in the test group). Measured on the full suite: 11556 tests in ~5 min at-n 4, several times faster than the same run serial. Use-n 4for any run wide enough to be worth waiting on; keep focused single-file runs serial, where worker startup costs more than it saves. - Drop
-nwhen the failure itself is what you are reading: xdist interleaves worker output, hides-xordering, and breaks--pdb. Reproduce a failure serially before diagnosing it. - A test that passes serially and fails only under
-nis a real defect, not an xdist artifact — usually shared machine-global state (a${TMPDIR}sentinel without its-${CSID}suffix, a fixed port, a written file outsidetmp_path). Fix the isolation; never pin that suite serial to hide it. - Lint/format edits via the pinned pre-commit hooks, never the bare tool:
pre-commit run ruff-check --files <changed-python-paths>,pre-commit run ruff-format --files <changed-python-paths>, andpre-commit run mdformat --files <changed-markdown-paths>; directrufformdformatinvocation drifts from the version/config pinned in.pre-commit-config.yaml. - Use
pre-commit run --all-filesonly when the task requires the repository-wide gate; preserve unrelated working-tree changes. - Release and build entry points are plugin-specific; follow
plugins/AGENTS.mdand the owning plugin's scripts and README. Remote publication remains human-owned.