Repository navigation
Bound Autopilot's verify rounds and state the program size before fan-out - #226
Merged
Merged
Conversation
Step 4 sent every proven finding back to the owner and gave each new head a fresh swarm, so a PR only reached a clean verdict when a reviewer found nothing. Send back only findings the diff causes, run the full swarm in the first round only, stop at two fix-forwards, and have the root state the program's size before the first owner starts.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
playbooks/autopilot-full.mdsent every proven finding to the owner. A reviewer can always find one more defect in the files a PR touches, so a PR reached a clean verdict only when a reviewer found nothing. The root now files a defect that trunk already has as an issue, or lists it as a declared gap, and does not hold the merge for it.tests/skill-rules.test.mjspins the three sentences, and thetools/forks.jsonentry describes the wider fork.Why
One seven-PR program ran under the old step 4. Each PR went through up to two audits and two fix rounds. The session transcripts record 67 subagents and about 815M cache-read tokens. After the first round, most findings were defects that
mainalready had.Verification
bun test tests/skill-rules.test.mjsgives 20 pass, 3 fail againstmain's playbook and 23 pass, 0 fail with the edit.bun tools/generate.mjs --checkpasses.tools/sync.mjsforpstackat the pinned SHA lists the file as apolicyfork and reports no failure.bun test tests/gives 953 pass, 35 skip, 24 fail on macOS. All 24 are intests/worktree-audit.test.mjs, which fails the same 24 onmainon this machine. The other test files together give 701 pass, 22 skip, 0 fail.Open risks
playbooks/orchestrate.md,autopilot-stack.mdorshipping.mdfor the same loop.Proposed CHANGES entry
Autopilot-full sends an owner only the findings its diff causes, runs the full swarm in the first round only, stops a PR at two fix-forwards, and states the program's size before the first owner starts.