Skip to content

chore(sync): pstack-claude 0.9.74 - #1

Merged
chhoumann merged 30 commits into
mainfrom
upstream-sync/552b1c8
Oct 7, 2026
Merged

chhoumann merged 30 commits into
mainfrom
upstream-sync/552b1c8

Conversation

@chhoumann

Copy link
Copy Markdown
Owner

Merges pstack-claude 0.9.70 through 0.9.74 (upstream 552b1c8, tracking pstack 0.15.13 / cursor/plugins 2cbf585) into t3-pstack as 1.0.3. Model defaults are unchanged. Upstream pstack 0.15.14+ is not in this release.

What came in

New skill

  • /poteto-help (0.9.72) maps a question about pstack to the skill, playbook, or principle that answers it and hands back a prompt to send.

Autopilot and watch-pr behavior

  • Autopilot-full bounds its verify rounds (#226). An owner gets only the findings its own diff causes, the full swarm runs in round one only, a PR stops at two fix-forwards, and the program's size is stated before the first owner starts.
  • watch-pr fails closed (#223). Unsettled mergeability, unregistered checks, and changed PR facts wait instead of reporting READY, and a branch behind its base stops at a new behind-base gate.

Tooling fixes

  • Generator release gate (#220). The first version heading in CHANGES must be the VERSION heading, and the validators are tighter on bare agent dispatches, model IDs, hook paths, missing tiers, agent names, and lead lines.
  • tools/sync.mjs rejects stray arguments, pins --no-diff3, and fails on path collisions (#221).
  • Fail-closed fixes in the fork check, log.sh, check-playbooks.mjs, and resume.mjs (#225). The security workflow also drops --include-git-root.
  • worktree-audit holds untracked and locked worktrees and resolves more symlink spellings (#224, #213, #206, #208). orch serializes stale-lock takeovers (#219).
  • A bare bun test is scoped to tests/ (#212), the generator keeps an ordered list of lead lines per file, and more.

Other runtimes

  • GitHub Copilot runtime (#164): a Copilot manifest, hooks, sheet checker, copilot-tools.md, and a Copilot preamble on each skill in its table. It only runs on Copilot.
  • Pi fixes: /loop stop, print-mode /loop, agent restore (#222), multi-select twins (#211), bun-installed Pi lookup (#215).

Adaptations

  • /poteto-help. t3-tools.md gets its Per-skill notes row. The row says the user types it as /t3-pstack:poteto-help, its model and routing check reads t3-pstack-models.md, and setup is /t3-pstack:setup-pstack. It also says install commands and public-copy links point at chhoumann/t3-pstack rather than pstack-claude. The skill's frontmatter is upstream's, unchanged.
  • Copilot manifest identity. .github/plugin/plugin.json takes the t3-pstack name, author, homepage, and repository. The generator's validateCopilotManifest and the packaging test require these to match the Claude Code manifest. Its displayName and description stay upstream's.
  • Copilot sheet name. The claude and codex arms of session-start.sh keep reading t3-pstack-models.md. The new copilot arm keeps upstream's pstack-models.md. session-start.sh claude and codex print the T3 context byte for byte as on main, including "This runs on T3". The overlay's hook-test helper now names the sheet per runtime.
  • Lead lines. Upstream now stamps an ordered list of preambles per skill. The T3 row displaces Codex: T3's preamble replaces the Codex one as in 1.0.0, and follows the Copilot one. stampLeadLine drops any generator-owned lead line a file no longer owns, which is what the overlay's earlier replace-in-place did. A Codex prompt stub still skips the codex-tools.md pointer on a skill with the T3 preamble.
  • Upstream tests adapted. Two generate.test.mjs tests used the Codex preamble on how and the "name": "pstack" Copilot manifest. They now use the Copilot preamble and t3-pstack. The sync fixture copies only t3-sheet.mjs rather than the whole setup-pstack/scripts directory, so upstream's Copilot scripts don't count as port-only.
  • Funding. Dropped .github/FUNDING.yml, the Buy Me a Coffee badge, and the "buy the maintainer a coffee" line. They fund pstack-claude's maintainer. The License section keeps the credit.
  • #220's checks. They flagged nothing in t3-tools.md or the T3 sheet example, so no overlay text changed for them.
  • VERSION 1.0.3 with a CHANGES entry. The overlay ledger in docs/t3-pstack-design.md now names the files this merge touches, and its stale "plugin name stays pstack" line is fixed.

Verification

  • GIT_CONFIG_GLOBAL=/dev/null bun test tests/: 1197 pass, 36 skip, 0 fail.
  • bun tools/generate.mjs --check: 83 generated files current, all checks pass.
  • node plugins/pstack/skills/setup-pstack/scripts/t3-sheet.mjs ~/.claude/t3-pstack-models.md ~/.config/t3-pstack/t3-catalog.json: ok: every role line parses and dispatches on this T3.

Commits: aba1d67 is the merge with conflict resolutions. d9ebc3b holds the separable adaptations.

michael-denyer and others added 30 commits October 5, 2026 19:59
* Keep each agent's lifecycle in one tagged state map

An agent is running here with its run, running under the pi process that
restore left it to, or ended. Two maps keyed by id held those three states
in combination, and a comment on finish() enforced that they changed
together. A single Map<id, AgentState> makes a launch and a finish one
write each, and the readers switch on the kind instead of inferring it
from which map has the id. Persisted entries are unchanged.

* Run one SIGTERM-then-SIGKILL ladder per child

PiChild.terminate and reapOrphan each spelled out the same ladder, and
stop() armed two on one child: close() scheduled one after exitGraceMs
and terminate() started another at once. terminateGroup in child.ts is
now the only ladder. PiChild.end closes stdin, cancels the one close()
would have armed, and starts it once.

* Test that a stop runs one SIGTERM-then-SIGKILL ladder

A stop used to leave close()'s exit-grace timer armed beside its own kill-grace ladder, so a child that ignored SIGTERM got a second SIGTERM and was killed at twice the exit grace instead of after the kill grace.
* Describe each fork's full diff in its forks.json why

Fourteen entries named only the change that first forked the file and missed later ones: the no-checks verdict in the watch-pr reader, policy, renderer, types, and tests; the stack pairing, thread paging, and deadline in github.ts; the fake gt option, colour, and branch-ref tests in orch.test.ts; and the standing objective, dynamic /loop, driver skill, and landing rules in poteto-mode and four playbooks. Each why now matches its diff against the derived upstream file at the pin.

* Drop a stale /goal clause and name two omitted fork changes

Upstream at e43c7ee never mentions /goal, so the multi-phase-plan why claimed a substitution that no longer exists. The autopilot-stack why now names the recorded head SHA and the policy.ts why the concurrent thread and check reads.
Upstream added this step in the same change that brought create-verification-skill, and the 0.9.10 sync ported the skill but not the step. The port's generator names the skill verify, so the check looks for verify or verify-*, and the invocation links the sibling skill instead of Cursor's install locations.
macOS du -sh pads the size to four columns, so a size shorter than that came back as an empty string, leaving the SIZE column blank and the size sort wrong.
…nscriptNeedles (#203)

* test(worktree-audit): pin shell boundaries and symlinked spellings of a worktree

* fix(poteto-mode): match shell boundaries and symlink spellings in transcriptNeedles

A path followed by a backtick, colon, semicolon, close paren, comma, pipe, ampersand, or angle bracket was missed, so a worktree whose only recent mention was in a command like `cd /x/wt;ls` could be classed safe. Git reports a worktree by its resolved path while a session names it through a symlink in an ancestor directory, as macOS does for /tmp, so those spellings are derived from the filesystem and matched too.

* test(worktree-audit): pin sentence-final paths, closers, alias failures, and the spellings contract

lastChats now takes the spellings of each worktree instead of resolving them, so its table never touches the host filesystem. pathSpellings is tested on its own with a real symlink fixture, skipped where symlinkSync is denied. An ancestor that cannot be listed must leave the chat fact unknown and warn, like every other discovery failure.

* perf(worktree-audit): scan one needle per spelling and check the byte after it

The previous head multiplied boundaries by spellings into up to 36 needles per worktree and ran one includes() pass per needle, which made a scan of unmentioned worktrees about four times slower than main. One indexOf needle per spelling with a boundary check on the following byte brings the scan under main's time. A period counts only before another boundary, and ] and } join the set.

pathSpellings moves out of the scanner into audit, where a failure other than a missing worktree makes that worktree's chat fact unknown through discover instead of being swallowed, and each ancestor directory's links are read once per run. A dangling or looping link is skipped because it cannot spell a path that exists.

* Name node as the worktree audit's runtime in the Codex and Pi notes

Under bun 1.3.14, realpath on a macOS cloud-storage link in the home directory fails with EPERM, so every worktree under it gets an unknown chat fact. Node resolves the same links.
…th tests (#207)

* Pin the at-pin SHA check, paths: parsing, and bun's flat Pi layout with tests

Each behaviour was checked by hand when it merged but nothing in the
suite exercised it. The typecheck's compiler paths move into
tools/pi-package.mjs so a test can call them without running tsc.

* fix(pi): name the skill file when a paths: line is not JSON

parsePaths named the file only for valid JSON of the wrong shape; a comma-separated line threw a bare JSON parse error. Both now throw the same file-naming fault.
#208 released 0.9.70 on top of #186-#207 but its entry covers only #208. Add the earlier fixes to the same entry and widen its title; #208's paragraphs are unchanged.
* Test that multi-select keeps a same-label twin selectable

Picking one of two choices that share a label removes the other from the next multi-select prompt, because picks are tracked by label. The 0.9.70 changelog says identical display text stays selectable; #208 tested only single-select.

* Track multi-select picks by choice, not label

Each choice's numbered display text is unique, so remembering picked display text hides only the choice the user picked. Filtering by the bare label hid every choice sharing that label.
* Run the worktree audit tests on windows-latest

* Make the worktree audit tests portable to Windows paths

* Report chmod-dependent audit cases as skipped on Windows and make the broken worktree portable

The three chmod failure cases now run through test.skipIf, so Windows CI
counts them as skips instead of dropping them. The broken worktree
corrupts its index instead of chmod-ing it, so its unknown-dirty row is
asserted on every platform.
* Scope a bare bun test to tests/

A bare bun test at the root also loaded the vendored poteto-mode scripts tests. watch-pr/cli.test.ts imports commander, which only the scripts directory's own bun install provides, so the run failed until a bootstrap call elsewhere in the suite had installed it, and passed on the next run.

* Let bun count both sides of the discovery test

A regex copy of bun's file patterns counted backups such as .orig files and nested node_modules that bun skips, so a stray file failed the test with bunfig.toml in place.
* Compose ancestor symlinks in worktree-audit's path spellings

pathSpellings applied each ancestor symlink to the worktree's resolved path alone, so a spelling that passes through two links, such as /tmp/link/x for /private/tmp/real/x when /private/tmp/link points at /private/tmp/real, was never produced. A transcript that named the worktree that way was missed and the worktree could be suggested as safe. Each directory on the path is now spelled from the root down, from its parent's spellings and from every ancestor link that resolves to it, so links compose. A link back up to an ancestor of its own directory is still applied through that directory's resolved spelling only, which keeps the set finite.

* Leave the released 0.9.70 changelog entry unchanged
…DIR (#215)

* Look for a bun-installed Pi under BUN_INSTALL and BUN_INSTALL_GLOBAL_DIR

findPiPackage only checked ~/.bun and ~/.cache/.bun, so a Pi added with BUN_INSTALL set was reported missing by tools/typecheck-pi.mjs and skipped by the tests under tests/pi/. It now follows bun's own order for the global directory: BUN_INSTALL_GLOBAL_DIR, then BUN_INSTALL/install/global, then the home defaults.

* Look under XDG_CACHE_HOME for a bun-installed Pi

bun puts its global directory under XDG_CACHE_HOME before HOME when the variable is set, which a Homebrew bun relies on. CONTRIBUTING.md now names every location searched and says CI installs Pi in the npm root.
* Release 0.9.71

* Say what single-select already allowed
* chore(sync): add substitution rules for the poteto-help skill

Upstream's poteto-help skill links the excluded docs/ tree and the
upstream README, lists the excluded make-bot-ui skill, and names Cursor's
skills docs, cloud agents, the Auto model, the driver skills, the
built-ins, Plan Mode, verify-* skills, and a bare /orchestrate. The
denylist rejects three of those lines, and the link validator and the
prose reference test reject others. Fifteen rules rewrite each shape so
the sync writes the file clean and the next sync applies them unattended.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* feat(pstack): sync to upstream 2cbf585 (v0.15.13)

The pin moves from e43c7ee to 2cbf585. The one change inside the sync
boundary is the poteto-help skill with its prompting and recipes
references; the other three upstream commits edit the guide and README
the port excludes.

poteto-help is a port-feature fork. Upstream's copy explains Cursor's
install, Custom Modes, and typed-only skill loading. The port's copy
explains the marketplace install, the SessionStart hook that keeps
routing on, model-invocable skills, the hidden principle leaves, and the
runtimes the port supports with their tool mappings. The reference table
gains its row and counts, NOTICE its attribution, and the generator its
Codex stub.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* chore(release): 0.9.71

Plugin auto-update installs by version number, so the sync ships under a
new one. The changelog entry records the sync report and the fork.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Keep help behavior in runtime sources and the declared fork

* Point poteto-help's babysit fallback at pstack's own skill

Claude Code has no bundled babysit skill. Without /poteto-mode, "babysit this pr" reaches the port's standalone /babysit.

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Michael Denyer <97485362+michael-denyer@users.noreply.github.com>
* Add the GitHub Copilot runtime

Copilot CLI and the Copilot app install pstack from the existing Claude
marketplace and manifest. This adds what the runtime needs beyond that.

- hooks/session-start.sh detects COPILOT_PLUGIN_ROOT and prints stamped
  JSON additionalContext, with the saved model choices escaped in POSIX awk
  so a session never reads the sheet outside the path sandbox.
- hooks/pre-tool-use.sh approves plugin-dir views and a strict vendored
  script form, and denies an off-sheet pstack:* task model or a bad sheet
  write. It is silent outside Copilot and always exits 0.
- poteto-mode/references/copilot-tools.md maps Claude tools, models,
  paths, and skills to Copilot, and runtimes.mjs stamps a Copilot preamble
  on each skill it covers.
- models.json gains a copilot block with no default IDs.
  setup-pstack/copilot.md asks each tier as an ask_user choice list.
- find-transcript.mjs and worktree-audit.mjs read Copilot session-state.
- tests/copilot.test.mjs, tests/pre-tool-use.test.mjs, session-hook cases,
  and the opt-in tests/copilot-smoke.sh.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* Document the GitHub Copilot runtime and release 0.9.59

README and docs/reference.md cover installing on the Copilot CLI and app,
the hook, and the tested CLI range 1.0.87 through 1.0.92. CONTEXT.md and
CONTRIBUTING.md name the Copilot build, its sheet, and the smoke test.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* Read copilot help config to EOF in the smoke test

On CLI 1.0.92-2 the help text outgrew the pipe buffer, so awk's early
exit gave copilot an EPIPE, it exited 1, and pipefail stopped the script.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* Clear the Copilot experiment cache before each smoke probe

Copilot CLI 1.0.92 drops plugin skills from -p sessions once it caches
the computer-use experiment assignment. The setup check now requires the
setup-pstack skill call to succeed, not only to be attempted.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* Test that a command after the sheet heredoc is left alone

A mutation check removed each deny and guard in the PreToolUse rules in
turn. Only the heredoc terminator check survived, so add the case it
guards and name the check in CHANGES.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* Let Copilot setup run every role on the session's model

A plan that exposes one model, or only automatic selection, has nothing to
choose per tier. Setup now opens with one question that sets every role,
and three entries per panel list, to inherit-parent. The PreToolUse sheet
check already accepts that shape without a vendor opt-out.

* Check the Copilot sheet when it is read and give Copilot its own hooks

Review of #164 asked for sheet validation at read time, an allowlisted
context, and no hand-typed role lists. One POSIX awk validator,
skills/setup-pstack/scripts/sheet.awk, now serves both readers:
session-start.sh checks the sheet before it adds role lines, and
setup-pstack runs check-sheet.sh after it writes the sheet. Its role
list is stamped from models.json. The context gets only known role
lines, rebuilt from checked values, so other sheet text never reaches
it. An invalid sheet yields a short "sheet invalid" line and the setup
note. The write-time enforcement in pre-tool-use goes away; the task
model deny and the vendored-script approval stay, and awk errors now
reach stderr.

Copilot reads .github/plugin/plugin.json before .claude-plugin, so it
gets its own manifest and hooks/copilot-hooks.json, which passes a
copilot argument. hooks/hooks.json is upstream's again. Hook commands
use literal ${COPILOT_PLUGIN_ROOT} paths, not $(dirname "$0").
The context is escaped at run time, so the stamped JSON copies and the
empty models.json copilot block are gone.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* Keep the Copilot smoke run's evidence when it fails

The cleanup trap deleted the temp dir, and events.jsonl with it, when
a probe errored under set -e. It now keeps and prints the dir when the
exit status is nonzero or any check failed. Role lists come from
models.json. New checks cover the allowlisted context, an invalid
sheet, check-sheet.sh running without a prompt, and each hook firing
once from Copilot's own hooks file.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* Describe the read-time Copilot sheet check in the docs

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* Record what verified the 0.9.70 Copilot runtime

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* Grant the sheet's directory in the Copilot smoke setup probes

The hook no longer approves a sheet write, so writing $COPILOT_HOME/pstack-models.md is the path permission the user grants at setup. A -p probe cannot grant it; --add-dir does, for that directory only.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* Record the 0.9.70 Copilot smoke run

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* Fix Copilot permission boundaries and transcript scoping

* Keep the session mandate when the Copilot sheet cannot be decoded

session-start.sh names each runtime's sheet reader through its plugin root instead of dirname, and treats a sheet read-sheet.sh cannot decode as missing, so set -e no longer drops the mandate for Claude and Codex. The smoke harness keeps its evidence on SIGTERM and SIGHUP and prints Copilot's output when install or list fails; the TUI driver stops at a permission request instead of approving it and exits 1 when the turn never ends; the Copilot transcript CLI test skips without node.

* Approve Copilot script path operands only inside the workspace

The PreToolUse hook accepted path operands under the plugin root, so an approved log.sh call could append to hooks/session-start-copilot.md and inject text into every later session without a permission prompt. Path operands now resolve under the workspace only, and a workspace inside the plugin gets no approvals.

* Spell the Copilot transcript-root test's paths through the Windows-safe helper

---------

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Michael Denyer <97485362+michael-denyer@users.noreply.github.com>
Adds .github/FUNDING.yml and one README line scoped to maintenance of the port, matching jamma. Nothing under plugins/pstack changes, so the directory listing and installed plugin are unaffected and no release is needed.
Uses the same shields.io badge as pyLocusZoom, placed under the title. Root README only; nothing under plugins/pstack changes.
… and resume (#225)

* test(resume): expect a drive-letter link to need registering

verifyLinks accepts one letter as a URL scheme, so it reads `C:/w.md` as a URL and publish does not ask for the link to be registered with --artifact. The test is marked test.failing until the next commit changes the filter.

* fix(resume): treat a drive-letter path as a local link

A URL scheme now needs two or more letters before the colon. `C:/proj/f.txt` then goes through the same registration check as every other local link.

* test(check-playbooks): pin the root lookup and the extends stems

Four cases print the success line today. They are a run with no argument from a subdirectory, a run with no argument outside a git repository, a root that does not exist, and `extends: ../SKILL`. Each test is marked test.failing until the next commit.

* fix(check-playbooks): resolve the root with git and reject a missing one

With no argument, the root is the output of `git rev-parse --show-toplevel`, so the check reads the project playbooks from any subdirectory. A root argument that is not a directory is an error. An extends stem that contains a path separator is not a playbook.

* test(show-me-your-work): race 40 writers on one log

log.sh hands a row longer than the stdio buffer of the shell to the kernel in several writes, and parallel writers interleave the pieces. The test is marked test.failing until the next commit.

* fix(show-me-your-work): append each log row in one write

log.sh builds the row first and passes it to perl on stdin, and perl appends it with one syswrite. One write to a file opened for append does not interleave with another writer. macOS has no flock(1). The row goes through stdin because Linux limits one argv string to 128 KiB.

* fix(ci): fail the fork check when jq fails or lists no component

The loop read its component list from a process substitution, so a jq error ran the loop zero times and the step passed. The step now captures the list before the loop, stops on a jq error, and fails on an empty list.

* fix(ci): drop the osv-scanner grep that could not match

osv-scanner v2.5.1 prints `No package sources found` on stderr and exits 128, so the step fails before the grep runs. The grep read only stdout. A run of the pinned image on an empty directory returned exit 128 with the message on stderr.

* test(check-playbooks): expect the working directory as root and a no-playbooks line

A repository with no `.agents/playbooks` directory still prints the success line. The tests expect one line that names the directory the check read, from the root and from a subdirectory.

The tests also expect the root to be the working directory again when no argument is given. The `git rev-parse` lookup made a run in a project that is not a git repository exit 1, and the security review flagged the subprocess. The three tests are marked test.failing until the next commit.

* fix(check-playbooks): drop the git lookup and report a missing playbooks directory

The root is the argument or the working directory, as it was before the lookup. The script starts no subprocess again. A bare `git` resolves through the search path, and on Windows the working directory of the child comes first, so a `git.exe` in an untrusted repository would have run. A project that is not a git repository also runs the check from its root again.

When the root has no `.agents/playbooks` directory, the check prints one line that names that directory and exits 0. A run from a subdirectory prints the same line and no longer reports a match.

* test(show-me-your-work): expect a row under PERL_UNICODE and without perl

The perl append fails in two environments where the printf append wrote the row. With `PERL_UNICODE` set, the handles of perl carry a `:utf8` layer and syswrite refuses them. Without perl on the PATH, the pipeline exits 127. Both tests are marked test.failing until the next commit.

* fix(show-me-your-work): append under PERL_UNICODE and without perl

perl now sets binmode on both handles before it reads the row, so a `:utf8` layer from `PERL_UNICODE` or `PERL5OPT` no longer makes syswrite refuse the handle. When perl is not on the PATH, log.sh appends the row with printf as it did before. That row can interleave with a parallel writer when it is longer than the stdio buffer of the shell.

* test(resume): pin the drive-letter filter with a table row

The dedicated test made a directory named `C:`, which a Windows checkout cannot create. One row in the existing table detects the same filter change. Before the change publish exits 0 for the row, and after it the link is rejected as a local link that does not resolve.

* test(check-playbooks): pin a root that is a file

Only a missing root was tested, so a check for existence in place of the directory check passed every test. The new case fails on main, where a file root prints the success line. Two test names now say what the code keys on, and the fixture helper states its null case once.

* docs(show-me-your-work): say why log.sh sets binmode on both handles

The comment gave a reason for STDOUT only, so the STDIN call looked removable. Without it, `PERL_UNICODE=SDA` turns the two bytes of a non-ASCII character into one.

* test(show-me-your-work): pin a row longer than one Linux argument

Two 100 KB cells make a 200 KB row. Linux limits one argument to 128 KiB, so a log.sh that passed the row to perl as an argument stops with `Argument list too long` there, while the row on stdin is appended. The test replaces a comment in log.sh that said so. The race test also names its cell for what makes it large enough.

* docs(ci,show-me-your-work): cut the new comments down to platform facts

The comment review removed sentences that restated the code or argued against alternatives the scripts do not use. Two facts stay, reworded, because the code cannot show them. bash ignores the exit status of a process substitution, and osv-scanner exits 128 when it finds no lockfile.

* fix(ci): fail the osv-scanner step when it finds no lockfile

--include-git-root makes the checkout's .git count as a package, so a scan
that found no lockfile exited 0. Without the flag that scan exits 128, and
both bun.lock files are still scanned.

* test(show-me-your-work): expect a row when perl is present but does not work

Three cases lose the row today: a perl shim that exits 3, a shim that exits 0
and writes nothing, and a PERL5OPT that names a module perl cannot load.

* fix(show-me-your-work): append with printf when perl does not work

log.sh asked only whether perl was on the path. A shim that fails, a shim
that does nothing, or a PERL5OPT perl cannot honour then lost the row. It now
asks perl for one syswrite first and appends with printf when that fails.

* test(show-me-your-work): expect a short write to say what was appended

Under a file size limit the row is cut short and log.sh prints an empty reason.

* fix(show-me-your-work): say how much of a row a short write appended

A short write sets no errno, so the message after a cut-off row was empty.
log.sh now prints the bytes appended and the row's length, and keeps the
system's reason for a write that appends nothing.

* test(show-me-your-work): race 400 KB rows so the unfixed script interleaves every run

With 20 KB rows the race test passed on the unfixed script in 8 of 20 Linux
runs. With four 100 KB cells per row it failed in 100 of 100.

* test: pin the mutations the suite let through and replay the fork check step

log.sh: a row with % cells, quiet stderr when perl does not work, and the
caller's stdin left unread. check-playbooks: a backslash stem where that file
exists. resume: a note that links URLs and an anchor. The fork check step is
read from ci.yml and run against a stub bun for each tools/upstream.json shape.

* chore(forks): state the log.sh fork in one sentence

The why now also names the short-write report and the printf fallback for a
perl that is missing or does not work.

* refactor(show-me-your-work): probe perl with a print and tidy the new tests

The probe only has to show that perl runs a program here, so it prints
instead of repeating the append's binmode and syswrite. The fork check tests
share one jq precondition, the short-write test drops a row it did not need,
and the log.sh why reads as four parallel verbs.
…ks bun refuses (#224)

* test(worktree-audit): a link into an unsearchable directory is not a failure

* fix(worktree-audit): skip links the process cannot resolve

bun's realpathSync throws EPERM on macOS's autofs /home, a sibling in an ancestor of every worktree, so each LAST_CHAT fact fell to unknown and the suite was red on macOS. A link that fails with EPERM or EACCES is dropped like a dangling one; the test for it landed in the previous commit.

* test(worktree-audit): hold worktrees with untracked files

Pins hold-untracked for a worktree whose only dirt is untracked files, counts them one per file through a directory, and keeps the count under a status.showUntrackedFiles=no config.

* fix(worktree-audit): hold worktrees with untracked files

Plain git status folded an untracked directory into one line and honoured status.showUntrackedFiles=no, and classify ignored the count, so a worktree full of never-added source files read safe. The audit now lists every untracked file and gives such a worktree the hold-untracked bucket. The DIRTY label reads untracked:N, since scratch wrongly implied throwaway. The test landed in the previous commit.

* test(worktree-audit): hold locked worktrees with their reason

Pins hold-locked ahead of every other bucket, a LOCKED column after BUCKET that carries the lock reason on one line or the word locked when none was given, and a dash elsewhere.

* fix(worktree-audit): hold locked worktrees with their reason

parseWorktrees dropped the locked field, so a worktree another tool had locked read safe, and the playbook would then delete by hand the directory that git worktree remove refuses. The audit now parses locked and its reason, puts hold-locked ahead of every other bucket, and prints the reason in a LOCKED column after BUCKET. The test landed in the previous commit.

* fix(worktree-audit): end a path at ? and !

A chat that asked "still using /x/wt?" went unseen because neither byte closed the path. Both now do; ? can glob one byte, but a wrong match there holds a worktree rather than deleting one.

* docs(poteto-mode): hold untracked and locked worktrees in the cleanup playbook

Step 4 no longer calls untracked files safe to drop. It names untracked:N and hold-locked, points at the LOCKED column for the reason, and keeps hands off a locked tree. The playbook becomes a declared fork. The Pi and Codex references describe which paths bun fails to resolve and what the audit does with each.

* ci: run the worktree audit tests on macOS

BSD du padding and bun EPERM on the autofs /home link both broke the audit only on macOS, where ubuntu-only CI could not see them.

* test(worktree-audit): name the hidden-untracked case and pin ? before a URL query

The end-to-end fixture set status.showUntrackedFiles=no for every worktree, behind a comment. That case now has its own test, which first shows that plain git status prints nothing for the worktree. A new row pins that ? ends a path before a URL query, which a rule that counts ? only at the end of a sentence would miss.

* refactor(worktree-audit): format the lock cell where the row is built

parseWorktrees returned the lock reason already flattened for display. It now returns the reason as git prints it. lockedLabel, beside dirtyLabel, flattens the reason to one line and supplies the dash for an unlocked worktree, so both rows that print the cell share one rule.

* docs(worktree-audit): compact three comments

The symlinkTargets note gives each tolerated error code its reason in four lines. The boundary note says in one sentence why ? ends a path although it can glob. The note on --untracked-files=all goes, because two tests fail without the flag: the per-file count in the end-to-end test and the status.showUntrackedFiles=no test.

* ci: report how the macOS runner resolves its root links

Temporary. The macOS job's comment says it guards the bun failure on the
autofs home link, and nobody has seen whether a hosted runner has that
failure. This step prints what node and bun answer there. A later commit
in this branch removes it and corrects the comment.

* fix(worktree-audit): keep a worktree out of safe when a link cannot be followed

The branch skipped any ancestor link that failed with EPERM or EACCES.
A link whose route crosses a directory the audit cannot search may still
land on the worktree, and a session that could search the route may have
named the worktree through it. Such a row read safe with no warning,
where origin/main reads review.

A link realpathSync cannot name is now identified by stat and matched to
the worktree's own directories by device and inode. bun's EPERM on the
macOS autofs home link is that case: stat follows it, so it is ruled out
by proof and the last-chat fact stays known. When stat fails too, the
failure propagates and the row reads review with a warning that names
the link. No permission code is tolerated on its own.

symlinkTargets takes its resolver as a parameter, so a test refuses
every link with EPERM on any machine.

* fix(worktree-audit): count submodule work that diff.ignoreSubmodules hides

With diff.ignoreSubmodules=all in the repository config, git status
prints nothing for a submodule that holds an edit, a never-added file or
a commit no remote has. The audit read such a worktree as clean and
safe, and git worktree remove --force then deletes the submodule's
objects with it. The status call now passes --ignore-submodules=none,
which overrides the config and any ignore setting in .gitmodules.

* fix(worktree-audit): show untracked files beside tracked edits in the DIRTY cell

A worktree with one tracked edit and five never-added files printed
wip:1. The cleanup playbook says to show the diff for such a row, and a
diff leaves untracked files out, so a reader could approve deleting
files they never saw. The cell now prints both counts when both are
nonzero, as wip:1,untracked:5. The bucket stays hold-wip.

The playbook's step 4 names the combined cell and says the diff omits
untracked files. The fork entry's why covers the added sentence.

* fix(worktree-audit): count untracked files past the default output buffer

execFileSync caps child output at 1 MiB. A worktree with 14,000
never-added files prints 2 MiB of git status, so the call failed and the
row read unknown and review with no reason, where it should read
hold-untracked. Every git call now runs with maxBuffer: Infinity.

The limit that remains is the longest string the runtime builds: 512 MiB
under node 24 and 2 GiB under bun 1.3.14. Past it the row reads unknown
and review again.

* test(worktree-audit): close the surviving mutations

Four single defects left every test green. The lock reason in the
end-to-end fixture now carries a tab and a trailing form feed, which git
does not trim, so the flatten and the trim in the LOCKED cell are both
pinned. A locked worktree whose directory is gone is pinned as
hold-locked: git never reports a locked worktree as prunable, so the
lock cell on the prunable row could only print a dash and is now that
constant. A link whose route is closed is pinned as an error.

The resolver test refuses only the fixture's own links and asserts that
both were refused. With every link refused, the Windows runner found six
more spellings by identity than by name, through junctions that realpath
answers with a long name where the test path holds an 8.3 name, and the
equality assertion failed there.

The macOS job comment no longer says the job covers the bun failure on
the autofs home link. The hosted runner resolves that link under bun,
and the temporary step that showed it is gone. dirtyLabel keeps its
labels as literals.

* fix(worktree-audit): match every ancestor link by the directory it lands on

Under node a link was matched only by the name realpathSync gave its target. node works that name out from the link's text: it resolves .. before the links in the target and keeps a route through the macOS data volume as written. A worktree named through such a link read safe with a chat from today, where bun read verify-recent-chat.

Every link is now followed with stat and matched by device and inode. The name is kept beside the identity, so no spelling an earlier version matched is lost. A dangling or looping link is read from stat's ENOENT or ELOOP. Any other stat failure still propagates.

The resolver parameter of symlinkTargets becomes a stat parameter. Its test gives way to a link to a directory the user cannot list, which bun cannot name on macOS, and to three rows read under node.

* fix(worktree-audit): keep every directory that shares a link's identity

The identity match kept the first directory on the worktree's path that shared a link's device and inode. On a filesystem that reports one identity for several directories, that could be an ancestor of the directory the link lands on, and the true spelling was lost.

Every directory that matches now counts. An extra spelling can only hold a worktree.

* docs(poteto-mode): show held work with the audit's own status call in the cleanup playbook

Step 4 said to show the diff. A file added inside a submodule reads wip:1 while git diff and git diff HEAD print nothing, and with diff.ignoreSubmodules=all so does a commit made inside one. An agent could read the empty diff as nothing to lose.

Step 4 now names the status call the audit counts from and a diff command that covers submodules, and says an empty diff never means nothing is at stake. Step 5 no longer says no commits are lost: a submodule's commits are stored with the worktree and go with it.
…ing cells (#219)

* test(orch): race two writers for a stale lock and round-trip formula-leading cells

* fix(orch): serialize stale-lock takeovers behind a takeover file

* fix(orch): round-trip cells that start with a spreadsheet formula character

* fix(orch): bound gt calls with a timeout so a hung gt releases the store lock

* test(orch): cover a takeover in progress and simplify the guard write

* fix(orch): make a forced writer wait out a stale-lock takeover

A writer run with --force overwrote the takeover file, so it could replace the lock while an unforced writer was between its re-read and its unlink. Both then held the lock. The takeover file is now exclusive for every writer, and the refusal says to retry.

* fix(orch): clear a takeover that a killed writer left behind

A writer killed between claiming a takeover and replacing the lock left the takeover file in place, and every later unforced writer was refused for good. The claim is now a directory that holds one file named for the claimant's pid. rename refuses a directory that holds a file, so one claim still excludes the next. The next writer removes a dead claimant's file by name, which cannot remove a live claimant's, and then claims as usual.

* fix(orch): carry on when the lock vanishes before the takeover removes it

A lock that its holder released between the takeover's re-read and its unlink surfaced as a raw ENOENT. A missing lock is the state the takeover wants, so it now goes on to create its own.

* fix(orch): keep a leading quote in rows written before cells were unquoted

The read side stripped any leading quote. Rows from before this branch never doubled one, so a brief such as 'quoted brief lost its first character and the ids 'foo and foo collided. Only a quote followed by a quote or a formula character is stripped now, which is the only shape the write side adds.

* fix(orch): kill a hung gt that ignores SIGTERM when its timeout expires

The timeout sent SIGTERM, so a gt that ignored it still held the store lock until it exited by itself. Both gt calls only read, so the timeout now sends SIGKILL. The test covers gt log and gt info separately.

* test(orch): round-trip inbox pointers that start with a formula or quote character

No test failed when the inbox reader stopped unquoting cells or the inbox writer stopped quoting them. This one pins the values read back and the pointer file on disk.

* refactor(orch): reuse the stale-lock test helper and correct the claim comment

The planted-claim test set up its stale lock by hand beside a helper that does the same. The claim comment said rename refuses a directory that holds a file, but the claim itself is such a rename. What rename cannot do is replace one.

* chore(forks): describe the orch store and test forks in full

The store entry now names the takeover claim directory, the quote added to a cell that starts with a quote, the refusal when the holder changed, and the SIGKILL on a timed-out gt. The test entry now names every takeover test, the inbox and older-row tests, and the per-call gt timeout tests.

* fix(orch): create the store lock with its pid already in it

A writer killed between the exclusive open and the pid write left an empty lock. An empty lock names no pid, so no later writer could judge it dead and every unforced writer was refused until someone passed --force. The pid now goes into a private file that is hard-linked onto the lock path, which fails with EEXIST when the lock exists. A filesystem without hard links falls back to the exclusive open and keeps the old window.

The new test kills a writer after every lock call, with and without a stale lock, and checks that the next writer gets in.

* test(orch): give the hung-gt tests a budget a loaded machine can meet

The gt info case needs one gt log call to finish inside the gt timeout. On a saturated machine a single start of the fake gt has taken over a second, so the 200 ms budget failed about one run in four. The budget is now 2000 ms, the fake gt hangs for 20 s, and the wall-clock bound is 10 s, so a gt that is not killed still fails the test.

* fix(orch): create the lock again when its holder releases before its pid is read

A writer whose create failed read the holder's pid next. When the holder released in between, the takeover threw a raw ENOENT and a plain open reported the lock as held by pid unknown. Both sites now share one helper that creates the lock again when it has vanished, twice at most, so a lock path that exists but never reads cannot loop.

* test(orch): pin removal of a dead claimant by name

A writer that listed the claim, saw only a dead claimant, and then paused must not remove a claim another writer has made since. Removing the whole claim directory at that point let both writers replace the lock, and every test still passed.

* chore(forks): describe the orch forks after the lock creation change

States which rows written before the quote was stripped keep a leading quote and which lose it, and covers the hard-linked lock create, the bounded retry, and the tests added for them.
…e unsettled (#223)

* test(watch-pr): pin unknown mergeability as a retry and BEHIND as a gate

* fix(watch-pr): retry unknown mergeability and gate a branch behind its base

* test(watch-pr): retry when PR facts change between the facts and checks reads

* fix(watch-pr): compare every PR fact after the checks read instead of head and base only

* test(watch-pr): wait for a second sighting before reporting a PR has no checks

* fix(watch-pr): confirm a no-checks reading across two polls before calling it ci-none

* test(watch-pr): pin the rollup cursor guard, a missing gh, and the trunk stop in stacks

* fix(watch-pr): refuse a check rollup cursor that does not advance

A repeating cursor paged until the deadline, and without end when no deadline was set. The review-thread loop already refuses one, so the rollup loop now raises the same retryable query failure.

* fix(watch-pr): report a binary that cannot be spawned as a query failure

A missing gh rejected with a plain Error, so the CLI exited 1 with an empty stdout. The spawn error is now a non-retryable spawn-failed query failure, and the CLI emits its status-query verdict with exit 7 on the first attempt.

* fix(watch-pr): stop both stack walks at the default branch

A PR whose head is the default branch, such as a backport or a release PR, became the parent of every PR that targets that branch. Stack discovery now reads the default branch and never walks through it. A long-lived branch that is not the default still links its PRs, and --stack-prs remains the way to name such a stack.

* test(watch-pr): pin the first no-checks sighting under a deadline, in a stack, and in a queue

A run whose deadline is shorter than the interval ends in TIMEOUT with the checks-unreported reason. A stack confirms each PR on its own sightings. A queued frontier is not reported blocker-free until the reading persists.

* refactor(watch-pr): shrink the watcher's diff against upstream

orderStack takes the default branch before the open PR list, so each existing call changes one line. The behind-base and checks-unreported branches sit after the upstream branches they join, which leaves those lines as upstream wrote them. The no-checks tests share one ticking clock.

* refactor(watch-pr): move what the comments explained into names

The second facts read is factsAfterChecks, the absent confirmer is neverConfirms, the deferred gate set is DEFERRED_WHILE_WAITING, orderStack takes trunk, and the fake reader option is factsOnReread. The comments those names replace are gone.

* fix(watch-pr): confirm a no-checks reading only after 60 seconds, whatever the interval

The second sighting came one --interval after the first, and --interval accepts any positive number, so a short interval still reported READY or merge-blocked before CI registered. The confirmation now needs 60 seconds of wall time since the first sighting of that head. The first check appeared within 9 seconds of 67 pushes measured in six repositories, and 60 seconds is the default interval, so a default caller answers no later than before.

* fix(watch-pr): wait on unknown mergeability instead of failing the query

GitHub computes mergeability on the first read that asks for it, so the first read after a push is often UNKNOWN. That threw a retryable query error, which spent --max-query-errors and slept at least 60 seconds, so a healthy PR was slow or exited 7 where it used to answer at once. UNKNOWN is now a wait reason re-polled at the interval. It is never READY, it ends at the caller's deadline as TIMEOUT, and --status-only prints the row and exits 0. Threads, failing checks, and requested changes still stop at once. The re-read no longer counts mergeability resolving from UNKNOWN as the PR changing.

* fix(watch-pr): report a branch behind its base ahead of a required review, and give it a route

BEHIND already stopped at once, but the fork registry said it waited for checks and no test covered either reading. It stays immediate: the branch cannot merge until it is updated, and the update restarts its checks. It now also outranks a required review, which deferred while checks were pending and so hid it. In github/docs six of eight BEHIND pull requests also need a review. Updating after the approval can dismiss it. The babysit and shipping playbooks route behind-base to the branch owner, as they route a conflict.

* fix(watch-pr): keep a PR headed at the default branch out of every stack

The trunk stop kept such a PR from being a parent, but it could still join as a child or seed a stack. A PR that brings main into a feature branch was ordered above the feature's own PR, where the base refused the pair as a cycle. Those PRs now leave the graph before it is built, so the rest of orderStack matches upstream again. A fork branch that happens to share the default branch's name still stacks.

* fix(watch-pr): say what a retry, a missing command, and --status-only's exit 0 mean

The RETRY line called a PR that changed mid-read a failed GitHub query. A status-query blocker told a caller with no gh on PATH to check authentication. The --status-only help promised exit 0 without saying that a row can still be waiting, and the --interval help did not say that a no-checks reading takes 60 seconds to confirm whatever the interval is.

* test(watch-pr): pin the readiness rules that mutation testing found unguarded

Adds a test for each surviving mutation in the audit: one sighting record shared across PRs, the default confirmer, a re-read that ignores state, mergedAt, or headRefName, a rollup cursor that returns to an earlier page, and a hard-coded trunk in stack discovery. The fake clocks now fail after an hour of polling, so a wait that never ends fails a test instead of spinning. The real-clock transport test asserts what each poll said and not how many polls fit, which a loaded machine changes.
* test(pi): cover /loop stop and a new /loop during a self-paced iteration

* fix(pi): let /loop stop end a self-paced loop while an iteration runs

* test(pi): cover a print-mode /loop whose prompt starts no run

* fix(pi): fail a one-shot /loop whose prompt would start no run

A print, JSON, or child run waits for the settle of the run /loop starts. Pi starts no run for an extension command or for a prompt without a model or credentials, and no settle follows, so the command never returned. It now checks those cases first, reports the reason, and marks the process failed.

* test(pi): cover a restored agent whose process is gone under a live parent pid

* fix(pi): orphan a restored agent whose own process no longer runs its session

Restore left any record with a live parent pid to that process. A parent pid has no identity to check, so after a reboot a reused pid kept the agent running forever, and stop_agent and send_message refused it. A record now stays with its launching process only while its own process is alive and carries its session id.

* test(pi): cover a resumed agent's second overflow file

* fix(pi): give each run of an agent its own overflow file

The full copy of output over 50 KB went to one file per agent, so a resumed run overwrote the file an earlier completion notice named. The file name now carries the time the run ended.

* test(pi): cover a resume whose system prompt file is gone

* fix(pi): write a lost system prompt file again before an agent resumes

A resume passed the path of the prompt file the first launch wrote without checking it. Pi appends a path it cannot find as the prompt text, so the agent ran without its agent file. Every launch now writes the file when it is not there, and fails when the agent's type no longer provides one.

* docs(pi): say when a session needs a pi models line

On a provider without its own column each alias resolves to an anthropic ID. Without Anthropic credentials the agent call then fails, and the mapping did not say so.

* docs(setup-pstack): name Pi where the session hook line applies

Step 4 called the line inert outside Claude Code and Codex, and the sheet text and the runtimes paragraph left Pi out, while the Pi row and the extension both honour it.

* fix(pi): say a refused loop wakeup may never have started

schedule_wakeup told the model the loop was stopped or replaced even when no /loop had started it.

* test(pi): simplify the fake's system prompt and model defaults

* refactor(pi): drop comments that restate the code

Keeps the comments that give a reason the code cannot show. Renames two test fixtures so their names carry what a comment used to explain.

* fix(pi): identify a restored agent's process by its parent and start time

Pi sets its process title, so ps shows the single word pi for every child and the --session-id match never held on the real program. A second pi on the same session then listed every live agent as stopped.

A running record now keeps the start time of its pid. Restore leaves the record to its launcher while the process is still that launcher's child, and signals a process only when its start time matches the record. A record without a start time is never signalled.

The fake pi sets its title as pi does, and the tests that stood in an unrelated sleep for the launching pi now use a second process that launches the agent itself.

* fix(pi): write the system prompt file before the worktree is made

A prompt file that could not be written left the agent's worktree and branch behind, because the worktree was created first. The file is written first, so the launch fails with nothing to clean up.

* fix(pi): seal the wakeup slot against the run a /loop command cut off

The re-arm check compared prompt text. It refused a live loop's re-arm worded any other way, in silence, and it let a stopped loop come back under a prompt without the /loop prefix or in the interval form.

A /loop command that ends or replaces the loop while a run is in flight now seals the wakeup slot against that run until it settles. The run is refused every wakeup, whatever the prompt, and the user is told. With no such command every wakeup is scheduled, as before the check. A loop asked for during a run starts once that run settles, so its own re-arm comes from a later run.
…220)

* fix(generate): require the newest CHANGES heading to equal VERSION

* fix(generate): catch a bare agent dispatch in any quoting or spacing

* fix(generate): flag any claude-* model ID, not only the available families

* fix(generate): require a hook path to resolve inside the plugin

* fix(generate): require the default, strongest, and panel tiers in models.json

* fix(generate): require an agent's frontmatter name to match its file name

* fix(generate): converge duplicated or separator-less lead lines when stamping

* fix(typecheck-pi): report the spawn error when bunx cannot start

* refactor(generate): simplify the new checks after review

* fix(generate): catch a bare dispatch behind a backticked, bold, or escaped key

The check read only a plain or quoted key followed by a colon. The tree's
own bullet spelling puts the key in backticks, and a call-style dispatch
uses an equals sign, so both passed. The key and the value may now each
carry quotes, backticks, bold markers, or JSON escapes, and the separator
may be a colon or an equals sign. The test runs each spelling against
every shipped agent name.

* fix(generate): read the whole key and the whole agent name in a dispatch

The pattern had no left boundary, so a key that only ends in
subagent_type matched. Its capture stopped at an underscore or a dot, so
an agent named poteto-agent_v2 or poteto-agent.local read as the plugin's
poteto-agent. The key now starts at a word boundary and the name must end
where the token ends. A dot that closes a sentence still ends the name.

* fix(generate): stop reading a release, uid, date, or issue number as a model ID

The versioned pattern matched claude-, any number of words, and a digit.
That flagged claude-code-2.1.267, /tmp/claude-501/, a dated backup
directory, an SDK version, and an issue slug, and the check has no
per-line exemption. A model ID now needs a one- or two-digit version
after one family word, or a leading generation followed by a family.

* fix(generate): require the first heading in CHANGES to be the VERSION heading

The gate compared VERSION with the first heading shaped like a release.
An entry added without a bump still passed when its heading had another
shape: a v prefix, brackets, two spaces, three hashes, Unreleased, or the
old number reused. The first line that opens a heading below the title
must now be the VERSION heading, and no version may head two entries.

* fix(generate): return problems when the plugin directory is absent

problems() resolved the plugin root's real path outside any check, so a
root with no plugins/pstack threw ENOENT where it used to return one
problem per failed check. The hook resolver now resolves the root when it
is asked about a path, which only happens once a hooks file has been
read.

* fix(generate): converge a lead line glued to the text below it

Stamping repaired a lead line that had lost the blank line above it but
left one that had lost the blank line below it, so the line stayed in
the body's paragraph and --check passed. The stamp now puts a blank line
between the last lead line and the text that follows. A lead line that
ends the file gains nothing.

* test(models): run the missing-tier case for default, strongest, and panel

The test named three tiers and deleted only strongest, so narrowing the
required list to that one tier left the suite green. Each tier now gets
its own case.

* fix(generate): keep a longer agent name from reading as a shorter one

The name's lookahead rejected a word character but not a hyphen, so on
poteto-agent-high_v2 the capture backed up to poteto-agent and the line
was flagged. The lookahead now rejects a hyphen too, so only the whole
token can match.

* refactor(generate): tidy the new checks after review

Say why the gate reads any ## line and not only a release-shaped one,
since the narrower list three lines down invites the old comparison.
Shorten the blank-line test in the lead stamp, and write the dispatch
test's non-empty guard the way the generator suite already does.

* fix(generate): read a claude-* slug as a model ID only when it names a listed family

The versioned-ID pattern could not tell claude-mythos-1 from claude-wt-1. It flagged a Claude Code major (claude-code-2), a worktree, pane, issue, or backup name, and a two-digit temp directory, none of which origin/main flags, and it passed a dotted ID such as claude-3.7-sonnet.

The check now reads each claude-* slug once and flags it when a family listed in models.json follows one of its hyphens as a whole word. That adds the IDs that put the generation first (claude-3-opus-20240229, claude-3.7-sonnet, anthropic/claude-3.5-sonnet) to what origin/main catches and flags nothing else. A family models.json does not list is no longer guessed at.

* fix(generate): skip fenced code when reading CHANGES headings

A correct release failed the gate when an entry quoted a heading inside a code fence: an older heading read as a second entry with one version, and a sample heading above the first entry read as the newest. origin/main passes both.

The gate now reads only the lines outside fenced code blocks, backtick or tilde. A VERSION heading that sits only inside a fence no longer counts as the entry.

* fix(generate): read only a version-led heading as a CHANGES entry

The gate took the first line starting with ## as the newest entry, so a correct release failed when a section such as "## About this file" or "### Format" sat above the entries. origin/main passes those.

The newest entry is now the first heading whose first word carries a version, at any level: "# 0.9.74 - new" above the VERSION heading fails, which no earlier version caught. A heading with no version is not read as an entry. That gives up "## Unreleased" and "## Next" over forgotten work, which origin/main also passes, because the line has the shape of a preamble section.

* fix(generate): catch a bare dispatch whose key follows a literal backslash-n

The word boundary before subagent_type read the n of a literal "\n" as part of the key, so a bare dispatch inside an escaped prompt string passed. origin/main flags that line in its own spelling for all twelve agents.

The key now only has to be free of a leading underscore, which is how a longer key such as my_subagent_type joins. That line still passes.

* fix(generate): keep an author paragraph apart when a lead glued below it is removed

Stamping removes each lead line with the blank line above it. When the lead had lost the blank line below it, that joined the paragraph above to the text below: an author paragraph above the leads merged into the first body paragraph. origin/main leaves such a file alone.

A removed lead now takes the blank line above it only when a blank line or the end of the file follows.

* fix(generate): let only version numbers sit between claude- and the family

The previous commit flagged a listed family after any hyphen of a claude-* slug, so a name such as claude-wt-opus failed where origin/main passes it. Every model ID that puts something before the family puts version numbers there.

The pattern is origin/main's again with one group added: claude-, any run of number parts, then a listed family. Nothing origin/main flags can pass, and the only additions are IDs such as claude-3-opus-20240229 and claude-3.7-sonnet.

* fix(generate): keep a longer key without an underscore out of the dispatch check

The previous commit replaced the word boundary before subagent_type with a no-underscore rule, which flagged a key such as presubagent_type where origin/main and the earlier boundary pass it in a bare spelling.

The key is a whole word again. A letter that follows a backslash does not extend it, so the key after a literal "\n" or "\t" is still caught.

* fix(generate): read an indented fence in CHANGES as a fence

A fence indented by one to three spaces, or nested under a list item, is still a fenced block, and its text may start at the margin. The gate read a heading quoted there as a second entry with one version.

* test(generate): pin the boundaries a mutation run left open

Each case fails under one change to the checks that the suites did not notice: a dispatch key followed by a single quote or a space, a bold value, a longer agent name ending in a capital or a dotted digit, a newer version that extends VERSION as a prefix, a heading that leads with a two-part number, three backticks in the middle of a line, and a lead line glued to the text on both sides.

* perf(generate): match the text before the dispatch key without a lookbehind

A pattern that opens with a lookbehind scanned a 1 MB line in about 21 ms, where the word-boundary pattern it replaced took about 2 ms. Consuming the start of the line, a non-word character, or a backslash escape before the key gives the same answer on 400,000 generated lines and runs at the earlier speed.

* fix(generate): count only a letter after a backslash as an escape before the dispatch key

The escape exception accepted any word character after a backslash, so a Markdown-escaped underscore joined a longer key to the check: my\_subagent_type failed where the word boundary passed it. The exception now covers a letter, as in a literal "\n" or "\t".

* perf(generate): read a CHANGES heading's first word in one pass

The entry pattern let two unbounded runs overlap, so its cost grew with the square of a long run of hashes or digits: 3.6 s on a 100 kB heading line, where origin/main takes under 1 ms. The gate now takes the first word with a pattern that cannot fail and searches it from the start of each run of digits. The two forms agree on 400,000 generated lines, and the same heading line now takes under 1 ms.

* fix(generate): close a CHANGES fence only at a bare marker

A marker with an info string inside a fenced block closed the fence, and the plain marker after it then opened one that never closed. A correct release whose preamble shows a nested fence failed with no VERSION heading. A marker indented four or more spaces is a code block and no longer opens a fence either.

The tests also pin the two halves of the closing rule that nothing checked: the closing run is the same character as the opening run, and it may be longer.

* refactor(generate): tidy the new gate tests and one comment after review

Two rows repeated the case next to them, two titles did not describe every row, and the comment in stampLeadLine sat between a condition and its else.
…n path collisions (#221)

* test(sync): a stray CLI argument must not run a real sync

* fix(sync): reject stray CLI arguments instead of running a real sync

* test(sync): conflicts carry merge-style markers under any git config, and the generator names ||||||| lines

* fix(sync): pin merge.conflictStyle so ||||||| sections never reach the tree, and scan for them anyway

* test(sync): case-only and file/directory collisions must fail before any write

* fix(sync): fail before any write on case-only renames and file/directory collisions

* test(sync): compare git's file modes only, and fail a pin-time run on an upstream file the port lacks

* fix(sync): compare only git's file modes, and fail a pin-time run on an upstream file the port lacks

* fix(sync): pin the conflict style with --no-diff3 so it holds outside a repository

Run from a directory outside any repository, git 2.54 applies the global merge.conflictStyle to merge-file and ignores -c merge.conflictStyle=merge, so a zdiff3 user still got a ||||||| base section. --no-diff3 resets the style after git reads the config, inside a repository and outside one.

* fix(sync): ask the filesystem whether a write collides with a file or directory

The walked paths list no empty directory, so an upstream file at an empty port directory threw EISDIR after earlier files were written. On a case-insensitive filesystem a port file B also blocked an upstream directory b the same way. Both now land in the collisions list and fail the run before any write.

* fix(sync): count every walked port path when looking for case pairs

A path another component carries, or one exclude names, has no outcome. A new upstream file under another spelling of it was written over it on a case-insensitive filesystem, with no conflict reported. The case check now groups the walked tree with the paths to write.

* chore(sync): reword two collision comments and run the in-repository merge case from the repository root

* refactor(sync): name the paths a sync writes once

* docs(contributing): say which conflict markers a sync writes

The marker scan also names a ||||||| line, but a sync never writes one.

* chore(sync): cut comments that restate the code, and inline the case-pair set

* fix(sync): ask the filesystem which entry a write reaches, never model how it folds

The collision check grouped paths by toLowerCase(). APFS also treats NFC and NFD, sharp s and ss, and a ligature and its letters as one name, so an upstream path under one of those spellings was written through to a port file spelled another way: the port's edit was replaced, exit 0, pin advanced.

Each part of a written path must now be absent from disk, or listed in its directory under exactly that spelling with the type the write needs. A part lstat finds that the listing does not spell that way is another entry under an alias, and the run fails naming the port's own path. A deleted path comes from the walk, so it is exact already.

The lookup is a parameter, so the rule runs on Linux against a stand-in filesystem that aliases a pair of names. On a filesystem that holds both spellings of a case pair, the pair is two files and the sync carries a case-only rename out; the lowercase grouping that failed it everywhere is gone.

* test(sync): an upstream directory that reaches a port directory spelled another way collides

Port-only Scripts/x.sh with upstream adding scripts/y.sh wrote the file into Scripts/, exit 0, and the next dry run at the new pin failed naming a file the port held. The exact-spelling rule stops at the directory, before any write, and names the port's Scripts. Before that rule both forms of this test failed with collisions [].

* test(sync): an upstream directory that reaches a port link spelled another way collides

A port link B with upstream adding b/new.md wrote through a link to a directory, threw EEXIST mid-loop on a dangling link after a sibling was rewritten, and threw an uncaught ENOTDIR on a link to a file. The exact-spelling rule stops at the link without following it, so all three fail before any write and name B. Before that rule the six forms of this test failed with collisions [], EEXIST, and ENOTDIR.

* fix(sync): keep a port file's permission bits unless the executable bit changes

Comparing modes as git records them stopped a 0664 clone from reading as a fork, but the normalised mode was also written back, so a merged or updated port file at 0600 or 0700 came out 0644 or 0755. A write now changes the mode only when the executable bit has to change.

* test(sync): pin git's mode rule from the upstream side and for an owner-only executable

Two mutants survived the suite: comparing the upstream file's raw permission bits, and reading the group-execute bit as git's executable bit. The unchanged-mode test now also runs with a group-writable upstream clone against a 0644 port file, and with a 0700 port script against upstream's 0755. Each mutant fails one of the new rows.

* fix(sync): name the stray argument, say rename or delete, and document collisions

The usage error now prints each unexpected argument above the usage line. The collision failure says to rename or delete each port path where it said to restructure the port. CONTRIBUTING describes path collisions, the strict argument list, and that a write leaves a port file's other permission bits alone.

* test(sync): spell the aliasing names as escapes, and pin the unidentified-entry collision

The four aliasing test names were literal characters, so the NFC and NFD constants read as the same string in a diff. They are JavaScript escapes now. A new test pins that an entry the lookup finds, which no listed name matches by inode, still fails the run. Two comments in the tool are reworded.

* refactor(sync): give each type collision its own guard line

The type check compared a boolean with isDirectory() and then re-tested it in a ternary. Two guards put each reason beside its condition. The mode paragraph in the header names its noun again, and the permission-bit test names the port's edit once.

* fix(sync): grant no permission the file's bits or the umask withhold

Under umask 077 a sync created a new executable file as 0755, flipped an existing 0600 file to 0755, and flipped a 0700 file to 0644. The same inputs gave 0700, 0700 and 0600 before this branch. A new file is now created through the umask. An existing file changes only its execute bits, and only when git's mode has to change. They are set where the file is readable, or all cleared.
Step 4 sent every proven finding back to the owner and gave each new head a fresh swarm, so a PR only reached a clean verdict when a reviewer found nothing. Send back only findings the diff causes, run the full swarm in the first round only, stop at two fix-forwards, and have the root state the program's size before the first owner starts.
Bump VERSION, add the CHANGES entry for #219 to #226, and stamp the manifests with bun tools/generate.mjs.
# Conflicts:
#	.claude-plugin/marketplace.json
#	CHANGES.md
#	README.md
#	VERSION
#	package.json
#	plugins/pstack/.claude-plugin/plugin.json
#	plugins/pstack/.codex-plugin/plugin.json
#	plugins/pstack/hooks/session-start.sh
#	plugins/pstack/skills/recall/SKILL.md
#	plugins/pstack/skills/show-me-your-work/SKILL.md
#	tests/generate.test.mjs
#	tools/generate.mjs
#	tools/runtimes.mjs
- t3-tools.md gets a /poteto-help row: its model and routing check reads
  t3-pstack-models.md, setup is /t3-pstack:setup-pstack, and install and
  public-copy links point at chhoumann/t3-pstack.
- The Copilot manifest takes the t3-pstack name, author, homepage, and
  repository, which the generator and the packaging test require to
  match the Claude Code manifest.
- Drop .github/FUNDING.yml and the Buy Me a Coffee README line; they fund
  pstack-claude's maintainer. The License section keeps the credit.
- VERSION 1.0.3 with its CHANGES entry, and the overlay ledger in
  docs/t3-pstack-design.md names the files this merge touched.
@chhoumann
chhoumann merged commit 966cdd3 into main Oct 7, 2026
14 checks passed
@chhoumann
chhoumann deleted the upstream-sync/552b1c8 branch October 7, 2026 07:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants