feat(skill): add cargento-release, which reads surfaces rather than commit prefixes - #233
Merged
Conversation
Adversarial review of the new skill. The measured claims all held: 27 tags, 26 decisions, 15 minor / 11 patch / 0 major, prefix classification right 23 of 26, the three named anomalies, and zero uses of BREAKING CHANGE or the ! suffix. The command blocks did not. - The route grep pinned on `path == "/` matched the GET ladder only. POST routes are dict entries (`"/api/answer": self._answer,`), so the release that added /api/ask, /api/answer, /api/dismiss and /api/ask/withdraw would have scanned clean. Widened to any quoted absolute path in http_api.py. - `$LAST` was set in step 1 and used in step 2, but each block is its own shell. An empty `$LAST` makes `"$LAST"..HEAD` resolve to `HEAD..HEAD`: zero commits, empty diffs, exit 0, indistinguishable from a release with nothing in it. Every block now re-derives it and checks it is non-empty. - The green check used `--limit 1`, which returns one of the three workflows that run on every main push. It reported a green Plugin Compatibility here. Filters on the head commit and lists all three instead. - The scan had no command for hook events or the stdio MCP tool, both of which step 3 and step 4 name as surfaces. v0.13.0 added mcp_server.py, 826 lines, invisible to the old list. Top-level payload keys live in aggregate.py, not sessions.py, so that pathspec gained aggregate.py too. - "The last four commands in step 2 cover it" named the wrong four: it excluded CLI flags and routes, the two most break-prone surfaces, and included a --stat that shows no removals. Replaced with a positional reference that cannot drift. - `$TAG` was used in step 1 before any version existed; the check now runs in step 4 where the proposal names one. `$VERSION` was never assigned in steps 5 and 6. - Step 5 re-pulls main, which can move the tip past the commit step 1 cleared. The workflow releases the tip, not the tagged commit, so the green check now repeats there. - The census sed left unprefixed merge subjects intact: eight full subject lines for v0.7.0 rather than one count. AGENTS.md said the Release workflow runs "behavior-focused dashboard test modules". It runs `unittest discover` over the whole dashboard suite. Corrected in both places. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Jared Scott <jared.scott@variable.team>
Three items from the adversarial review's left-alone list, acted on because each
is cheap and no CI cycle is being spent to get them.
The two cross-document references were plain text, matching `burndown`'s style.
That style is why nothing checked them: `validate_plugins.py` resolves links and
anchors in canonical skill bodies, and a file with no links passes that check
vacuously. Renaming `## Parallel Work` or `## Voice and tone` would have broken
the skill silently with CI still green. Both are now links, and the probe
confirms the check bites: renaming the AGENTS.md heading fails the build with
`Markdown anchor does not exist`. The divergence from `burndown` is deliberate
and is the safer direction.
AGENTS.md restated a figure the skill owns ("across 26 measured version
decisions the prefixes would have been wrong three times"). It was accurate, but
the doc-ownership rule exists to stop exactly this: a second copy nothing would
remind you to update when the history section is re-measured. It now points at
the skill's history section instead of carrying the number.
Step 6 asserted the `chore(release)` bump commit unconditionally. The workflow
skips it when the manifests already carry the tagged version, which is the
initial-release path and the resume path, so the sentence was false on two real
paths. Step 1's parity check is what makes it unconditional on a normal release,
and that is now said rather than assumed.
Left alone deliberately, with the reviewer's reasoning accepted: the plural in
"store relocation environment variables" (four are honored, so the plural is
more accurate than the singular would be), `quality-gate.yml` outside the
Python-floor scan (a CI pin moving alone is not a floor raise, and pyproject is
the declaration site), and `fail_under` riding along in that grep (a threshold
ratchet is worth seeing at release time, and it is now labelled as a CI knob
rather than as evidence of a bump).
Signed-off-by: Jared Scott <jared.scott@variable.team>
Contributor
CoverageThreshold: |
gcko
added a commit
that referenced
this pull request
Aug 28, 2026
…sync marker (#234) Two loose ends from cutting v0.17.0, the first real run of `cargento-release`. Step 5 said `git tag "v$VERSION"`. On a machine with `tag.gpgsign true` that tries to create a signed tag, finds no message, and dies with `no tag message?` before anything reaches the remote. It failed on the release it was written for. The fix is `git -c tag.gpgsign=false`, which forces the lightweight form. That is what this repository's release tags are (`git cat-file -t v0.16.0` prints `commit`, not `tag`) and it is what the workflow's own `git tag -f` re-creates when it moves the tag onto the bump commit, so the override matches the end state rather than working around it. Where signing is off it does nothing. A `git cat-file -t` check rides along because the wrong outcome is silent in the other direction: an annotated tag pushes fine and the mismatch only shows up later. Rehearsed this time rather than reasoned about: created a throwaway tag locally, confirmed `commit`, deleted it, and confirmed the bare form still fails. The comment in the skill records why that rehearsal is possible and why the original command was the one that went untested. Also stamps `docs-synced-through` from `0f9d086` to `ec10918`, per AGENTS.md's rule to advance it once from main after merges rather than per branch. The range covers four commits: the previous stamp (#231), the sessions-table title label (#232), the `cargento-release` skill itself (#233), and the v0.17.0 release bump. Signed-off-by: Jared Scott <jared.scott@variable.team>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds
cargento-release, a repository development skill that decides which version to cut from what changed since the last tag, cuts it, verifies it landed, and optionally posts release notes to Slack.Why a skill rather than a paragraph in AGENTS.md
AGENTS.md already owns how to release: tag
main, the workflow does the rest. It said nothing about which number, and that turns out to be the part with a wrong intuition sitting under it.Measured over the whole tag history: 27 release tags, so 26 version decisions after
0.1.0. Fifteen minor, eleven patch, no major ever. Reading the bump off commit prefixes would have agreed 23 times and been wrong 3 times:v0.4.3was a patch containing afeat— listing Cargento in the shared marketplace. Packaging.v0.8.1was a patch containing afeat—feat(bench), a tool for measuring a session mix you do not have. A developer benchmark is not a product surface.v0.8.0was a minor containing only afix— to the release workflow itself. This one does not dissolve and the skill says so.The first two stop being exceptions the moment the question is "can a user see it?" rather than "what prefix did the author use?". So the skill reads surfaces (
web/, CLI flags, HTTP routes, the harness registry, payload keys, hook events, the stdio MCP tool, the Python floor, the shipped skill body) and treats the commit census as evidence rather than verdict.Major is deliberately not inferable. Cargento is
0.x, and semver holds the major at 0 until the interface is declared stable, so1.0.0is a decision to commit to that interface rather than a consequence of any diff. The skill refuses to infer one and instead lists what1.0.0would freeze. A skill that classified majors from a changelog would be inventing an answer.The Slack half
Gated, and in that order deliberately: ask during the release approval, then draft, then verify every claim against the commit range, then run
humanizer:humanizer, then show the operator the text, then send. Verification precedes humanising because humanising a false claim only makes it read better. The channel is resolved by search at run time rather than hardcoded, since this repository is public.Adversarial review
One reviewer, with authority to fix, found 8 defects — every one a silently-empty scan, which is the worst failure mode here because a grep that returns nothing reads as "no breaking changes". The one worth naming: the HTTP-route grep matched only the GET ladder, because POST routes are dict entries. On the range that added the ask lane it found 7 changes and missed 9, including
/api/ask,/api/answer,/api/dismissand/api/ask/withdraw. A release adding four public routes would have scanned clean.sync-docsdocuments that exact trap; the skill shipped the bug it warns about.It also found
$LASTunset across shell blocks (making"$LAST"..HEADresolve toHEAD..HEAD: zero commits, empty diffs, exit 0),gh run list --limit 1returning one of three workflows that run on every main push, and two whole surfaces missing from the scan — hook events andmcp_server.py, which landed 826 lines inv0.13.0.Every claimed number survived re-derivation. It tested competing classification rules and found none that fits as well, and it reported the honest limit:
cargento_runtime/web/did not exist beforev0.6.0, so five of the fifteen minors predate the layout the scan pins to and a full backtest of the surface rule is not available. The skill never claims one.AGENTS.md
Two additions (the doc-ownership row, and a pointer from "Versioning and Releases" to the skill), plus one correction the review surfaced: that section claimed the Release workflow runs "behavior-focused dashboard test modules". It runs
unittest discoverover the whole dashboard suite.Verification
validate_plugins.pygreen, and proven to cover the new skill: removing the Codex alias fails withmissing Codex alias for canonical Claude skill.## Parallel Workfails withMarkdown anchor does not exist.short_description38 chars (25 to 64),default_promptcontains$cargento-release, symlink relative and correctly targeted, no bannedlocalhostspelling.ruff format --check,mypy --strict(108 files),lint_embedded.py(node present, JS syntax genuinely checked), 1,935 dashboard tests + 192 scripts tests OK,bump_version.py --current0.16.0 with no manifest in the diff.