Skip to content

ci(publish): cancel superseded runs for pull requests only - #145

Merged
chrisuthe merged 1 commit into
Sendspin:mainfrom
chrisuthe:chrisuthe/task/stop-pages-deploys-failing-after-a-cancelled
Oct 7, 2026
Merged

chrisuthe merged 1 commit into
Sendspin:mainfrom
chrisuthe:chrisuthe/task/stop-pages-deploys-failing-after-a-cancelled

Conversation

@chrisuthe

Copy link
Copy Markdown
Member

Problem

Since #138, two main publish runs have built the report successfully and then failed to deploy it, so the published report is stale (192 cases, revisions [3,4,5], against 16 scenarios and revisions up to 6 on main). scripts/detect_regressions.py compares every pull request against that published baseline, so the regression gate is weakened for as long as it stays stale.

What actually happens

The failed deploy-pages jobs never start. GitHub rejects them before dispatch:

run commit deploy-pages annotation
37652436774 55c4d0b Internal server error. Correlation ID: 6e0591d9-56c8-44d2-b32f-7578e7ef5deb
37639144952 935b5ae Internal server error. Correlation ID: 6af99b7e-2315-4ba1-b23b-2de8ff9748d6

Both jobs have no runner, no steps and no log, and both completed exactly 56 seconds after their publish-report finished. The annotation is only visible on the job page; the check-run annotations API returns an empty list for both.

It is not leftover environment state. The deployments API shows no deployment was created for either failed run, and none for the cancelled runs before them (4c9c79a, 72dcdc4), whose deploy-pages was skipped. The github-pages environment has a single rule, the main branch policy, and no other run held the pages group at either failure.

What the two failed runs share, and no successful deploy does: their publish-report job is the one that cancelled an in-progress predecessor through the per-ref cancel-in-progress group. Runs since #138 that cancelled nothing deployed normally (98f963c, 4a5236c, dd3b9ed), as did every run before it.

This is a correlation, two of two against three of three. Why GitHub errors is not observable from outside.

Change

Only pull request runs share a per-ref concurrency group. Every other event gets a group keyed on github.run_id, so a run that deploys never cancels another through it.

  • Pull request runs never deploy (deploy-pages is gated on github.event_name != 'pull_request'), so cancelling them cannot interact with a deployment. They are also where the superseded macOS minutes were spent, so that saving is kept in full.
  • Superseded main runs run to completion again, as they did before ci(publish): cancel superseded runs per ref without cancelling Pages deployments #138. Two can overlap; their deploys are still serialised by the pages group.

The steps, their conditions and the job outputs are untouched.

Alternatives

  • cancel-in-progress: ${{ github.event_name == 'pull_request' }} with the group unchanged. main runs would then queue in the shared group, and a newer arrival cancels the pending one. That is still a deploying run displacing another through this group, so there is no evidence it avoids the error.
  • Deploy from a separate workflow_run workflow. A second workflow, a cross-run artifact download and default-branch semantics, for a one-line problem.
  • Revert ci(publish): cancel superseded runs per ref without cancelling Pages deployments #138. Gives up a measured saving on pull requests for no reason.

Known limits

Verification

Verified locally:

  • actionlint .github/workflows/publish.yml exits 0 with no findings (actionlint 1.7.12), before and after the change.
  • python -m unittest discover -s tests: 344 tests, OK, 2 skipped.

Only a live merge can show:

  • That deploys succeed again. This pull request is cross-fork, so its own check runs the base branch's workflow and says nothing about the change.
  • That a main push arriving while another main run is in progress leaves both running and both deploying.
  • That a second push to a pull request still cancels the first run.

A run whose publish-report cancelled an in-progress predecessor has its
deploy-pages job rejected by GitHub with an internal server error before
it is dispatched, so the report is not published.

Only pull request runs now share a per-ref concurrency group. Every
other event gets a group of its own, so a run that deploys never cancels
another through it. Pull request runs never deploy and are where the
superseded macOS minutes were being spent.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Approval recommended

The focused workflow change matches the stated deployment-safety goal while preserving PR cancellation.

0 open findings

What changed in this PR

Limits superseded-run cancellation to pull requests, preventing deploy-capable runs from cancelling one another.

Changes:

  • Uses the PR ref for pull-request concurrency groups.
  • Uses the unique run ID for all other events.
File Description
.github/​workflows/​publish.yml Adjusts publish-job concurrency grouping by event type.

🧠 Review effort: Balanced


💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@chrisuthe
chrisuthe marked this pull request as ready for review October 7, 2026 18:51
@chrisuthe
chrisuthe merged commit 5099eb6 into Sendspin:main Oct 7, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants