Repository navigation
Auto-retry fuzz-daily runs killed by hosted-runner shutdown (#3471) - #3472
Conversation
Fresh-eyes reviewReviewer verdict: APPROVE. The important item and nits 1, 2, and 4 are addressed in 234f73b; nit 3 is covered by the new 'deliberate gap' comment. Review: PR #3472 — auto-retry fuzz-daily runner killsI verified the load-bearing GitHub Actions semantics against the docs and real-world precedent before judging the code: Blocking issuesNone found. Important issues
Nits / questions
Conventions check: PR title carries VerdictAPPROVE — the mechanism is sound and every GitHub-Actions assumption it rests on checks out; the one Important item is a diagnosability fix, not a correctness bug. |
6 of the 10 daily fuzz runs between Sep 10 and Sep 19, 2026 died the same way: minutes of total silence from both server and fuzzer, then "The runner has received a shutdown signal" and step exit 143 — the hosted ubuntu-latest VM reclaimed mid-fuzz, well under the 40-minute job timeout and before the fuzzer's own deadline. The daily fuzz signal was red ~60% of the time with no actionable product cause. fuzz-daily-retry.yml listens for completions of "Fuzz API: daily run" and re-runs the fuzz job when the kill signature is present: the fuzz job failed and its if: always() "Upload server log" step ended skipped (a runner shutdown skips post-steps). A genuine fuzz-anomaly failure runs the upload step and is never retried — a rerun draws a fresh random seed and could mask a real finding with a passing attempt. The gate allows the original attempt plus two reruns (run_attempt < 3); after that the failure stands for triage. A retry job inside fuzz-daily.yml cannot work: rerunFailedJobs requires a completed run, and the run is not complete while one of its own jobs is executing — hence the separate workflow_run-triggered workflow, with actions: write scoped to exactly that call. scripts/tests/test_fuzz_daily_retry.py pins the couplings that would otherwise drift silently: the trigger name matching fuzz-daily.yml's name: key, the attempt bound, the job/step names the signature keys on, and the narrow permission grant.
- The 422 catch logs err.message and the API's message, so any rerun refusal that is not a duplicate-delivery race is visible in the retry run's log instead of a fixed 'run already re-running'. - The contract test pins 'run_attempt < 3' exactly and requires workflow_run.workflows to stay a list (a bare string would turn the membership check into a substring match). - Comments and the test docstring describe the skip mechanism accurately: the runner VM is gone with the job, so every remaining step — the if: always() upload step included — never starts and is reported skipped; and the inverse gap (runner dies during the upload step: failure/cancelled, no retry, red run stays visible) is stated as deliberate.
234f73b to
bcbd48a
Compare
Python 3.14.8 wraps WatchedFileHandler.emit's reopenIfNeeded in its own try/except and routes failures to handleError, so a hostile log path (file replaced by a directory, rotated volume gone) no longer raises out of super().emit(). RotationSafeFileHandler's suspend path therefore never engaged on 3.14.8: _sink_broken stayed False, the one-shot suspension warning was never emitted, and every record retried the failing reopen with a stderr traceback — the exact behavior the suspend design exists to prevent. CI hit this as a deterministic failure of test_reopen_failure_suspends_sink_not_the_ call_site (Python 3.14.8 on runners vs 3.14.7 in the devenv). emit() now calls reopenIfNeeded() itself, inside its own try, and delegates to FileHandler.emit directly (the stdlib's check inside WatchedFileHandler.emit becomes a no-op after a successful reopen, and on failure ours raises first). The recursion-contract test moves its patch target from WatchedFileHandler.emit to FileHandler.emit — the watched wrapper is no longer on the call path — and a new test installs the 3.14.8 swallowing shape of the stdlib emit to pin the suspend behavior on any toolchain.
|
Created backport PR for
Please cherry-pick the changes locally and resolve any conflicts. git fetch origin backport-3472-to-stable/2.0
git worktree add --checkout .worktree/backport-3472-to-stable/2.0 backport-3472-to-stable/2.0
cd .worktree/backport-3472-to-stable/2.0
git reset --hard HEAD^
git cherry-pick -x 0d7490a92cd7210da08a88d91463074b0af86435 |
Summary
Closes #3471.
Adds
fuzz-daily-retry.yml, a small companion workflow that automatically re-runs the daily fuzz job when the failure is the hosted-runner shutdown signature from the issue — 6 of the 10 daily runs between Sep 10 and Sep 19, 2026 died with minutes of total silence, then "The runner has received a shutdown signal" and step exit 143, leaving the daily fuzz signal red with no actionable product cause.How it works:
workflow_runcompletions of "Fuzz API: daily run". A retry job insidefuzz-daily.ymlitself cannot work:rerunFailedJobsrequires the run to be completed, and the run is not complete while one of its own jobs is still executing.fuzzjob failed and itsif: always()"Upload server log" step ended skipped — a runner shutdown skips post-steps (verified against the Sep 10–19 failure logs). A genuine fuzz-anomaly failure runs the upload step and is never retried, because a rerun draws a fresh random seed and could mask a real finding with a passing attempt.run_attempt < 3(original attempt plus two reruns); after that the failure stands for human triage.actions: writefor thererunFailedJobscall.scripts/tests/test_fuzz_daily_retry.pypins the couplings that would otherwise drift silently: the trigger name matchingfuzz-daily.yml'sname:key, the attempt bound, the job/step names the signature keys on, and the narrow permission grant.Testing
python -m pytest scripts/tests -v— 168 passed (4 new).pre-commit run xenon --all-files— passed.actionlint .github/workflows/fuzz-daily-retry.yml— clean.CI-only change with no operator-visible effect, so no changelog entry (matches #3416).
Rider: file-sink suspend fix (#3551)
CI on this PR hit the Python 3.14.8 stdlib change twice (3.14.8 swallows
WatchedFileHandlerreopen failures intohandleError), which had disabledRotationSafeFileHandler's suspend-on-failure path entirely.emit()now runs the reopen check inside its own try; a regression test installs the 3.14.8 swallowing shape to pin the behavior on any toolchain. Closes #3551.