Skip to content

ci: compare benchmarks on the same runner - #2271

Merged
cpunion merged 3 commits into
xgo-dev:mainfrom
cpunion:codex/benchmark-same-runner-baseline-20260803
Aug 3, 2026
Merged

cpunion merged 3 commits into
xgo-dev:mainfrom
cpunion:codex/benchmark-same-runner-baseline-20260803

Conversation

@cpunion

@cpunion cpunion commented Aug 3, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Check out the pull request base and head sequentially at the same source path and measure both in each platform job.
  • Feed the paired result to setup-benchmark-go-action@v1, so the PR comment compares vs base from the same runner instead of historical data from another machine.
  • Share dependency setup and Go build caches, warm program builds before timing, and report the median of five Go/LLGo microbenchmark samples.
  • Keep main pushes at one suite and retain the existing 20-minute timeout.
  • Consolidate the benchmark commands in benchmark/baseline/run.sh so both revisions use the same current-checkout harness.

The paired artifact/publisher support was added in xgo-dev/setup-benchmark-go-action#3.

Runtime impact

Across the latest 15 successful benchmark runs before this change:

Platform Median Range P90
Linux 3m25s 3m09s–3m34s 3m32s
macOS 3m44s 3m19s–5m29s 4m53s

The first hosted paired run completed in 5m07s on Linux and 5m38s on macOS, about 50% above the previous medians rather than twice as long. After the review stability changes, Linux completed in 5m01s and macOS in 7m03s.

A complete local five-sample suite on macOS/Apple M4 Max took 87.72s, versus 69.93–73s with one sample. Thus main remains a single suite with roughly 15–20s added for warm-up and sampling, while PR jobs remain around 5–6 minutes with wide headroom under 20 minutes.

Stability checks

The first paired report exposed false changes even though compiler code was unchanged: macOS fmtprintf build showed -48.6%, several microbenchmarks moved 20–35%, and binary sizes differed. The review update now:

  • uses the exact same source path for base and head, eliminating path-dependent size differences;
  • performs an unmeasured build warm-up before the three timed program builds;
  • records five independent 250ms microbenchmark samples and reports their median;
  • documents that very small changes still require confirmation across workflow runs.

On the refreshed Linux artifact, all binary size metrics are identical, microbenchmark deltas are 0–1.23%, and the largest program build/run delta is 2.50%.

Validation

  • bash -n benchmark/baseline/run.sh
  • shellcheck benchmark/baseline/run.sh
  • actionlint .github/workflows/benchmark.yml
  • go test ./benchmark/baseline -count=1
  • complete local relative-path benchmark suite and export with five samples
  • action ingestion verification: five stored samples with the correct median
  • end-to-end paired artifact rendering with real llgo output (vs base, trusted same-runner footer)

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review summary

Solid refactor: extracting the benchmark steps into benchmark/baseline/run.sh and running the base + current suites sequentially on the same runner is a clean way to eliminate cross-machine variance. Shell hygiene is good — set -euo pipefail, per-stage subshells for cd isolation, quoted expansions, arg-count validation, and pipefail correctly makes the | tee pipelines fail-fast. The README changes accurately match the workflow and script behavior, and the CI security posture is correct (pull_request trigger + contents: read + no secrets + persist-credentials: false); this PR does not change that posture since PR code was already built/executed here.

The findings below are about CI cost, measurement bias, and a couple of robustness nits — none are blocking.

CI cost & measurement (not inline — spans the workflow)

  • Doubled runtime vs. the 20-minute cap (.github/workflows/benchmark.yml:28, :52-65). The new "Measure pull request base" step adds a full second run.sh pass on every PR: another go build -p=1 ./cmd/llgo of the compiler plus all collector runs (3 workloads × 3 builds + 7 runs) and three go test/llgo test benchmark stages. With -p=1 and mostly GOMAXPROCS=1 there is little parallelism to absorb the doubling, and macos-latest is slower. Worth confirming both runners have comfortable headroom under timeout-minutes: 20, or bumping it / caching the base compiler build (build time is pure overhead, not a measured metric).

  • Fixed base→current ordering bias (.github/workflows/benchmark.yml:52-65). Running base then current back-to-back removes cross-machine variance but introduces a systematic warm-up bias: the first suite runs cold (disk/thermal/toolchain caches), the second warm. Since the order is always base-first, sub-few-percent deltas may be dominated by ordering noise (short -benchtime=250ms, small sample counts give little power to average it out). Consider documenting that small deltas are within noise, or alternating order.

Comment thread benchmark/baseline/run.sh
Comment thread benchmark/baseline/run.sh
Comment thread benchmark/baseline/run.sh
@codecov

codecov Bot commented Aug 3, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@github-actions

github-actions Bot commented Aug 3, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

df0bb22bfe8e | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Build vs base Run vs base
Linux cprintf 18544 B +0.0% 282.914 ms -1.5% (better) 1.323 ms +6.5% (worse)
Linux fmtprintf 2218040 B +0.0% 3.018 s +1.0% (worse) 2.437 ms -7.4% (better)
Linux println 71504 B +0.0% 285.540 ms +1.1% (worse) 1.567 ms -0.1% (better)
macOS cprintf 84672 B +0.0% 437.333 ms -36.2% (better) 4.119 ms -47.5% (better)
macOS fmtprintf 2361520 B +0.0% 3.681 s -2.4% (better) 19.908 ms -2.7% (better)
macOS println 125712 B +0.0% 397.735 ms -47.2% (better) 4.716 ms -37.3% (better)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 13.390 ns/op +0.1% (worse)
Linux BenchmarkMergeCompilerFlags 151 ns/op +0.5% (worse)
Linux BenchmarkMergeLinkerFlags 94.280 ns/op +0.1% (worse)
Linux BenchmarkChannelBuffered 34.310 ns/op +0.1% (worse)
Linux BenchmarkChannelHandoff 26118 ns/op -1.7% (better)
Linux BenchmarkDefer 46.620 ns/op +1.6% (worse)
Linux BenchmarkDirectCall 1.556 ns/op -0.1% (better)
Linux BenchmarkGlobalRead 1.556 ns/op -0.1% (better)
Linux BenchmarkGlobalWrite 2.487 ns/op +0.0%
Linux BenchmarkGoroutine 29641 ns/op -0.8% (better)
Linux BenchmarkInterfaceCall 7.783 ns/op -0.0% (better)
Linux BenchmarkRuntimeGetG 2.492 ns/op +0.1% (worse)
macOS BenchmarkLookupPCRandom 11.840 ns/op -5.2% (better)
macOS BenchmarkMergeCompilerFlags 120.900 ns/op -41.6% (better)
macOS BenchmarkMergeLinkerFlags 75.680 ns/op -25.3% (better)
macOS BenchmarkChannelBuffered 23.510 ns/op +4.6% (worse)
macOS BenchmarkChannelHandoff 8325 ns/op +7.0% (worse)
macOS BenchmarkDefer 32.740 ns/op +0.6% (worse)
macOS BenchmarkDirectCall 1.065 ns/op -3.5% (better)
macOS BenchmarkGlobalRead 1.066 ns/op +0.5% (worse)
macOS BenchmarkGlobalWrite 1.116 ns/op -7.2% (better)
macOS BenchmarkGoroutine 46387 ns/op +11.0% (worse)
macOS BenchmarkInterfaceCall 4.520 ns/op -13.4% (better)
macOS BenchmarkRuntimeGetG 2.192 ns/op +4.9% (worse)

Compared with 11125dce30e7 measured in the same runner job.

@cpunion

cpunion commented Aug 3, 2026

Copy link
Copy Markdown
Collaborator Author

Review follow-up:

  • The first hosted paired run completed in 5m07s on Linux and 5m38s on macOS, versus the previous medians of 3m25s and 3m44s. That is about a 50% increase, not a doubling, with substantial headroom under the unchanged 20-minute timeout. Base/current measurement steps were 2m46s/1m48s on Linux and 2m36s/1m30s on macOS.
  • The first report also confirmed an ordering/noise problem: this PR did not change compiler code, but single-sample results included macOS fmtprintf build -48.6% and several 20–35% microbenchmark changes.
  • 49ba5ee now checks base and head out into the same source path, adds an unmeasured program-build warm-up, explicitly resets the result stream, normalizes local paths, and documents the stable-harness coupling/noise boundary.
  • df0bb22 increases each Go/LLGo microbenchmark from one 250ms sample to five samples; the action stores all five and reports their median. I kept 250ms rather than raising both duration and count to limit runner cost.

A complete local relative-path run with warm-up and three samples took 83.49s versus 69.93–73s before the review changes. Moving from three to five samples should add only a few more seconds per suite; the refreshed hosted jobs will provide the final cost and variance check.

@cpunion

cpunion commented Aug 3, 2026

Copy link
Copy Markdown
Collaborator Author

Five-sample local follow-up: the complete suite finished in 87.72s and exported exactly five samples per microbenchmark. Compared with the 69.93s one-sample warm-cache run, this adds 17.79s per suite, so a paired PR job should add roughly 36s—not minutes—over the first hosted paired run. The refreshed estimate is about 5m43s on Linux and 6m14s on macOS, still with wide headroom under 20 minutes.

@cpunion cpunion left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Direct response to the review on 2326692:

  1. CI cost / timeout: confirmed on both hosted runners. The original paired run was 5m07s on Linux and 5m38s on macOS. The five-sample revision was 5m01s on Linux and 7m03s on macOS, so both retain substantial headroom under the unchanged 20-minute timeout. No timeout increase is needed.

  2. Fixed base→current ordering: confirmed as a real limitation. The original single-sample run produced false 20–48% changes. Commits 49ba5ee and df0bb22 now use the same source path, warm program builds before timing, record five microbenchmark samples, and explicitly document that small deltas require repeated workflow confirmation. Linux false deltas fell to 0–1.23% and all size metrics became identical. macOS still shows a phase/core-scheduling effect across the two whole-suite processes, so increasing benchtime alone would not solve it; the PR now documents this remaining noise rather than treating every delta as conclusive. Alternating at sample granularity would require a larger paired-runner refactor and materially more runner time, so it is not added here.

  3. Inline robustness comments: go.txt reset plus append-only tee, absolute path normalization with a complete relative-path test, and the current-harness/base compatibility comment are implemented in 49ba5ee. Each original inline thread has a direct reply and is resolved.

@cpunion
cpunion merged commit 91d416b into xgo-dev:main Aug 3, 2026
44 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant