Skip to content

build: report final Wasm, ELF and PE sizes - #2719

Merged
cpunion merged 12 commits into
xgo-dev:mainfrom
cpunion:codex/fix-size-report-2679-20261003
Oct 4, 2026
Merged

cpunion merged 12 commits into
xgo-dev:mainfrom
cpunion:codex/fix-size-report-2679-20261003

Conversation

@cpunion

@cpunion cpunion commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

llgo build -size measures final Wasm, ELF and PE artifacts directly, preserves Go method names, and avoids duplicate ELF symbol ranges. PE reports exclude raw-file alignment padding and distinguish stored data from virtual zero-fill. Mach-O retains its native tool fallback.

Move parsing, aggregation and text/JSON output into the standard-library-only internal/sizereport package. The build layer selects final artifacts after post-link processing and supplies package metadata. Requested-report failures propagate to the caller.

Part of #2679.

Validation: standalone tests with CGO_ENABLED=0 (97.7% statement coverage), vet, build integration tests, PE32/PE32+ and malformed-artifact cases, real Windows executables, and native/WASI final-artifact smoke tests. PE regressions cover section-relative symbols and zero-fill at a nonzero virtual address. Real-artifact acceptance lives under test/sizereport.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Size-report rework review

Solid, well-tested change. The ELF interval sweep, LEB128 decoding (signed + unsigned with overflow/bounds checks), and Wasm byte-accounting are correct and backed by strong tests, including a fuzz harness that enforces code + data + custom_bytes + structure_bytes == file_size. Documentation matches the implemented behavior. Findings below are a few cleanups and one design decision worth confirming; nothing blocking.

Design decision to confirm — report failure now aborts a successful build

build.go propagates reportFinalSize errors instead of printing Warning: size report failed, and the doc states this intentionally. The format/level are pre-validated in ensureSizeReporting, but reader/os.Stat/stdout-write failures in the diagnostic path are not — so a transient failure while reporting can now discard the exit status of an otherwise-successful compile (the binary is still on disk, but the command exits non-zero). If that is the intended contract ("a requested report that cannot be produced fails the build"), no change needed; if purely diagnostic I/O failures should not fail a good build, consider distinguishing them.

Performance — symbol attribution is O(symbols x packages) on large binaries

nameResolver.matchModule/matchPackage (internal/build/resolver.go, unchanged here but now exercised per-symbol by both the ELF and Wasm readers) linearly scan every package in link.allPkgs (the full transitive dependency set — often hundreds to thousands) and only cache positive matches. Symbols that match no package (runtime/C/asm/mangled symbols — frequently a large share of a linked binary) re-scan all packages on every call and are never memoized, because "" doubles as both "absent" and the stored value. On a large binary this is the dominant cost. Consider caching misses (sentinel distinct from a real value) and/or using the precomputed pkgPrefixes with a sorted-prefix binary search instead of a per-symbol linear scan. Not blocking, but worth addressing given this PR makes the resolver the hot path.

Comment thread internal/build/size_report.go Outdated
Comment thread internal/build/size_report.go Outdated
Comment thread doc/size-report.md Outdated
@codecov

codecov Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.67675% with 7 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
internal/build/size_report_build.go 90.38% 5 Missing ⚠️
internal/sizereport/size_report.go 98.37% 2 Missing ⚠️

📢 Thoughts on this report? Let us know!

@github-actions

github-actions Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

34acf063ad36 | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 154193 B 0 B / +0.0% 88742 B 0 B / +0.0%
cprintf/j32-goos-js/LLGo 152295 B 0 B / +0.0% 73165 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 141248 B 0 B / +0.0% 92630 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 152848 B 0 B / +0.0% 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 153560 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3213561 B 0 B / +0.0% 132442 B 0 B / +0.0%
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3192660 B 0 B / +0.0% 101635 B 0 B / +0.0%
fmtprintf/j64-emscripten-memory64/LLGo 2950168 B 0 B / +0.0% 139285 B 0 B / +0.0%
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2347479 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2344141 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-emscripten/LLGo 153428 B 0 B / +0.0% 88742 B 0 B / +0.0%
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 151768 B 0 B / +0.0% 73165 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 140582 B 0 B / +0.0% 92630 B 0 B / +0.0%
reflectcall/j32-emscripten/LLGo 1543898 B 0 B / +0.0% 105908 B 0 B / +0.0%
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1546371 B 0 B / +0.0% 90331 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1428570 B 0 B / +0.0% 111641 B 0 B / +0.0%
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1276821 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1274429 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 152495 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-wasi/LLGo 153207 B 0 B / +0.0% 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 6.373 s +322.2 ms / +5.3% (worse)
j32-goos-js 6.200 s +109.4 ms / +1.8% (worse)
j64-emscripten-memory64 5.792 s +517.9 ms / +9.8% (worse)
reflectcall/w32-wasi 24.465 s -1.475 s / -5.7% (better)
w32-goos-wasip1 4.559 s +273.3 ms / +6.4% (worse)
w32-wasi 4.517 s +437.2 ms / +10.7% (worse)

Compared with 5f1f13897af7 measured in the same runner job.

@github-actions

github-actions Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

34acf063ad36 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 8568 B 0 B / +0.0% 387 B 0 B / +0.0% 1.019 s +35.96 ms / +3.7% (worse) 1.243 ms -31.21 us / -2.5% (better)
Linux cprintf-lto 8408 B 0 B / +0.0% 368 B 0 B / +0.0% 962.091 ms +25.88 ms / +2.8% (worse) 1.250 ms +12.58 us / +1.0% (worse)
Linux fmtprintf 4880944 B 0 B / +0.0% 501619 B 0 B / +0.0% 7.832 s +425.7 ms / +5.7% (worse) 3.097 ms +77.65 us / +2.6% (worse)
Linux fmtprintf-lto 3600808 B 0 B / +0.0% 438538 B 0 B / +0.0% 17.956 s +1.794 s / +11.1% (worse) 3.348 ms +362.9 us / +12.2% (worse)
Linux println 674952 B 0 B / +0.0% 16855 B 0 B / +0.0% 975.365 ms +30.92 ms / +3.3% (worse) 1.737 ms +199.7 us / +13.0% (worse)
Linux println-lto 190024 B 0 B / +0.0% 14273 B 0 B / +0.0% 1.318 s +94.61 ms / +7.7% (worse) 1.595 ms +1.834 us / +0.1% (worse)
macOS cprintf 50736 B 0 B / +0.0% 4409 B 0 B / +0.0% 881.753 ms -53.69 ms / -5.7% (better) 3.429 ms -157.5 us / -4.4% (better)
macOS cprintf-lto 50496 B 0 B / +0.0% 161 B 0 B / +0.0% 807.964 ms -54.92 ms / -6.4% (better) 2.255 ms +71.75 us / +3.3% (worse)
macOS fmtprintf 1773456 B 0 B / +0.0% 877780 B 0 B / +0.0% 3.975 s -3.394 s / -46.1% (better) 8.090 ms -1.348 ms / -14.3% (better)
macOS fmtprintf-lto 1360928 B 0 B / +0.0% 762428 B 0 B / +0.0% 8.885 s -3.578 s / -28.7% (better) 3.748 ms -4.084 ms / -52.1% (better)
macOS println 99344 B 0 B / +0.0% 24216 B 0 B / +0.0% 906.784 ms +12.35 ms / +1.4% (worse) 5.482 ms +1.768 ms / +47.6% (worse)
macOS println-lto 83664 B 0 B / +0.0% 21457 B 0 B / +0.0% 1.139 s -81.08 ms / -6.6% (better) 3.417 ms -950.3 us / -21.8% (better)
Windows MinGW cprintf 651264 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.626 s -11.04 ms / -0.7% (better) 3.641 ms +127.7 us / +3.6% (worse)
Windows MinGW cprintf-lto 43520 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.674 s +28.55 ms / +1.7% (worse) 3.699 ms +243.8 us / +7.1% (worse)
Windows MinGW fmtprintf 5432832 B 0 B / +0.0% 600310 B 0 B / +0.0% 7.216 s +30.66 ms / +0.4% (worse) 7.973 ms -85.6 us / -1.1% (better)
Windows MinGW fmtprintf-lto 4127232 B 0 B / +0.0% 546582 B 0 B / +0.0% 14.454 s +140.9 ms / +1.0% (worse) 7.594 ms -227.9 us / -2.9% (better)
Windows MinGW println 706560 B 0 B / +0.0% 25190 B 0 B / +0.0% 1.652 s +29.73 ms / +1.8% (worse) 7.328 ms +985.2 us / +15.5% (worse)
Windows MinGW println-lto 208896 B 0 B / +0.0% 22054 B 0 B / +0.0% 1.963 s +58.94 ms / +3.1% (worse) 6.618 ms -29.2 us / -0.4% (better)
Windows MinGW 386 cprintf 601600 B 0 B / +0.0% 5326 B 0 B / +0.0% 1.657 s -26.38 ms / -1.6% (better) 4.926 ms -364.9 us / -6.9% (better)
Windows MinGW 386 cprintf-lto 103424 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.686 s -32.2 ms / -1.9% (better) 5.088 ms +25.3 us / +0.5% (worse)
Windows MinGW 386 fmtprintf 4745216 B 0 B / +0.0% 472478 B 0 B / +0.0% 7.459 s -47.36 ms / -0.6% (better) 10.301 ms +28.8 us / +0.3% (worse)
Windows MinGW 386 fmtprintf-lto 4148736 B 0 B / +0.0% 451114 B 0 B / +0.0% 14.561 s +132.6 ms / +0.9% (worse) 10.580 ms +473.8 us / +4.7% (worse)
Windows MinGW 386 println 653312 B 0 B / +0.0% 21490 B 0 B / +0.0% 1.674 s +16.72 ms / +1.0% (worse) 8.487 ms +232.7 us / +2.8% (worse)
Windows MinGW 386 println-lto 258560 B 0 B / +0.0% 19306 B 0 B / +0.0% 1.994 s +47.89 ms / +2.5% (worse) 8.751 ms -334.7 us / -3.7% (better)
Windows MinGW ARM64 cprintf 661504 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.883 s -52.57 ms / -2.7% (better) 6.314 ms -653.8 us / -9.4% (better)
Windows MinGW ARM64 cprintf-lto 43520 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.905 s -70.15 ms / -3.6% (better) 6.274 ms -773 us / -11.0% (better)
Windows MinGW ARM64 fmtprintf 5343744 B 0 B / +0.0% 510876 B 0 B / +0.0% 7.210 s -16.89 ms / -0.2% (better) 13.504 ms -495.9 us / -3.5% (better)
Windows MinGW ARM64 fmtprintf-lto 4302336 B 0 B / +0.0% 477264 B 0 B / +0.0% 13.917 s +90.45 ms / +0.7% (worse) 13.240 ms +1.2 us / +0.009064% (worse)
Windows MinGW ARM64 println 714240 B 0 B / +0.0% 23884 B 0 B / +0.0% 1.882 s -80.63 ms / -4.1% (better) 10.773 ms -1.377 ms / -11.3% (better)
Windows MinGW ARM64 println-lto 215552 B 0 B / +0.0% 21232 B 0 B / +0.0% 2.154 s -59.72 ms / -2.7% (better) 11.183 ms -335.9 us / -2.9% (better)
Windows MSVC cprintf 893952 B 0 B / +0.0% 65798 B 0 B / +0.0% 1.199 s +52.28 ms / +4.6% (worse) 2.238 ms +240.6 us / +12.0% (worse)
Windows MSVC cprintf-lto 289792 B 0 B / +0.0% 65734 B 0 B / +0.0% 1.118 s +13.8 ms / +1.2% (worse) 2.206 ms -62.8 us / -2.8% (better)
Windows MSVC fmtprintf 5732864 B 0 B / +0.0% 695862 B 0 B / +0.0% 5.524 s +57.18 ms / +1.0% (worse) 5.771 ms -327.4 us / -5.4% (better)
Windows MSVC fmtprintf-lto 4449792 B 0 B / +0.0% 646134 B 0 B / +0.0% 10.233 s -81.04 ms / -0.8% (better) 5.521 ms -259.8 us / -4.5% (better)
Windows MSVC println 1017344 B 0 B / +0.0% 120854 B 0 B / +0.0% 1.189 s +34.33 ms / +3.0% (worse) 4.476 ms +20.1 us / +0.5% (worse)
Windows MSVC println-lto 528384 B 0 B / +0.0% 118390 B 0 B / +0.0% 1.247 s +11.14 ms / +0.9% (worse) 4.469 ms -340.3 us / -7.1% (better)
Windows MSVC 386 cprintf 513536 B 0 B / +0.0% 3931 B 0 B / +0.0% 1.225 s +88.36 ms / +7.8% (worse) 4.881 ms +891.3 us / +22.3% (worse)
Windows MSVC 386 cprintf-lto 44032 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.135 s -2.85 ms / -0.3% (better) 4.258 ms +306.3 us / +7.8% (worse)
Windows MSVC 386 fmtprintf 4479488 B 0 B / +0.0% 455868 B 0 B / +0.0% 5.459 s -8.776 ms / -0.2% (better) 8.522 ms -51.9 us / -0.6% (better)
Windows MSVC 386 fmtprintf-lto 3899392 B 0 B / +0.0% 426651 B 0 B / +0.0% 10.626 s +35.46 ms / +0.3% (worse) 9.589 ms +1.132 ms / +13.4% (worse)
Windows MSVC 386 println 567296 B 0 B / +0.0% 20340 B 0 B / +0.0% 1.132 s -83.18 ms / -6.8% (better) 7.111 ms -1.486 ms / -17.3% (better)
Windows MSVC 386 println-lto 199168 B 0 B / +0.0% 18501 B 0 B / +0.0% 1.329 s -24.21 ms / -1.8% (better) 7.058 ms -174 us / -2.4% (better)
Windows MSVC ARM64 cprintf 662528 B 0 B / +0.0% 4192 B 0 B / +0.0% 1.616 s -33.11 ms / -2.0% (better) 6.970 ms -565.6 us / -7.5% (better)
Windows MSVC ARM64 cprintf-lto 47616 B 0 B / +0.0% 4084 B 0 B / +0.0% 1.609 s -50.92 ms / -3.1% (better) 6.822 ms -374.6 us / -5.2% (better)
Windows MSVC ARM64 fmtprintf 5339648 B 0 B / +0.0% 510808 B 0 B / +0.0% 6.782 s -327.5 ms / -4.6% (better) 15.347 ms -318.9 us / -2.0% (better)
Windows MSVC ARM64 fmtprintf-lto 4309504 B 0 B / +0.0% 477924 B 0 B / +0.0% 13.114 s -282 ms / -2.1% (better) 14.559 ms -664.7 us / -4.4% (better)
Windows MSVC ARM64 println 715264 B 0 B / +0.0% 23908 B 0 B / +0.0% 1.600 s -76.61 ms / -4.6% (better) 12.086 ms -1.519 ms / -11.2% (better)
Windows MSVC ARM64 println-lto 220672 B 0 B / +0.0% 21380 B 0 B / +0.0% 1.842 s -49.39 ms / -2.6% (better) 12.882 ms -248 us / -1.9% (better)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.710 ns/op +0.22 ns/op / +1.5% (worse)
Linux BenchmarkMergeCompilerFlags 219.900 ns/op +25 ns/op / +12.8% (worse)
Linux BenchmarkMergeLinkerFlags 143.800 ns/op +6.2 ns/op / +4.5% (worse)
Linux BenchmarkChannelBuffered 56 ns/op -1.18 ns/op / -2.1% (better)
Linux BenchmarkChannelHandoff 14428 ns/op -518 ns/op / -3.5% (better)
Linux BenchmarkDefer 55.020 ns/op +7.11 ns/op / +14.8% (worse)
Linux BenchmarkDirectCall 1.185 ns/op +0.014 ns/op / +1.2% (worse)
Linux BenchmarkGlobalRead 1.173 ns/op +0.007 ns/op / +0.6% (worse)
Linux BenchmarkGlobalWrite 7.905 ns/op +0.132 ns/op / +1.7% (worse)
Linux BenchmarkGoroutine 30933 ns/op -6702 ns/op / -17.8% (better)
Linux BenchmarkInterfaceCall 5.839 ns/op -0.055 ns/op / -0.9% (better)
Linux BenchmarkRuntimeGetG 2.916 ns/op -0.159 ns/op / -5.2% (better)
macOS BenchmarkLookupPCRandom 10.990 ns/op -4.15 ns/op / -27.4% (better)
macOS BenchmarkMergeCompilerFlags 85.170 ns/op -24.13 ns/op / -22.1% (better)
macOS BenchmarkMergeLinkerFlags 54.720 ns/op -36.96 ns/op / -40.3% (better)
macOS BenchmarkChannelBuffered 22.480 ns/op -6.31 ns/op / -21.9% (better)
macOS BenchmarkChannelHandoff 6632 ns/op -3755 ns/op / -36.2% (better)
macOS BenchmarkDefer 29.010 ns/op -8.24 ns/op / -22.1% (better)
macOS BenchmarkDirectCall 1.003 ns/op -0.098 ns/op / -8.9% (better)
macOS BenchmarkGlobalRead 0.943 ns/op -0.1544 ns/op / -14.1% (better)
macOS BenchmarkGlobalWrite 0.944 ns/op -0.1479 ns/op / -13.5% (better)
macOS BenchmarkGoroutine 37146 ns/op +1384 ns/op / +3.9% (worse)
macOS BenchmarkInterfaceCall 3.490 ns/op -0.565 ns/op / -13.9% (better)
macOS BenchmarkRuntimeGetG 1.882 ns/op -0.327 ns/op / -14.8% (better)
Windows MinGW BenchmarkLookupPCRandom 13.210 ns/op +0.05 ns/op / +0.4% (worse)
Windows MinGW BenchmarkMergeCompilerFlags 657.400 ns/op +36.7 ns/op / +5.9% (worse)
Windows MinGW BenchmarkMergeLinkerFlags 587.500 ns/op +40.1 ns/op / +7.3% (worse)
Windows MinGW BenchmarkChannelBuffered 32.150 ns/op +1.82 ns/op / +6.0% (worse)
Windows MinGW BenchmarkChannelHandoff 933.400 ns/op +67.1 ns/op / +7.7% (worse)
Windows MinGW BenchmarkDefer 55.830 ns/op -0.66 ns/op / -1.2% (better)
Windows MinGW BenchmarkDirectCall 1.538 ns/op -0.01 ns/op / -0.6% (better)
Windows MinGW BenchmarkGlobalRead 1.540 ns/op -0.009 ns/op / -0.6% (better)
Windows MinGW BenchmarkGlobalWrite 2.447 ns/op -0.021 ns/op / -0.9% (better)
Windows MinGW BenchmarkGoroutine 90004 ns/op -346 ns/op / -0.4% (better)
Windows MinGW BenchmarkInterfaceCall 8.370 ns/op +0.001 ns/op / +0.01195% (worse)
Windows MinGW BenchmarkRuntimeGetG 2.481 ns/op -0.005 ns/op / -0.2% (better)
Windows MinGW 386 BenchmarkLookupPCRandom 26.570 ns/op -0.03 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkMergeCompilerFlags 756.400 ns/op +41.1 ns/op / +5.7% (worse)
Windows MinGW 386 BenchmarkMergeLinkerFlags 718.900 ns/op +18.4 ns/op / +2.6% (worse)
Windows MinGW 386 BenchmarkChannelBuffered 41.510 ns/op -0.39 ns/op / -0.9% (better)
Windows MinGW 386 BenchmarkChannelHandoff 894.100 ns/op +38.9 ns/op / +4.5% (worse)
Windows MinGW 386 BenchmarkDefer 44.630 ns/op -2.11 ns/op / -4.5% (better)
Windows MinGW 386 BenchmarkDirectCall 1.551 ns/op -0.014 ns/op / -0.9% (better)
Windows MinGW 386 BenchmarkGlobalRead 1.553 ns/op 0 ns/op / +0.0%
Windows MinGW 386 BenchmarkGlobalWrite 7.770 ns/op -0.013 ns/op / -0.2% (better)
Windows MinGW 386 BenchmarkGoroutine 104700 ns/op -741 ns/op / -0.7% (better)
Windows MinGW 386 BenchmarkInterfaceCall 8.404 ns/op -0.034 ns/op / -0.4% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 1.930 ns/op +0.004 ns/op / +0.2% (worse)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.050 ns/op -0.08 ns/op / -0.7% (better)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 575.400 ns/op +4 ns/op / +0.7% (worse)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 535.600 ns/op +6.2 ns/op / +1.2% (worse)
Windows MinGW ARM64 BenchmarkChannelBuffered 37.520 ns/op -1.95 ns/op / -4.9% (better)
Windows MinGW ARM64 BenchmarkChannelHandoff 1598 ns/op -100 ns/op / -5.9% (better)
Windows MinGW ARM64 BenchmarkDefer 56.250 ns/op +0.71 ns/op / +1.3% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.663 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkGlobalRead 0.664 ns/op +0.0006 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.737 ns/op +0.0003 ns/op / +0.0407% (worse)
Windows MinGW ARM64 BenchmarkGoroutine 59186 ns/op +218 ns/op / +0.4% (worse)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.144 ns/op +0.004 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.769 ns/op -0.038 ns/op / -2.1% (better)
Windows MSVC BenchmarkLookupPCRandom 7.148 ns/op +0.199 ns/op / +2.9% (worse)
Windows MSVC BenchmarkMergeCompilerFlags 707.400 ns/op +31.6 ns/op / +4.7% (worse)
Windows MSVC BenchmarkMergeLinkerFlags 929.500 ns/op +286.3 ns/op / +44.5% (worse)
Windows MSVC BenchmarkChannelBuffered 28.130 ns/op +0.24 ns/op / +0.9% (worse)
Windows MSVC BenchmarkChannelHandoff 797.500 ns/op +98.8 ns/op / +14.1% (worse)
Windows MSVC BenchmarkDefer 30.030 ns/op +0.29 ns/op / +1.0% (worse)
Windows MSVC BenchmarkDirectCall 0.921 ns/op +0.0286 ns/op / +3.2% (worse)
Windows MSVC BenchmarkGlobalRead 0.914 ns/op +0.0193 ns/op / +2.2% (worse)
Windows MSVC BenchmarkGlobalWrite 4.605 ns/op +0.12 ns/op / +2.7% (worse)
Windows MSVC BenchmarkGoroutine 41797 ns/op +3398 ns/op / +8.8% (worse)
Windows MSVC BenchmarkInterfaceCall 4.929 ns/op +0.047 ns/op / +1.0% (worse)
Windows MSVC BenchmarkRuntimeGetG 1.024 ns/op +0.006 ns/op / +0.6% (worse)
Windows MSVC 386 BenchmarkLookupPCRandom 21.550 ns/op 0 ns/op / +0.0%
Windows MSVC 386 BenchmarkMergeCompilerFlags 553.900 ns/op +1.6 ns/op / +0.3% (worse)
Windows MSVC 386 BenchmarkMergeLinkerFlags 526.100 ns/op +8.5 ns/op / +1.6% (worse)
Windows MSVC 386 BenchmarkChannelBuffered 33.990 ns/op +0.08 ns/op / +0.2% (worse)
Windows MSVC 386 BenchmarkChannelHandoff 691.500 ns/op +15.1 ns/op / +2.2% (worse)
Windows MSVC 386 BenchmarkDefer 37.820 ns/op -1.14 ns/op / -2.9% (better)
Windows MSVC 386 BenchmarkDirectCall 1.356 ns/op -0.002 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGlobalRead 1.358 ns/op 0 ns/op / +0.0%
Windows MSVC 386 BenchmarkGlobalWrite 6.981 ns/op -0.004 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGoroutine 71826 ns/op -2407 ns/op / -3.2% (better)
Windows MSVC 386 BenchmarkInterfaceCall 7.340 ns/op +0.004 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkRuntimeGetG 1.630 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.060 ns/op -0.08 ns/op / -0.7% (better)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 640 ns/op +56 ns/op / +9.6% (worse)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 588.100 ns/op +48.5 ns/op / +9.0% (worse)
Windows MSVC ARM64 BenchmarkChannelBuffered 38.620 ns/op +0.07 ns/op / +0.2% (worse)
Windows MSVC ARM64 BenchmarkChannelHandoff 2312 ns/op +87 ns/op / +3.9% (worse)
Windows MSVC ARM64 BenchmarkDefer 61.230 ns/op -0.84 ns/op / -1.4% (better)
Windows MSVC ARM64 BenchmarkDirectCall 0.663 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 BenchmarkGlobalRead 0.663 ns/op -0.0003 ns/op / -0.04521% (better)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.796 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 BenchmarkGoroutine 58985 ns/op -1096 ns/op / -1.8% (better)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.139 ns/op -0.007 ns/op / -0.2% (better)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.809 ns/op +0.004 ns/op / +0.2% (worse)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 912.400 ns/op +12.5 ns/op / +1.4% (worse)
Linux AfterFuncZeroDelivery/LLGo 48539 ns/op +11743 ns/op / +31.9% (worse)
Linux CreateStop/Go 294.400 ns/op +5.9 ns/op / +2.0% (worse)
Linux CreateStop/LLGo 1822 ns/op -21 ns/op / -1.1% (better)
Linux RearmStopped/Go 117.900 ns/op +1.7 ns/op / +1.5% (worse)
Linux RearmStopped/LLGo 1385 ns/op +7 ns/op / +0.5% (worse)
Linux ResetActive/Go 69.440 ns/op +0.99 ns/op / +1.4% (worse)
Linux ResetActive/LLGo 867.800 ns/op +49.4 ns/op / +6.0% (worse)
Linux ResetHeap1024/Go 67.340 ns/op +0.34 ns/op / +0.5% (worse)
Linux ResetHeap1024/LLGo 179.900 ns/op +5.9 ns/op / +3.4% (worse)
macOS AfterFuncZeroDelivery/Go 384.700 ns/op -199.3 ns/op / -34.1% (better)
macOS AfterFuncZeroDelivery/LLGo 66132 ns/op -14923 ns/op / -18.4% (better)
macOS CreateStop/Go 115.200 ns/op -30.8 ns/op / -21.1% (better)
macOS CreateStop/LLGo 465.100 ns/op -16.7 ns/op / -3.5% (better)
macOS RearmStopped/Go 53.460 ns/op -6.97 ns/op / -11.5% (better)
macOS RearmStopped/LLGo 321.200 ns/op -264.3 ns/op / -45.1% (better)
macOS ResetActive/Go 37.650 ns/op -10.76 ns/op / -22.2% (better)
macOS ResetActive/LLGo 152.900 ns/op -71.9 ns/op / -32.0% (better)
macOS ResetHeap1024/Go 37.760 ns/op -10.42 ns/op / -21.6% (better)
macOS ResetHeap1024/LLGo 74.450 ns/op -12.37 ns/op / -14.2% (better)
Windows MinGW AfterFuncZeroDelivery/Go 554 ns/op -0.5 ns/op / -0.1% (better)
Windows MinGW AfterFuncZeroDelivery/LLGo 177506 ns/op +843 ns/op / +0.5% (worse)
Windows MinGW CreateStop/Go 116.500 ns/op +1.7 ns/op / +1.5% (worse)
Windows MinGW CreateStop/LLGo 418.700 ns/op -8.4 ns/op / -2.0% (better)
Windows MinGW RearmStopped/Go 31.290 ns/op +0.08 ns/op / +0.3% (worse)
Windows MinGW RearmStopped/LLGo 385.700 ns/op +113.3 ns/op / +41.6% (worse)
Windows MinGW ResetActive/Go 20.050 ns/op +0.07 ns/op / +0.4% (worse)
Windows MinGW ResetActive/LLGo 166.100 ns/op -3.3 ns/op / -1.9% (better)
Windows MinGW ResetHeap1024/Go 20.310 ns/op -0.04 ns/op / -0.2% (better)
Windows MinGW ResetHeap1024/LLGo 124.200 ns/op -5.1 ns/op / -3.9% (better)
Windows MinGW 386 AfterFuncZeroDelivery/Go 950.100 ns/op +4 ns/op / +0.4% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 197133 ns/op +1083 ns/op / +0.6% (worse)
Windows MinGW 386 CreateStop/Go 193.200 ns/op +2.1 ns/op / +1.1% (worse)
Windows MinGW 386 CreateStop/LLGo 532.300 ns/op +8.8 ns/op / +1.7% (worse)
Windows MinGW 386 RearmStopped/Go 63.660 ns/op +0.13 ns/op / +0.2% (worse)
Windows MinGW 386 RearmStopped/LLGo 370.200 ns/op +3.6 ns/op / +1.0% (worse)
Windows MinGW 386 ResetActive/Go 39.370 ns/op +0.1 ns/op / +0.3% (worse)
Windows MinGW 386 ResetActive/LLGo 941.100 ns/op +49.1 ns/op / +5.5% (worse)
Windows MinGW 386 ResetHeap1024/Go 39.760 ns/op +0.28 ns/op / +0.7% (worse)
Windows MinGW 386 ResetHeap1024/LLGo 192.300 ns/op +3.6 ns/op / +1.9% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 666.500 ns/op -0.9 ns/op / -0.1% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 151916 ns/op +6612 ns/op / +4.6% (worse)
Windows MinGW ARM64 CreateStop/Go 195.100 ns/op -1.8 ns/op / -0.9% (better)
Windows MinGW ARM64 CreateStop/LLGo 424.500 ns/op -1.9 ns/op / -0.4% (better)
Windows MinGW ARM64 RearmStopped/Go 70.610 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 RearmStopped/LLGo 257.800 ns/op -0.5 ns/op / -0.2% (better)
Windows MinGW ARM64 ResetActive/Go 31.130 ns/op +0.06 ns/op / +0.2% (worse)
Windows MinGW ARM64 ResetActive/LLGo 140.100 ns/op +8.4 ns/op / +6.4% (worse)
Windows MinGW ARM64 ResetHeap1024/Go 31.240 ns/op +0.12 ns/op / +0.4% (worse)
Windows MinGW ARM64 ResetHeap1024/LLGo 125.400 ns/op +0.7 ns/op / +0.6% (worse)
Windows MSVC AfterFuncZeroDelivery/Go 407.500 ns/op +3.7 ns/op / +0.9% (worse)
Windows MSVC AfterFuncZeroDelivery/LLGo 84406 ns/op +1728 ns/op / +2.1% (worse)
Windows MSVC CreateStop/Go 117.200 ns/op +0.4 ns/op / +0.3% (worse)
Windows MSVC CreateStop/LLGo 311.900 ns/op -5.1 ns/op / -1.6% (better)
Windows MSVC RearmStopped/Go 46.050 ns/op +0.12 ns/op / +0.3% (worse)
Windows MSVC RearmStopped/LLGo 191.600 ns/op +12.5 ns/op / +7.0% (worse)
Windows MSVC ResetActive/Go 20.400 ns/op +0.18 ns/op / +0.9% (worse)
Windows MSVC ResetActive/LLGo 441.700 ns/op -77 ns/op / -14.8% (better)
Windows MSVC ResetHeap1024/Go 20.110 ns/op 0 ns/op / +0.0%
Windows MSVC ResetHeap1024/LLGo 90.830 ns/op +1.14 ns/op / +1.3% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/Go 767.500 ns/op +6.2 ns/op / +0.8% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 134767 ns/op +3334 ns/op / +2.5% (worse)
Windows MSVC 386 CreateStop/Go 167.200 ns/op +0.1 ns/op / +0.1% (worse)
Windows MSVC 386 CreateStop/LLGo 363.300 ns/op -15.3 ns/op / -4.0% (better)
Windows MSVC 386 RearmStopped/Go 57 ns/op +0.42 ns/op / +0.7% (worse)
Windows MSVC 386 RearmStopped/LLGo 262.100 ns/op +0.8 ns/op / +0.3% (worse)
Windows MSVC 386 ResetActive/Go 32.570 ns/op +0.01 ns/op / +0.03071% (worse)
Windows MSVC 386 ResetActive/LLGo 857.600 ns/op +24.1 ns/op / +2.9% (worse)
Windows MSVC 386 ResetHeap1024/Go 32.850 ns/op +0.02 ns/op / +0.1% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 141.900 ns/op -2 ns/op / -1.4% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 664.400 ns/op -21 ns/op / -3.1% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 173030 ns/op -2265 ns/op / -1.3% (better)
Windows MSVC ARM64 CreateStop/Go 205.600 ns/op +5.8 ns/op / +2.9% (worse)
Windows MSVC ARM64 CreateStop/LLGo 499.100 ns/op -97.9 ns/op / -16.4% (better)
Windows MSVC ARM64 RearmStopped/Go 70.550 ns/op -0.01 ns/op / -0.01417% (better)
Windows MSVC ARM64 RearmStopped/LLGo 282.700 ns/op -8.4 ns/op / -2.9% (better)
Windows MSVC ARM64 ResetActive/Go 31.050 ns/op -0.02 ns/op / -0.1% (better)
Windows MSVC ARM64 ResetActive/LLGo 162.900 ns/op +9.4 ns/op / +6.1% (worse)
Windows MSVC ARM64 ResetHeap1024/Go 31.160 ns/op +0.01 ns/op / +0.0321% (worse)
Windows MSVC ARM64 ResetHeap1024/LLGo 140.200 ns/op +2.5 ns/op / +1.8% (worse)

Compared with 5f1f13897af7 measured in the same runner job.

@cpunion
cpunion force-pushed the codex/fix-size-report-2679-20261003 branch from adad928 to 5c7c7aa Compare October 3, 2026 05:29
@cpunion cpunion changed the title build: measure final Wasm and ELF sizes accurately build: report final Wasm, ELF and PE sizes Oct 4, 2026
@cpunion

cpunion commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator Author

@fennoai review

The size parser, attribution and output now live in the independently tested internal/sizereport package, including final PE artifacts. The failing Wasm SIMD CI path now executes the full suite with Emscripten native Wasm EH; WASI compilation and the Go/Wasmtime oracle remain checked. EH capabilities and cache isolation have regression tests, and the full SIMD suite passes with both cold and cached packages. Report failures still propagate; no CI failure tolerance was added.

@fennoai

fennoai Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Review complete

Reviewed the latest revision (783ec4f...dc9a06c2) and submitted a PR review: #pullrequestreview-5404395670.

All three findings from the prior review are resolved — resolver negative-cache, unbounded io.ReadAll, and the double CollectArtifacts snapshot.

The new work (the internal/sizereport package extraction, PE/COFF reader, Emscripten native Wasm EH feature injection + EMCC_CFLAGS cache keying, and the SIMD CI changes) is correct and well-tested. Security and documentation passes found no issues. Posted only two optional micro-nits on the PE reader; nothing blocking.

View job run

@zhouguangyuan0718 zhouguangyuan0718 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the final-artifact size accounting and the latest Emscripten EH/cache changes. Focused size-report and build-integration tests passed, including the added EH/cache tests. Checked real Wasm, ELF, and PE artifacts against their file/section sizes. No blocking findings in the reviewed changes. The latest CI run is still in progress.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Follow-up review — sizereport extraction, PE support, Wasm EH

Reviewed the latest revision (783ec4f...dc9a06c2). The three findings from the previous review are all resolved:

  • Resolver O(symbols x packages) — fixed: matchPackage/matchModule now cache misses (map[...] = "" + comma-ok lookup), so each distinct symbol scans packages at most once.
  • Unbounded io.ReadAll — fixed: Collect rejects non-regular files and readWasmSizeBytes presizes the buffer to the stat size with a single io.ReadFull, plus a grown/shrunken-file guard.
  • Double CollectArtifacts — fixed: reportBuildOutputs collects one artifact snapshot and shares it between the size report and the artifact listing (covered by TestReportBuildOutputsSharesArtifacts).

The new work is solid and well-tested:

  • The internal/sizereport extraction is clean — stdlib-only, exported Package/Artifact/Report, tool invocation (llvm-readelf) kept in the build layer behind ErrUnsupportedFormat.
  • The new PE/COFF reader correctly excludes raw-file alignment padding, models virtual zero-fill as BSS via a shadow-section index, bounds-checks SectionNumber, and uses uint32-widened arithmetic that cannot overflow. Malformed PE/ELF/Wasm inputs all fail closed (no panics, no unbounded allocation).
  • applyEmscriptenEHFeature correctly replaces (not duplicates) the target-features attribute, skips declarations, is idempotent, and only materializes cost on Emscripten J32/J64 builds. EMCC_CFLAGS cache-key isolation is scoped to js/wasm and test-verified.
  • Documentation (doc/size-report.md, test/simd/README.md) matches the implemented behavior, including the JSON schema, the ram = data + zero-fill definition, and all three validation commands.

Security and documentation passes found nothing. Only two optional micro-nits below; nothing blocking.

Comment thread internal/sizereport/size_report_pe.go Outdated
Comment thread internal/sizereport/size_report_pe.go Outdated
@cpunion
cpunion force-pushed the codex/fix-size-report-2679-20261003 branch from dc9a06c to 34acf06 Compare October 4, 2026 06:47
@cpunion
cpunion merged commit d92347c into xgo-dev:main Oct 4, 2026
92 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants