Skip to content

wasm: honor WASI LTO and preserve threaded exception lowering - #2726

Merged
cpunion merged 3 commits into
mainfrom
codex/dev-wasm-full-lto-fix-20261003
Oct 4, 2026
Merged

cpunion merged 3 commits into
mainfrom
codex/dev-wasm-full-lto-fix-20261003

Conversation

@zhouguangyuan0718

@zhouguangyuan0718 zhouguangyuan0718 commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

WASI accepts -lto=thin/full without passing -flto to Clang. Ordinary builds silently produce native Wasm objects, while dev Full LTO emits llvm.type.checked.load for Go GlobalDCE and crashes during premature object code generation (Do not know how to promote this operator!).

Pass the selected LTO mode through WASI compilation and linking. Forward the requested optimization level to wasm-ld with --lto-O0 through --lto-O3; Os/Oz use --lto-O2 with size attributes retained in bitcode. Without this forwarding, wasm-ld uses its default O2 optimizer even when the user selects O0 or O3. No LTO optimizer option is added when LTO is disabled.

Also pass the Wasm exception model and atomics/bulk-memory/exception-handling defaults to the link-time backend: otherwise Full LTO drops SjLj catch pads, and independent ThinLTO modules lose shared-memory/TLS support or fail to select EH instructions. Keep standard Wasm EH and Go GlobalDCE method-retention metadata enabled.

Add off/thin/full flag coverage across O0/O1/O2/O3/Os/Oz and a CI regression using a dev compiler and the shared Wasmer runner. The regression verifies actual LTO backend objects, merged Full LTO bitcode, and dev type-checked-load metadata. Under both Thin and Full LTO, it runs interface dispatch, SIMD tests, and GC/nogc Goexit, defer, panic/recover and C longjmp across pthreads, including expected init/main Goexit and uncaught panic failures. Wasmer chooses the backend automatically, its module cache stays enabled, and engine tracing is suppressed for test-output checks. The multiline CI step explicitly enables set -euo pipefail.

Scope and related work:

  • Based on main 3a8054a87, including the merged Wasmer migration (wasm: switch threaded WASI runner to Wasmer #2725) and browser coverage OIDC fix (fix(ci): allow OIDC for browser coverage uploads #2728).
  • -mattr supplies backend defaults; it does not merge required features into an existing function target-features attribute. An LLVM 22 TLS/atomic probe with only +simd128 on its definition fails shared-memory linking in both Thin and Full LTO despite these defaults; merging atomics/bulk-memory/exception-handling into that attribute fixes the probe. This is a confirmed backend boundary, not a reproduced failure of the ordinary LLGo matrix below.
  • build: fix C-only DCE and WASI LTO function features #2721 still contains separate C-only deadcode linking, missing-metadata diagnostics, function-attribute handling, and threaded-GC validation. After this PR lands, its duplicate flag changes can be removed; attribute handling should retain existing SIMD features and fill required features on definitions that already have an attribute. This PR does not replace that work.

Validation on macOS arm64 with Go 1.27.0, LLVM/LLD 22.1.8 and Wasmer 7.5.0:

  • Focused WASI flags, LTO optimization mapping and dev GlobalDCE tests passed, as did the full internal/crosscompile short suite.
  • Actual dev-compiler builds at O0/O3/Os/Oz passed under both Thin and Full LTO. Clang's verbose output confirmed the expected --lto-O* in the wasm-ld command; each build produced a nonempty LTO backend object and ran under Wasmer. The previous compiler's O0 ThinLTO invocation lacked --lto-O0.
  • dev/test_wasm_wasi_lto.py passed the complete Thin/Full LTO matrix, including SIMD, verified LTO artifacts, dev metadata, and all GC/nogc exception cases.
  • The existing-attribute LLVM probe failed without feature merging and linked successfully with merging in both modes.
  • Workflow YAML parsing and git diff --check passed.

The dev Full LTO promotion crash was also reproduced on unchanged main 783ec4fd3. The earlier ordinary Thin/Full LTO results reported in #2725 demonstrated flag acceptance only; the checks above exercise actual LTO. CI for the updated head is pending.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review summary

The change is well-targeted. The LTO-mode gating in crosscompile.go is correct, and TestUseWASILTOEnablesSjLjAtLink now cleanly asserts the new link/clang flags appear exactly when lto.Mode.Enabled() is true across Off/Thin/Full. The new Python harness builds both thin and full modes and validates the LTO backend object and merged bitcode, giving strong coverage of the fix. Security and performance passes found nothing noteworthy (build-time flags only, no runtime hot path, list-form subprocess calls with check=True and timeouts).

One comment-accuracy finding is inlined below. A minor CI nit (the new Test dev WASI Thin and Full LTO step omits the set -euo pipefail guard used by every other multiline bash step in this workflow, so a go build failure surfaces as a confusing "file not found" from the next line rather than the real build error) is optional to address.

Comment thread internal/crosscompile/crosscompile.go Outdated
// Without an explicit exception model, codegen drops SjLj catch pads.
export.LDFLAGS = append(export.LDFLAGS, "-Wl,--mllvm=-exception-model=wasm")
// ThinLTO compiles Go modules independently of the C modules that
// carry these features. Preserve shared memory, TLS and Wasm EH.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] Comment scopes flags to ThinLTO and misnames -mattr features

The comment says "ThinLTO compiles Go modules independently ... Preserve shared memory, TLS and Wasm EH", but the enclosing guard is ltoMode.Enabled(), which is true for both lto.Thin and lto.Full (internal/lto/lto.go), and the harness added in this PR exercises both modes — so these flags matter under full LTO too, not just thin. The feature list is also inaccurate: the actual -mattr value is +atomics,+bulk-memory,+exception-handling; there is no +tls attribute (TLS on wasm is emergent from atomics+bulk-memory). Consider rewording to cover both LTO modes and to list the real attributes (atomics / bulk-memory / exception-handling) so a future reader does not assume the flags are thin-only or inert under -lto=full.

@codecov

codecov Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@github-actions

github-actions Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

db70b67550c0 | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 154193 B 0 B / +0.0% 88742 B 0 B / +0.0%
cprintf/j32-goos-js/LLGo 152295 B 0 B / +0.0% 73165 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 141248 B 0 B / +0.0% 92630 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 153508 B 0 B / +0.0% 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 154198 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3213612 B 0 B / +0.0% 132442 B 0 B / +0.0%
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3192706 B 0 B / +0.0% 101635 B 0 B / +0.0%
fmtprintf/j64-emscripten-memory64/LLGo 2950218 B 0 B / +0.0% 139285 B 0 B / +0.0%
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2349336 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2345976 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-emscripten/LLGo 153428 B 0 B / +0.0% 88742 B 0 B / +0.0%
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 151768 B 0 B / +0.0% 73165 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 140582 B 0 B / +0.0% 92630 B 0 B / +0.0%
reflectcall/j32-emscripten/LLGo 1543965 B 0 B / +0.0% 105908 B 0 B / +0.0%
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1546420 B 0 B / +0.0% 90331 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1428634 B 0 B / +0.0% 111641 B 0 B / +0.0%
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1277776 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1275362 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 153155 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-wasi/LLGo 153845 B 0 B / +0.0% 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 6.107 s -224.3 ms / -3.5% (better)
j32-goos-js 5.982 s -198.1 ms / -3.2% (better)
j64-emscripten-memory64 5.319 s -28.01 ms / -0.5% (better)
reflectcall/w32-wasi 23.065 s -473.8 ms / -2.0% (better)
w32-goos-wasip1 4.356 s -133.6 ms / -3.0% (better)
w32-wasi 4.025 s -66.05 ms / -1.6% (better)

Compared with 3a8054a87f76 measured in the same runner job.

@github-actions

github-actions Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

db70b67550c0 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 8568 B 0 B / +0.0% 387 B 0 B / +0.0% 926.695 ms -366.5 us / -0.03953% (better) 1.233 ms +3.999 us / +0.3% (worse)
Linux cprintf-lto 8408 B 0 B / +0.0% 368 B 0 B / +0.0% 952.042 ms +15.36 ms / +1.6% (worse) 1.267 ms +18.74 us / +1.5% (worse)
Linux fmtprintf 4880952 B +8 B / +0.0001639% (worse) 501619 B 0 B / +0.0% 7.317 s +288.4 ms / +4.1% (worse) 3.101 ms +8.161 us / +0.3% (worse)
Linux fmtprintf-lto 3600808 B 0 B / +0.0% 438538 B 0 B / +0.0% 15.865 s -147.8 ms / -0.9% (better) 3.044 ms +145.4 us / +5.0% (worse)
Linux println 674952 B 0 B / +0.0% 16855 B 0 B / +0.0% 956.210 ms +20.93 ms / +2.2% (worse) 1.645 ms +57.28 us / +3.6% (worse)
Linux println-lto 190024 B 0 B / +0.0% 14273 B 0 B / +0.0% 1.243 s +1.388 ms / +0.1% (worse) 1.603 ms +30.16 us / +1.9% (worse)
macOS cprintf 50736 B 0 B / +0.0% 4409 B 0 B / +0.0% 1.109 s -274.6 ms / -19.8% (better) 4.646 ms +396.8 us / +9.3% (worse)
macOS cprintf-lto 50496 B 0 B / +0.0% 161 B 0 B / +0.0% 824.193 ms -486.2 ms / -37.1% (better) 2.602 ms -2.054 ms / -44.1% (better)
macOS fmtprintf 1773456 B 0 B / +0.0% 877780 B 0 B / +0.0% 6.805 s +306.1 ms / +4.7% (worse) 11.020 ms +1.733 ms / +18.7% (worse)
macOS fmtprintf-lto 1360928 B 0 B / +0.0% 762428 B 0 B / +0.0% 13.429 s -1.615 s / -10.7% (better) 4.462 ms -644.3 us / -12.6% (better)
macOS println 99344 B 0 B / +0.0% 24216 B 0 B / +0.0% 855.610 ms -475.2 ms / -35.7% (better) 4.189 ms -661.3 us / -13.6% (better)
macOS println-lto 83664 B 0 B / +0.0% 21457 B 0 B / +0.0% 1.425 s -181.2 ms / -11.3% (better) 5.530 ms -2.619 ms / -32.1% (better)
Windows MinGW cprintf 651264 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.764 s +23.05 ms / +1.3% (worse) 3.615 ms +90.9 us / +2.6% (worse)
Windows MinGW cprintf-lto 43520 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.866 s +87.45 ms / +4.9% (worse) 4.222 ms +231.7 us / +5.8% (worse)
Windows MinGW fmtprintf 5432832 B 0 B / +0.0% 600310 B 0 B / +0.0% 7.623 s -64.63 ms / -0.8% (better) 8.916 ms -622.4 us / -6.5% (better)
Windows MinGW fmtprintf-lto 4127232 B 0 B / +0.0% 546582 B 0 B / +0.0% 15.259 s -374.2 ms / -2.4% (better) 8.694 ms +209.8 us / +2.5% (worse)
Windows MinGW println 706560 B 0 B / +0.0% 25190 B 0 B / +0.0% 1.780 s +18.79 ms / +1.1% (worse) 7.415 ms -210.9 us / -2.8% (better)
Windows MinGW println-lto 208896 B 0 B / +0.0% 22054 B 0 B / +0.0% 2.101 s +76.24 ms / +3.8% (worse) 7.248 ms -393.4 us / -5.1% (better)
Windows MinGW 386 cprintf 601600 B 0 B / +0.0% 5326 B 0 B / +0.0% 1.629 s +12.94 ms / +0.8% (worse) 5.042 ms +31.2 us / +0.6% (worse)
Windows MinGW 386 cprintf-lto 103424 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.674 s +15.76 ms / +1.0% (worse) 5.139 ms +121.8 us / +2.4% (worse)
Windows MinGW 386 fmtprintf 4745216 B 0 B / +0.0% 472478 B 0 B / +0.0% 7.380 s +137.3 ms / +1.9% (worse) 10.648 ms +481.4 us / +4.7% (worse)
Windows MinGW 386 fmtprintf-lto 4148736 B 0 B / +0.0% 451114 B 0 B / +0.0% 14.300 s +109.9 ms / +0.8% (worse) 9.859 ms -481.6 us / -4.7% (better)
Windows MinGW 386 println 653312 B 0 B / +0.0% 21490 B 0 B / +0.0% 1.630 s +3.307 ms / +0.2% (worse) 8.300 ms -3.022 ms / -26.7% (better)
Windows MinGW 386 println-lto 258560 B 0 B / +0.0% 19306 B 0 B / +0.0% 1.916 s +23.51 ms / +1.2% (worse) 8.330 ms +43.7 us / +0.5% (worse)
Windows MinGW ARM64 cprintf 661504 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.898 s +18.37 ms / +1.0% (worse) 6.441 ms +15.9 us / +0.2% (worse)
Windows MinGW ARM64 cprintf-lto 43520 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.921 s -7.098 ms / -0.4% (better) 6.431 ms -17.3 us / -0.3% (better)
Windows MinGW ARM64 fmtprintf 5343744 B 0 B / +0.0% 510876 B 0 B / +0.0% 7.035 s +97.29 ms / +1.4% (worse) 12.797 ms +131.6 us / +1.0% (worse)
Windows MinGW ARM64 fmtprintf-lto 4302336 B 0 B / +0.0% 477264 B 0 B / +0.0% 13.582 s -327 us / -0.002408% (better) 12.142 ms +139.2 us / +1.2% (worse)
Windows MinGW ARM64 println 714240 B 0 B / +0.0% 23884 B 0 B / +0.0% 1.919 s +38.34 ms / +2.0% (worse) 10.945 ms -157.7 us / -1.4% (better)
Windows MinGW ARM64 println-lto 215552 B 0 B / +0.0% 21232 B 0 B / +0.0% 2.160 s +20.28 ms / +0.9% (worse) 10.670 ms -164.3 us / -1.5% (better)
Windows MSVC cprintf 893952 B 0 B / +0.0% 65798 B 0 B / +0.0% 1.577 s -188.8 ms / -10.7% (better) 3.751 ms +430.3 us / +13.0% (worse)
Windows MSVC cprintf-lto 289792 B 0 B / +0.0% 65734 B 0 B / +0.0% 1.602 s +24.15 ms / +1.5% (worse) 3.410 ms +47.8 us / +1.4% (worse)
Windows MSVC fmtprintf 5732864 B 0 B / +0.0% 695862 B 0 B / +0.0% 7.023 s +56.7 ms / +0.8% (worse) 8.869 ms +92.6 us / +1.1% (worse)
Windows MSVC fmtprintf-lto 4449792 B 0 B / +0.0% 646134 B 0 B / +0.0% 13.643 s +17.32 ms / +0.1% (worse) 9.129 ms +702.3 us / +8.3% (worse)
Windows MSVC println 1017344 B 0 B / +0.0% 120854 B 0 B / +0.0% 1.538 s -34.97 ms / -2.2% (better) 7.047 ms +259.2 us / +3.8% (worse)
Windows MSVC println-lto 528384 B 0 B / +0.0% 118390 B 0 B / +0.0% 1.793 s +53.1 us / +0.002961% (worse) 6.917 ms -191.2 us / -2.7% (better)
Windows MSVC 386 cprintf 513536 B 0 B / +0.0% 3931 B 0 B / +0.0% 1.619 s -17.64 ms / -1.1% (better) 5.680 ms -105.2 us / -1.8% (better)
Windows MSVC 386 cprintf-lto 44032 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.543 s -100.8 ms / -6.1% (better) 5.434 ms -371.1 us / -6.4% (better)
Windows MSVC 386 fmtprintf 4479488 B 0 B / +0.0% 455868 B 0 B / +0.0% 7.150 s -76.22 ms / -1.1% (better) 11.403 ms -877.6 us / -7.1% (better)
Windows MSVC 386 fmtprintf-lto 3899392 B 0 B / +0.0% 426651 B 0 B / +0.0% 13.563 s -58.81 ms / -0.4% (better) 11.037 ms -865.4 us / -7.3% (better)
Windows MSVC 386 println 567296 B 0 B / +0.0% 20340 B 0 B / +0.0% 1.527 s -73.06 ms / -4.6% (better) 9.729 ms -989 us / -9.2% (better)
Windows MSVC 386 println-lto 199168 B 0 B / +0.0% 18501 B 0 B / +0.0% 1.802 s -65.33 ms / -3.5% (better) 9.513 ms -788.2 us / -7.7% (better)
Windows MSVC ARM64 cprintf 662528 B 0 B / +0.0% 4192 B 0 B / +0.0% 1.668 s +22.37 ms / +1.4% (worse) 7.515 ms +308.5 us / +4.3% (worse)
Windows MSVC ARM64 cprintf-lto 47616 B 0 B / +0.0% 4084 B 0 B / +0.0% 1.701 s +41.33 ms / +2.5% (worse) 7.330 ms +395.1 us / +5.7% (worse)
Windows MSVC ARM64 fmtprintf 5339648 B 0 B / +0.0% 510808 B 0 B / +0.0% 6.701 s -251.5 ms / -3.6% (better) 14.271 ms -768.2 us / -5.1% (better)
Windows MSVC ARM64 fmtprintf-lto 4309504 B 0 B / +0.0% 477924 B 0 B / +0.0% 13.020 s -26.9 ms / -0.2% (better) 14.158 ms -1.497 ms / -9.6% (better)
Windows MSVC ARM64 println 715264 B 0 B / +0.0% 23908 B 0 B / +0.0% 1.641 s +33.15 ms / +2.1% (worse) 12.117 ms -1.174 ms / -8.8% (better)
Windows MSVC ARM64 println-lto 220672 B 0 B / +0.0% 21380 B 0 B / +0.0% 1.862 s -13.61 ms / -0.7% (better) 12.256 ms -167.4 us / -1.3% (better)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.490 ns/op -0.04 ns/op / -0.3% (better)
Linux BenchmarkMergeCompilerFlags 201.300 ns/op +7.8 ns/op / +4.0% (worse)
Linux BenchmarkMergeLinkerFlags 134 ns/op +6.3 ns/op / +4.9% (worse)
Linux BenchmarkChannelBuffered 55.050 ns/op -0.06 ns/op / -0.1% (better)
Linux BenchmarkChannelHandoff 12679 ns/op -2900 ns/op / -18.6% (better)
Linux BenchmarkDefer 46.670 ns/op -0.58 ns/op / -1.2% (better)
Linux BenchmarkDirectCall 1.168 ns/op -0.004 ns/op / -0.3% (better)
Linux BenchmarkGlobalRead 1.166 ns/op +0.001 ns/op / +0.1% (worse)
Linux BenchmarkGlobalWrite 7.757 ns/op -0.006 ns/op / -0.1% (better)
Linux BenchmarkGoroutine 24225 ns/op -1466 ns/op / -5.7% (better)
Linux BenchmarkInterfaceCall 5.850 ns/op +0.026 ns/op / +0.4% (worse)
Linux BenchmarkRuntimeGetG 3.010 ns/op +0.029 ns/op / +1.0% (worse)
macOS BenchmarkLookupPCRandom 17.220 ns/op -1.52 ns/op / -8.1% (better)
macOS BenchmarkMergeCompilerFlags 168.200 ns/op +4.8 ns/op / +2.9% (worse)
macOS BenchmarkMergeLinkerFlags 99.210 ns/op -6.79 ns/op / -6.4% (better)
macOS BenchmarkChannelBuffered 34.360 ns/op -10.2 ns/op / -22.9% (better)
macOS BenchmarkChannelHandoff 12710 ns/op -789 ns/op / -5.8% (better)
macOS BenchmarkDefer 47.320 ns/op -9.46 ns/op / -16.7% (better)
macOS BenchmarkDirectCall 1.246 ns/op -0.178 ns/op / -12.5% (better)
macOS BenchmarkGlobalRead 1.185 ns/op -0.133 ns/op / -10.1% (better)
macOS BenchmarkGlobalWrite 1.405 ns/op -0.315 ns/op / -18.3% (better)
macOS BenchmarkGoroutine 64815 ns/op +16965 ns/op / +35.5% (worse)
macOS BenchmarkInterfaceCall 5.245 ns/op -1.028 ns/op / -16.4% (better)
macOS BenchmarkRuntimeGetG 2.878 ns/op -0.236 ns/op / -7.6% (better)
Windows MinGW BenchmarkLookupPCRandom 13.200 ns/op +0.11 ns/op / +0.8% (worse)
Windows MinGW BenchmarkMergeCompilerFlags 665.300 ns/op +61.4 ns/op / +10.2% (worse)
Windows MinGW BenchmarkMergeLinkerFlags 594.300 ns/op +44.7 ns/op / +8.1% (worse)
Windows MinGW BenchmarkChannelBuffered 30.380 ns/op -0.83 ns/op / -2.7% (better)
Windows MinGW BenchmarkChannelHandoff 879.200 ns/op -74.3 ns/op / -7.8% (better)
Windows MinGW BenchmarkDefer 59.280 ns/op +1.7 ns/op / +3.0% (worse)
Windows MinGW BenchmarkDirectCall 1.548 ns/op -0.003 ns/op / -0.2% (better)
Windows MinGW BenchmarkGlobalRead 1.549 ns/op -0.001 ns/op / -0.1% (better)
Windows MinGW BenchmarkGlobalWrite 2.471 ns/op -0.001 ns/op / -0.04045% (better)
Windows MinGW BenchmarkGoroutine 89369 ns/op +708 ns/op / +0.8% (worse)
Windows MinGW BenchmarkInterfaceCall 8.371 ns/op +0.005 ns/op / +0.1% (worse)
Windows MinGW BenchmarkRuntimeGetG 2.477 ns/op -0.004 ns/op / -0.2% (better)
Windows MinGW 386 BenchmarkLookupPCRandom 26.950 ns/op +0.35 ns/op / +1.3% (worse)
Windows MinGW 386 BenchmarkMergeCompilerFlags 718.600 ns/op +4.8 ns/op / +0.7% (worse)
Windows MinGW 386 BenchmarkMergeLinkerFlags 670.500 ns/op +6.7 ns/op / +1.0% (worse)
Windows MinGW 386 BenchmarkChannelBuffered 41.160 ns/op -0.11 ns/op / -0.3% (better)
Windows MinGW 386 BenchmarkChannelHandoff 822 ns/op +21.3 ns/op / +2.7% (worse)
Windows MinGW 386 BenchmarkDefer 41.290 ns/op -1.09 ns/op / -2.6% (better)
Windows MinGW 386 BenchmarkDirectCall 1.551 ns/op +0.005 ns/op / +0.3% (worse)
Windows MinGW 386 BenchmarkGlobalRead 1.554 ns/op +0.003 ns/op / +0.2% (worse)
Windows MinGW 386 BenchmarkGlobalWrite 7.779 ns/op +0.005 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkGoroutine 106738 ns/op +1218 ns/op / +1.2% (worse)
Windows MinGW 386 BenchmarkInterfaceCall 8.369 ns/op -0.031 ns/op / -0.4% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 1.930 ns/op +0.004 ns/op / +0.2% (worse)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.100 ns/op +0.03 ns/op / +0.2% (worse)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 573.700 ns/op +4.4 ns/op / +0.8% (worse)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 536.100 ns/op -1.2 ns/op / -0.2% (better)
Windows MinGW ARM64 BenchmarkChannelBuffered 37.570 ns/op -1.34 ns/op / -3.4% (better)
Windows MinGW ARM64 BenchmarkChannelHandoff 2294 ns/op -368 ns/op / -13.8% (better)
Windows MinGW ARM64 BenchmarkDefer 55.040 ns/op -0.02 ns/op / -0.03632% (better)
Windows MinGW ARM64 BenchmarkDirectCall 0.664 ns/op +0.0006 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkGlobalRead 0.663 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkGlobalWrite 0.737 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkGoroutine 61013 ns/op -2662 ns/op / -4.2% (better)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.143 ns/op +0.003 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.803 ns/op -0.006 ns/op / -0.3% (better)
Windows MSVC BenchmarkLookupPCRandom 12.780 ns/op -0.35 ns/op / -2.7% (better)
Windows MSVC BenchmarkMergeCompilerFlags 622.300 ns/op -1.7 ns/op / -0.3% (better)
Windows MSVC BenchmarkMergeLinkerFlags 549.400 ns/op -1.2 ns/op / -0.2% (better)
Windows MSVC BenchmarkChannelBuffered 29.760 ns/op +0.16 ns/op / +0.5% (worse)
Windows MSVC BenchmarkChannelHandoff 1076 ns/op +37 ns/op / +3.6% (worse)
Windows MSVC BenchmarkDefer 53.990 ns/op -0.52 ns/op / -1.0% (better)
Windows MSVC BenchmarkDirectCall 1.547 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC BenchmarkGlobalRead 1.548 ns/op -0.002 ns/op / -0.1% (better)
Windows MSVC BenchmarkGlobalWrite 2.471 ns/op -0.003 ns/op / -0.1% (better)
Windows MSVC BenchmarkGoroutine 88471 ns/op -1687 ns/op / -1.9% (better)
Windows MSVC BenchmarkInterfaceCall 8.993 ns/op -0.022 ns/op / -0.2% (better)
Windows MSVC BenchmarkRuntimeGetG 2.486 ns/op +0.008 ns/op / +0.3% (worse)
Windows MSVC 386 BenchmarkLookupPCRandom 26.540 ns/op 0 ns/op / +0.0%
Windows MSVC 386 BenchmarkMergeCompilerFlags 768.800 ns/op +20.2 ns/op / +2.7% (worse)
Windows MSVC 386 BenchmarkMergeLinkerFlags 667.900 ns/op +17.6 ns/op / +2.7% (worse)
Windows MSVC 386 BenchmarkChannelBuffered 38.780 ns/op +0.1 ns/op / +0.3% (worse)
Windows MSVC 386 BenchmarkChannelHandoff 851.600 ns/op -41.3 ns/op / -4.6% (better)
Windows MSVC 386 BenchmarkDefer 45.580 ns/op -1.98 ns/op / -4.2% (better)
Windows MSVC 386 BenchmarkDirectCall 1.548 ns/op +0.001 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkGlobalRead 1.553 ns/op +0.004 ns/op / +0.3% (worse)
Windows MSVC 386 BenchmarkGlobalWrite 7.782 ns/op +0.006 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkGoroutine 109809 ns/op -4025 ns/op / -3.5% (better)
Windows MSVC 386 BenchmarkInterfaceCall 8.393 ns/op -0.11 ns/op / -1.3% (better)
Windows MSVC 386 BenchmarkRuntimeGetG 2.166 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.130 ns/op +0.04 ns/op / +0.3% (worse)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 573.900 ns/op +2.3 ns/op / +0.4% (worse)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 539.800 ns/op +14.6 ns/op / +2.8% (worse)
Windows MSVC ARM64 BenchmarkChannelBuffered 38.580 ns/op +0.98 ns/op / +2.6% (worse)
Windows MSVC ARM64 BenchmarkChannelHandoff 3066 ns/op -584 ns/op / -16.0% (better)
Windows MSVC ARM64 BenchmarkDefer 61.820 ns/op -3.08 ns/op / -4.7% (better)
Windows MSVC ARM64 BenchmarkDirectCall 0.664 ns/op +0.0002 ns/op / +0.03015% (worse)
Windows MSVC ARM64 BenchmarkGlobalRead 0.664 ns/op +0.0003 ns/op / +0.04519% (worse)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.797 ns/op -0.001 ns/op / -0.02633% (better)
Windows MSVC ARM64 BenchmarkGoroutine 59759 ns/op +151 ns/op / +0.3% (worse)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.147 ns/op +0.009 ns/op / +0.2% (worse)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.771 ns/op +0.001 ns/op / +0.1% (worse)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 905.500 ns/op +4.6 ns/op / +0.5% (worse)
Linux AfterFuncZeroDelivery/LLGo 36092 ns/op -4924 ns/op / -12.0% (better)
Linux CreateStop/Go 288.200 ns/op -2.5 ns/op / -0.9% (better)
Linux CreateStop/LLGo 1870 ns/op +179 ns/op / +10.6% (worse)
Linux RearmStopped/Go 115.800 ns/op -0.4 ns/op / -0.3% (better)
Linux RearmStopped/LLGo 1173 ns/op -77 ns/op / -6.2% (better)
Linux ResetActive/Go 68.500 ns/op -0.11 ns/op / -0.2% (better)
Linux ResetActive/LLGo 857.200 ns/op +68 ns/op / +8.6% (worse)
Linux ResetHeap1024/Go 67.320 ns/op +0.12 ns/op / +0.2% (worse)
Linux ResetHeap1024/LLGo 175.400 ns/op +0.7 ns/op / +0.4% (worse)
macOS AfterFuncZeroDelivery/Go 613.800 ns/op -18.6 ns/op / -2.9% (better)
macOS AfterFuncZeroDelivery/LLGo 114215 ns/op +2203 ns/op / +2.0% (worse)
macOS CreateStop/Go 240.600 ns/op +48 ns/op / +24.9% (worse)
macOS CreateStop/LLGo 721.600 ns/op -22.5 ns/op / -3.0% (better)
macOS RearmStopped/Go 78.500 ns/op +5.69 ns/op / +7.8% (worse)
macOS RearmStopped/LLGo 560.600 ns/op -151 ns/op / -21.2% (better)
macOS ResetActive/Go 57.330 ns/op -2.65 ns/op / -4.4% (better)
macOS ResetActive/LLGo 253.500 ns/op -41.5 ns/op / -14.1% (better)
macOS ResetHeap1024/Go 56.640 ns/op -5.24 ns/op / -8.5% (better)
macOS ResetHeap1024/LLGo 88.560 ns/op -11.25 ns/op / -11.3% (better)
Windows MinGW AfterFuncZeroDelivery/Go 574.600 ns/op +7.1 ns/op / +1.3% (worse)
Windows MinGW AfterFuncZeroDelivery/LLGo 182716 ns/op -3028 ns/op / -1.6% (better)
Windows MinGW CreateStop/Go 116.200 ns/op +2.1 ns/op / +1.8% (worse)
Windows MinGW CreateStop/LLGo 421.700 ns/op -23.7 ns/op / -5.3% (better)
Windows MinGW RearmStopped/Go 31.350 ns/op -0.03 ns/op / -0.1% (better)
Windows MinGW RearmStopped/LLGo 270.400 ns/op -2.9 ns/op / -1.1% (better)
Windows MinGW ResetActive/Go 20.040 ns/op -0.04 ns/op / -0.2% (better)
Windows MinGW ResetActive/LLGo 158.300 ns/op +5.6 ns/op / +3.7% (worse)
Windows MinGW ResetHeap1024/Go 20.430 ns/op -0.06 ns/op / -0.3% (better)
Windows MinGW ResetHeap1024/LLGo 124.200 ns/op -2.5 ns/op / -2.0% (better)
Windows MinGW 386 AfterFuncZeroDelivery/Go 967.700 ns/op +12.1 ns/op / +1.3% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 197180 ns/op +1211 ns/op / +0.6% (worse)
Windows MinGW 386 CreateStop/Go 189.900 ns/op -2.7 ns/op / -1.4% (better)
Windows MinGW 386 CreateStop/LLGo 530.700 ns/op +27 ns/op / +5.4% (worse)
Windows MinGW 386 RearmStopped/Go 63.290 ns/op -0.18 ns/op / -0.3% (better)
Windows MinGW 386 RearmStopped/LLGo 351.200 ns/op -5.1 ns/op / -1.4% (better)
Windows MinGW 386 ResetActive/Go 38.970 ns/op -0.14 ns/op / -0.4% (better)
Windows MinGW 386 ResetActive/LLGo 986.700 ns/op +39 ns/op / +4.1% (worse)
Windows MinGW 386 ResetHeap1024/Go 39.430 ns/op +0.03 ns/op / +0.1% (worse)
Windows MinGW 386 ResetHeap1024/LLGo 187.200 ns/op -0.6 ns/op / -0.3% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 667.900 ns/op -6.6 ns/op / -1.0% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 145993 ns/op -254 ns/op / -0.2% (better)
Windows MinGW ARM64 CreateStop/Go 199.400 ns/op -1.7 ns/op / -0.8% (better)
Windows MinGW ARM64 CreateStop/LLGo 359.900 ns/op +7.8 ns/op / +2.2% (worse)
Windows MinGW ARM64 RearmStopped/Go 70.580 ns/op +0.01 ns/op / +0.01417% (worse)
Windows MinGW ARM64 RearmStopped/LLGo 248.200 ns/op -0.4 ns/op / -0.2% (better)
Windows MinGW ARM64 ResetActive/Go 31.040 ns/op +0.05 ns/op / +0.2% (worse)
Windows MinGW ARM64 ResetActive/LLGo 118.800 ns/op -10.9 ns/op / -8.4% (better)
Windows MinGW ARM64 ResetHeap1024/Go 31.120 ns/op +0.13 ns/op / +0.4% (worse)
Windows MinGW ARM64 ResetHeap1024/LLGo 124.500 ns/op -0.7 ns/op / -0.6% (better)
Windows MSVC AfterFuncZeroDelivery/Go 576.400 ns/op +31.6 ns/op / +5.8% (worse)
Windows MSVC AfterFuncZeroDelivery/LLGo 175531 ns/op -427 ns/op / -0.2% (better)
Windows MSVC CreateStop/Go 116.200 ns/op +0.2 ns/op / +0.2% (worse)
Windows MSVC CreateStop/LLGo 411.800 ns/op -12.8 ns/op / -3.0% (better)
Windows MSVC RearmStopped/Go 31.700 ns/op +0.49 ns/op / +1.6% (worse)
Windows MSVC RearmStopped/LLGo 254.900 ns/op +5.4 ns/op / +2.2% (worse)
Windows MSVC ResetActive/Go 20.130 ns/op +0.08 ns/op / +0.4% (worse)
Windows MSVC ResetActive/LLGo 140.700 ns/op -3.7 ns/op / -2.6% (better)
Windows MSVC ResetHeap1024/Go 20.540 ns/op +0.14 ns/op / +0.7% (worse)
Windows MSVC ResetHeap1024/LLGo 123.600 ns/op -3.1 ns/op / -2.4% (better)
Windows MSVC 386 AfterFuncZeroDelivery/Go 965.600 ns/op +5.7 ns/op / +0.6% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 201090 ns/op -849 ns/op / -0.4% (better)
Windows MSVC 386 CreateStop/Go 191.300 ns/op -1 ns/op / -0.5% (better)
Windows MSVC 386 CreateStop/LLGo 591.200 ns/op -66.9 ns/op / -10.2% (better)
Windows MSVC 386 RearmStopped/Go 63.610 ns/op +0.26 ns/op / +0.4% (worse)
Windows MSVC 386 RearmStopped/LLGo 330.700 ns/op +2.1 ns/op / +0.6% (worse)
Windows MSVC 386 ResetActive/Go 39.160 ns/op +0.12 ns/op / +0.3% (worse)
Windows MSVC 386 ResetActive/LLGo 954.600 ns/op +8.4 ns/op / +0.9% (worse)
Windows MSVC 386 ResetHeap1024/Go 39.560 ns/op +0.01 ns/op / +0.02528% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 170.800 ns/op +1.4 ns/op / +0.8% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 664.500 ns/op -11.9 ns/op / -1.8% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 175114 ns/op -3453 ns/op / -1.9% (better)
Windows MSVC ARM64 CreateStop/Go 206.100 ns/op +7.2 ns/op / +3.6% (worse)
Windows MSVC ARM64 CreateStop/LLGo 382.300 ns/op +5.5 ns/op / +1.5% (worse)
Windows MSVC ARM64 RearmStopped/Go 70.670 ns/op +0.07 ns/op / +0.1% (worse)
Windows MSVC ARM64 RearmStopped/LLGo 270.100 ns/op -1.6 ns/op / -0.6% (better)
Windows MSVC ARM64 ResetActive/Go 31.100 ns/op +0.03 ns/op / +0.1% (worse)
Windows MSVC ARM64 ResetActive/LLGo 134.400 ns/op +3 ns/op / +2.3% (worse)
Windows MSVC ARM64 ResetHeap1024/Go 31.140 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 ResetHeap1024/LLGo 137.100 ns/op +0.4 ns/op / +0.3% (worse)

Compared with 3a8054a87f76 measured in the same runner job.

@zhouguangyuan0718
zhouguangyuan0718 force-pushed the codex/dev-wasm-full-lto-fix-20261003 branch from 1906f8e to c98d59e Compare October 4, 2026 00:59
@zhouguangyuan0718
zhouguangyuan0718 force-pushed the codex/dev-wasm-full-lto-fix-20261003 branch from c98d59e to 68f6f92 Compare October 4, 2026 09:40

@cpunion cpunion left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已独立审查当前 head 68f6f92715cd,并对照 #2721。建议补齐下面的 LTO 优化等级遗漏后,先合并本 PR,再让 #2721 rebase、去除重复的参数修复。

本 PR 的参数修复及验收测试值得采用,但不能完整替代 #2721:后者还修复 C-only deadcode 链接时错误纳入未构建 runtime 包、缺失 metadata 的诊断,并有 Full LTO 多线程 GC 验收。函数属性处理也不能仅凭 -mattr 就全部删除;下面区分了已经确认的 LLVM 后端边界与未复现的普通 LLGo 回归。

独立验证环境:macOS arm64、Go 1.27.0、LLVM/LLD 22.1.8、Wasmer 7.3.0。

  • TestUseWASILTOEnablesSjLjAtLink、TestLTOLinkerOptFlag 和完整 internal/crosscompile short 测试通过。
  • 本 PR 新增的 dev/test_wasm_wasi_lto.py 完整通过,包括 Thin/Full 实际 LTO 产物、dev Full-LTO 接口 metadata、GC/nogc 的 Goexit/defer/C longjmp/recover、init/main Goexit、uncaught panic,以及两种 LTO 下的 SIMD 测试。
  • 额外构建并运行 wasm-wasi-threaded-gc 的 Full LTO 用例,通过 wasi threaded gc ok,覆盖 worker 根、STW、C 阻塞及 arena 扩容。
  • 实际 LLGo SIMD + atomic 样例,以及 deadcodedrop、nogc/PCLN-none 变体通过。

本地运行使用 Wasmer 7.3.0;CI 中仓库固定 Wasmer 7.5.0 的相关 WASI 任务也已通过,二者不混同。

建议收敛为:后端提供默认特性,已有 target-features 的函数合并必要特性并保留 SIMD;这样可以缩小 #2721 的属性处理范围,而不是逐函数重复写全部默认属性。

if ltoMode.Enabled() {
// Clang does not forward -fwasm-exceptions to the LTO backend.
// Without an explicit exception model, codegen drops SjLj catch pads.
export.LDFLAGS = append(export.LDFLAGS, "-Wl,--mllvm=-exception-model=wasm")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] 请把选择的优化等级传给 wasm-ld 的 LTO 优化器

这里把 -O0 / -O3 复制进 clang 的最终链接参数,但没有生成 --lto-O0 / --lto-O3。我用当前 dev 编译器分别构建了 -O0 -lto=thin 和 -O3 -lto=full 的 println:最终 clang 链接参数有对应 -O,实际 wasm-ld 参数没有 --lto-O*。LLVM 22 wasm-ld 的 LTO 默认等级仍是 2,因此选择的优化等级在链接阶段没有生效。

默认值见 LLVM Driver.cpp。#2721 已有这部分修复,可直接复用本文件现有的映射:

if optFlag := ltoLinkerOptFlag(level); optFlag != "" {
    export.LDFLAGS = append(export.LDFLAGS, "-Wl,"+optFlag)
}

建议在 Thin/Full 两种模式下补上 O0/O3 及 Os/Oz 的参数覆盖,并断言关闭 LTO 时不加入该参数。当前 WASI 新测试只使用 O2,无法发现这个遗漏。

export.LDFLAGS = append(export.LDFLAGS, "-Wl,--mllvm=-exception-model=wasm")
// ThinLTO compiles Go modules independently of the C modules that
// carry these features. Preserve shared memory, TLS and Wasm EH.
export.LDFLAGS = append(export.LDFLAGS, "-Xlinker", "--mllvm=-mattr=+atomics,+bulk-memory,+exception-handling")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

选型边界:-mattr 不是已有函数 target-features 的合并或强制覆盖

这不是已经复现的普通 LLGo SIMD 回归;当前新增矩阵及我补的真实 Go SIMD/atomic 样例都通过。不过,不能据此认为这一行与 #2721 的属性补齐完全等价。

LLVM 22 的 getSubtargetImpl(Function) 会用已有函数属性代替 TargetFS;随后 coalesceFeatures 对模块内定义函数的特性取并集。因此,如果 LTO 后端模块中的定义函数都已有属性,仅传 -mattr 不保证 atomics/bulk-memory/EH 被保留。

我在 LLVM/LLD 22.1.8 上验证的最小输入如下:

target datalayout = "e-m:e-p:32:32-i64:64-n32:64-S128"
target triple = "wasm32-unknown-wasip1"
@state = thread_local global i32 0, align 4

define i32 @inc() #0 {
  %p = call ptr @llvm.threadlocal.address.p0(ptr @state)
  %old = atomicrmw add ptr %p, i32 1 seq_cst
  ret i32 %old
}
declare ptr @llvm.threadlocal.address.p0(ptr)
attributes #0 = { "target-features"="+simd128" }

用 LLVM 22 工具运行:

clang -c probe.ll -target wasm32-wasip1-threads -O2 -flto=thin \
  -matomics -mbulk-memory -fwasm-exceptions -o probe.o
wasm-ld probe.o --no-entry --export=inc --shared-memory \
  --max-memory=16777216 \
  --mllvm=-mattr=+atomics,+bulk-memory,+exception-handling \
  --mllvm=-exception-model=wasm --mllvm=-wasm-enable-sjlj \
  --mllvm=-wasm-use-legacy-eh=false -o probe.wasm

会报 --shared-memory is disallowed ... because it was not compiled with 'atomics' or 'bulk-memory' features;换成 Full LTO 也复现。去掉 shared-memory 限制再检查后端产物,可以看到 TLS 被去掉,atomicrmw 变成普通 load/add/store。把该函数属性改成 +atomics,+bulk-memory,+exception-handling,+simd128 后,shared-memory 链接通过,并保留 TLS/原子指令。

实际 LLGo 样例中的 init 等无该属性的定义让模块能够取得后端默认特性,所以那些样例通过,不代表这条 LLVM 边界不存在。我的建议是采用本 PR 的后端默认参数,同时把 #2721 收窄为只补齐已有属性的函数,并增加上述已有属性 + TLS/atomic 的实际后端回归测试。

@cpunion cpunion left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

复核新提交 db70b67550c0:未发现新增的合入阻塞,建议按原顺序先合并本 PR,再精简 #2721。

上轮意见的处理情况:

  • LTO 优化等级遗漏已修复。 新实现直接复用 ltoLinkerOptFlag,只在 LTO 开启时加入链接参数;关闭 LTO、Thin/Full × O0/O1/O2/O3/Os/Oz 的 18 组配置测试均通过,且断言不会加入重复或错误的 optimizer flag。
  • 函数属性的替代边界已说明清楚。 新注释和 PR 描述准确区分后端默认特性与已有函数 target-features 的合并。本 PR 没有宣称替代 #2721 的属性处理;这一项继续由 #2721 补齐已有属性并添加后端回归测试,不作为当前参数修复的合入阻塞。
  • 原注释的 Thin/Full 范围及特性名称已修正,新增 CI 步骤也补了显式 shell guard。新增 diff 集中,没有发现不必要的功能改动。

本轮独立验证(Go 1.27.0、LLVM/LLD 22.1.8、Wasmer 7.3.0):

  • 构建新 head 的 dev 编译器,相关 flags 测试及完整 internal/crosscompile short 测试通过。
  • 对 globaldce_interface_matrix 分别运行 Thin/Full × O0/O3 四组实际构建及执行;在 clang -v 输出的实际 wasm-ld 命令中确认 --lto-O0/O3,验证 LTO 后端产物非空,全部接口输出正确。两组 Full LTO 还检查了 merged preopt bitcode 中的 llvm.type.checked.load,确认保留 dev GlobalDCE 路径。
  • 最新提交的 CI 已完成:97 项成功,1 项 release 跳过,无失败;包括使用仓库固定 Wasmer 7.5.0 的实际 LTO/SIMD/EH 验收。当前 PR 无合并冲突。

#2721 仍保留 C-only deadcode 链接与 metadata 诊断、已有函数特性合并和 Full LTO 多线程 GC 验收;本 PR 合并后删除其中重复的参数修复即可。

@cpunion
cpunion merged commit 1b01cb9 into main Oct 4, 2026
98 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants