Skip to content

wasm: support SIMD in the WAMR classic interpreter - #2723

Closed
zhouguangyuan0718 wants to merge 2 commits into
xgo-dev:mainfrom
zhouguangyuan0718:codex/wamr-simd-runner-20261003
Closed

zhouguangyuan0718 wants to merge 2 commits into
xgo-dev:mainfrom
zhouguangyuan0718:codex/wamr-simd-runner-20261003

Conversation

@zhouguangyuan0718

Copy link
Copy Markdown
Collaborator

The main-branch WASI SIMD suite fails before execution with unknown value type: WAMR 2.4.5's classic interpreter rejects v128, while the fast interpreter cannot execute the legacy exception handling used by W32.

Add a WAMR patch that reuses its existing SIMDe-backed SIMD operations in the classic interpreter. The adapter reads original Wasm immediates, uses the classic value stack and memory checks, preserves all four vector cells through locals/globals/select/drop/calls, and extends block scanning for the rewritten opcodes. Correct overlapping catch-payload copies and the saved-payload offset used by rethrow, with vector regressions for both.

The build keeps the existing classic/EH/pthread configuration, enables SIMD explicitly, includes the new patch in the local build identity, and handles an empty generator-argument array under macOS Bash 3.2. CI's existing WAMR patch hash invalidates its cache automatically.

A raw Wasm regression now runs at the start of the existing WASI acceptance driver. It covers high vector lanes, unaligned memory, indirect calls, select/drop, block scanning, cross-call catch/rethrow, and a trapping out-of-bounds vector load.

Validation on macOS:

  • Fresh bash dev/build_iwasm.sh build and cache reuse passed.
  • Main-branch WASI SIMD suite passed at O2; the expanded suite from simd: expand SIMD128 operations and runtime integration #2722 passed against the same runner.
  • The raw SIMD/EH regression passed on arm64 and x86_64 (Rosetta) WAMR builds.
  • 57 applicable WebAssembly SIMD spec files passed on each architecture: 24,335 assertions per run, using spec revision 2e44bf79e68cc4fe689f43eea8e7d25b6fdfbce1. Multi-module linking and multi-memory cases are outside this runner configuration. Two input files required whitespace-only normalization for WAMR's Python harness; their assertions were retained.
  • The full WASI pthread acceptance driver passed, including exception/Goexit isolation, threaded GC, heap-growth boundaries, finalizers, reflection/GC races, filesystem operations, the focused standard-library slice, and the GOROOT sentinel.

Unblocks WASI execution for #2722. Linux and Windows validation will run in CI.

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: classic-interpreter SIMD support

This PR adds a WAMR 2.4.5 patch that runs WebAssembly SIMD (v128) in the classic interpreter so the legacy-EH / pthreads profile can execute SIMD, plus build wiring, a .wat regression fixture, a Python boundary test, and validation docs. The port of the fast-interpreter SIMD semantics onto the classic value stack looks faithful and I found no merge-blockers. Findings below are minor/optional.

Verified correct (no change needed):

  • v128.store pops data then address; CHECK_MEMORY_OVERFLOW guards the right maddr with correct widths (16 for full load/store, 8 for extend-loads, width/8 for splat/lane/zero). Loads/stores check before dereferencing.
  • 4-cell (v128 = 4×32-bit) stack discipline is consistent across POP_V128/PUSH_V128, DROP_128/SELECT_128, typed select, and local/global get/set/tee.
  • Lane-index immediates (extract/replace/load_lane/store_lane/shuffle) are not re-checked at runtime, which is safe only because the loader validation (WASM_OP_SIMD_PREFIX case) is enabled under the new || WASM_ENABLE_SIMDE guard. Keep the execution and loader guards in lock-step — if a future refactor drops || WASM_ENABLE_SIMDE from the loader's SIMD validation, these handlers become OOB read/write primitives driven by untrusted bytecode.
  • EH changes: catch now copies from the saved payload (frame_sp - cell_num_to_copy) avoiding source/dest overlap for a multi-cell v128; rethrow resolves the tag type via module->e->tags[...] with the is_import_tag branch (correctly handling imported tags) and reads the payload after the tag word.
  • V128 / GET_V128_FROM_ADDR / PUT_V128_TO_ADDR / LOAD_V128 / STORE_V128 resolve from shared headers (wasm.h, wasm_runtime_common.h) that the classic TU already includes, so the patch compiles without adding them locally.

See inline comments for the minor items.

+ SIMD_CASE(SIMD_i64x2_extmul_high_i32x4_u,
+ SIMD_DOUBLE_OP(simde_wasm_u64x2_extmul_high_u32x4));
+
+ /* f32x4 opertions */

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Typo: f32x4 opertions → operations (the i64x2/f64x2 sections nearby spell it correctly).

+ V128 v2 = POP_V128();
+ V128 v3 = POP_V128();
+
+ simde_v128_t simde_result = simde_wasm_v128_bitselect(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

v128.bitselect operand naming is confusing although the behavior is correct. Here v1 is popped first (top of stack = the mask/condition) and passed as SIMDe's 3rd arg, while v3 is the first spec operand. Naming the condition v1 and operand-1 v3 inverts the Wasm spec's conventional names and makes this hard to audit. Consider renaming to c / v2 / v1 matching the spec so the order is self-evident. Non-blocking.

+#define SIMD_SPLAT_OP_F64(simde_func) \
+ SIMD_SPLAT_OP(simde_func, POP_F64, float64)
+
+ case SIMD_i8x16_splat:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SIMD_i8x16_splat is hand-expanded while every other splat uses the SIMD_SPLAT_OP_* macros defined just above. It also reuses the shared val variable (declared as tbl_elem_idx_t, not uint32) rather than a local. SIMD_CASE(SIMD_i8x16_splat, SIMD_SPLAT_OP_I32(simde_wasm_i8x16_splat)); would be consistent and avoid touching val. Non-blocking.

+ SIMD_SPLAT_OP_F32(simde_wasm_f32x4_splat));
+ SIMD_CASE(SIMD_f64x2_splat,
+ SIMD_SPLAT_OP_F64(simde_wasm_f64x2_splat));
+#define SIMD_LANE_HANDLE_UNALIGNED_ACCESS()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SIMD_LANE_HANDLE_UNALIGNED_ACCESS() expands to nothing and is invoked in five places (carried over from the fast interpreter). If it's an intentional hook for future unaligned handling, a one-line comment would clarify; otherwise it's dead code. Non-blocking.

fast interpreter, with the classic value stack, original Wasm immediates,
and memory bounds checks. Preserve all four cells through locals, globals,
select/drop, calls, and exception payloads. Copy catch values from the
saved payload to avoid overlap; rethrow reads that payload after the tag.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Doc completeness: the rethrow hunk makes two changes — (a) reading the payload after the tag, and (b) a new tag-type resolution that handles imported tags via module->e->tags[...].is_import_tag. The patch header and wasm-wasi-validation.md describe only (a). Consider mentioning the imported-tag fix in the header so it fully reflects the hunk. (Only the inline code comment currently notes it.) Optional.

@codecov

codecov Bot commented Oct 3, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@github-actions

github-actions Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

db3bab76d8de | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 8568 B 0 B / +0.0% 387 B 0 B / +0.0% 989.592 ms -36.26 ms / -3.5% (better) 1.294 ms -20.78 us / -1.6% (better)
Linux cprintf-lto 8408 B 0 B / +0.0% 368 B 0 B / +0.0% 982.265 ms -72.12 ms / -6.8% (better) 1.314 ms -21.7 us / -1.6% (better)
Linux fmtprintf 4880952 B 0 B / +0.0% 501619 B 0 B / +0.0% 7.400 s +78.59 ms / +1.1% (worse) 3.233 ms +115.4 us / +3.7% (worse)
Linux fmtprintf-lto 3600808 B 0 B / +0.0% 438538 B 0 B / +0.0% 16.358 s -507.8 ms / -3.0% (better) 2.917 ms -72.92 us / -2.4% (better)
Linux println 674952 B 0 B / +0.0% 16855 B 0 B / +0.0% 1.009 s -51.32 ms / -4.8% (better) 1.694 ms -67.51 us / -3.8% (better)
Linux println-lto 190024 B 0 B / +0.0% 14273 B 0 B / +0.0% 1.285 s -91.69 ms / -6.7% (better) 1.627 ms -18.93 us / -1.2% (better)
macOS cprintf 50736 B 0 B / +0.0% 4409 B 0 B / +0.0% 896.207 ms -29.04 ms / -3.1% (better) 2.593 ms -441.8 us / -14.6% (better)
macOS cprintf-lto 50496 B 0 B / +0.0% 161 B 0 B / +0.0% 849.285 ms -438.9 ms / -34.1% (better) 2.561 ms -243.3 us / -8.7% (better)
macOS fmtprintf 1773456 B 0 B / +0.0% 877780 B 0 B / +0.0% 7.060 s +2.402 s / +51.6% (worse) 14.534 ms +6.367 ms / +78.0% (worse)
macOS fmtprintf-lto 1360928 B 0 B / +0.0% 762428 B 0 B / +0.0% 15.364 s +6.034 s / +64.7% (worse) 11.733 ms +7.942 ms / +209.5% (worse)
macOS println 99344 B 0 B / +0.0% 24216 B 0 B / +0.0% 861.186 ms -222.5 ms / -20.5% (better) 3.573 ms +222.5 us / +6.6% (worse)
macOS println-lto 83664 B 0 B / +0.0% 21457 B 0 B / +0.0% 1.024 s -27.21 ms / -2.6% (better) 3.442 ms -365.9 us / -9.6% (better)
Windows MinGW cprintf 651264 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.677 s -9.898 ms / -0.6% (better) 3.540 ms -484.6 us / -12.0% (better)
Windows MinGW cprintf-lto 43520 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.699 s -100.5 ms / -5.6% (better) 3.372 ms -334.9 us / -9.0% (better)
Windows MinGW fmtprintf 5432832 B 0 B / +0.0% 600310 B 0 B / +0.0% 7.219 s -48.66 ms / -0.7% (better) 8.181 ms +130.6 us / +1.6% (worse)
Windows MinGW fmtprintf-lto 4127232 B 0 B / +0.0% 546582 B 0 B / +0.0% 14.609 s +23.68 ms / +0.2% (worse) 8.316 ms +170.2 us / +2.1% (worse)
Windows MinGW println 706560 B 0 B / +0.0% 25190 B 0 B / +0.0% 1.650 s -21.06 ms / -1.3% (better) 6.446 ms +154.2 us / +2.5% (worse)
Windows MinGW println-lto 208896 B 0 B / +0.0% 22054 B 0 B / +0.0% 1.953 s -15.94 ms / -0.8% (better) 6.503 ms -491.6 us / -7.0% (better)
Windows MinGW 386 cprintf 601600 B 0 B / +0.0% 5326 B 0 B / +0.0% 1.742 s +19.98 ms / +1.2% (worse) 5.697 ms +656.3 us / +13.0% (worse)
Windows MinGW 386 cprintf-lto 103424 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.829 s +103.8 ms / +6.0% (worse) 4.962 ms -520.9 us / -9.5% (better)
Windows MinGW 386 fmtprintf 4745216 B 0 B / +0.0% 472478 B 0 B / +0.0% 7.438 s -30.37 ms / -0.4% (better) 10.557 ms +443.4 us / +4.4% (worse)
Windows MinGW 386 fmtprintf-lto 4148736 B 0 B / +0.0% 451114 B 0 B / +0.0% 14.721 s -129.7 ms / -0.9% (better) 10.289 ms -286.2 us / -2.7% (better)
Windows MinGW 386 println 653312 B 0 B / +0.0% 21490 B 0 B / +0.0% 1.714 s -7.716 ms / -0.4% (better) 8.499 ms -620.1 us / -6.8% (better)
Windows MinGW 386 println-lto 258560 B 0 B / +0.0% 19306 B 0 B / +0.0% 1.960 s -23.95 ms / -1.2% (better) 8.469 ms -525.8 us / -5.8% (better)
Windows MinGW ARM64 cprintf 661504 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.923 s -20.3 ms / -1.0% (better) 6.430 ms -437.6 us / -6.4% (better)
Windows MinGW ARM64 cprintf-lto 43520 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.946 s +22.16 ms / +1.2% (worse) 6.246 ms -219.2 us / -3.4% (better)
Windows MinGW ARM64 fmtprintf 5343744 B 0 B / +0.0% 510876 B 0 B / +0.0% 7.042 s +21.22 ms / +0.3% (worse) 12.903 ms -206.2 us / -1.6% (better)
Windows MinGW ARM64 fmtprintf-lto 4302336 B 0 B / +0.0% 477264 B 0 B / +0.0% 13.792 s -99.32 ms / -0.7% (better) 14.163 ms +631.5 us / +4.7% (worse)
Windows MinGW ARM64 println 714240 B 0 B / +0.0% 23884 B 0 B / +0.0% 1.915 s +36.39 ms / +1.9% (worse) 11.470 ms +135.6 us / +1.2% (worse)
Windows MinGW ARM64 println-lto 215552 B 0 B / +0.0% 21232 B 0 B / +0.0% 2.194 s +51.14 ms / +2.4% (worse) 11.383 ms +256.8 us / +2.3% (worse)
Windows MSVC cprintf 893952 B 0 B / +0.0% 65798 B 0 B / +0.0% 1.876 s +333.2 ms / +21.6% (worse) 4.130 ms +13.9 us / +0.3% (worse)
Windows MSVC cprintf-lto 289792 B 0 B / +0.0% 65734 B 0 B / +0.0% 1.679 s +115.9 ms / +7.4% (worse) 3.510 ms +107.2 us / +3.2% (worse)
Windows MSVC fmtprintf 5732864 B 0 B / +0.0% 695862 B 0 B / +0.0% 7.516 s +540.5 ms / +7.7% (worse) 9.842 ms +1.236 ms / +14.4% (worse)
Windows MSVC fmtprintf-lto 4449792 B 0 B / +0.0% 646134 B 0 B / +0.0% 14.955 s +1.436 s / +10.6% (worse) 9.899 ms +1.087 ms / +12.3% (worse)
Windows MSVC println 1017344 B 0 B / +0.0% 120854 B 0 B / +0.0% 1.697 s +141.5 ms / +9.1% (worse) 8.104 ms +1.027 ms / +14.5% (worse)
Windows MSVC println-lto 528384 B 0 B / +0.0% 118390 B 0 B / +0.0% 1.936 s +160.2 ms / +9.0% (worse) 11.709 ms +4.648 ms / +65.8% (worse)
Windows MSVC 386 cprintf 513536 B 0 B / +0.0% 3931 B 0 B / +0.0% 1.361 s +55.21 ms / +4.2% (worse) 6.258 ms +535.9 us / +9.4% (worse)
Windows MSVC 386 cprintf-lto 44032 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.294 s -15.27 ms / -1.2% (better) 5.056 ms -599.5 us / -10.6% (better)
Windows MSVC 386 fmtprintf 4479488 B 0 B / +0.0% 455868 B 0 B / +0.0% 6.054 s -26.29 ms / -0.4% (better) 11.494 ms +733.1 us / +6.8% (worse)
Windows MSVC 386 fmtprintf-lto 3899392 B 0 B / +0.0% 426651 B 0 B / +0.0% 11.510 s -69.55 ms / -0.6% (better) 10.604 ms -1.926 ms / -15.4% (better)
Windows MSVC 386 println 567296 B 0 B / +0.0% 20340 B 0 B / +0.0% 1.298 s -304.4 us / -0.02344% (better) 8.863 ms -941.5 us / -9.6% (better)
Windows MSVC 386 println-lto 199168 B 0 B / +0.0% 18501 B 0 B / +0.0% 1.506 s -19.58 ms / -1.3% (better) 8.802 ms -742.1 us / -7.8% (better)
Windows MSVC ARM64 cprintf 662528 B 0 B / +0.0% 4192 B 0 B / +0.0% 1.172 s +27.48 ms / +2.4% (worse) 4.655 ms +20.1 us / +0.4% (worse)
Windows MSVC ARM64 cprintf-lto 47616 B 0 B / +0.0% 4084 B 0 B / +0.0% 1.173 s +30.44 ms / +2.7% (worse) 4.656 ms +35.6 us / +0.8% (worse)
Windows MSVC ARM64 fmtprintf 5339648 B 0 B / +0.0% 510808 B 0 B / +0.0% 5.022 s +51.21 ms / +1.0% (worse) 9.717 ms -166.9 us / -1.7% (better)
Windows MSVC ARM64 fmtprintf-lto 4309504 B 0 B / +0.0% 477924 B 0 B / +0.0% 9.043 s -283 ms / -3.0% (better) 9.486 ms -485.9 us / -4.9% (better)
Windows MSVC ARM64 println 715264 B 0 B / +0.0% 23908 B 0 B / +0.0% 1.162 s +18.43 ms / +1.6% (worse) 8.828 ms +531.6 us / +6.4% (worse)
Windows MSVC ARM64 println-lto 220672 B 0 B / +0.0% 21380 B 0 B / +0.0% 1.352 s +45.79 ms / +3.5% (worse) 8.467 ms +263.7 us / +3.2% (worse)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.480 ns/op 0 ns/op / +0.0%
Linux BenchmarkMergeCompilerFlags 195.500 ns/op +1.4 ns/op / +0.7% (worse)
Linux BenchmarkMergeLinkerFlags 127.600 ns/op +1.2 ns/op / +0.9% (worse)
Linux BenchmarkChannelBuffered 55.080 ns/op +0.03 ns/op / +0.1% (worse)
Linux BenchmarkChannelHandoff 13235 ns/op -1245 ns/op / -8.6% (better)
Linux BenchmarkDefer 48.980 ns/op +1.14 ns/op / +2.4% (worse)
Linux BenchmarkDirectCall 1.166 ns/op +0.001 ns/op / +0.1% (worse)
Linux BenchmarkGlobalRead 1.164 ns/op -0.007 ns/op / -0.6% (better)
Linux BenchmarkGlobalWrite 7.762 ns/op +0.002 ns/op / +0.02577% (worse)
Linux BenchmarkGoroutine 28264 ns/op -7497 ns/op / -21.0% (better)
Linux BenchmarkInterfaceCall 5.818 ns/op -0.021 ns/op / -0.4% (better)
Linux BenchmarkRuntimeGetG 2.988 ns/op -0.021 ns/op / -0.7% (better)
macOS BenchmarkLookupPCRandom 13.520 ns/op +1.06 ns/op / +8.5% (worse)
macOS BenchmarkMergeCompilerFlags 112.900 ns/op +3.3 ns/op / +3.0% (worse)
macOS BenchmarkMergeLinkerFlags 69.130 ns/op +0.41 ns/op / +0.6% (worse)
macOS BenchmarkChannelBuffered 27.070 ns/op +2.01 ns/op / +8.0% (worse)
macOS BenchmarkChannelHandoff 7923 ns/op -2116 ns/op / -21.1% (better)
macOS BenchmarkDefer 40.560 ns/op +8.26 ns/op / +25.6% (worse)
macOS BenchmarkDirectCall 1.137 ns/op +0.06 ns/op / +5.6% (worse)
macOS BenchmarkGlobalRead 1.110 ns/op +0.048 ns/op / +4.5% (worse)
macOS BenchmarkGlobalWrite 1.137 ns/op +0.086 ns/op / +8.2% (worse)
macOS BenchmarkGoroutine 43246 ns/op +9466 ns/op / +28.0% (worse)
macOS BenchmarkInterfaceCall 4.384 ns/op +0.48 ns/op / +12.3% (worse)
macOS BenchmarkRuntimeGetG 2.187 ns/op +0.05 ns/op / +2.3% (worse)
Windows MinGW BenchmarkLookupPCRandom 12.810 ns/op -0.43 ns/op / -3.2% (better)
Windows MinGW BenchmarkMergeCompilerFlags 656.200 ns/op +46.7 ns/op / +7.7% (worse)
Windows MinGW BenchmarkMergeLinkerFlags 601.900 ns/op +70.7 ns/op / +13.3% (worse)
Windows MinGW BenchmarkChannelBuffered 30.050 ns/op -0.16 ns/op / -0.5% (better)
Windows MinGW BenchmarkChannelHandoff 860.200 ns/op +5.6 ns/op / +0.7% (worse)
Windows MinGW BenchmarkDefer 56.060 ns/op -0.14 ns/op / -0.2% (better)
Windows MinGW BenchmarkDirectCall 1.547 ns/op -0.003 ns/op / -0.2% (better)
Windows MinGW BenchmarkGlobalRead 1.550 ns/op +0.002 ns/op / +0.1% (worse)
Windows MinGW BenchmarkGlobalWrite 2.471 ns/op -0.004 ns/op / -0.2% (better)
Windows MinGW BenchmarkGoroutine 87438 ns/op -1103 ns/op / -1.2% (better)
Windows MinGW BenchmarkInterfaceCall 8.374 ns/op -0.056 ns/op / -0.7% (better)
Windows MinGW BenchmarkRuntimeGetG 2.476 ns/op -0.195 ns/op / -7.3% (better)
Windows MinGW 386 BenchmarkLookupPCRandom 26.520 ns/op +0.03 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkMergeCompilerFlags 755.500 ns/op -6.7 ns/op / -0.9% (better)
Windows MinGW 386 BenchmarkMergeLinkerFlags 695.900 ns/op -7.4 ns/op / -1.1% (better)
Windows MinGW 386 BenchmarkChannelBuffered 41.410 ns/op -0.04 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkChannelHandoff 885.100 ns/op +74.6 ns/op / +9.2% (worse)
Windows MinGW 386 BenchmarkDefer 42.620 ns/op +0.11 ns/op / +0.3% (worse)
Windows MinGW 386 BenchmarkDirectCall 1.547 ns/op -0.003 ns/op / -0.2% (better)
Windows MinGW 386 BenchmarkGlobalRead 1.551 ns/op 0 ns/op / +0.0%
Windows MinGW 386 BenchmarkGlobalWrite 7.786 ns/op +0.011 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkGoroutine 105316 ns/op +2225 ns/op / +2.2% (worse)
Windows MinGW 386 BenchmarkInterfaceCall 8.385 ns/op -0.018 ns/op / -0.2% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 1.931 ns/op +0.006 ns/op / +0.3% (worse)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.100 ns/op -0.06 ns/op / -0.5% (better)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 627.600 ns/op +43.4 ns/op / +7.4% (worse)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 572.700 ns/op +33.3 ns/op / +6.2% (worse)
Windows MinGW ARM64 BenchmarkChannelBuffered 37.590 ns/op -1.99 ns/op / -5.0% (better)
Windows MinGW ARM64 BenchmarkChannelHandoff 2545 ns/op +214 ns/op / +9.2% (worse)
Windows MinGW ARM64 BenchmarkDefer 57.190 ns/op +0.14 ns/op / +0.2% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.663 ns/op -0.0006 ns/op / -0.1% (better)
Windows MinGW ARM64 BenchmarkGlobalRead 0.664 ns/op -0.0028 ns/op / -0.4% (better)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.737 ns/op -0.0042 ns/op / -0.6% (better)
Windows MinGW ARM64 BenchmarkGoroutine 65886 ns/op +582 ns/op / +0.9% (worse)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.144 ns/op -0.001 ns/op / -0.02413% (better)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.771 ns/op -0.042 ns/op / -2.3% (better)
Windows MSVC BenchmarkLookupPCRandom 13.170 ns/op +0.05 ns/op / +0.4% (worse)
Windows MSVC BenchmarkMergeCompilerFlags 649.200 ns/op +13 ns/op / +2.0% (worse)
Windows MSVC BenchmarkMergeLinkerFlags 568.200 ns/op +13.1 ns/op / +2.4% (worse)
Windows MSVC BenchmarkChannelBuffered 29.740 ns/op +0.05 ns/op / +0.2% (worse)
Windows MSVC BenchmarkChannelHandoff 1029 ns/op -23 ns/op / -2.2% (better)
Windows MSVC BenchmarkDefer 55.050 ns/op -0.72 ns/op / -1.3% (better)
Windows MSVC BenchmarkDirectCall 1.549 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC BenchmarkGlobalRead 1.551 ns/op +0.003 ns/op / +0.2% (worse)
Windows MSVC BenchmarkGlobalWrite 2.470 ns/op -0.001 ns/op / -0.04047% (better)
Windows MSVC BenchmarkGoroutine 93077 ns/op +2055 ns/op / +2.3% (worse)
Windows MSVC BenchmarkInterfaceCall 9.016 ns/op -0.009 ns/op / -0.1% (better)
Windows MSVC BenchmarkRuntimeGetG 2.478 ns/op -0.003 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkLookupPCRandom 60.080 ns/op -0.33 ns/op / -0.5% (better)
Windows MSVC 386 BenchmarkMergeCompilerFlags 638.400 ns/op -8.6 ns/op / -1.3% (better)
Windows MSVC 386 BenchmarkMergeLinkerFlags 626.900 ns/op +18.4 ns/op / +3.0% (worse)
Windows MSVC 386 BenchmarkChannelBuffered 46.800 ns/op +0.21 ns/op / +0.5% (worse)
Windows MSVC 386 BenchmarkChannelHandoff 1522 ns/op +302 ns/op / +24.8% (worse)
Windows MSVC 386 BenchmarkDefer 41.040 ns/op +1.53 ns/op / +3.9% (worse)
Windows MSVC 386 BenchmarkDirectCall 0.879 ns/op +0.0214 ns/op / +2.5% (worse)
Windows MSVC 386 BenchmarkGlobalRead 0.885 ns/op +0.0115 ns/op / +1.3% (worse)
Windows MSVC 386 BenchmarkGlobalWrite 14.790 ns/op -0.01 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGoroutine 214503 ns/op +3796 ns/op / +1.8% (worse)
Windows MSVC 386 BenchmarkInterfaceCall 4.648 ns/op +0.04 ns/op / +0.9% (worse)
Windows MSVC 386 BenchmarkRuntimeGetG 1.396 ns/op -0.009 ns/op / -0.6% (better)
Windows MSVC ARM64 BenchmarkLookupPCRandom 8.872 ns/op -0.01 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 459.100 ns/op +11.8 ns/op / +2.6% (worse)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 441.700 ns/op +11.8 ns/op / +2.7% (worse)
Windows MSVC ARM64 BenchmarkChannelBuffered 39.650 ns/op +2.07 ns/op / +5.5% (worse)
Windows MSVC ARM64 BenchmarkChannelHandoff 1797 ns/op -33 ns/op / -1.8% (better)
Windows MSVC ARM64 BenchmarkDefer 55.990 ns/op -3.63 ns/op / -6.1% (better)
Windows MSVC ARM64 BenchmarkDirectCall 0.626 ns/op -0.0004 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkGlobalRead 0.660 ns/op -0.0237 ns/op / -3.5% (better)
Windows MSVC ARM64 BenchmarkGlobalWrite 4.032 ns/op +0.002 ns/op / +0.04963% (worse)
Windows MSVC ARM64 BenchmarkGoroutine 35760 ns/op +44 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkInterfaceCall 3.144 ns/op -0.04 ns/op / -1.3% (better)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.879 ns/op -0.001 ns/op / -0.1% (better)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 902.600 ns/op +3 ns/op / +0.3% (worse)
Linux AfterFuncZeroDelivery/LLGo 36799 ns/op -85 ns/op / -0.2% (better)
Linux CreateStop/Go 289.900 ns/op +0.6 ns/op / +0.2% (worse)
Linux CreateStop/LLGo 1892 ns/op -64 ns/op / -3.3% (better)
Linux RearmStopped/Go 115.800 ns/op +1.1 ns/op / +1.0% (worse)
Linux RearmStopped/LLGo 1381 ns/op +110 ns/op / +8.7% (worse)
Linux ResetActive/Go 68.720 ns/op +1.2 ns/op / +1.8% (worse)
Linux ResetActive/LLGo 817.300 ns/op +24.5 ns/op / +3.1% (worse)
Linux ResetHeap1024/Go 67.050 ns/op -0.05 ns/op / -0.1% (better)
Linux ResetHeap1024/LLGo 176.400 ns/op -0.3 ns/op / -0.2% (better)
macOS AfterFuncZeroDelivery/Go 431 ns/op -7.7 ns/op / -1.8% (better)
macOS AfterFuncZeroDelivery/LLGo 86050 ns/op +15266 ns/op / +21.6% (worse)
macOS CreateStop/Go 142.700 ns/op +3.2 ns/op / +2.3% (worse)
macOS CreateStop/LLGo 575.800 ns/op +143.3 ns/op / +33.1% (worse)
macOS RearmStopped/Go 56.420 ns/op -0.42 ns/op / -0.7% (better)
macOS RearmStopped/LLGo 343.700 ns/op +26.3 ns/op / +8.3% (worse)
macOS ResetActive/Go 51.910 ns/op +10.06 ns/op / +24.0% (worse)
macOS ResetActive/LLGo 172.900 ns/op +30.1 ns/op / +21.1% (worse)
macOS ResetHeap1024/Go 41.630 ns/op -1.25 ns/op / -2.9% (better)
macOS ResetHeap1024/LLGo 85.660 ns/op -2.53 ns/op / -2.9% (better)
Windows MinGW AfterFuncZeroDelivery/Go 566.300 ns/op -2.9 ns/op / -0.5% (better)
Windows MinGW AfterFuncZeroDelivery/LLGo 179272 ns/op +272 ns/op / +0.2% (worse)
Windows MinGW CreateStop/Go 115.300 ns/op +1.1 ns/op / +1.0% (worse)
Windows MinGW CreateStop/LLGo 408.900 ns/op +0.9 ns/op / +0.2% (worse)
Windows MinGW RearmStopped/Go 31.500 ns/op +0.21 ns/op / +0.7% (worse)
Windows MinGW RearmStopped/LLGo 268.100 ns/op +2.2 ns/op / +0.8% (worse)
Windows MinGW ResetActive/Go 20.080 ns/op +0.07 ns/op / +0.3% (worse)
Windows MinGW ResetActive/LLGo 149 ns/op -13.6 ns/op / -8.4% (better)
Windows MinGW ResetHeap1024/Go 20.640 ns/op +0.09 ns/op / +0.4% (worse)
Windows MinGW ResetHeap1024/LLGo 125.500 ns/op -0.7 ns/op / -0.6% (better)
Windows MinGW 386 AfterFuncZeroDelivery/Go 958.800 ns/op -74.2 ns/op / -7.2% (better)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 198230 ns/op -23059 ns/op / -10.4% (better)
Windows MinGW 386 CreateStop/Go 193.200 ns/op +0.9 ns/op / +0.5% (worse)
Windows MinGW 386 CreateStop/LLGo 499 ns/op -161.5 ns/op / -24.5% (better)
Windows MinGW 386 RearmStopped/Go 63.440 ns/op +0.03 ns/op / +0.04731% (worse)
Windows MinGW 386 RearmStopped/LLGo 356.400 ns/op -46.4 ns/op / -11.5% (better)
Windows MinGW 386 ResetActive/Go 39.070 ns/op +0.09 ns/op / +0.2% (worse)
Windows MinGW 386 ResetActive/LLGo 939.700 ns/op -72.3 ns/op / -7.1% (better)
Windows MinGW 386 ResetHeap1024/Go 39.590 ns/op -0.48 ns/op / -1.2% (better)
Windows MinGW 386 ResetHeap1024/LLGo 187.600 ns/op -4.2 ns/op / -2.2% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 657.800 ns/op -6.7 ns/op / -1.0% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 150802 ns/op -2602 ns/op / -1.7% (better)
Windows MinGW ARM64 CreateStop/Go 212.700 ns/op +14.4 ns/op / +7.3% (worse)
Windows MinGW ARM64 CreateStop/LLGo 364.300 ns/op -13.3 ns/op / -3.5% (better)
Windows MinGW ARM64 RearmStopped/Go 70.540 ns/op -0.11 ns/op / -0.2% (better)
Windows MinGW ARM64 RearmStopped/LLGo 249.200 ns/op -0.9 ns/op / -0.4% (better)
Windows MinGW ARM64 ResetActive/Go 31.090 ns/op +0.05 ns/op / +0.2% (worse)
Windows MinGW ARM64 ResetActive/LLGo 121.100 ns/op +0.5 ns/op / +0.4% (worse)
Windows MinGW ARM64 ResetHeap1024/Go 31.150 ns/op +0.06 ns/op / +0.2% (worse)
Windows MinGW ARM64 ResetHeap1024/LLGo 124.400 ns/op -0.2 ns/op / -0.2% (better)
Windows MSVC AfterFuncZeroDelivery/Go 566.800 ns/op +0.2 ns/op / +0.0353% (worse)
Windows MSVC AfterFuncZeroDelivery/LLGo 176881 ns/op +1185 ns/op / +0.7% (worse)
Windows MSVC CreateStop/Go 118.300 ns/op +2.3 ns/op / +2.0% (worse)
Windows MSVC CreateStop/LLGo 461.100 ns/op +52.7 ns/op / +12.9% (worse)
Windows MSVC RearmStopped/Go 31.440 ns/op +0.11 ns/op / +0.4% (worse)
Windows MSVC RearmStopped/LLGo 255.500 ns/op +3.3 ns/op / +1.3% (worse)
Windows MSVC ResetActive/Go 20.410 ns/op +0.41 ns/op / +2.0% (worse)
Windows MSVC ResetActive/LLGo 145.200 ns/op +4 ns/op / +2.8% (worse)
Windows MSVC ResetHeap1024/Go 20.380 ns/op -0.13 ns/op / -0.6% (better)
Windows MSVC ResetHeap1024/LLGo 124.100 ns/op -0.1 ns/op / -0.1% (better)
Windows MSVC 386 AfterFuncZeroDelivery/Go 865.900 ns/op -26.5 ns/op / -3.0% (better)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 309785 ns/op -8476 ns/op / -2.7% (better)
Windows MSVC 386 CreateStop/Go 220.300 ns/op -0.1 ns/op / -0.04537% (better)
Windows MSVC 386 CreateStop/LLGo 835.500 ns/op +98.5 ns/op / +13.4% (worse)
Windows MSVC 386 RearmStopped/Go 81.160 ns/op +0.51 ns/op / +0.6% (worse)
Windows MSVC 386 RearmStopped/LLGo 331.800 ns/op -96.8 ns/op / -22.6% (better)
Windows MSVC 386 ResetActive/Go 38.910 ns/op +0.05 ns/op / +0.1% (worse)
Windows MSVC 386 ResetActive/LLGo 269.700 ns/op +9.7 ns/op / +3.7% (worse)
Windows MSVC 386 ResetHeap1024/Go 39.150 ns/op +0.11 ns/op / +0.3% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 143.600 ns/op +1.8 ns/op / +1.3% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 628.400 ns/op -1.4 ns/op / -0.2% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 84402 ns/op -442 ns/op / -0.5% (better)
Windows MSVC ARM64 CreateStop/Go 185.300 ns/op -1.7 ns/op / -0.9% (better)
Windows MSVC ARM64 CreateStop/LLGo 376.900 ns/op -7 ns/op / -1.8% (better)
Windows MSVC ARM64 RearmStopped/Go 72.640 ns/op +0.01 ns/op / +0.01377% (worse)
Windows MSVC ARM64 RearmStopped/LLGo 265.800 ns/op -2.6 ns/op / -1.0% (better)
Windows MSVC ARM64 ResetActive/Go 33.670 ns/op -0.01 ns/op / -0.02969% (better)
Windows MSVC ARM64 ResetActive/LLGo 133.300 ns/op -15.5 ns/op / -10.4% (better)
Windows MSVC ARM64 ResetHeap1024/Go 33.190 ns/op -0.5 ns/op / -1.5% (better)
Windows MSVC ARM64 ResetHeap1024/LLGo 139.500 ns/op +0.5 ns/op / +0.4% (worse)

Compared with 783ec4fd3d57 measured in the same runner job.

@github-actions

github-actions Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

db3bab76d8de | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 154193 B 0 B / +0.0% 88742 B 0 B / +0.0%
cprintf/j32-goos-js/LLGo 152295 B 0 B / +0.0% 73165 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 141248 B 0 B / +0.0% 92630 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 152848 B 0 B / +0.0% 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 153560 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3213561 B 0 B / +0.0% 132442 B 0 B / +0.0%
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3192660 B 0 B / +0.0% 101635 B 0 B / +0.0%
fmtprintf/j64-emscripten-memory64/LLGo 2950168 B 0 B / +0.0% 139285 B 0 B / +0.0%
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2347479 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2344141 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-emscripten/LLGo 153428 B 0 B / +0.0% 88742 B 0 B / +0.0%
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 151768 B 0 B / +0.0% 73165 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 140582 B 0 B / +0.0% 92630 B 0 B / +0.0%
reflectcall/j32-emscripten/LLGo 1543898 B 0 B / +0.0% 105908 B 0 B / +0.0%
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1546371 B 0 B / +0.0% 90331 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1428570 B 0 B / +0.0% 111641 B 0 B / +0.0%
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1276821 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1274429 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 152495 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-wasi/LLGo 153207 B 0 B / +0.0% 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 6.048 s -159.5 ms / -2.6% (better)
j32-goos-js 6.093 s +121.3 ms / +2.0% (worse)
j64-emscripten-memory64 5.235 s -35.62 ms / -0.7% (better)
reflectcall/w32-wasi 23.104 s -6.55 ms / -0.02834% (better)
w32-goos-wasip1 4.249 s -93.19 ms / -2.1% (better)
w32-wasi 4.012 s -117.6 ms / -2.8% (better)

Compared with 783ec4fd3d57 measured in the same runner job.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant