Skip to content

abi: lower multi-element array copies like cmd/compile CanSSA - #2692

Open
visualfc wants to merge 7 commits into
xgo-dev:mainfrom
visualfc:fix/can-ssa-array-copies
Open

visualfc wants to merge 7 commits into
xgo-dev:mainfrom
visualfc:fix/can-ssa-array-copies

Conversation

@visualfc

@visualfc visualfc commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Depends on #2690 (host 4KiB aggregate copy pass). This PR adds cmd/compile-style handling for multi-element arrays larger than the CanSSA size limit.

Summary

cmd/compile CanSSA never represents arrays with more than one element as SSA values once they exceed 4 pointer words (32 bytes on 64-bit). Copies stay in memory (OpMove / LoweredMoveLoop). LLGo emitted first-class load [N x T] / store [N x T], and LLVM default<Os> then scalarized them (SLP + ISel).

crypto/internal/fips140/mldsa is the sharp case: ringElement / nttElement are [256]uint32 (1KiB), under the 4KiB copy threshold. ntt starts with:

%v = load [256 x i32], ptr %src
store [256 x i32] %v, ptr %tmp

After C ABI lowering, this PR rewrites those array copies to memmove, same pass as #2690, gated on size > 4*PtrSize. Smaller arrays such as [2]int64 stay first-class so they can remain in registers. The frontend safepoint predicate uses the same ShouldSnapshotAggregateLoad helper as the backend.

C ABI is unchanged:

  • Function signatures, sret, and MaxImplicitStackVarSize (64KiB) are untouched
  • A load used as a call argument is left as a first-class value
  • LowerLargeAggregates (pre-C-ABI sret) does not rewrite 1KiB array copies

Performance (-Os backend, isolated package)

package main #2690 this PR
crypto/internal/fips140/mldsa 6.06s 5.98s 1.35s
crypto/internal/fips140/mlkem 8.68s 3.77s 1.34s
golang.org/x/tools/go/ast/edge 3.82s 0.86s 0.86s
compress/flate 1.41s 0.85s 0.73s

go/ast/edge and most of compress/flate are already covered by the 4KiB copy pass in #2690. This PR is the remaining 1KiB polynomial copies in mldsa / mlkem.

Test plan

  • go test ./internal/abi -count=1 -run 'ShouldLower|MultiElement|AggregateCopies|LargeAggregateThreshold'
  • go test ./cl -count=1 -run TestCompileLargeSnapshotGCRoots
  • go test ./internal/cabi -count=1
  • [256 x i32] and [5 x i64] → memmove; [2 x i64] / [1 x i64] unchanged; call-argument arrays unchanged
  • 40-byte array multi-store snapshot publishes AllocU GC roots

@codecov

codecov Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 94.28571% with 8 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
internal/abi/large.go 93.70% 8 Missing ⚠️

📢 Thoughts on this report? Let us know!

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: extend aggregate-copy lowering to all architectures

This PR makes LowerWasmAggregateCopies run on every target (not just wasm) and adds a copyMultiElementArrays path that lowers any array with more than one element, regardless of size. The core motivation (LLVM scalarizes first-class loads of multi-element arrays like NTT polynomials) is sound and the memmove/memcpy selection is correct.

One correctness concern stands out and is worth confirming before merge: the frontend GC-safepoint predicate in cl/gcroot.go and the ABI lowering predicate in internal/abi/large.go now disagree about which loads become heap-allocating snapshots. See the inline comments.

Summary of findings

  • P1 (correctness): isGCSafepoint still uses a size-only threshold (≥4 KiB) while isLargeCopy now snapshots multi-element arrays of any size — the snapshot path heap-allocates via runtime.AllocU (a new safepoint) that the root plan may not account for.
  • P2 (performance): small fixed-size arrays (e.g. [2]int, [4]byte) that previously stayed in registers on native targets are now forced through memmove or, in the multi-store case, a heap allocation.
  • P3 (clarity): Wasm-prefixed names/filename now describe an architecture-independent pass; dead goarch parameter in wasm_copies.go.
  • Tests: the new multi-element-array cases only exercise the single-adjacent-store memmove fast path (no GCRoots, no AllocU); the small-array + safepoint interaction is untested.

Comment thread cl/gcroot.go Outdated
size := p.prog.SizeOf(p.type_(load.Type(), llssa.InGo))
if size > llabi.MaxImplicitStackVarSize ||
(p.prog.Target().GOARCH == "wasm" && size >= llabi.MinWasmAggregateCopySize) {
if size > llabi.MaxImplicitStackVarSize || size >= llabi.MinWasmAggregateCopySize {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 — safepoint predicate no longer mirrors the lowering predicate.

This condition treats an aggregate load as a safepoint only when size > MaxImplicitStackVarSize (64 KiB) or size >= MinWasmAggregateCopySize (4 KiB) — i.e. effectively size >= 4 KiB.

But isLargeCopy in internal/abi/large.go now returns true for any multi-element array via the new copyMultiElementArrays path, with no size floor. When such a small array (e.g. [2]*T = 16 bytes) has multiple stores/extracts, transformStoredLoad takes the snapshot path and calls allocResult → runtime.AllocU, a real heap allocation and therefore a new GC safepoint that was absent from the frontend root plan. The pass roots its own snapshot and the copy source, but unrelated Go pointers live only across that newly-inserted alloc get no root slot here, since functionHasGCSafepoint/isGCSafepoint won't flag a sub-4 KiB array. A collection during AllocU could then reclaim a still-referenced object (use-after-free; collector is non-moving so it's premature-free rather than relocation).

The comment just above states the intent is to "account for that added allocation now," so this branch should mirror isMultiElementArray: also return true for *types.Array with length > 1 regardless of size. Consider factoring a single shared predicate so the two passes cannot drift again.

Comment thread internal/abi/large.go Outdated
}

func (l largeAggregateLowerer) isMultiElementArray(typ llvm.Type) bool {
return typ.TypeKind() == llvm.ArrayTypeKind && typ.ArrayLength() > 1

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This predicate has no size floor, so [2 x i64] (16 bytes) and similar tiny arrays now qualify as "large copies" on every architecture. Two consequences:

  1. Correctness: in the multi-store/extract case these go through the allocating snapshot path (AllocU), which the size-only safepoint check in cl/gcroot.go does not recognize — see the inline comment there.
  2. Performance: small fixed-size arrays that previously stayed as register-resident first-class SSA values on native targets are now forced through memmove (single-store) or a heap allocation (multi-store), defeating SROA/value-forwarding with no scalarization benefit for arrays that fit in a few registers.

Consider gating copyMultiElementArrays on a minimum size (or a small register-width threshold) rather than ArrayLength() > 1, which preserves the fix for large NTT-style arrays while leaving small arrays in registers.

Comment thread internal/build/wasm_copies.go Outdated
if goarch != "wasm" {
return 0
}
func lowerWasmAggregateCopies(_ string, td llvm.TargetData, mod llvm.Module, config abi.AggregateLoweringConfig) int {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3 — dead parameter. With the goarch != "wasm" early return removed, the first parameter is now unused (_ string) yet both call sites still pass ctx.buildConf.Goarch. It misleadingly suggests the pass still branches on architecture. Consider dropping the parameter and updating call sites.

Relatedly, the Wasm-prefixed names (lowerWasmAggregateCopies, LowerWasmAggregateCopies, MinWasmAggregateCopySize, and the wasm_copies.go filename) now describe an architecture-independent pass. Renaming to arch-neutral names — or at least a doc comment noting the pass now runs on all targets and "Wasm" is retained for historical reasons — would prevent a future reader from reintroducing a wasm guard. (The Wasm bool config field is still legitimately wasm-specific.)

@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

875c1e699147 | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 146973 B 0 B / +0.0% 74890 B 0 B / +0.0%
cprintf/j32-goos-js/LLGo 145407 B 0 B / +0.0% 73165 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 134516 B 0 B / +0.0% 78778 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 141956 B 0 B / +0.0% 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 141704 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3203377 B 0 B / +0.0% 118590 B 0 B / +0.0%
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3182480 B -296 B / -0.0093% (better) 101635 B 0 B / +0.0%
fmtprintf/j64-emscripten-memory64/LLGo 2940319 B 0 B / +0.0% 125433 B 0 B / +0.0%
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2835788 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2702221 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-emscripten/LLGo 146205 B 0 B / +0.0% 74890 B 0 B / +0.0%
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 144876 B 0 B / +0.0% 73165 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 133847 B 0 B / +0.0% 78778 B 0 B / +0.0%
reflectcall/j32-emscripten/LLGo 1534085 B 0 B / +0.0% 92056 B 0 B / +0.0%
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1536830 B 0 B / +0.0% 90331 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1419257 B 0 B / +0.0% 97789 B 0 B / +0.0%
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1540668 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1466211 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 141170 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-wasi/LLGo 140989 B 0 B / +0.0% 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 5.720 s +66.34 ms / +1.2% (worse)
j32-goos-js 5.890 s +96.77 ms / +1.7% (worse)
j64-emscripten-memory64 5.043 s +125.1 ms / +2.5% (worse)
reflectcall/w32-wasi 25.119 s -228.6 ms / -0.9% (better)
w32-goos-wasip1 4.512 s +4.88 ms / +0.1% (worse)
w32-wasi 4.482 s +111.2 ms / +2.5% (worse)

Compared with 07a0059ce2e6 measured in the same runner job.

@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

875c1e699147 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 7160 B 0 B / +0.0% 387 B 0 B / +0.0% 526.446 ms -6.137 ms / -1.2% (better) 1.308 ms -31.37 us / -2.3% (better)
Linux cprintf-lto 6912 B 0 B / +0.0% 368 B 0 B / +0.0% 531.839 ms +6.808 ms / +1.3% (worse) 1.258 ms -46.69 us / -3.6% (better)
Linux fmtprintf 1662272 B 0 B / +0.0% 498359 B 0 B / +0.0% 3.683 s +55.44 ms / +1.5% (worse) 2.979 ms -198.9 us / -6.3% (better)
Linux fmtprintf-lto 1500384 B 0 B / +0.0% 436113 B 0 B / +0.0% 10.673 s -71.39 ms / -0.7% (better) 2.859 ms -34.19 us / -1.2% (better)
Linux println 68992 B 0 B / +0.0% 16783 B 0 B / +0.0% 543.962 ms +7.968 ms / +1.5% (worse) 1.594 ms +25.64 us / +1.6% (worse)
Linux println-lto 59848 B 0 B / +0.0% 14199 B 0 B / +0.0% 791.661 ms +6.219 ms / +0.8% (worse) 1.612 ms +19.34 us / +1.2% (worse)
macOS cprintf 68064 B 0 B / +0.0% 4429 B 0 B / +0.0% 722.508 ms -179.9 ms / -19.9% (better) 2.474 ms -524.3 us / -17.5% (better)
macOS cprintf-lto 84288 B 0 B / +0.0% 193 B 0 B / +0.0% 854.481 ms -56.48 ms / -6.2% (better) 2.667 ms -2.827 ms / -51.5% (better)
macOS fmtprintf 1504672 B 0 B / +0.0% 874572 B 0 B / +0.0% 3.418 s -854.7 ms / -20.0% (better) 4.922 ms -1.522 ms / -23.6% (better)
macOS fmtprintf-lto 1192704 B 0 B / +0.0% 848216 B 0 B / +0.0% 9.490 s +264 ms / +2.9% (worse) 4.770 ms +432.1 us / +10.0% (worse)
macOS println 117136 B 0 B / +0.0% 37421 B 0 B / +0.0% 850.748 ms -166.6 ms / -16.4% (better) 4.677 ms +574.8 us / +14.0% (worse)
macOS println-lto 119472 B 0 B / +0.0% 34824 B 0 B / +0.0% 1.117 s -111.1 ms / -9.0% (better) 4.784 ms -503.4 us / -9.5% (better)
Windows MinGW cprintf 19456 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.285 s +4.246 ms / +0.3% (worse) 3.404 ms -158.2 us / -4.4% (better)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.316 s +13.48 ms / +1.0% (worse) 3.432 ms -73.6 us / -2.1% (better)
Windows MinGW fmtprintf 1932800 B 0 B / +0.0% 598950 B 0 B / +0.0% 4.108 s -40.97 ms / -1.0% (better) 8.666 ms +572 us / +7.1% (worse)
Windows MinGW fmtprintf-lto 1957376 B 0 B / +0.0% 547318 B 0 B / +0.0% 10.373 s +284.8 ms / +2.8% (worse) 7.815 ms -925.8 us / -10.6% (better)
Windows MinGW println 75776 B 0 B / +0.0% 25142 B 0 B / +0.0% 1.293 s -14.29 ms / -1.1% (better) 7.157 ms +607.9 us / +9.3% (worse)
Windows MinGW println-lto 69120 B 0 B / +0.0% 21990 B 0 B / +0.0% 1.530 s +11.57 ms / +0.8% (worse) 6.924 ms +241.5 us / +3.6% (worse)
Windows MinGW 386 cprintf 43520 B 0 B / +0.0% 5326 B 0 B / +0.0% 1.219 s -2.26 ms / -0.2% (better) 5.152 ms -307.5 us / -5.6% (better)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.250 s -11.99 ms / -0.9% (better) 5.072 ms -389.2 us / -7.1% (better)
Windows MinGW 386 fmtprintf 1896448 B 0 B / +0.0% 472414 B 0 B / +0.0% 4.115 s +15.56 ms / +0.4% (worse) 12.054 ms +1.66 ms / +16.0% (worse)
Windows MinGW 386 fmtprintf-lto 2180608 B 0 B / +0.0% 451258 B 0 B / +0.0% 10.346 s +842.5 ms / +8.9% (worse) 10.629 ms +220 us / +2.1% (worse)
Windows MinGW 386 println 96256 B 0 B / +0.0% 21458 B 0 B / +0.0% 1.265 s +29.11 ms / +2.4% (worse) 8.666 ms -423.9 us / -4.7% (better)
Windows MinGW 386 println-lto 74240 B 0 B / +0.0% 19314 B 0 B / +0.0% 1.482 s +9.261 ms / +0.6% (worse) 8.893 ms -2.074 ms / -18.9% (better)
Windows MinGW ARM64 cprintf 18944 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.610 s -22.46 ms / -1.4% (better) 7.153 ms +180.5 us / +2.6% (worse)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.653 s +827.9 us / +0.1% (worse) 7.018 ms +152.6 us / +2.2% (worse)
Windows MinGW ARM64 fmtprintf 1819136 B 0 B / +0.0% 509608 B 0 B / +0.0% 4.285 s -175.5 ms / -3.9% (better) 13.360 ms -734.2 us / -5.2% (better)
Windows MinGW ARM64 fmtprintf-lto 1879040 B 0 B / +0.0% 476152 B 0 B / +0.0% 9.646 s -334.9 ms / -3.4% (better) 12.464 ms -1.361 ms / -9.8% (better)
Windows MinGW ARM64 println 72192 B 0 B / +0.0% 23876 B 0 B / +0.0% 1.625 s -14.95 ms / -0.9% (better) 11.851 ms -637.7 us / -5.1% (better)
Windows MinGW ARM64 println-lto 68608 B 0 B / +0.0% 21224 B 0 B / +0.0% 1.810 s -60.68 ms / -3.2% (better) 10.971 ms -1.097 ms / -9.1% (better)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65798 B 0 B / +0.0% 1.146 s -8.612 ms / -0.7% (better) 3.774 ms -36.1 us / -0.9% (better)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65734 B 0 B / +0.0% 1.169 s -14.15 ms / -1.2% (better) 4.051 ms -144.7 us / -3.4% (better)
Windows MSVC fmtprintf 1643008 B 0 B / +0.0% 694502 B 0 B / +0.0% 3.993 s -82.95 ms / -2.0% (better) 9.562 ms -437 us / -4.4% (better)
Windows MSVC fmtprintf-lto 1634816 B 0 B / +0.0% 647062 B 0 B / +0.0% 9.266 s -236.3 ms / -2.5% (better) 10.256 ms -680.5 us / -6.2% (better)
Windows MSVC println 194560 B 0 B / +0.0% 120822 B 0 B / +0.0% 1.159 s -9.044 ms / -0.8% (better) 7.992 ms -427.8 us / -5.1% (better)
Windows MSVC println-lto 192512 B 0 B / +0.0% 118342 B 0 B / +0.0% 1.412 s +27.81 ms / +2.0% (worse) 10.507 ms +2.559 ms / +32.2% (worse)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 1.278 s +222.9 ms / +21.1% (worse) 5.738 ms +208.9 us / +3.8% (worse)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.150 s -142.7 ms / -11.0% (better) 5.832 ms +615 us / +11.8% (worse)
Windows MSVC 386 fmtprintf 1204224 B 0 B / +0.0% 455804 B 0 B / +0.0% 3.785 s +4.897 ms / +0.1% (worse) 11.313 ms -61.7 us / -0.5% (better)
Windows MSVC 386 fmtprintf-lto 1241088 B 0 B / +0.0% 427195 B 0 B / +0.0% 8.793 s +247.2 ms / +2.9% (worse) 11.657 ms +272.4 us / +2.4% (worse)
Windows MSVC 386 println 36352 B 0 B / +0.0% 20324 B 0 B / +0.0% 1.124 s +61.98 ms / +5.8% (worse) 9.218 ms +197.2 us / +2.2% (worse)
Windows MSVC 386 println-lto 35840 B 0 B / +0.0% 18549 B 0 B / +0.0% 1.479 s +202.9 ms / +15.9% (worse) 10.209 ms -1.868 ms / -15.5% (better)
Windows MSVC ARM64 cprintf 11776 B 0 B / +0.0% 4192 B 0 B / +0.0% 1.229 s +5.941 ms / +0.5% (worse) 6.692 ms -238.5 us / -3.4% (better)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 4084 B 0 B / +0.0% 1.224 s -17.23 ms / -1.4% (better) 6.621 ms -219.1 us / -3.2% (better)
Windows MSVC ARM64 fmtprintf 1386496 B 0 B / +0.0% 509544 B 0 B / +0.0% 3.809 s +75.05 ms / +2.0% (worse) 13.856 ms -503.8 us / -3.5% (better)
Windows MSVC ARM64 fmtprintf-lto 1404928 B 0 B / +0.0% 476820 B 0 B / +0.0% 8.599 s -41.37 ms / -0.5% (better) 14.383 ms +266.4 us / +1.9% (worse)
Windows MSVC ARM64 println 45056 B 0 B / +0.0% 23908 B 0 B / +0.0% 1.226 s -1.523 ms / -0.1% (better) 12.850 ms +1.295 ms / +11.2% (worse)
Windows MSVC ARM64 println-lto 42496 B 0 B / +0.0% 21380 B 0 B / +0.0% 1.404 s -1.406 ms / -0.1% (better) 11.934 ms -11.4 us / -0.1% (better)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.640 ns/op +0.16 ns/op / +1.1% (worse)
Linux BenchmarkMergeCompilerFlags 196.400 ns/op +3.5 ns/op / +1.8% (worse)
Linux BenchmarkMergeLinkerFlags 142.400 ns/op +3.7 ns/op / +2.7% (worse)
Linux BenchmarkChannelBuffered 55.070 ns/op +0.09 ns/op / +0.2% (worse)
Linux BenchmarkChannelHandoff 14993 ns/op +1962 ns/op / +15.1% (worse)
Linux BenchmarkDefer 51.310 ns/op +2.38 ns/op / +4.9% (worse)
Linux BenchmarkDirectCall 1.579 ns/op +0.026 ns/op / +1.7% (worse)
Linux BenchmarkGlobalRead 1.557 ns/op +0.391 ns/op / +33.5% (worse)
Linux BenchmarkGlobalWrite 7.758 ns/op -0.002 ns/op / -0.02577% (better)
Linux BenchmarkGoroutine 23852 ns/op -11147 ns/op / -31.8% (better)
Linux BenchmarkInterfaceCall 5.871 ns/op +0.029 ns/op / +0.5% (worse)
Linux BenchmarkRuntimeGetG 2.719 ns/op -0.265 ns/op / -8.9% (better)
macOS BenchmarkLookupPCRandom 21.850 ns/op +0.55 ns/op / +2.6% (worse)
macOS BenchmarkMergeCompilerFlags 212.700 ns/op +14.1 ns/op / +7.1% (worse)
macOS BenchmarkMergeLinkerFlags 158.100 ns/op +38.2 ns/op / +31.9% (worse)
macOS BenchmarkChannelBuffered 33.580 ns/op +2.57 ns/op / +8.3% (worse)
macOS BenchmarkChannelHandoff 8164 ns/op -3598 ns/op / -30.6% (better)
macOS BenchmarkDefer 39.910 ns/op +1.34 ns/op / +3.5% (worse)
macOS BenchmarkDirectCall 1.192 ns/op +0.088 ns/op / +8.0% (worse)
macOS BenchmarkGlobalRead 1.104 ns/op +0.02 ns/op / +1.8% (worse)
macOS BenchmarkGlobalWrite 1.189 ns/op +0.111 ns/op / +10.3% (worse)
macOS BenchmarkGoroutine 51971 ns/op +5324 ns/op / +11.4% (worse)
macOS BenchmarkInterfaceCall 4.648 ns/op +0.611 ns/op / +15.1% (worse)
macOS BenchmarkRuntimeGetG 2.502 ns/op -0.145 ns/op / -5.5% (better)
Windows MinGW BenchmarkLookupPCRandom 12.970 ns/op -0.21 ns/op / -1.6% (better)
Windows MinGW BenchmarkMergeCompilerFlags 659.100 ns/op +6.4 ns/op / +1.0% (worse)
Windows MinGW BenchmarkMergeLinkerFlags 590.600 ns/op +21.2 ns/op / +3.7% (worse)
Windows MinGW BenchmarkChannelBuffered 29.190 ns/op -1.05 ns/op / -3.5% (better)
Windows MinGW BenchmarkChannelHandoff 943.200 ns/op -60.8 ns/op / -6.1% (better)
Windows MinGW BenchmarkDefer 56.450 ns/op -1.37 ns/op / -2.4% (better)
Windows MinGW BenchmarkDirectCall 1.549 ns/op -0.002 ns/op / -0.1% (better)
Windows MinGW BenchmarkGlobalRead 1.546 ns/op -0.31 ns/op / -16.7% (better)
Windows MinGW BenchmarkGlobalWrite 2.471 ns/op +0.014 ns/op / +0.6% (worse)
Windows MinGW BenchmarkGoroutine 88972 ns/op -5763 ns/op / -6.1% (better)
Windows MinGW BenchmarkInterfaceCall 8.365 ns/op -0.014 ns/op / -0.2% (better)
Windows MinGW BenchmarkRuntimeGetG 2.169 ns/op +0.002 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkLookupPCRandom 26.520 ns/op 0 ns/op / +0.0%
Windows MinGW 386 BenchmarkMergeCompilerFlags 768.300 ns/op +29.7 ns/op / +4.0% (worse)
Windows MinGW 386 BenchmarkMergeLinkerFlags 709.600 ns/op +42.9 ns/op / +6.4% (worse)
Windows MinGW 386 BenchmarkChannelBuffered 39.540 ns/op +0.34 ns/op / +0.9% (worse)
Windows MinGW 386 BenchmarkChannelHandoff 804.600 ns/op -30.1 ns/op / -3.6% (better)
Windows MinGW 386 BenchmarkDefer 43.220 ns/op -0.61 ns/op / -1.4% (better)
Windows MinGW 386 BenchmarkDirectCall 1.548 ns/op 0 ns/op / +0.0%
Windows MinGW 386 BenchmarkGlobalRead 1.862 ns/op +0.312 ns/op / +20.1% (worse)
Windows MinGW 386 BenchmarkGlobalWrite 7.777 ns/op +0.007 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkGoroutine 106732 ns/op +2773 ns/op / +2.7% (worse)
Windows MinGW 386 BenchmarkInterfaceCall 8.365 ns/op -0.007 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 2.167 ns/op -0.005 ns/op / -0.2% (better)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.150 ns/op +0.04 ns/op / +0.3% (worse)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 581.200 ns/op -1.1 ns/op / -0.2% (better)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 556.100 ns/op +11.7 ns/op / +2.1% (worse)
Windows MinGW ARM64 BenchmarkChannelBuffered 38.460 ns/op +1.06 ns/op / +2.8% (worse)
Windows MinGW ARM64 BenchmarkChannelHandoff 2609 ns/op -911 ns/op / -25.9% (better)
Windows MinGW ARM64 BenchmarkDefer 53.010 ns/op -4.15 ns/op / -7.3% (better)
Windows MinGW ARM64 BenchmarkDirectCall 0.589 ns/op -0.0008 ns/op / -0.1% (better)
Windows MinGW ARM64 BenchmarkGlobalRead 0.884 ns/op +0.2209 ns/op / +33.3% (worse)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.590 ns/op -0.0736 ns/op / -11.1% (better)
Windows MinGW ARM64 BenchmarkGoroutine 64062 ns/op -2401 ns/op / -3.6% (better)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.204 ns/op +0.064 ns/op / +1.5% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.770 ns/op +0.001 ns/op / +0.1% (worse)
Windows MSVC BenchmarkLookupPCRandom 13.210 ns/op +0.04 ns/op / +0.3% (worse)
Windows MSVC BenchmarkMergeCompilerFlags 619.300 ns/op -3 ns/op / -0.5% (better)
Windows MSVC BenchmarkMergeLinkerFlags 544.800 ns/op -9.8 ns/op / -1.8% (better)
Windows MSVC BenchmarkChannelBuffered 31.360 ns/op -0.75 ns/op / -2.3% (better)
Windows MSVC BenchmarkChannelHandoff 1108 ns/op +24 ns/op / +2.2% (worse)
Windows MSVC BenchmarkDefer 55.090 ns/op -0.04 ns/op / -0.1% (better)
Windows MSVC BenchmarkDirectCall 1.550 ns/op -0.003 ns/op / -0.2% (better)
Windows MSVC BenchmarkGlobalRead 1.856 ns/op +0.308 ns/op / +19.9% (worse)
Windows MSVC BenchmarkGlobalWrite 2.460 ns/op -0.012 ns/op / -0.5% (better)
Windows MSVC BenchmarkGoroutine 88941 ns/op -74 ns/op / -0.1% (better)
Windows MSVC BenchmarkInterfaceCall 8.373 ns/op -0.007 ns/op / -0.1% (better)
Windows MSVC BenchmarkRuntimeGetG 1.862 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkLookupPCRandom 26.610 ns/op +0.09 ns/op / +0.3% (worse)
Windows MSVC 386 BenchmarkMergeCompilerFlags 711.500 ns/op -38.5 ns/op / -5.1% (better)
Windows MSVC 386 BenchmarkMergeLinkerFlags 696.200 ns/op +20.5 ns/op / +3.0% (worse)
Windows MSVC 386 BenchmarkChannelBuffered 39.440 ns/op -5.37 ns/op / -12.0% (better)
Windows MSVC 386 BenchmarkChannelHandoff 862.900 ns/op -37.7 ns/op / -4.2% (better)
Windows MSVC 386 BenchmarkDefer 46.520 ns/op -1 ns/op / -2.1% (better)
Windows MSVC 386 BenchmarkDirectCall 1.548 ns/op +0.002 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkGlobalRead 1.550 ns/op -0.31 ns/op / -16.7% (better)
Windows MSVC 386 BenchmarkGlobalWrite 7.771 ns/op -0.023 ns/op / -0.3% (better)
Windows MSVC 386 BenchmarkGoroutine 109623 ns/op -899 ns/op / -0.8% (better)
Windows MSVC 386 BenchmarkInterfaceCall 8.358 ns/op +0.275 ns/op / +3.4% (worse)
Windows MSVC 386 BenchmarkRuntimeGetG 2.477 ns/op -0.004 ns/op / -0.2% (better)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.060 ns/op -0.02 ns/op / -0.2% (better)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 548.800 ns/op -23.4 ns/op / -4.1% (better)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 515.600 ns/op -14.6 ns/op / -2.8% (better)
Windows MSVC ARM64 BenchmarkChannelBuffered 37.440 ns/op -1.41 ns/op / -3.6% (better)
Windows MSVC ARM64 BenchmarkChannelHandoff 1817 ns/op +151 ns/op / +9.1% (worse)
Windows MSVC ARM64 BenchmarkDefer 62.380 ns/op +1.66 ns/op / +2.7% (worse)
Windows MSVC ARM64 BenchmarkDirectCall 0.590 ns/op -0.0001 ns/op / -0.01696% (better)
Windows MSVC ARM64 BenchmarkGlobalRead 0.885 ns/op +0.2214 ns/op / +33.4% (worse)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.753 ns/op -0.008 ns/op / -0.2% (better)
Windows MSVC ARM64 BenchmarkGoroutine 55306 ns/op -127 ns/op / -0.2% (better)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.145 ns/op +0.007 ns/op / +0.2% (worse)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.803 ns/op +0.034 ns/op / +1.9% (worse)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 906 ns/op +0.9 ns/op / +0.1% (worse)
Linux AfterFuncZeroDelivery/LLGo 38776 ns/op -8672 ns/op / -18.3% (better)
Linux CreateStop/Go 290.900 ns/op +1.7 ns/op / +0.6% (worse)
Linux CreateStop/LLGo 1794 ns/op -66 ns/op / -3.5% (better)
Linux RearmStopped/Go 114.600 ns/op -1.3 ns/op / -1.1% (better)
Linux RearmStopped/LLGo 1650 ns/op +357 ns/op / +27.6% (worse)
Linux ResetActive/Go 67.530 ns/op -1.17 ns/op / -1.7% (better)
Linux ResetActive/LLGo 760.400 ns/op +123.1 ns/op / +19.3% (worse)
Linux ResetHeap1024/Go 67.100 ns/op 0 ns/op / +0.0%
Linux ResetHeap1024/LLGo 177.200 ns/op +2.2 ns/op / +1.3% (worse)
macOS AfterFuncZeroDelivery/Go 633.900 ns/op -48.9 ns/op / -7.2% (better)
macOS AfterFuncZeroDelivery/LLGo 102436 ns/op +2681 ns/op / +2.7% (worse)
macOS CreateStop/Go 271.600 ns/op +64.2 ns/op / +31.0% (worse)
macOS CreateStop/LLGo 814.500 ns/op -336.5 ns/op / -29.2% (better)
macOS RearmStopped/Go 90.690 ns/op +8.43 ns/op / +10.2% (worse)
macOS RearmStopped/LLGo 378.200 ns/op +3.3 ns/op / +0.9% (worse)
macOS ResetActive/Go 69.880 ns/op +6.45 ns/op / +10.2% (worse)
macOS ResetActive/LLGo 170.600 ns/op -1.5 ns/op / -0.9% (better)
macOS ResetHeap1024/Go 64.490 ns/op -1.13 ns/op / -1.7% (better)
macOS ResetHeap1024/LLGo 103.800 ns/op +8.03 ns/op / +8.4% (worse)
Windows MinGW AfterFuncZeroDelivery/Go 587.400 ns/op +18 ns/op / +3.2% (worse)
Windows MinGW AfterFuncZeroDelivery/LLGo 192516 ns/op +6439 ns/op / +3.5% (worse)
Windows MinGW CreateStop/Go 114.800 ns/op -0.4 ns/op / -0.3% (better)
Windows MinGW CreateStop/LLGo 587.400 ns/op +155.1 ns/op / +35.9% (worse)
Windows MinGW RearmStopped/Go 33.960 ns/op +2.38 ns/op / +7.5% (worse)
Windows MinGW RearmStopped/LLGo 281 ns/op +10 ns/op / +3.7% (worse)
Windows MinGW ResetActive/Go 20.040 ns/op +0.02 ns/op / +0.1% (worse)
Windows MinGW ResetActive/LLGo 165.100 ns/op +7.8 ns/op / +5.0% (worse)
Windows MinGW ResetHeap1024/Go 22.550 ns/op +2.2 ns/op / +10.8% (worse)
Windows MinGW ResetHeap1024/LLGo 124.800 ns/op +2.1 ns/op / +1.7% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/Go 965.500 ns/op +12.4 ns/op / +1.3% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 197953 ns/op +781 ns/op / +0.4% (worse)
Windows MinGW 386 CreateStop/Go 193.600 ns/op -3.9 ns/op / -2.0% (better)
Windows MinGW 386 CreateStop/LLGo 518.600 ns/op +33.1 ns/op / +6.8% (worse)
Windows MinGW 386 RearmStopped/Go 63.390 ns/op -0.1 ns/op / -0.2% (better)
Windows MinGW 386 RearmStopped/LLGo 348.200 ns/op +8.2 ns/op / +2.4% (worse)
Windows MinGW 386 ResetActive/Go 38.910 ns/op -0.18 ns/op / -0.5% (better)
Windows MinGW 386 ResetActive/LLGo 593 ns/op -366.8 ns/op / -38.2% (better)
Windows MinGW 386 ResetHeap1024/Go 39.480 ns/op +0.08 ns/op / +0.2% (worse)
Windows MinGW 386 ResetHeap1024/LLGo 187.500 ns/op -1.5 ns/op / -0.8% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 680 ns/op +18.5 ns/op / +2.8% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 149079 ns/op -7898 ns/op / -5.0% (better)
Windows MinGW ARM64 CreateStop/Go 202.700 ns/op -11.8 ns/op / -5.5% (better)
Windows MinGW ARM64 CreateStop/LLGo 360.600 ns/op -2.6 ns/op / -0.7% (better)
Windows MinGW ARM64 RearmStopped/Go 70.550 ns/op -0.02 ns/op / -0.02834% (better)
Windows MinGW ARM64 RearmStopped/LLGo 251.100 ns/op -3.5 ns/op / -1.4% (better)
Windows MinGW ARM64 ResetActive/Go 30.990 ns/op -0.07 ns/op / -0.2% (better)
Windows MinGW ARM64 ResetActive/LLGo 121.200 ns/op +1.4 ns/op / +1.2% (worse)
Windows MinGW ARM64 ResetHeap1024/Go 31.070 ns/op +0.07 ns/op / +0.2% (worse)
Windows MinGW ARM64 ResetHeap1024/LLGo 127.600 ns/op +0.3 ns/op / +0.2% (worse)
Windows MSVC AfterFuncZeroDelivery/Go 547.800 ns/op -28.2 ns/op / -4.9% (better)
Windows MSVC AfterFuncZeroDelivery/LLGo 174229 ns/op -136 ns/op / -0.1% (better)
Windows MSVC CreateStop/Go 117.400 ns/op +2.2 ns/op / +1.9% (worse)
Windows MSVC CreateStop/LLGo 426.500 ns/op -24.2 ns/op / -5.4% (better)
Windows MSVC RearmStopped/Go 31.530 ns/op +0.29 ns/op / +0.9% (worse)
Windows MSVC RearmStopped/LLGo 261.300 ns/op -9.2 ns/op / -3.4% (better)
Windows MSVC ResetActive/Go 20.060 ns/op -0.08 ns/op / -0.4% (better)
Windows MSVC ResetActive/LLGo 160.200 ns/op +11.3 ns/op / +7.6% (worse)
Windows MSVC ResetHeap1024/Go 20.660 ns/op +0.25 ns/op / +1.2% (worse)
Windows MSVC ResetHeap1024/LLGo 124.400 ns/op -4.5 ns/op / -3.5% (better)
Windows MSVC 386 AfterFuncZeroDelivery/Go 938.300 ns/op -11 ns/op / -1.2% (better)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 206404 ns/op +7061 ns/op / +3.5% (worse)
Windows MSVC 386 CreateStop/Go 191.100 ns/op +0.2 ns/op / +0.1% (worse)
Windows MSVC 386 CreateStop/LLGo 458.100 ns/op +11.2 ns/op / +2.5% (worse)
Windows MSVC 386 RearmStopped/Go 63.370 ns/op -0.07 ns/op / -0.1% (better)
Windows MSVC 386 RearmStopped/LLGo 326.300 ns/op +3.2 ns/op / +1.0% (worse)
Windows MSVC 386 ResetActive/Go 39 ns/op +0.01 ns/op / +0.02565% (worse)
Windows MSVC 386 ResetActive/LLGo 940.300 ns/op -35.2 ns/op / -3.6% (better)
Windows MSVC 386 ResetHeap1024/Go 39.330 ns/op -0.18 ns/op / -0.5% (better)
Windows MSVC 386 ResetHeap1024/LLGo 169.900 ns/op -2 ns/op / -1.2% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 670.800 ns/op -2.3 ns/op / -0.3% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 141165 ns/op +2908 ns/op / +2.1% (worse)
Windows MSVC ARM64 CreateStop/Go 201.800 ns/op +4.5 ns/op / +2.3% (worse)
Windows MSVC ARM64 CreateStop/LLGo 451.900 ns/op -7.6 ns/op / -1.7% (better)
Windows MSVC ARM64 RearmStopped/Go 70.630 ns/op +0.04 ns/op / +0.1% (worse)
Windows MSVC ARM64 RearmStopped/LLGo 284 ns/op +2.5 ns/op / +0.9% (worse)
Windows MSVC ARM64 ResetActive/Go 31.110 ns/op +0.01 ns/op / +0.03215% (worse)
Windows MSVC ARM64 ResetActive/LLGo 145.300 ns/op -3.9 ns/op / -2.6% (better)
Windows MSVC ARM64 ResetHeap1024/Go 31.220 ns/op +0.13 ns/op / +0.4% (worse)
Windows MSVC ARM64 ResetHeap1024/LLGo 139.300 ns/op +2.3 ns/op / +1.7% (worse)

Compared with 07a0059ce2e6 measured in the same runner job.

@visualfc
visualfc force-pushed the fix/can-ssa-array-copies branch from f1bffcb to 7c34fc2 Compare September 29, 2026 00:11
@visualfc

Copy link
Copy Markdown
Collaborator Author

Addressed the P1/P2 notes: multi-element array copies are now gated on cmd/compile's CanSSA size limit (4 pointer words, 32 bytes on 64-bit). [2]int64 stays first-class; NTT-sized arrays still lower to memmove. isGCSafepoint uses the same ShouldSnapshotAggregateLoad predicate as the backend so an AllocU snapshot cannot appear without a matching root plan.

@visualfc

Copy link
Copy Markdown
Collaborator Author

Rebased onto #2690 so AllocU snapshots inherit the debug-location fix.

LLVM default<Os> scalarizes first-class loads of mid-size array and
struct values. Packages such as go/ast/edge and crypto/internal/fips140/mlkem
spend seconds in SLP/ISel on copies a few KiB above the old host threshold.

Reuse the existing Wasm 4KiB copy pass on every target: load/store of
aggregates >= 4KiB become memmove. Return sret and the native C ABI stay
at the 64KiB MaxImplicitStackVarSize limit.
Drop the unused goarch argument now that the copy pass runs on every
target, and rename the pass to LowerAggregateCopies so the old Wasm
prefix does not imply it is wasm-only. config.Wasm still selects the
GC-root frame layout.

Simplify isGCSafepoint to size >= 4KiB. Run the same post-C-ABI copy
pass on native C export wrappers, which previously skipped it.
LowerAggregateCopies inserts runtime.AllocU for multi-use aggregate
loads. In functions with debug info, LLVM requires inlinable calls to
have a !dbg location. Copy the load/call's debug loc onto AllocU, matching
the existing memcpy/memmove handling.
cmd/compile never represents arrays with more than one element as SSA
values; copies stay in memory as OpMove. LLGo emitted first-class
load/store of types such as [256 x i32], and LLVM default<Os> then spent
seconds in SLP and AArch64 ISel on ML-DSA NTT.

After C ABI lowering, rewrite those array copies to memmove. Loads used
as call arguments are left unchanged so register passing of small arrays
stays C-compatible. Return sret and MaxImplicitStackVarSize are unchanged.
cmd/compile's CanSSA limit is 4 pointer words. Arrays larger than that
are copied in memory; smaller multi-element arrays stay first-class so
they can remain in registers and do not allocate snapshots.

Share ShouldSnapshotAggregateLoad with the frontend safepoint predicate
so a heap snapshot cannot appear without a matching GC root plan.
Multi-use copies of arrays larger than 4 pointer words were lowered with
AllocU on every target. On Wasm that inserted extra heap safepoints into
the precise collector, which stalled runtime.GC and JS value finalizers.

Keep lowering those arrays on Wasm. Snapshots smaller than 4KiB use an
entry alloca; AllocU and GC roots remain only for copies at the existing
4KiB threshold, matching the host/Wasm split that already passed CI.
@visualfc
visualfc force-pushed the fix/can-ssa-array-copies branch from c4dd95e to d028c5a Compare September 30, 2026 02:22

@cpunion cpunion left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One additional finding in the #2692-only changes; not repeating the #2690 review.

Comment thread internal/abi/large.go
} else {
b.SetInsertPointBefore(first)
}
return b.CreateAlloca(typ, "")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Allow disjoint stack snapshots to reuse storage

Each sub-4 KiB snapshot receives a separate entry-block alloca without lifetime markers, so even mutually exclusive branches retain the sum of their snapshot slots.

Reproducer: eight switch cases, each doing v := Source; Mutate(&Source, caseID); Dest = v, where Source/Dest are [256]uint32 and the non-inlined mutator modifies the array. Ordinary Go source reproduces native snapshot stack reservation growing from 1,024 to 8,192 bytes. In the corresponding LLVM fixture at -Os (LLVM 22), the ARM64 frame grows from 1,056 to 8,256 bytes, and the Wasm32 backend reserves 8,192 linear-stack bytes instead of zero on the base.

Copies remain correct, but stack usage scales with the total number of snapshots rather than their peak simultaneous lifetime, increasing fixed-stack exhaustion risk. Please retain loop-safe entry allocation while providing lifetimes/slot reuse or a frame-budget fallback, and add a mutually exclusive-branch stack-usage regression test.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed: sub-4 KiB snapshots still use a loop-safe entry alloca. Disjoint live ranges of the same type now share one slot and get llvm.lifetime.start / llvm.lifetime.end (LLVM 22 pointer-only form).

An 8-way switch of [256]uint32 copies reserves one 1 KiB slot instead of eight. Overlapping snapshots still get separate slots; a loop still keeps a single entry alloca. Covered by TestLowerMultiElementArrayCopyStackReuse.

Sub-4KiB aggregate snapshots each got a dedicated entry alloca, so
mutually exclusive branches reserved the sum of every copy. Keep the
loop-safe entry allocation, share one slot across non-overlapping live
ranges of the same type, and mark occupants with llvm.lifetime.start/end.

An 8-way switch of [256]uint32 copies now reserves 1KiB instead of 8KiB.
Overlapping snapshots still get separate slots.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants