Skip to content

ssa: emit static itabs and box constant iface data - #2704

Open
visualfc wants to merge 1 commit into
xgo-dev:mainfrom
visualfc:fix/static-iface-box
Open

visualfc wants to merge 1 commit into
xgo-dev:mainfrom
visualfc:fix/static-iface-box

Conversation

@visualfc

@visualfc visualfc commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Known concrete-to-interface (T2I) conversions emit a read-only _llgo_itab$ global ({Inter, Type, Hash, Fun[]}) with weak_odr linkage, matching cmd/compile's go:itab symbols.

Ordinary builds use that global as the runtime vtable. The T2I site is the only reference, so --gc-sections / -dead_strip can drop itabs that belong to dead functions. LTO and deadcode-drop keep calling NewItab so unused interface methods can be dropped (Fun[] would pin them). LTO still emits the static template for de-virt and erases unused templates after the plugin runs.

Non-direct interface values (integers, bools, small aggregates) stay IfaceIndir, matching Go: data is a pointer to a copy. Compile-time constants go in a WeakODR _llgo_ifacebox$ global instead of AllocU, like cmd/compile's static temps / staticuint64s.

Any LLVM-constant payload is eligible. Box names hash the constant's bits. Zero-sized constants stay AllocU so boxing does not mutate the shared sentinel.

Together, a constant T2I is {itab, box} with no runtime call or heap allocation. interequal compares the (inter, _type) pair so a static itab and a dynamically allocated itab for the same conversion compare equal.

Supersedes #2703 (static itabs are included in this commit).

Test plan

  • go test ./ssa
  • cl/_testrt/staticitab (including any(T).(I) equality)
  • cl/_testrt/staticifacebox (any(int64(42)) has no AllocU)
  • cl/_testrt/eface
  • Local Darwin check: unreferenced _llgo_ifacebox$ dropped by -dead_strip; LTO also drops them

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: static iface box / static itab for known T2I

This is a well-targeted change. MakeInterface now boxes compile-time-constant non-direct interface data into deduplicated read-only globals (_llgo_ifacebox$...) instead of AllocU+Store, and staticItab emits a runtime static itab global for known type-to-interface conversions even without LTO plugin markers. The runtime interequal fix correctly restores the invariant that a T2I value (static itab global) and an I2I value (NewItab) of the same interface compare equal, and it is memory-safe (both tabs are nil-guarded before dereference). The verified behaviors:

  • The new comments are accurate — staticItab's hash matches abiCommonFields (sha256(typeName) → little-endian uint32), and integers/bools do stay IfaceIndir.
  • interequal's fallback short-circuits on _type mismatch, so the extra ifaceeq only runs when the type pair actually matches.
  • Test renames/assertions in ssa_test.go are consistent with the new emit-always behavior.

Findings below are refinements, not blockers. The main items are two compile-time hot-path inefficiencies (expensive dedup keys computed before the cache lookup) and a couple of correctness invariants worth documenting.

Semantics note (not inline-placeable)

Boxing all constant non-direct types (structs, strings, slices, arrays, big ints) into shared GlobalConstant globals is broader than cmd/compile, which only shares addresses for the small-integer staticuint64s range and read-only static temps. Two consequences worth confirming are intended:

  • Identical (type, constant) pairs now resolve to a single shared address. Code that extracts the iface data word (via reflect internals, //go:linkname, or unsafe) and relies on distinct addresses will observe shared identity.
  • Because the box is GlobalConstant, mutating the pointee through the interface via unsafe would now fault on a read-only page, whereas the previous AllocU+Store path produced writable storage.

Neither is guaranteed by the Go spec, so this is likely acceptable — flagging so it's a conscious decision.

Comment thread ssa/interface.go Outdated
Comment on lines +134 to +136
sum := sha256.Sum256([]byte(typeName + "\x00" + x.impl.String()))
name := "_llgo_ifacebox$" + base64.RawURLEncoding.EncodeToString(sum[:])
if g := b.Pkg.VarOf(name); g != nil {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] staticIfaceBox: compute the expensive dedup key after the cache check

The dedup lookup b.Pkg.VarOf(name) runs after the key is built, so every constant interface conversion — including repeated conversions of a constant that is already boxed — pays x.impl.String() (serializes the full LLVM constant to textual IR, O(constant size)) plus a sha256 and base64 on every call. For programs that box the same large aggregate constant in N places, N-1 of those are cache hits that still do the full IR serialization + hash.

Consider a cheaper first-level key (e.g. the llvm.Value pointer or a small structural key) and only fall back to String()+sha256 on a miss, or memoize the computed name on the constant.

Comment thread ssa/interface.go
Comment on lines +49 to +52
if rawIntf.NumMethods() == 0 || concrete == nil {
return Expr{}, false
}
rawIntf = rawIntf.Complete()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] staticItab: move the VarOf cache check before NewMethodSet/AssignableTo

rawIntf.Complete(), types.AssignableTo, types.NewMethodSet(concrete) and the per-method Lookup loop all run before the VarOf(name) cache check further down. The global name derives only from intfName/typeName (via abi.TypeName), so repeated T2I conversions of the same (interface, concrete) pair rebuild the full method set on every call only to find the itab already exists. Computing the name and checking VarOf first would make repeated conversions O(1).

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two additional inline notes (continuation of the prior review).

Comment thread ssa/interface.go Outdated
}
prog := b.Prog
typeName, _ := prog.abi.TypeName(typ.raw.Type)
sum := sha256.Sum256([]byte(typeName + "\x00" + x.impl.String()))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] Fragile constant identity key

The box name hashes x.impl.String() (the LLVM IR text). For constants that reference other globals — notably ConstString, whose data pointer points at an anonymous private global — the printed form can depend on module-local numbering (@0, @1, …). Distinct byte payloads printing identically is extremely unlikely today (referenced string globals are content-addressed via p.strs), but the correctness of the ODR-merged box relies on that implicit invariant. A short comment noting the assumption, or keying on the constant's actual bytes, would make this robust against future changes to anonymous/private global naming.

Comment thread ssa/interface.go
if rawIntf.NumMethods() == 0 || concrete == nil {
return Expr{}, false
}
rawIntf = rawIntf.Complete()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] Dedup trusts abi.TypeName injectivity

Both the itab and the iface box are deduplicated by sha256(abi.TypeName(...)). If two distinct concrete types.Type values can ever map to the same abi.TypeName string, the first-emitted itab/box would be silently reused for the second, wrong type — its fun slots would point at the wrong methods, or wrong data would be shared. This is pre-existing key style, but since the itab global is now emitted in all builds (not just LTO), the blast radius is larger. Worth confirming abi.TypeName is injective for every concrete type reaching here, or folding in a stronger discriminator.

@codecov

codecov Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 96.33028% with 4 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
ssa/interface.go 96.29% 4 Missing ⚠️

📢 Thoughts on this report? Let us know!

@visualfc
visualfc force-pushed the fix/static-iface-box branch from 48eba88 to 4b3a724 Compare September 30, 2026 05:18
@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

3fed6ed0c3b2 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 8568 B 0 B / +0.0% 387 B 0 B / +0.0% 1.001 s -3.087 ms / -0.3% (better) 1.233 ms -92.85 us / -7.0% (better)
Linux cprintf-lto 8408 B 0 B / +0.0% 368 B 0 B / +0.0% 979.913 ms -15.74 ms / -1.6% (better) 1.255 ms +14.17 us / +1.1% (worse)
Linux fmtprintf 4886688 B +6752 B / +0.1% (worse) 492638 B -8675 B / -1.7% (better) 7.424 s -23.14 ms / -0.3% (better) 3.058 ms -129.9 us / -4.1% (better)
Linux fmtprintf-lto 3591984 B -6304 B / -0.2% (better) 428037 B -10120 B / -2.3% (better) 16.603 s +18.17 ms / +0.1% (worse) 2.929 ms +14.06 us / +0.5% (worse)
Linux println 675144 B +776 B / +0.1% (worse) 16649 B -206 B / -1.2% (better) 986.914 ms -40.37 ms / -3.9% (better) 1.587 ms -64.59 us / -3.9% (better)
Linux println-lto 190048 B +56 B / +0.02947% (worse) 14089 B -184 B / -1.3% (better) 1.293 s -35.31 ms / -2.7% (better) 1.612 ms -7.476 us / -0.5% (better)
macOS cprintf 50736 B 0 B / +0.0% 4409 B 0 B / +0.0% 881.082 ms -291.5 ms / -24.9% (better) 2.693 ms -1.043 ms / -27.9% (better)
macOS cprintf-lto 50496 B 0 B / +0.0% 161 B 0 B / +0.0% 891.702 ms -114.1 ms / -11.3% (better) 2.365 ms +58.58 us / +2.5% (worse)
macOS fmtprintf 1827296 B +53840 B / +3.0% (worse) 870528 B -6888 B / -0.8% (better) 5.042 s -298.9 ms / -5.6% (better) 9.751 ms +674.1 us / +7.4% (worse)
macOS fmtprintf-lto 1390144 B +29216 B / +2.1% (worse) 753824 B -8092 B / -1.1% (better) 10.449 s -3.111 s / -22.9% (better) 4.846 ms -54.75 us / -1.1% (better)
macOS println 101152 B +1808 B / +1.8% (worse) 24006 B -216 B / -0.9% (better) 907.409 ms -249.8 ms / -21.6% (better) 4.020 ms -5.282 ms / -56.8% (better)
macOS println-lto 84032 B +368 B / +0.4% (worse) 21257 B -200 B / -0.9% (better) 1.018 s -306.1 ms / -23.1% (better) 3.415 ms -640.4 us / -15.8% (better)
Windows MinGW cprintf 650240 B +512 B / +0.1% (worse) 4550 B 0 B / +0.0% 1.727 s +66.21 ms / +4.0% (worse) 3.458 ms -45.9 us / -1.3% (better)
Windows MinGW cprintf-lto 43520 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.750 s +87.14 ms / +5.2% (worse) 4.734 ms +1.207 ms / +34.2% (worse)
Windows MinGW fmtprintf 5454336 B +23040 B / +0.4% (worse) 595670 B -4336 B / -0.7% (better) 7.427 s +235.2 ms / +3.3% (worse) 8.301 ms +10.3 us / +0.1% (worse)
Windows MinGW fmtprintf-lto 4127232 B +3072 B / +0.1% (worse) 539254 B -6992 B / -1.3% (better) 14.962 s +298.1 ms / +2.0% (worse) 8.412 ms +89.6 us / +1.1% (worse)
Windows MinGW println 706560 B +1024 B / +0.1% (worse) 24998 B -192 B / -0.8% (better) 1.731 s +50.51 ms / +3.0% (worse) 6.515 ms +236.9 us / +3.8% (worse)
Windows MinGW println-lto 209408 B +512 B / +0.2% (worse) 21846 B -208 B / -0.9% (better) 1.995 s +11.88 ms / +0.6% (worse) 6.900 ms +158.8 us / +2.4% (worse)
Windows MinGW 386 cprintf 600576 B -512 B / -0.1% (better) 5326 B 0 B / +0.0% 1.631 s -28.48 ms / -1.7% (better) 5.124 ms -176.9 us / -3.3% (better)
Windows MinGW 386 cprintf-lto 103424 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.653 s -44.24 ms / -2.6% (better) 5.916 ms +325.9 us / +5.8% (worse)
Windows MinGW 386 fmtprintf 4752384 B +8704 B / +0.2% (worse) 465966 B -6336 B / -1.3% (better) 7.317 s +23.6 ms / +0.3% (worse) 10.081 ms +143.6 us / +1.4% (worse)
Windows MinGW 386 fmtprintf-lto 4167680 B +21504 B / +0.5% (worse) 445314 B -5504 B / -1.2% (better) 14.120 s -7.936 ms / -0.1% (better) 11.281 ms +1.221 ms / +12.1% (worse)
Windows MinGW 386 println 653824 B +2048 B / +0.3% (worse) 21282 B -208 B / -1.0% (better) 1.635 s -30.98 ms / -1.9% (better) 8.455 ms -1.489 ms / -15.0% (better)
Windows MinGW 386 println-lto 259072 B +512 B / +0.2% (worse) 19114 B -192 B / -1.0% (better) 1.902 s -38.14 ms / -2.0% (better) 9.033 ms -565.7 us / -5.9% (better)
Windows MinGW ARM64 cprintf 660480 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.935 s +12.82 ms / +0.7% (worse) 6.756 ms +182.8 us / +2.8% (worse)
Windows MinGW ARM64 cprintf-lto 43520 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.982 s -45.29 ms / -2.2% (better) 6.548 ms -728.3 us / -10.0% (better)
Windows MinGW ARM64 fmtprintf 5348352 B +5632 B / +0.1% (worse) 504616 B -6072 B / -1.2% (better) 7.202 s -104.9 ms / -1.4% (better) 13.518 ms -1.028 ms / -7.1% (better)
Windows MinGW ARM64 fmtprintf-lto 4319232 B +19968 B / +0.5% (worse) 471408 B -5532 B / -1.2% (better) 13.924 s +119.4 ms / +0.9% (worse) 13.693 ms +770.2 us / +6.0% (worse)
Windows MinGW ARM64 println 713728 B +1024 B / +0.1% (worse) 23644 B -240 B / -1.0% (better) 1.923 s -94.94 ms / -4.7% (better) 11.195 ms -1.063 ms / -8.7% (better)
Windows MinGW ARM64 println-lto 216064 B +512 B / +0.2% (worse) 21036 B -196 B / -0.9% (better) 2.196 s -84.89 ms / -3.7% (better) 11.701 ms -708.9 us / -5.7% (better)
Windows MSVC cprintf 893440 B +1536 B / +0.2% (worse) 65798 B 0 B / +0.0% 1.805 s +255.8 ms / +16.5% (worse) 3.683 ms +357.4 us / +10.7% (worse)
Windows MSVC cprintf-lto 289792 B 0 B / +0.0% 65734 B 0 B / +0.0% 1.709 s -166 ms / -8.9% (better) 3.546 ms +139.7 us / +4.1% (worse)
Windows MSVC fmtprintf 5753344 B +20992 B / +0.4% (worse) 691190 B -4352 B / -0.6% (better) 7.139 s +59.24 ms / +0.8% (worse) 9.096 ms +432.4 us / +5.0% (worse)
Windows MSVC fmtprintf-lto 4449792 B +1536 B / +0.03453% (worse) 638806 B -6976 B / -1.1% (better) 13.916 s +86.67 ms / +0.6% (worse) 9.497 ms +25.8 us / +0.3% (worse)
Windows MSVC println 1017344 B +512 B / +0.1% (worse) 120662 B -192 B / -0.2% (better) 1.581 s -157.4 ms / -9.1% (better) 8.018 ms +793.8 us / +11.0% (worse)
Windows MSVC println-lto 528384 B 0 B / +0.0% 118182 B -208 B / -0.2% (better) 1.841 s +34.92 ms / +1.9% (worse) 8.297 ms +1.225 ms / +17.3% (worse)
Windows MSVC 386 cprintf 512512 B 0 B / +0.0% 3931 B 0 B / +0.0% 1.165 s -5.025 ms / -0.4% (better) 4.431 ms +115.8 us / +2.7% (worse)
Windows MSVC 386 cprintf-lto 44032 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.187 s -3.596 ms / -0.3% (better) 4.450 ms +123.8 us / +2.9% (worse)
Windows MSVC 386 fmtprintf 4486656 B +8192 B / +0.2% (worse) 449340 B -6352 B / -1.4% (better) 5.441 s -139 ms / -2.5% (better) 8.721 ms -917.8 us / -9.5% (better)
Windows MSVC 386 fmtprintf-lto 3910656 B +13824 B / +0.4% (worse) 420315 B -6080 B / -1.4% (better) 10.809 s +54.71 ms / +0.5% (worse) 8.887 ms -490.2 us / -5.2% (better)
Windows MSVC 386 println 566784 B +512 B / +0.1% (worse) 20132 B -208 B / -1.0% (better) 1.161 s -24.99 ms / -2.1% (better) 7.318 ms +46.4 us / +0.6% (worse)
Windows MSVC 386 println-lto 199168 B 0 B / +0.0% 18293 B -208 B / -1.1% (better) 1.343 s -49.16 ms / -3.5% (better) 7.352 ms -168.6 us / -2.2% (better)
Windows MSVC ARM64 cprintf 662016 B +512 B / +0.1% (worse) 4192 B 0 B / +0.0% 1.628 s -1.668 ms / -0.1% (better) 7.197 ms +71.7 us / +1.0% (worse)
Windows MSVC ARM64 cprintf-lto 47616 B 0 B / +0.0% 4084 B 0 B / +0.0% 1.651 s +20.84 ms / +1.3% (worse) 7.328 ms +218 us / +3.1% (worse)
Windows MSVC ARM64 fmtprintf 5344256 B +5632 B / +0.1% (worse) 504552 B -6080 B / -1.2% (better) 6.713 s +44.35 ms / +0.7% (worse) 14.417 ms -902.7 us / -5.9% (better)
Windows MSVC ARM64 fmtprintf-lto 4326912 B +21504 B / +0.5% (worse) 472084 B -5520 B / -1.2% (better) 12.913 s -44.26 ms / -0.3% (better) 14.321 ms -613.7 us / -4.1% (better)
Windows MSVC ARM64 println 715776 B +1024 B / +0.1% (worse) 23668 B -240 B / -1.0% (better) 1.628 s -4.591 ms / -0.3% (better) 12.318 ms -612.9 us / -4.7% (better)
Windows MSVC ARM64 println-lto 220160 B 0 B / +0.0% 21188 B -192 B / -0.9% (better) 1.870 s +9.63 ms / +0.5% (worse) 12.447 ms -749.3 us / -5.7% (better)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.660 ns/op +0.21 ns/op / +1.5% (worse)
Linux BenchmarkMergeCompilerFlags 192.100 ns/op -13.8 ns/op / -6.7% (better)
Linux BenchmarkMergeLinkerFlags 133.300 ns/op +5.7 ns/op / +4.5% (worse)
Linux BenchmarkChannelBuffered 55.280 ns/op +0.14 ns/op / +0.3% (worse)
Linux BenchmarkChannelHandoff 19494 ns/op +6692 ns/op / +52.3% (worse)
Linux BenchmarkDefer 50.780 ns/op +0.02 ns/op / +0.0394% (worse)
Linux BenchmarkDirectCall 1.164 ns/op -0.419 ns/op / -26.5% (better)
Linux BenchmarkGlobalRead 1.165 ns/op -0.003 ns/op / -0.3% (better)
Linux BenchmarkGlobalWrite 7.804 ns/op +0.034 ns/op / +0.4% (worse)
Linux BenchmarkGoroutine 30674 ns/op +6303 ns/op / +25.9% (worse)
Linux BenchmarkInterfaceCall 5.629 ns/op -0.226 ns/op / -3.9% (better)
Linux BenchmarkRuntimeGetG 2.991 ns/op +0.635 ns/op / +27.0% (worse)
macOS BenchmarkLookupPCRandom 13.490 ns/op -2.37 ns/op / -14.9% (better)
macOS BenchmarkMergeCompilerFlags 118.900 ns/op +0.7 ns/op / +0.6% (worse)
macOS BenchmarkMergeLinkerFlags 70.790 ns/op -9.14 ns/op / -11.4% (better)
macOS BenchmarkChannelBuffered 26.690 ns/op -7.55 ns/op / -22.1% (better)
macOS BenchmarkChannelHandoff 7230 ns/op -5423 ns/op / -42.9% (better)
macOS BenchmarkDefer 31.300 ns/op -13.01 ns/op / -29.4% (better)
macOS BenchmarkDirectCall 1.105 ns/op -0.093 ns/op / -7.8% (better)
macOS BenchmarkGlobalRead 1.215 ns/op -0.076 ns/op / -5.9% (better)
macOS BenchmarkGlobalWrite 1.093 ns/op -0.405 ns/op / -27.0% (better)
macOS BenchmarkGoroutine 51803 ns/op -17251 ns/op / -25.0% (better)
macOS BenchmarkInterfaceCall 3.856 ns/op -1.22 ns/op / -24.0% (better)
macOS BenchmarkRuntimeGetG 2.121 ns/op -0.659 ns/op / -23.7% (better)
Windows MinGW BenchmarkLookupPCRandom 12.860 ns/op -0.24 ns/op / -1.8% (better)
Windows MinGW BenchmarkMergeCompilerFlags 619.900 ns/op +23.1 ns/op / +3.9% (worse)
Windows MinGW BenchmarkMergeLinkerFlags 549 ns/op +6.2 ns/op / +1.1% (worse)
Windows MinGW BenchmarkChannelBuffered 30.840 ns/op -0.37 ns/op / -1.2% (better)
Windows MinGW BenchmarkChannelHandoff 913.900 ns/op -22.3 ns/op / -2.4% (better)
Windows MinGW BenchmarkDefer 55.440 ns/op -1.43 ns/op / -2.5% (better)
Windows MinGW BenchmarkDirectCall 1.547 ns/op 0 ns/op / +0.0%
Windows MinGW BenchmarkGlobalRead 1.862 ns/op +0.314 ns/op / +20.3% (worse)
Windows MinGW BenchmarkGlobalWrite 2.446 ns/op -0.023 ns/op / -0.9% (better)
Windows MinGW BenchmarkGoroutine 90695 ns/op -902 ns/op / -1.0% (better)
Windows MinGW BenchmarkInterfaceCall 8.378 ns/op +0.011 ns/op / +0.1% (worse)
Windows MinGW BenchmarkRuntimeGetG 1.862 ns/op -0.614 ns/op / -24.8% (better)
Windows MinGW 386 BenchmarkLookupPCRandom 26.490 ns/op -0.15 ns/op / -0.6% (better)
Windows MinGW 386 BenchmarkMergeCompilerFlags 768.500 ns/op +22.1 ns/op / +3.0% (worse)
Windows MinGW 386 BenchmarkMergeLinkerFlags 672.700 ns/op -32.3 ns/op / -4.6% (better)
Windows MinGW 386 BenchmarkChannelBuffered 38.920 ns/op -2.34 ns/op / -5.7% (better)
Windows MinGW 386 BenchmarkChannelHandoff 898 ns/op -21.2 ns/op / -2.3% (better)
Windows MinGW 386 BenchmarkDefer 41.730 ns/op -1.8 ns/op / -4.1% (better)
Windows MinGW 386 BenchmarkDirectCall 1.551 ns/op +0.001 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkGlobalRead 1.859 ns/op +0.31 ns/op / +20.0% (worse)
Windows MinGW 386 BenchmarkGlobalWrite 7.776 ns/op -0.008 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGoroutine 103818 ns/op -3311 ns/op / -3.1% (better)
Windows MinGW 386 BenchmarkInterfaceCall 8.048 ns/op -0.324 ns/op / -3.9% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 1.927 ns/op -0.004 ns/op / -0.2% (better)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.070 ns/op +0.05 ns/op / +0.4% (worse)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 570.800 ns/op +15.2 ns/op / +2.7% (worse)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 540.700 ns/op +7.2 ns/op / +1.3% (worse)
Windows MinGW ARM64 BenchmarkChannelBuffered 38.600 ns/op +0.93 ns/op / +2.5% (worse)
Windows MinGW ARM64 BenchmarkChannelHandoff 2136 ns/op -263 ns/op / -11.0% (better)
Windows MinGW ARM64 BenchmarkDefer 58.570 ns/op +3.14 ns/op / +5.7% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.589 ns/op -0.0739 ns/op / -11.1% (better)
Windows MinGW ARM64 BenchmarkGlobalRead 0.885 ns/op +0.2213 ns/op / +33.3% (worse)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.663 ns/op -0.0732 ns/op / -9.9% (better)
Windows MinGW ARM64 BenchmarkGoroutine 61161 ns/op -2843 ns/op / -4.4% (better)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.183 ns/op +0.038 ns/op / +0.9% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.769 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC BenchmarkLookupPCRandom 13.150 ns/op -0.07 ns/op / -0.5% (better)
Windows MSVC BenchmarkMergeCompilerFlags 631.300 ns/op -9.7 ns/op / -1.5% (better)
Windows MSVC BenchmarkMergeLinkerFlags 552.600 ns/op -3 ns/op / -0.5% (better)
Windows MSVC BenchmarkChannelBuffered 29.520 ns/op -0.26 ns/op / -0.9% (better)
Windows MSVC BenchmarkChannelHandoff 1157 ns/op +64 ns/op / +5.9% (worse)
Windows MSVC BenchmarkDefer 55.450 ns/op +0.37 ns/op / +0.7% (worse)
Windows MSVC BenchmarkDirectCall 1.546 ns/op -0.003 ns/op / -0.2% (better)
Windows MSVC BenchmarkGlobalRead 1.547 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC BenchmarkGlobalWrite 2.457 ns/op -0.014 ns/op / -0.6% (better)
Windows MSVC BenchmarkGoroutine 90104 ns/op +4333 ns/op / +5.1% (worse)
Windows MSVC BenchmarkInterfaceCall 9.335 ns/op +0.327 ns/op / +3.6% (worse)
Windows MSVC BenchmarkRuntimeGetG 2.174 ns/op -0.306 ns/op / -12.3% (better)
Windows MSVC 386 BenchmarkLookupPCRandom 21.550 ns/op -0.02 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkMergeCompilerFlags 551.800 ns/op +6.1 ns/op / +1.1% (worse)
Windows MSVC 386 BenchmarkMergeLinkerFlags 521.200 ns/op +3.2 ns/op / +0.6% (worse)
Windows MSVC 386 BenchmarkChannelBuffered 33.950 ns/op 0 ns/op / +0.0%
Windows MSVC 386 BenchmarkChannelHandoff 687.700 ns/op -86.6 ns/op / -11.2% (better)
Windows MSVC 386 BenchmarkDefer 41.090 ns/op +2.66 ns/op / +6.9% (worse)
Windows MSVC 386 BenchmarkDirectCall 1.358 ns/op +0.001 ns/op / +0.1% (worse)
Windows MSVC 386 BenchmarkGlobalRead 1.358 ns/op 0 ns/op / +0.0%
Windows MSVC 386 BenchmarkGlobalWrite 6.981 ns/op 0 ns/op / +0.0%
Windows MSVC 386 BenchmarkGoroutine 74083 ns/op +134 ns/op / +0.2% (worse)
Windows MSVC 386 BenchmarkInterfaceCall 7.354 ns/op +0.015 ns/op / +0.2% (worse)
Windows MSVC 386 BenchmarkRuntimeGetG 1.632 ns/op -0.19 ns/op / -10.4% (better)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.090 ns/op +0.03 ns/op / +0.2% (worse)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 588.700 ns/op +12.7 ns/op / +2.2% (worse)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 550.800 ns/op -0.4 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkChannelBuffered 39.010 ns/op +0.53 ns/op / +1.4% (worse)
Windows MSVC ARM64 BenchmarkChannelHandoff 2887 ns/op -708 ns/op / -19.7% (better)
Windows MSVC ARM64 BenchmarkDefer 62.610 ns/op +0.72 ns/op / +1.2% (worse)
Windows MSVC ARM64 BenchmarkDirectCall 0.589 ns/op -0.0737 ns/op / -11.1% (better)
Windows MSVC ARM64 BenchmarkGlobalRead 0.589 ns/op -0.0745 ns/op / -11.2% (better)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.794 ns/op -0.001 ns/op / -0.02635% (better)
Windows MSVC ARM64 BenchmarkGoroutine 60139 ns/op +488 ns/op / +0.8% (worse)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.140 ns/op +0.001 ns/op / +0.02416% (worse)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.770 ns/op -0.04 ns/op / -2.2% (better)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 913.600 ns/op +11.6 ns/op / +1.3% (worse)
Linux AfterFuncZeroDelivery/LLGo 37507 ns/op +730 ns/op / +2.0% (worse)
Linux CreateStop/Go 291.400 ns/op +2.8 ns/op / +1.0% (worse)
Linux CreateStop/LLGo 1706 ns/op -1 ns/op / -0.1% (better)
Linux RearmStopped/Go 116.600 ns/op +0.6 ns/op / +0.5% (worse)
Linux RearmStopped/LLGo 1255 ns/op +8 ns/op / +0.6% (worse)
Linux ResetActive/Go 69.030 ns/op +0.06 ns/op / +0.1% (worse)
Linux ResetActive/LLGo 709.200 ns/op +20.6 ns/op / +3.0% (worse)
Linux ResetHeap1024/Go 67.160 ns/op +0.04 ns/op / +0.1% (worse)
Linux ResetHeap1024/LLGo 179.100 ns/op +2.8 ns/op / +1.6% (worse)
macOS AfterFuncZeroDelivery/Go 536.100 ns/op +37 ns/op / +7.4% (worse)
macOS AfterFuncZeroDelivery/LLGo 95174 ns/op +14371 ns/op / +17.8% (worse)
macOS CreateStop/Go 148.500 ns/op -28.8 ns/op / -16.2% (better)
macOS CreateStop/LLGo 461.400 ns/op -10.6 ns/op / -2.2% (better)
macOS RearmStopped/Go 58.210 ns/op -4.79 ns/op / -7.6% (better)
macOS RearmStopped/LLGo 391.200 ns/op -282.3 ns/op / -41.9% (better)
macOS ResetActive/Go 49.240 ns/op -0.73 ns/op / -1.5% (better)
macOS ResetActive/LLGo 194 ns/op -10.6 ns/op / -5.2% (better)
macOS ResetHeap1024/Go 43.080 ns/op -5.32 ns/op / -11.0% (better)
macOS ResetHeap1024/LLGo 85.610 ns/op -7.84 ns/op / -8.4% (better)
Windows MinGW AfterFuncZeroDelivery/Go 566.800 ns/op -2.9 ns/op / -0.5% (better)
Windows MinGW AfterFuncZeroDelivery/LLGo 178799 ns/op -1434 ns/op / -0.8% (better)
Windows MinGW CreateStop/Go 113 ns/op -0.5 ns/op / -0.4% (better)
Windows MinGW CreateStop/LLGo 443.300 ns/op +11.9 ns/op / +2.8% (worse)
Windows MinGW RearmStopped/Go 31.540 ns/op +0.24 ns/op / +0.8% (worse)
Windows MinGW RearmStopped/LLGo 275.900 ns/op +3.7 ns/op / +1.4% (worse)
Windows MinGW ResetActive/Go 20.120 ns/op -0.02 ns/op / -0.1% (better)
Windows MinGW ResetActive/LLGo 154.600 ns/op -6.5 ns/op / -4.0% (better)
Windows MinGW ResetHeap1024/Go 20.510 ns/op +0.09 ns/op / +0.4% (worse)
Windows MinGW ResetHeap1024/LLGo 125.200 ns/op +1.2 ns/op / +1.0% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/Go 959.100 ns/op -1.5 ns/op / -0.2% (better)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 198635 ns/op -613 ns/op / -0.3% (better)
Windows MinGW 386 CreateStop/Go 190 ns/op -1.7 ns/op / -0.9% (better)
Windows MinGW 386 CreateStop/LLGo 521.700 ns/op +9.1 ns/op / +1.8% (worse)
Windows MinGW 386 RearmStopped/Go 63.540 ns/op +0.05 ns/op / +0.1% (worse)
Windows MinGW 386 RearmStopped/LLGo 350.800 ns/op -16.2 ns/op / -4.4% (better)
Windows MinGW 386 ResetActive/Go 39.170 ns/op +0.08 ns/op / +0.2% (worse)
Windows MinGW 386 ResetActive/LLGo 1004 ns/op -27 ns/op / -2.6% (better)
Windows MinGW 386 ResetHeap1024/Go 39.670 ns/op +0.17 ns/op / +0.4% (worse)
Windows MinGW 386 ResetHeap1024/LLGo 188.100 ns/op -1.1 ns/op / -0.6% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 662.700 ns/op -1.9 ns/op / -0.3% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 151414 ns/op +2265 ns/op / +1.5% (worse)
Windows MinGW ARM64 CreateStop/Go 197.200 ns/op -4.5 ns/op / -2.2% (better)
Windows MinGW ARM64 CreateStop/LLGo 361.800 ns/op +1.6 ns/op / +0.4% (worse)
Windows MinGW ARM64 RearmStopped/Go 70.630 ns/op +0.05 ns/op / +0.1% (worse)
Windows MinGW ARM64 RearmStopped/LLGo 250.900 ns/op -9.5 ns/op / -3.6% (better)
Windows MinGW ARM64 ResetActive/Go 31.090 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 ResetActive/LLGo 117.500 ns/op -4.9 ns/op / -4.0% (better)
Windows MinGW ARM64 ResetHeap1024/Go 31.150 ns/op +0.06 ns/op / +0.2% (worse)
Windows MinGW ARM64 ResetHeap1024/LLGo 125.800 ns/op +1.5 ns/op / +1.2% (worse)
Windows MSVC AfterFuncZeroDelivery/Go 546.300 ns/op -24.6 ns/op / -4.3% (better)
Windows MSVC AfterFuncZeroDelivery/LLGo 176715 ns/op +91 ns/op / +0.1% (worse)
Windows MSVC CreateStop/Go 116.700 ns/op -0.4 ns/op / -0.3% (better)
Windows MSVC CreateStop/LLGo 411.800 ns/op -7.8 ns/op / -1.9% (better)
Windows MSVC RearmStopped/Go 31.500 ns/op +0.02 ns/op / +0.1% (worse)
Windows MSVC RearmStopped/LLGo 263.900 ns/op -18.4 ns/op / -6.5% (better)
Windows MSVC ResetActive/Go 20.160 ns/op -0.25 ns/op / -1.2% (better)
Windows MSVC ResetActive/LLGo 150.300 ns/op -0.3 ns/op / -0.2% (better)
Windows MSVC ResetHeap1024/Go 20.490 ns/op -0.01 ns/op / -0.04878% (better)
Windows MSVC ResetHeap1024/LLGo 127.100 ns/op -1.9 ns/op / -1.5% (better)
Windows MSVC 386 AfterFuncZeroDelivery/Go 799.300 ns/op +37.7 ns/op / +5.0% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 133212 ns/op -607 ns/op / -0.5% (better)
Windows MSVC 386 CreateStop/Go 168.300 ns/op +1.2 ns/op / +0.7% (worse)
Windows MSVC 386 CreateStop/LLGo 401.900 ns/op +29.3 ns/op / +7.9% (worse)
Windows MSVC 386 RearmStopped/Go 56.600 ns/op -0.03 ns/op / -0.1% (better)
Windows MSVC 386 RearmStopped/LLGo 260.900 ns/op -4.8 ns/op / -1.8% (better)
Windows MSVC 386 ResetActive/Go 32.590 ns/op +0.01 ns/op / +0.03069% (worse)
Windows MSVC 386 ResetActive/LLGo 911.300 ns/op +21.8 ns/op / +2.5% (worse)
Windows MSVC 386 ResetHeap1024/Go 32.750 ns/op -0.05 ns/op / -0.2% (better)
Windows MSVC 386 ResetHeap1024/LLGo 143.100 ns/op -1.4 ns/op / -1.0% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 686.300 ns/op +18.9 ns/op / +2.8% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 146145 ns/op -27776 ns/op / -16.0% (better)
Windows MSVC ARM64 CreateStop/Go 198.800 ns/op +1.4 ns/op / +0.7% (worse)
Windows MSVC ARM64 CreateStop/LLGo 388.600 ns/op -2.3 ns/op / -0.6% (better)
Windows MSVC ARM64 RearmStopped/Go 70.550 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 RearmStopped/LLGo 276.900 ns/op +5 ns/op / +1.8% (worse)
Windows MSVC ARM64 ResetActive/Go 31.060 ns/op -0.02 ns/op / -0.1% (better)
Windows MSVC ARM64 ResetActive/LLGo 135.800 ns/op +5.9 ns/op / +4.5% (worse)
Windows MSVC ARM64 ResetHeap1024/Go 31.090 ns/op +0.03 ns/op / +0.1% (worse)
Windows MSVC ARM64 ResetHeap1024/LLGo 140.300 ns/op +0.6 ns/op / +0.4% (worse)

Compared with d98b43f90a3d measured in the same runner job.

@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

bb2f4d3cbcc5 | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 153510 B -683 B / -0.4% (better) 88742 B 0 B / +0.0%
cprintf/j32-goos-js/LLGo 151500 B -795 B / -0.5% (better) 73165 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 141143 B -105 B / -0.1% (better) 92630 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 152668 B -180 B / -0.1% (better) 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 153380 B -180 B / -0.1% (better) 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3189496 B -23932 B / -0.7% (better) 132442 B 0 B / +0.0%
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3168726 B -23797 B / -0.7% (better) 101635 B 0 B / +0.0%
fmtprintf/j64-emscripten-memory64/LLGo 2929079 B -20954 B / -0.7% (better) 139285 B 0 B / +0.0%
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2333963 B -13388 B / -0.6% (better) 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2330582 B -13431 B / -0.6% (better) 0 B 0 B / 0.0%
j32-emscripten/LLGo 152744 B -684 B / -0.4% (better) 88742 B 0 B / +0.0%
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 150972 B -796 B / -0.5% (better) 73165 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 140468 B -114 B / -0.1% (better) 92629 B -1 B / -0.00108% (better)
reflectcall/j32-emscripten/LLGo 1537919 B -5852 B / -0.4% (better) 105908 B 0 B / +0.0%
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1540185 B -6037 B / -0.4% (better) 90331 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1424269 B -4157 B / -0.3% (better) 111641 B 0 B / +0.0%
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1275225 B -1468 B / -0.1% (better) 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1272820 B -1481 B / -0.1% (better) 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 152315 B -180 B / -0.1% (better) 0 B 0 B / 0.0%
w32-wasi/LLGo 153027 B -180 B / -0.1% (better) 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 6.996 s +529.5 ms / +8.2% (worse)
j32-goos-js 6.836 s +514.4 ms / +8.1% (worse)
j64-emscripten-memory64 6.068 s +502.5 ms / +9.0% (worse)
reflectcall/w32-wasi 27.022 s +1.159 s / +4.5% (worse)
w32-goos-wasip1 4.877 s +458.8 ms / +10.4% (worse)
w32-wasi 4.805 s +562.3 ms / +13.3% (worse)

Compared with d98b43f90a3d measured in the same runner job.

@visualfc
visualfc force-pushed the fix/static-iface-box branch 9 times, most recently from db6ed63 to 3fed6ed Compare October 2, 2026 02:20
@visualfc visualfc changed the title ssa: box constant iface data in read-only globals ssa: emit static itabs and box constant iface data Oct 2, 2026
visualfc added a commit to visualfc/llgo that referenced this pull request Oct 2, 2026
Codecov patch coverage on xgo-dev#2704 was 89.9% (target 95%). The misses were
constantFingerprint's non-int kinds and staticIfaceBox's VarOf hit after
the per-value cache is cleared.
Known concrete-to-interface conversions emit a read-only _llgo_itab$
global (Inter, Type, Hash, Fun) with weak_odr linkage, matching
cmd/compile's go:itab symbols.

Ordinary builds use that global as the runtime vtable. The T2I site is
the only reference, so --gc-sections/-dead_strip can drop itabs that
belong to dead functions. LTO and deadcode-drop keep calling NewItab so
unused interface methods can be dropped; Fun[] would pin them. LTO still
emits the static template for de-virt and erases unused templates after
the plugin runs.

Non-direct interface values (integers, bools, small aggregates) still
use IfaceIndir, matching Go: data is a pointer to a copy. Compile-time
constants go in a WeakODR _llgo_ifacebox$ global instead of AllocU,
like cmd/compile's static temps.

Box names hash the constant's bits, not LLVM's printed form, so
same-length strings in different packages do not collide under
Windows COMDAT. Repeated boxing of the same LLVM value reuses a
per-package pointer cache. Zero-sized constants stay AllocU so boxing
does not mutate the shared sentinel.

Together, a constant T2I is {itab, box} with no runtime call or heap
allocation. Hash is copied from the type descriptor when present.
interequal compares the (inter, _type) pair so a static itab and a
dynamically allocated itab for the same conversion compare equal.
@visualfc
visualfc force-pushed the fix/static-iface-box branch from 4b91a17 to bb2f4d3 Compare October 2, 2026 06:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant