Skip to content

cl: fold trivial constant iface constructors - #2705

Open
visualfc wants to merge 2 commits into
xgo-dev:mainfrom
visualfc:fix/trivial-iface-fold
Open

visualfc wants to merge 2 commits into
xgo-dev:mainfrom
visualfc:fix/trivial-iface-fold

Conversation

@visualfc

@visualfc visualfc commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

A call whose SSA body is MakeInterface of a Convert/ChangeType of a constant parameter is lowered at the call site to MakeInterface of that concrete value.

The callee is recognized by that SSA shape (single block, no //go:noinline), not by package path or function name. constant.MakeInt64 / MakeBool match because they are return int64Val(x) / return boolVal(b). User helpers such as func boxMyInt(x int64) any { return myInt(x) } use the same path. Functions with extra control flow (MakeString) are left as real calls.

Depends on: #2704 (static itabs + iface boxes). Without those, this fold would emit AllocU+NewItab per call. Merge after #2704. This branch contains #2704 plus this commit; only the latest commit is in scope.

Test plan

  • cl/_testrt/constiface (MakeInt64, MakeBool, and a local boxMyInt all become {itab, box})

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: fold trivial constant interface boxing

Solid optimization overall — the static itab / static ifacebox globals replace runtime NewItab and AllocU for compile-time constants, and the interequal fix correctly keeps T2I (static itab) and I2I (NewItab) representations of the same iface equal by comparing the (inter, _type) pair with proper nil guards. Verified: the removed LTO/compiler.used template comments are no longer stale, the staticItab hash matches the concrete type descriptor, and boxed constants are safely immutable under Go semantics (assertions copy out; no addressable path reaches the read-only box).

Two correctness concerns worth addressing before merge, plus one minor doc nit — inline.

Comment thread cl/constiface.go Outdated
case *ssa.ChangeType:
chain[t] = true
v = t.X
case *ssa.Convert:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Convert in fold chain reinterprets the raw constant under the wrong type

analyzeTrivialIfaceBox accepts *ssa.Convert in the chain and returns concrete = mi.X.Type() (the type after the conversion), while foldConstantMakeValue then applies the callee argument's raw constant via b.Const(c.Value, concrete) at cl/constiface.go:50 — skipping the conversion arithmetic.

This is correct only for representation-preserving conversions. For a value-changing Convert it is unsound. Example: func box(x int32) any { return any(string(x)) } called as box(65). Here concrete is string but c.Value is the integer 65; b.Const dispatches on the target kind (ssa/expr.go), hits the types.String branch, and calls constant.StringVal on an integer constant — producing a wrong result or a panic, where real Go yields any("A"). Numeric narrowing/widening and int↔float Convert have the same hazard.

*ssa.Convert in go/ssa is exactly the set of value-changing conversions, so including it here is unsafe. Suggest dropping *ssa.Convert from the recognized chain (keeping only ChangeType/ChangeInterface, which are representation-identical), or skip folding whenever a Convert is present.

Comment thread ssa/interface.go
if g.impl.IsNil() {
return llvm.Value{}, false
}
g.impl.SetGlobalConstant(true)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] staticIfaceBox mutates the shared zero-sized sentinel for zero-sized types

For a zero-sized value type (e.g. struct{}, [0]T, or a named type over them — none are directIfaceType, so MakeInterface routes them here), NewVarEx → doNewVarEx does not create a dedicated global. It returns the shared module zero-sized sentinel (__llgo.moduleZeroSizedAlloc$) as an isZeroSizedAlias Global (ssa/decl.go:209-219).

g.Init(x) is a no-op for that alias, but the following lines are not guarded and run on the shared sentinel:

g.impl.SetGlobalConstant(true)
g.impl.SetUnnamedAddr(true)
b.Pkg.setODRLinkage(g.impl, llvm.WeakODRLinkage)

Boxing any zero-sized constant thus flips the shared sentinel to GlobalConstant and rewrites its linkage (from LinkOnceODRLinkage, or re-COMDATs it on Windows) — a cross-cutting side effect on the address handed out for all zero-sized allocations in the module. It also pollutes the _llgo_ifacebox$ dedup cache with an entry pointing at the sentinel.

Suggest bailing out (return false) before the constant/linkage mutations when the resulting Global is a zero-sized alias — the isZeroSizedAlias field already exists on aGlobal, or guard on TypeAllocSize(storageType)==0. Zero-sized boxes don't benefit from this optimization anyway.

@codecov

codecov Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 93.37349% with 11 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
cl/constiface.go 87.03% 7 Missing ⚠️
ssa/interface.go 96.29% 4 Missing ⚠️

📢 Thoughts on this report? Let us know!

@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

4f1be2eba561 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 8568 B 0 B / +0.0% 387 B 0 B / +0.0% 1.007 s +34.07 ms / +3.5% (worse) 1.277 ms -48.47 us / -3.7% (better)
Linux cprintf-lto 8408 B 0 B / +0.0% 368 B 0 B / +0.0% 1.044 s +89.17 ms / +9.3% (worse) 1.313 ms +26.66 us / +2.1% (worse)
Linux fmtprintf 4886680 B +6744 B / +0.1% (worse) 492638 B -8675 B / -1.7% (better) 7.436 s -2.215 ms / -0.02979% (better) 3.087 ms -31 us / -1.0% (better)
Linux fmtprintf-lto 3591984 B -6304 B / -0.2% (better) 428037 B -10120 B / -2.3% (better) 16.450 s -679.2 ms / -4.0% (better) 2.933 ms +83.1 us / +2.9% (worse)
Linux println 675144 B +776 B / +0.1% (worse) 16649 B -206 B / -1.2% (better) 1.039 s +61.02 ms / +6.2% (worse) 1.761 ms +63.74 us / +3.8% (worse)
Linux println-lto 190048 B +56 B / +0.02947% (worse) 14089 B -184 B / -1.3% (better) 1.401 s +146.1 ms / +11.6% (worse) 1.730 ms +110.5 us / +6.8% (worse)
macOS cprintf 50736 B 0 B / +0.0% 4409 B 0 B / +0.0% 1.305 s -50.01 ms / -3.7% (better) 6.475 ms -354.3 us / -5.2% (better)
macOS cprintf-lto 50496 B 0 B / +0.0% 161 B 0 B / +0.0% 1.194 s -423.7 ms / -26.2% (better) 3.912 ms +588.5 us / +17.7% (worse)
macOS fmtprintf 1827296 B +53840 B / +3.0% (worse) 870528 B -6888 B / -0.8% (better) 8.366 s +1.439 s / +20.8% (worse) 17.258 ms +3.721 ms / +27.5% (worse)
macOS fmtprintf-lto 1390144 B +29216 B / +2.1% (worse) 753824 B -8092 B / -1.1% (better) 13.755 s -5.365 s / -28.1% (better) 5.340 ms -142.8 us / -2.6% (better)
macOS println 101152 B +1808 B / +1.8% (worse) 24006 B -216 B / -0.9% (better) 1.180 s -145.4 ms / -11.0% (better) 7.244 ms +1.411 ms / +24.2% (worse)
macOS println-lto 84032 B +368 B / +0.4% (worse) 21257 B -200 B / -0.9% (better) 1.521 s -12.43 ms / -0.8% (better) 4.667 ms +317.7 us / +7.3% (worse)
Windows MinGW cprintf 650240 B +512 B / +0.1% (worse) 4550 B 0 B / +0.0% 1.755 s +83.63 ms / +5.0% (worse) 4.083 ms +681.3 us / +20.0% (worse)
Windows MinGW cprintf-lto 43520 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.722 s +24.06 ms / +1.4% (worse) 4.022 ms +502.5 us / +14.3% (worse)
Windows MinGW fmtprintf 5454336 B +23040 B / +0.4% (worse) 595670 B -4336 B / -0.7% (better) 7.305 s +70.14 ms / +1.0% (worse) 8.361 ms +37.8 us / +0.5% (worse)
Windows MinGW fmtprintf-lto 4127232 B +3072 B / +0.1% (worse) 539254 B -6992 B / -1.3% (better) 14.690 s +154.8 ms / +1.1% (worse) 8.769 ms +558.7 us / +6.8% (worse)
Windows MinGW println 706560 B +1024 B / +0.1% (worse) 24998 B -192 B / -0.8% (better) 1.914 s +233.7 ms / +13.9% (worse) 7.512 ms +982.3 us / +15.0% (worse)
Windows MinGW println-lto 209408 B +512 B / +0.2% (worse) 21846 B -208 B / -0.9% (better) 2.214 s +277.9 ms / +14.4% (worse) 7.895 ms +1.307 ms / +19.8% (worse)
Windows MinGW 386 cprintf 600576 B -512 B / -0.1% (better) 5326 B 0 B / +0.0% 1.735 s -13.41 ms / -0.8% (better) 5.671 ms -362.4 us / -6.0% (better)
Windows MinGW 386 cprintf-lto 103424 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.801 s +26.42 ms / +1.5% (worse) 6.265 ms +865.9 us / +16.0% (worse)
Windows MinGW 386 fmtprintf 4752384 B +8704 B / +0.2% (worse) 465966 B -6336 B / -1.3% (better) 7.646 s +479.9 us / +0.006277% (worse) 12.809 ms +901.9 us / +7.6% (worse)
Windows MinGW 386 fmtprintf-lto 4167680 B +21504 B / +0.5% (worse) 445314 B -5504 B / -1.2% (better) 15.378 s +420.1 ms / +2.8% (worse) 13.657 ms +1.118 ms / +8.9% (worse)
Windows MinGW 386 println 653824 B +2048 B / +0.3% (worse) 21282 B -208 B / -1.0% (better) 1.810 s +74.38 ms / +4.3% (worse) 10.187 ms +700.4 us / +7.4% (worse)
Windows MinGW 386 println-lto 259072 B +512 B / +0.2% (worse) 19114 B -192 B / -1.0% (better) 2.070 s +19.84 ms / +1.0% (worse) 9.812 ms -84.7 us / -0.9% (better)
Windows MinGW ARM64 cprintf 660480 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.830 s -85.06 ms / -4.4% (better) 6.068 ms -329.8 us / -5.2% (better)
Windows MinGW ARM64 cprintf-lto 43520 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.880 s -68.87 ms / -3.5% (better) 6.490 ms -153.2 us / -2.3% (better)
Windows MinGW ARM64 fmtprintf 5348352 B +5632 B / +0.1% (worse) 504616 B -6072 B / -1.2% (better) 6.980 s -31.09 ms / -0.4% (better) 13.011 ms +220.6 us / +1.7% (worse)
Windows MinGW ARM64 fmtprintf-lto 4319232 B +19968 B / +0.5% (worse) 471408 B -5532 B / -1.2% (better) 13.774 s +363.4 ms / +2.7% (worse) 12.152 ms -487.2 us / -3.9% (better)
Windows MinGW ARM64 println 713728 B +1024 B / +0.1% (worse) 23644 B -240 B / -1.0% (better) 1.908 s -8.484 ms / -0.4% (better) 10.541 ms -483.7 us / -4.4% (better)
Windows MinGW ARM64 println-lto 216064 B +512 B / +0.2% (worse) 21036 B -196 B / -0.9% (better) 2.192 s +40.12 ms / +1.9% (worse) 10.681 ms +94.2 us / +0.9% (worse)
Windows MSVC cprintf 893440 B +1536 B / +0.2% (worse) 65798 B 0 B / +0.0% 1.366 s -11.01 ms / -0.8% (better) 3.693 ms +141.4 us / +4.0% (worse)
Windows MSVC cprintf-lto 289792 B 0 B / +0.0% 65734 B 0 B / +0.0% 1.423 s +29.1 ms / +2.1% (worse) 3.396 ms -63.7 us / -1.8% (better)
Windows MSVC fmtprintf 5753344 B +20992 B / +0.4% (worse) 691190 B -4352 B / -0.6% (better) 6.273 s +127.8 ms / +2.1% (worse) 8.483 ms -164.5 us / -1.9% (better)
Windows MSVC fmtprintf-lto 4449792 B +1536 B / +0.03453% (worse) 638806 B -6976 B / -1.1% (better) 12.280 s -412 ms / -3.2% (better) 9.162 ms -651.3 us / -6.6% (better)
Windows MSVC println 1017344 B +512 B / +0.1% (worse) 120662 B -192 B / -0.2% (better) 1.354 s -17.15 ms / -1.3% (better) 6.652 ms -1.262 ms / -15.9% (better)
Windows MSVC println-lto 528384 B 0 B / +0.0% 118182 B -208 B / -0.2% (better) 1.585 s +21.1 ms / +1.3% (worse) 6.660 ms +103.3 us / +1.6% (worse)
Windows MSVC 386 cprintf 512512 B 0 B / +0.0% 3931 B 0 B / +0.0% 1.395 s +114.8 ms / +9.0% (worse) 5.417 ms +129.6 us / +2.5% (worse)
Windows MSVC 386 cprintf-lto 44032 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.294 s -85.28 ms / -6.2% (better) 5.937 ms +757.9 us / +14.6% (worse)
Windows MSVC 386 fmtprintf 4486656 B +8192 B / +0.2% (worse) 449340 B -6352 B / -1.4% (better) 6.001 s -104.5 ms / -1.7% (better) 12.346 ms +1.895 ms / +18.1% (worse)
Windows MSVC 386 fmtprintf-lto 3910656 B +13824 B / +0.4% (worse) 420315 B -6080 B / -1.4% (better) 11.325 s -139.8 ms / -1.2% (better) 10.746 ms -632.7 us / -5.6% (better)
Windows MSVC 386 println 566784 B +512 B / +0.1% (worse) 20132 B -208 B / -1.0% (better) 1.252 s -29.6 ms / -2.3% (better) 8.802 ms +262.5 us / +3.1% (worse)
Windows MSVC 386 println-lto 199168 B 0 B / +0.0% 18293 B -208 B / -1.1% (better) 1.494 s -9.116 ms / -0.6% (better) 9.279 ms +152.8 us / +1.7% (worse)
Windows MSVC ARM64 cprintf 662016 B +512 B / +0.1% (worse) 4192 B 0 B / +0.0% 1.672 s +12.51 ms / +0.8% (worse) 7.276 ms -275.6 us / -3.6% (better)
Windows MSVC ARM64 cprintf-lto 47616 B 0 B / +0.0% 4084 B 0 B / +0.0% 1.658 s -5.79 ms / -0.3% (better) 7.578 ms -243.5 us / -3.1% (better)
Windows MSVC ARM64 fmtprintf 5344256 B +5632 B / +0.1% (worse) 504552 B -6080 B / -1.2% (better) 6.871 s +61.29 ms / +0.9% (worse) 15.925 ms +283.8 us / +1.8% (worse)
Windows MSVC ARM64 fmtprintf-lto 4326912 B +21504 B / +0.5% (worse) 472084 B -5520 B / -1.2% (better) 13.484 s +310.5 ms / +2.4% (worse) 15.344 ms +94.4 us / +0.6% (worse)
Windows MSVC ARM64 println 715776 B +1024 B / +0.1% (worse) 23668 B -240 B / -1.0% (better) 1.640 s +3.397 ms / +0.2% (worse) 13.112 ms -547.5 us / -4.0% (better)
Windows MSVC ARM64 println-lto 220160 B 0 B / +0.0% 21188 B -192 B / -0.9% (better) 1.869 s -14.54 ms / -0.8% (better) 13.913 ms +509.8 us / +3.8% (worse)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.630 ns/op +0.06 ns/op / +0.4% (worse)
Linux BenchmarkMergeCompilerFlags 202.200 ns/op -5.2 ns/op / -2.5% (better)
Linux BenchmarkMergeLinkerFlags 126.400 ns/op -8.9 ns/op / -6.6% (better)
Linux BenchmarkChannelBuffered 54.630 ns/op -0.76 ns/op / -1.4% (better)
Linux BenchmarkChannelHandoff 13097 ns/op -811 ns/op / -5.8% (better)
Linux BenchmarkDefer 51.580 ns/op -1.7 ns/op / -3.2% (better)
Linux BenchmarkDirectCall 1.170 ns/op -0.427 ns/op / -26.7% (better)
Linux BenchmarkGlobalRead 1.172 ns/op +0.004 ns/op / +0.3% (worse)
Linux BenchmarkGlobalWrite 7.778 ns/op +0.001 ns/op / +0.01286% (worse)
Linux BenchmarkGoroutine 24004 ns/op -875 ns/op / -3.5% (better)
Linux BenchmarkInterfaceCall 5.647 ns/op -0.218 ns/op / -3.7% (better)
Linux BenchmarkRuntimeGetG 2.970 ns/op +0.605 ns/op / +25.6% (worse)
macOS BenchmarkLookupPCRandom 13.540 ns/op -5.04 ns/op / -27.1% (better)
macOS BenchmarkMergeCompilerFlags 209.400 ns/op +79.2 ns/op / +60.8% (worse)
macOS BenchmarkMergeLinkerFlags 99.010 ns/op +2.84 ns/op / +3.0% (worse)
macOS BenchmarkChannelBuffered 28.070 ns/op -9.25 ns/op / -24.8% (better)
macOS BenchmarkChannelHandoff 7341 ns/op +2 ns/op / +0.02725% (worse)
macOS BenchmarkDefer 34.140 ns/op -22.27 ns/op / -39.5% (better)
macOS BenchmarkDirectCall 1.140 ns/op -0.143 ns/op / -11.1% (better)
macOS BenchmarkGlobalRead 1.090 ns/op -0.197 ns/op / -15.3% (better)
macOS BenchmarkGlobalWrite 1.063 ns/op -0.503 ns/op / -32.1% (better)
macOS BenchmarkGoroutine 40884 ns/op -34085 ns/op / -45.5% (better)
macOS BenchmarkInterfaceCall 4.294 ns/op -0.97 ns/op / -18.4% (better)
macOS BenchmarkRuntimeGetG 2.280 ns/op -0.629 ns/op / -21.6% (better)
Windows MinGW BenchmarkLookupPCRandom 13.060 ns/op +0.25 ns/op / +2.0% (worse)
Windows MinGW BenchmarkMergeCompilerFlags 618.100 ns/op -6.6 ns/op / -1.1% (better)
Windows MinGW BenchmarkMergeLinkerFlags 527.700 ns/op -14.3 ns/op / -2.6% (better)
Windows MinGW BenchmarkChannelBuffered 30.410 ns/op -0.12 ns/op / -0.4% (better)
Windows MinGW BenchmarkChannelHandoff 976.500 ns/op +6.6 ns/op / +0.7% (worse)
Windows MinGW BenchmarkDefer 55.470 ns/op -1.22 ns/op / -2.2% (better)
Windows MinGW BenchmarkDirectCall 1.548 ns/op 0 ns/op / +0.0%
Windows MinGW BenchmarkGlobalRead 1.860 ns/op +0.31 ns/op / +20.0% (worse)
Windows MinGW BenchmarkGlobalWrite 2.447 ns/op -0.024 ns/op / -1.0% (better)
Windows MinGW BenchmarkGoroutine 89989 ns/op -3361 ns/op / -3.6% (better)
Windows MinGW BenchmarkInterfaceCall 8.366 ns/op 0 ns/op / +0.0%
Windows MinGW BenchmarkRuntimeGetG 1.862 ns/op -0.617 ns/op / -24.9% (better)
Windows MinGW 386 BenchmarkLookupPCRandom 68.770 ns/op -0.52 ns/op / -0.8% (better)
Windows MinGW 386 BenchmarkMergeCompilerFlags 945.400 ns/op +28.8 ns/op / +3.1% (worse)
Windows MinGW 386 BenchmarkMergeLinkerFlags 852.200 ns/op -5.2 ns/op / -0.6% (better)
Windows MinGW 386 BenchmarkChannelBuffered 52.270 ns/op -1.61 ns/op / -3.0% (better)
Windows MinGW 386 BenchmarkChannelHandoff 6469 ns/op +3997 ns/op / +161.7% (worse)
Windows MinGW 386 BenchmarkDefer 59.990 ns/op +5.2 ns/op / +9.5% (worse)
Windows MinGW 386 BenchmarkDirectCall 0.869 ns/op -0.002 ns/op / -0.2% (better)
Windows MinGW 386 BenchmarkGlobalRead 1.166 ns/op +0.2962 ns/op / +34.1% (worse)
Windows MinGW 386 BenchmarkGlobalWrite 16.110 ns/op -0.07 ns/op / -0.4% (better)
Windows MinGW 386 BenchmarkGoroutine 216234 ns/op -8 ns/op / -0.0037% (better)
Windows MinGW 386 BenchmarkInterfaceCall 6.589 ns/op -0.098 ns/op / -1.5% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 2.028 ns/op +0.01 ns/op / +0.5% (worse)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.030 ns/op -0.05 ns/op / -0.4% (better)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 573.700 ns/op +5.5 ns/op / +1.0% (worse)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 537.600 ns/op +1.3 ns/op / +0.2% (worse)
Windows MinGW ARM64 BenchmarkChannelBuffered 37.480 ns/op -1.64 ns/op / -4.2% (better)
Windows MinGW ARM64 BenchmarkChannelHandoff 2361 ns/op -650 ns/op / -21.6% (better)
Windows MinGW ARM64 BenchmarkDefer 57.880 ns/op +0.21 ns/op / +0.4% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.590 ns/op -0.0736 ns/op / -11.1% (better)
Windows MinGW ARM64 BenchmarkGlobalRead 0.885 ns/op +0.2214 ns/op / +33.4% (worse)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.664 ns/op -0.0735 ns/op / -10.0% (better)
Windows MinGW ARM64 BenchmarkGoroutine 64499 ns/op +2889 ns/op / +4.7% (worse)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.192 ns/op +0.049 ns/op / +1.2% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.801 ns/op -0.007 ns/op / -0.4% (better)
Windows MSVC BenchmarkLookupPCRandom 9.749 ns/op -0.045 ns/op / -0.5% (better)
Windows MSVC BenchmarkMergeCompilerFlags 486.500 ns/op +8.3 ns/op / +1.7% (worse)
Windows MSVC BenchmarkMergeLinkerFlags 433.400 ns/op -10.2 ns/op / -2.3% (better)
Windows MSVC BenchmarkChannelBuffered 37.360 ns/op +0.14 ns/op / +0.4% (worse)
Windows MSVC BenchmarkChannelHandoff 1084 ns/op -358 ns/op / -24.8% (better)
Windows MSVC BenchmarkDefer 60.930 ns/op +2.24 ns/op / +3.8% (worse)
Windows MSVC BenchmarkDirectCall 0.915 ns/op -0.0426 ns/op / -4.4% (better)
Windows MSVC BenchmarkGlobalRead 1.001 ns/op +0.1125 ns/op / +12.7% (worse)
Windows MSVC BenchmarkGlobalWrite 7.284 ns/op +0.037 ns/op / +0.5% (worse)
Windows MSVC BenchmarkGoroutine 72319 ns/op +2672 ns/op / +3.8% (worse)
Windows MSVC BenchmarkInterfaceCall 5.676 ns/op +0.087 ns/op / +1.6% (worse)
Windows MSVC BenchmarkRuntimeGetG 1.396 ns/op -0.311 ns/op / -18.2% (better)
Windows MSVC 386 BenchmarkLookupPCRandom 61 ns/op +0.12 ns/op / +0.2% (worse)
Windows MSVC 386 BenchmarkMergeCompilerFlags 668.100 ns/op -2.1 ns/op / -0.3% (better)
Windows MSVC 386 BenchmarkMergeLinkerFlags 635.600 ns/op -14.6 ns/op / -2.2% (better)
Windows MSVC 386 BenchmarkChannelBuffered 46.620 ns/op -0.51 ns/op / -1.1% (better)
Windows MSVC 386 BenchmarkChannelHandoff 1625 ns/op +569 ns/op / +53.9% (worse)
Windows MSVC 386 BenchmarkDefer 41.040 ns/op +3.02 ns/op / +7.9% (worse)
Windows MSVC 386 BenchmarkDirectCall 0.948 ns/op +0.0175 ns/op / +1.9% (worse)
Windows MSVC 386 BenchmarkGlobalRead 1.226 ns/op +0.2788 ns/op / +29.4% (worse)
Windows MSVC 386 BenchmarkGlobalWrite 14.950 ns/op -0.03 ns/op / -0.2% (better)
Windows MSVC 386 BenchmarkGoroutine 221935 ns/op +1694 ns/op / +0.8% (worse)
Windows MSVC 386 BenchmarkInterfaceCall 4.428 ns/op -0.2 ns/op / -4.3% (better)
Windows MSVC 386 BenchmarkRuntimeGetG 1.511 ns/op +0.075 ns/op / +5.2% (worse)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.160 ns/op +0.08 ns/op / +0.7% (worse)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 570.900 ns/op -2.7 ns/op / -0.5% (better)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 535.400 ns/op +4 ns/op / +0.8% (worse)
Windows MSVC ARM64 BenchmarkChannelBuffered 37.650 ns/op -0.92 ns/op / -2.4% (better)
Windows MSVC ARM64 BenchmarkChannelHandoff 2597 ns/op +396 ns/op / +18.0% (worse)
Windows MSVC ARM64 BenchmarkDefer 62.080 ns/op -0.89 ns/op / -1.4% (better)
Windows MSVC ARM64 BenchmarkDirectCall 0.590 ns/op -0.0736 ns/op / -11.1% (better)
Windows MSVC ARM64 BenchmarkGlobalRead 0.590 ns/op -0.0793 ns/op / -11.9% (better)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.794 ns/op -0.003 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkGoroutine 59935 ns/op +1616 ns/op / +2.8% (worse)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.140 ns/op +0.001 ns/op / +0.02416% (worse)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.770 ns/op 0 ns/op / +0.0%
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 959.900 ns/op +51.7 ns/op / +5.7% (worse)
Linux AfterFuncZeroDelivery/LLGo 48999 ns/op +2971 ns/op / +6.5% (worse)
Linux CreateStop/Go 301.700 ns/op +13.7 ns/op / +4.8% (worse)
Linux CreateStop/LLGo 1709 ns/op -44 ns/op / -2.5% (better)
Linux RearmStopped/Go 121.200 ns/op +5.2 ns/op / +4.5% (worse)
Linux RearmStopped/LLGo 1371 ns/op +27 ns/op / +2.0% (worse)
Linux ResetActive/Go 72.050 ns/op +3.37 ns/op / +4.9% (worse)
Linux ResetActive/LLGo 766.900 ns/op +61.6 ns/op / +8.7% (worse)
Linux ResetHeap1024/Go 71.130 ns/op +3.89 ns/op / +5.8% (worse)
Linux ResetHeap1024/LLGo 175.700 ns/op -2.7 ns/op / -1.5% (better)
macOS AfterFuncZeroDelivery/Go 408.100 ns/op -193.1 ns/op / -32.1% (better)
macOS AfterFuncZeroDelivery/LLGo 106503 ns/op -1008 ns/op / -0.9% (better)
macOS CreateStop/Go 158.600 ns/op -5.3 ns/op / -3.2% (better)
macOS CreateStop/LLGo 515 ns/op -123.1 ns/op / -19.3% (better)
macOS RearmStopped/Go 58.700 ns/op -17.28 ns/op / -22.7% (better)
macOS RearmStopped/LLGo 563.200 ns/op +106 ns/op / +23.2% (worse)
macOS ResetActive/Go 49.250 ns/op -4.17 ns/op / -7.8% (better)
macOS ResetActive/LLGo 243.500 ns/op -5.5 ns/op / -2.2% (better)
macOS ResetHeap1024/Go 43.750 ns/op -8.21 ns/op / -15.8% (better)
macOS ResetHeap1024/LLGo 89.040 ns/op -59.36 ns/op / -40.0% (better)
Windows MinGW AfterFuncZeroDelivery/Go 565.800 ns/op +0.1 ns/op / +0.01768% (worse)
Windows MinGW AfterFuncZeroDelivery/LLGo 180759 ns/op +370 ns/op / +0.2% (worse)
Windows MinGW CreateStop/Go 122.700 ns/op +7.6 ns/op / +6.6% (worse)
Windows MinGW CreateStop/LLGo 440 ns/op +29.4 ns/op / +7.2% (worse)
Windows MinGW RearmStopped/Go 32.070 ns/op +0.54 ns/op / +1.7% (worse)
Windows MinGW RearmStopped/LLGo 274.100 ns/op -2.3 ns/op / -0.8% (better)
Windows MinGW ResetActive/Go 20.180 ns/op +0.1 ns/op / +0.5% (worse)
Windows MinGW ResetActive/LLGo 152.100 ns/op -6 ns/op / -3.8% (better)
Windows MinGW ResetHeap1024/Go 20.510 ns/op +0.09 ns/op / +0.4% (worse)
Windows MinGW ResetHeap1024/LLGo 125.800 ns/op -1.5 ns/op / -1.2% (better)
Windows MinGW 386 AfterFuncZeroDelivery/Go 1126 ns/op +11 ns/op / +1.0% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 197185 ns/op -1627 ns/op / -0.8% (better)
Windows MinGW 386 CreateStop/Go 275.600 ns/op -19.3 ns/op / -6.5% (better)
Windows MinGW 386 CreateStop/LLGo 619.100 ns/op +62 ns/op / +11.1% (worse)
Windows MinGW 386 RearmStopped/Go 92.760 ns/op +0.35 ns/op / +0.4% (worse)
Windows MinGW 386 RearmStopped/LLGo 392.800 ns/op -25.7 ns/op / -6.1% (better)
Windows MinGW 386 ResetActive/Go 46.130 ns/op -0.11 ns/op / -0.2% (better)
Windows MinGW 386 ResetActive/LLGo 353.300 ns/op +74.5 ns/op / +26.7% (worse)
Windows MinGW 386 ResetHeap1024/Go 46.360 ns/op +0.14 ns/op / +0.3% (worse)
Windows MinGW 386 ResetHeap1024/LLGo 185.200 ns/op -1.4 ns/op / -0.8% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 658.700 ns/op -17.4 ns/op / -2.6% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 151073 ns/op +7443 ns/op / +5.2% (worse)
Windows MinGW ARM64 CreateStop/Go 193.300 ns/op -5.9 ns/op / -3.0% (better)
Windows MinGW ARM64 CreateStop/LLGo 355.200 ns/op -12.3 ns/op / -3.3% (better)
Windows MinGW ARM64 RearmStopped/Go 70.560 ns/op -0.07 ns/op / -0.1% (better)
Windows MinGW ARM64 RearmStopped/LLGo 252.300 ns/op +4 ns/op / +1.6% (worse)
Windows MinGW ARM64 ResetActive/Go 31.100 ns/op +0.03 ns/op / +0.1% (worse)
Windows MinGW ARM64 ResetActive/LLGo 117.800 ns/op -1.7 ns/op / -1.4% (better)
Windows MinGW ARM64 ResetHeap1024/Go 31.180 ns/op +0.1 ns/op / +0.3% (worse)
Windows MinGW ARM64 ResetHeap1024/LLGo 124.600 ns/op +0.1 ns/op / +0.1% (worse)
Windows MSVC AfterFuncZeroDelivery/Go 609.800 ns/op +7.7 ns/op / +1.3% (worse)
Windows MSVC AfterFuncZeroDelivery/LLGo 131773 ns/op +1585 ns/op / +1.2% (worse)
Windows MSVC CreateStop/Go 165.300 ns/op +6.7 ns/op / +4.2% (worse)
Windows MSVC CreateStop/LLGo 758 ns/op +171.3 ns/op / +29.2% (worse)
Windows MSVC RearmStopped/Go 59.970 ns/op +0.39 ns/op / +0.7% (worse)
Windows MSVC RearmStopped/LLGo 283 ns/op +13.3 ns/op / +4.9% (worse)
Windows MSVC ResetActive/Go 26.480 ns/op -0.03 ns/op / -0.1% (better)
Windows MSVC ResetActive/LLGo 182.800 ns/op -23.1 ns/op / -11.2% (better)
Windows MSVC ResetHeap1024/Go 26.900 ns/op +0.15 ns/op / +0.6% (worse)
Windows MSVC ResetHeap1024/LLGo 111.100 ns/op -0.3 ns/op / -0.3% (better)
Windows MSVC 386 AfterFuncZeroDelivery/Go 875 ns/op -0.7 ns/op / -0.1% (better)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 319841 ns/op -4006 ns/op / -1.2% (better)
Windows MSVC 386 CreateStop/Go 224.100 ns/op -0.9 ns/op / -0.4% (better)
Windows MSVC 386 CreateStop/LLGo 776.600 ns/op -3.1 ns/op / -0.4% (better)
Windows MSVC 386 RearmStopped/Go 80.920 ns/op -1.78 ns/op / -2.2% (better)
Windows MSVC 386 RearmStopped/LLGo 330.100 ns/op -104.1 ns/op / -24.0% (better)
Windows MSVC 386 ResetActive/Go 39.140 ns/op -0.77 ns/op / -1.9% (better)
Windows MSVC 386 ResetActive/LLGo 284.800 ns/op +1.4 ns/op / +0.5% (worse)
Windows MSVC 386 ResetHeap1024/Go 40.190 ns/op +0.77 ns/op / +2.0% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 147.800 ns/op +1.4 ns/op / +1.0% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 667.200 ns/op +1.7 ns/op / +0.3% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 153947 ns/op -22827 ns/op / -12.9% (better)
Windows MSVC ARM64 CreateStop/Go 204.400 ns/op +3.6 ns/op / +1.8% (worse)
Windows MSVC ARM64 CreateStop/LLGo 407 ns/op -6.8 ns/op / -1.6% (better)
Windows MSVC ARM64 RearmStopped/Go 70.560 ns/op +0.02 ns/op / +0.02835% (worse)
Windows MSVC ARM64 RearmStopped/LLGo 275.300 ns/op +0.2 ns/op / +0.1% (worse)
Windows MSVC ARM64 ResetActive/Go 31.080 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 ResetActive/LLGo 132.400 ns/op -4.5 ns/op / -3.3% (better)
Windows MSVC ARM64 ResetHeap1024/Go 31.200 ns/op +0.09 ns/op / +0.3% (worse)
Windows MSVC ARM64 ResetHeap1024/LLGo 136.800 ns/op -0.9 ns/op / -0.7% (better)

Compared with d98b43f90a3d measured in the same runner job.

@visualfc
visualfc force-pushed the fix/trivial-iface-fold branch from b0f449e to aaa3c6e Compare September 30, 2026 05:18
@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

8c9553ca7cb4 | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 153510 B -683 B / -0.4% (better) 88742 B 0 B / +0.0%
cprintf/j32-goos-js/LLGo 151500 B -795 B / -0.5% (better) 73165 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 141143 B -105 B / -0.1% (better) 92630 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 152668 B -180 B / -0.1% (better) 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 153380 B -180 B / -0.1% (better) 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3189496 B -23932 B / -0.7% (better) 132442 B 0 B / +0.0%
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3168726 B -23797 B / -0.7% (better) 101635 B 0 B / +0.0%
fmtprintf/j64-emscripten-memory64/LLGo 2929079 B -20954 B / -0.7% (better) 139285 B 0 B / +0.0%
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2333963 B -13388 B / -0.6% (better) 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2330582 B -13431 B / -0.6% (better) 0 B 0 B / 0.0%
j32-emscripten/LLGo 152744 B -684 B / -0.4% (better) 88742 B 0 B / +0.0%
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 150972 B -796 B / -0.5% (better) 73165 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 140468 B -114 B / -0.1% (better) 92629 B -1 B / -0.00108% (better)
reflectcall/j32-emscripten/LLGo 1537919 B -5852 B / -0.4% (better) 105908 B 0 B / +0.0%
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1540185 B -6037 B / -0.4% (better) 90331 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1424269 B -4157 B / -0.3% (better) 111641 B 0 B / +0.0%
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1275225 B -1468 B / -0.1% (better) 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1272820 B -1481 B / -0.1% (better) 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 152315 B -180 B / -0.1% (better) 0 B 0 B / 0.0%
w32-wasi/LLGo 153027 B -180 B / -0.1% (better) 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 6.367 s -89.78 ms / -1.4% (better)
j32-goos-js 6.255 s +117.2 ms / +1.9% (worse)
j64-emscripten-memory64 5.577 s +175.5 ms / +3.2% (worse)
reflectcall/w32-wasi 23.055 s -873.5 ms / -3.7% (better)
w32-goos-wasip1 4.502 s +123.7 ms / +2.8% (worse)
w32-wasi 4.230 s +150.1 ms / +3.7% (worse)

Compared with d98b43f90a3d measured in the same runner job.

@visualfc
visualfc force-pushed the fix/trivial-iface-fold branch 4 times, most recently from d0f5383 to 8c9553c Compare October 2, 2026 02:23
Known concrete-to-interface conversions emit a read-only _llgo_itab$
global (Inter, Type, Hash, Fun) with weak_odr linkage, matching
cmd/compile's go:itab symbols.

Ordinary builds use that global as the runtime vtable. The T2I site is
the only reference, so --gc-sections/-dead_strip can drop itabs that
belong to dead functions. LTO and deadcode-drop keep calling NewItab so
unused interface methods can be dropped; Fun[] would pin them. LTO still
emits the static template for de-virt and erases unused templates after
the plugin runs.

Non-direct interface values (integers, bools, small aggregates) still
use IfaceIndir, matching Go: data is a pointer to a copy. Compile-time
constants go in a WeakODR _llgo_ifacebox$ global instead of AllocU,
like cmd/compile's static temps.

Box names hash the constant's bits, not LLVM's printed form, so
same-length strings in different packages do not collide under
Windows COMDAT. Repeated boxing of the same LLVM value reuses a
per-package pointer cache. Zero-sized constants stay AllocU so boxing
does not mutate the shared sentinel.

Together, a constant T2I is {itab, box} with no runtime call or heap
allocation. Hash is copied from the type descriptor when present.
interequal compares the (inter, _type) pair so a static itab and a
dynamically allocated itab for the same conversion compare equal.
A call whose SSA body is MakeInterface of a Convert/ChangeType of a
constant parameter is lowered at the call site to MakeInterface of
that concrete value. No function-name matching: constant.MakeInt64
and user helpers such as boxMyInt(x int64) any { return myInt(x) }
use the same path.

Depends on static itabs and iface boxes so the folded value is a
compile-time {itab, box} pair.
@visualfc
visualfc force-pushed the fix/trivial-iface-fold branch from 8c9553c to 4f1be2e Compare October 2, 2026 06:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant