Skip to content

ssa: use static itabs for known T2I conversions - #2703

Open
visualfc wants to merge 1 commit into
xgo-dev:mainfrom
visualfc:fix/static-itab
Open

visualfc wants to merge 1 commit into
xgo-dev:mainfrom
visualfc:fix/static-itab

Conversation

@visualfc

@visualfc visualfc commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Known concrete-to-interface (T2I) conversions emit a read-only _llgo_itab$ global ({Inter, Type, Hash, Fun[]}) with weak_odr linkage, matching cmd/compile's go:itab symbols. Identification is by type: the method set is complete and types.AssignableTo holds.

Ordinary builds use that global as the runtime vtable. The T2I site is the only reference, so --gc-sections / -dead_strip can drop itabs that belong to dead functions. LTO and deadcode-drop keep calling NewItab so unused interface methods can be dropped (Fun[] would pin them). LTO still emits the static template for de-virt and erases unused templates after the plugin runs.

NewItab remains for I2I and for types only known at runtime. Hash is copied from the type descriptor when present. interequal compares the (inter, _type) pair so a static itab and a dynamically allocated itab for the same conversion compare equal.

Stack: first of four. Independent of the later PRs.

Test plan

  • go test ./ssa
  • cl/_testrt/staticitab (including any(T).(I) equality)
  • go test ./cl -run TestRunAndTestFromTestmeta
  • go test -tags=dev LTO plugin tests (TestBuildAndCheckSymbolsFromTestltoLTOPlugin, TestLTOPluginAggregateStaticItabDevirt)
  • go test -tags=dev -run TestBuildAndCheckSymbolsFromTestdrop ./cl (including interface_match)
  • fmtprintf file size vs main: __TEXT unchanged; LTO drops remaining _llgo_itab$

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: static itab for known T2I conversions

This is a clean, well-scoped change. staticItab now emits a compile-time itab global for known T2I conversions (interface with methods + assignable concrete type) and unsafeInterface uses it directly instead of always calling runtime.NewItab. Because I2I still goes through NewItab, the two representations of the same dynamic type can hold different tab pointers, so interequal correctly gains a type-pair fallback.

Correctness verified:

  • interequal fast path (x.tab == y.tab) is preserved; the new branch runs only on pointer mismatch and is nil-safe (the x.tab == nil || y.tab == nil guard precedes the field access at alg.go:265).
  • The fallback compares canonical descriptor pointers (inter, _type). abiType/abiIntf globals are name-canonicalized with WeakODRLinkage, so the static itab's inter/_type are the same pointers NewItab stores — pointer equality holds within a binary. It still calls ifaceeq, so data comparison is retained.
  • interhash (alg.go:103) and TypeAssert (interface.go:422) key on tab._type, not the itab pointer, so map hashing and type assertions stay consistent across both representations.
  • The static itab hash field matches abiCommonFields (sha256(typeName)[:4]), so type-switch/hash paths agree — the added comment at interface.go:88 is accurate.

Findings (all non-blocking):

  1. Test coverage gap for the interequal fallback — inline on cl/_testrt/staticitab/in.go. The new fallback branch (alg.go:262-265) is the reason for this PR, but the new test only exercises the fast path.

  2. Stale comment in cl/_testlto/globaldce_static_itab_devirt/in.go:50-52 (outside the diff hunk, so noted here). The comment still reads:

    // Interface equality relies on canonical runtime itab identity. The static template is analysis-only; the direct conversion must still agree with an interface assembled through reflection.
    Both claims now describe the old behavior. Equality no longer relies on itab-pointer identity (the whole point of the interequal change), and the static itab is used directly for T2I rather than being "analysis-only". This is the same stale wording removed elsewhere in the PR ("compile-time certificate", "template disappears before GlobalDCE") — this occurrence was missed. The adjacent SYMBOL check on line 8 was correctly flipped from SYMBOL-NOT to SYMBOL-DAG: _llgo_itab, so only the prose is out of date.

  3. Minor (compile-time only): in staticItab, the two abi.TypeName lookups + sha256 + base64 that build the global name run on every call before the VarOf(name) dedup check (interface.go:67-73). For repeated T2I sites of the same (interface, type) pair this is redundant work per site. Optional: key a fast in-memory map on (rawIntf, concrete) before hashing. No runtime or binary-size impact.

No correctness or security defects found.

panic(i.M())
}
j := boxed(T(41))
if i != j {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The interequal change in runtime/internal/runtime/alg.go adds a fallback branch (comparing inter/_type when itab pointers differ) specifically to handle a T2I static-itab value compared against an I2I NewItab value. But this test never exercises that branch: both i (var i I = T(41)) and j (boxed(T(41))) are T2I conversions, and the static-itab global name is a hash of (intfName, typeName), so both resolve to the same global. x.tab == y.tab holds and the comparison takes the fast path — a regression in the fallback would still pass here.

Consider adding a case where one operand is produced via I2I (e.g. through ChangeInterface / assigning via a second interface type so unsafeInterface is reached with a nil concrete and routes to newItab) so the two operands have different tab pointers, plus an inequality check against a different concrete type.

@codecov

codecov Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 98.00000% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
ssa/interface.go 97.95% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

39c6afb6a9cb | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 7160 B 0 B / +0.0% 387 B 0 B / +0.0% 529.314 ms -26.03 ms / -4.7% (better) 1.266 ms -3.405 us / -0.3% (better)
Linux cprintf-lto 6912 B 0 B / +0.0% 368 B 0 B / +0.0% 554.420 ms -42.28 ms / -7.1% (better) 1.271 ms -28.71 us / -2.2% (better)
Linux fmtprintf 1662424 B -1000 B / -0.1% (better) 494539 B -3898 B / -0.8% (better) 3.807 s +139.1 ms / +3.8% (worse) 3.176 ms -1.083 us / -0.03408% (better)
Linux fmtprintf-lto 1494328 B -6840 B / -0.5% (better) 430732 B -5391 B / -1.2% (better) 11.129 s -89.7 ms / -0.8% (better) 3.437 ms +443.5 us / +14.8% (worse)
Linux println 69384 B +136 B / +0.2% (worse) 16826 B +21 B / +0.1% (worse) 572.107 ms -1.288 ms / -0.2% (better) 1.729 ms +63.81 us / +3.8% (worse)
Linux println-lto 59992 B +16 B / +0.02668% (worse) 14219 B +10 B / +0.1% (worse) 830.373 ms +9.062 ms / +1.1% (worse) 1.574 ms -82.54 us / -5.0% (better)
macOS cprintf 68064 B 0 B / +0.0% 4429 B 0 B / +0.0% 1.055 s +255.6 ms / +31.9% (worse) 3.877 ms -505.8 us / -11.5% (better)
macOS cprintf-lto 84288 B 0 B / +0.0% 193 B 0 B / +0.0% 1.230 s +374.7 ms / +43.8% (worse) 6.316 ms +3.073 ms / +94.7% (worse)
macOS fmtprintf 1506560 B +1728 B / +0.1% (worse) 872632 B -2852 B / -0.3% (better) 4.117 s +374.4 ms / +10.0% (worse) 5.762 ms +111.9 us / +2.0% (worse)
macOS fmtprintf-lto 1192688 B -16 B / -0.001341% (better) 845172 B -3808 B / -0.4% (better) 10.954 s -3.276 s / -23.0% (better) 5.792 ms -1.187 ms / -17.0% (better)
macOS println 117296 B +80 B / +0.1% (worse) 37609 B +44 B / +0.1% (worse) 981.547 ms +133.6 ms / +15.8% (worse) 5.467 ms +1.123 ms / +25.9% (worse)
macOS println-lto 119472 B 0 B / +0.0% 34976 B +32 B / +0.1% (worse) 1.200 s +113.8 ms / +10.5% (worse) 5.323 ms -2.259 ms / -29.8% (better)
Windows MinGW cprintf 19456 B 0 B / +0.0% 4550 B 0 B / +0.0% 1.321 s -29.32 ms / -2.2% (better) 3.527 ms +25.6 us / +0.7% (worse)
Windows MinGW cprintf-lto 17920 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.347 s +16.19 ms / +1.2% (worse) 3.633 ms +118.9 us / +3.4% (worse)
Windows MinGW fmtprintf 1935872 B +2560 B / +0.1% (worse) 598502 B -528 B / -0.1% (better) 4.175 s -43.23 ms / -1.0% (better) 8.433 ms -695.4 us / -7.6% (better)
Windows MinGW fmtprintf-lto 1953792 B -4096 B / -0.2% (better) 544758 B -2560 B / -0.5% (better) 10.394 s +101.8 ms / +1.0% (worse) 8.675 ms -195.1 us / -2.2% (better)
Windows MinGW println 75776 B 0 B / +0.0% 25158 B +16 B / +0.1% (worse) 1.331 s +18.76 ms / +1.4% (worse) 6.968 ms -68.2 us / -1.0% (better)
Windows MinGW println-lto 69632 B +512 B / +0.7% (worse) 22006 B +16 B / +0.1% (worse) 1.552 s -37.82 ms / -2.4% (better) 6.693 ms -322.5 us / -4.6% (better)
Windows MinGW 386 cprintf 43520 B 0 B / +0.0% 5326 B 0 B / +0.0% 916.871 ms -132.5 ms / -12.6% (better) 4.510 ms -138.9 us / -3.0% (better)
Windows MinGW 386 cprintf-lto 20992 B 0 B / +0.0% 5094 B 0 B / +0.0% 932.958 ms -98.29 ms / -9.5% (better) 4.225 ms -372.9 us / -8.1% (better)
Windows MinGW 386 fmtprintf 1896448 B 0 B / +0.0% 470238 B -2224 B / -0.5% (better) 3.114 s -133.5 ms / -4.1% (better) 9.754 ms -305.8 us / -3.0% (better)
Windows MinGW 386 fmtprintf-lto 2186752 B +5632 B / +0.3% (worse) 449358 B -1916 B / -0.4% (better) 7.162 s -329.2 ms / -4.4% (better) 9.188 ms -1.073 ms / -10.5% (better)
Windows MinGW 386 println 96256 B 0 B / +0.0% 21490 B +16 B / +0.1% (worse) 918.308 ms -25.63 ms / -2.7% (better) 7.593 ms +424.1 us / +5.9% (worse)
Windows MinGW 386 println-lto 74752 B +512 B / +0.7% (worse) 19338 B +4 B / +0.02069% (worse) 1.172 s +35.97 ms / +3.2% (worse) 7.586 ms -125.8 us / -1.6% (better)
Windows MinGW ARM64 cprintf 18944 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.553 s +26.9 us / +0.001732% (worse) 6.336 ms -292.4 us / -4.4% (better)
Windows MinGW ARM64 cprintf-lto 17920 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.590 s -1.852 ms / -0.1% (better) 6.543 ms -215.9 us / -3.2% (better)
Windows MinGW ARM64 fmtprintf 1820160 B +512 B / +0.02814% (worse) 507172 B -2544 B / -0.5% (better) 4.224 s -40.13 ms / -0.9% (better) 12.376 ms -329.5 us / -2.6% (better)
Windows MinGW ARM64 fmtprintf-lto 1878528 B -1536 B / -0.1% (better) 473552 B -2624 B / -0.6% (better) 9.506 s -147.7 ms / -1.5% (better) 12.867 ms -707.1 us / -5.2% (better)
Windows MinGW ARM64 println 72192 B 0 B / +0.0% 23912 B +28 B / +0.1% (worse) 1.555 s -23.14 ms / -1.5% (better) 11.413 ms +73.3 us / +0.6% (worse)
Windows MinGW ARM64 println-lto 68608 B 0 B / +0.0% 21248 B +16 B / +0.1% (worse) 1.745 s -40.96 ms / -2.3% (better) 11.140 ms -1.041 ms / -8.5% (better)
Windows MSVC cprintf 120320 B 0 B / +0.0% 65798 B 0 B / +0.0% 1.072 s -19.38 ms / -1.8% (better) 3.303 ms -43 us / -1.3% (better)
Windows MSVC cprintf-lto 119808 B 0 B / +0.0% 65734 B 0 B / +0.0% 1.089 s +15.4 ms / +1.4% (worse) 3.386 ms +45.5 us / +1.4% (worse)
Windows MSVC fmtprintf 1644032 B +512 B / +0.03115% (worse) 694054 B -528 B / -0.1% (better) 3.863 s +173.9 ms / +4.7% (worse) 9.706 ms -273.3 us / -2.7% (better)
Windows MSVC fmtprintf-lto 1632256 B -3072 B / -0.2% (better) 644454 B -2624 B / -0.4% (better) 9.526 s +83.51 ms / +0.9% (worse) 9.890 ms +440.9 us / +4.7% (worse)
Windows MSVC println 195072 B +512 B / +0.3% (worse) 120854 B +32 B / +0.02649% (worse) 1.114 s +57.06 ms / +5.4% (worse) 7.948 ms +938.7 us / +13.4% (worse)
Windows MSVC println-lto 192512 B 0 B / +0.0% 118358 B 0 B / +0.0% 1.351 s +99.37 ms / +7.9% (worse) 9.188 ms +2.021 ms / +28.2% (worse)
Windows MSVC 386 cprintf 9728 B 0 B / +0.0% 3931 B 0 B / +0.0% 889.564 ms +17.93 ms / +2.1% (worse) 5.217 ms -30.2 us / -0.6% (better)
Windows MSVC 386 cprintf-lto 9216 B 0 B / +0.0% 3853 B 0 B / +0.0% 889.180 ms -5.925 ms / -0.7% (better) 5.192 ms +260.4 us / +5.3% (worse)
Windows MSVC 386 fmtprintf 1202688 B -2048 B / -0.2% (better) 453628 B -2224 B / -0.5% (better) 3.285 s +124.5 ms / +3.9% (worse) 11.041 ms +695.8 us / +6.7% (worse)
Windows MSVC 386 fmtprintf-lto 1241088 B -512 B / -0.04124% (better) 425115 B -2096 B / -0.5% (better) 7.573 s +196.4 ms / +2.7% (worse) 10.762 ms -271.4 us / -2.5% (better)
Windows MSVC 386 println 36352 B 0 B / +0.0% 20340 B +16 B / +0.1% (worse) 1.091 s +109.4 ms / +11.1% (worse) 8.902 ms +315.1 us / +3.7% (worse)
Windows MSVC 386 println-lto 35328 B -512 B / -1.4% (better) 18565 B 0 B / +0.0% 1.072 s +28.9 ms / +2.8% (worse) 8.546 ms +55.1 us / +0.6% (worse)
Windows MSVC ARM64 cprintf 11776 B 0 B / +0.0% 4192 B 0 B / +0.0% 1.245 s -31.82 ms / -2.5% (better) 7.167 ms +86.2 us / +1.2% (worse)
Windows MSVC ARM64 cprintf-lto 10752 B 0 B / +0.0% 4084 B 0 B / +0.0% 1.275 s +830.2 us / +0.1% (worse) 7.062 ms +78.5 us / +1.1% (worse)
Windows MSVC ARM64 fmtprintf 1384960 B -2048 B / -0.1% (better) 507112 B -2544 B / -0.5% (better) 3.887 s -19.08 ms / -0.5% (better) 15.069 ms +524 us / +3.6% (worse)
Windows MSVC ARM64 fmtprintf-lto 1404416 B -1024 B / -0.1% (better) 474212 B -2624 B / -0.6% (better) 8.908 s -157.8 ms / -1.7% (better) 14.548 ms -1.567 ms / -9.7% (better)
Windows MSVC ARM64 println 45056 B 0 B / +0.0% 23940 B +32 B / +0.1% (worse) 1.267 s +11.55 ms / +0.9% (worse) 13.231 ms +192.6 us / +1.5% (worse)
Windows MSVC ARM64 println-lto 42496 B 0 B / +0.0% 21396 B +16 B / +0.1% (worse) 1.463 s +11.15 ms / +0.8% (worse) 12.235 ms -1.108 ms / -8.3% (better)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.520 ns/op +0.08 ns/op / +0.6% (worse)
Linux BenchmarkMergeCompilerFlags 199.500 ns/op -20.6 ns/op / -9.4% (better)
Linux BenchmarkMergeLinkerFlags 126 ns/op -18.6 ns/op / -12.9% (better)
Linux BenchmarkChannelBuffered 55.200 ns/op -0.51 ns/op / -0.9% (better)
Linux BenchmarkChannelHandoff 12812 ns/op -1107 ns/op / -8.0% (better)
Linux BenchmarkDefer 50.510 ns/op +4.09 ns/op / +8.8% (worse)
Linux BenchmarkDirectCall 1.167 ns/op -0.421 ns/op / -26.5% (better)
Linux BenchmarkGlobalRead 1.167 ns/op -0.031 ns/op / -2.6% (better)
Linux BenchmarkGlobalWrite 7.771 ns/op -0.011 ns/op / -0.1% (better)
Linux BenchmarkGoroutine 24787 ns/op +737 ns/op / +3.1% (worse)
Linux BenchmarkInterfaceCall 5.838 ns/op -0.027 ns/op / -0.5% (better)
Linux BenchmarkRuntimeGetG 2.334 ns/op -0.676 ns/op / -22.5% (better)
macOS BenchmarkLookupPCRandom 27.490 ns/op +8.5 ns/op / +44.8% (worse)
macOS BenchmarkMergeCompilerFlags 159.600 ns/op +12.1 ns/op / +8.2% (worse)
macOS BenchmarkMergeLinkerFlags 105 ns/op +11.7 ns/op / +12.5% (worse)
macOS BenchmarkChannelBuffered 37.210 ns/op -0.58 ns/op / -1.5% (better)
macOS BenchmarkChannelHandoff 7723 ns/op -5060 ns/op / -39.6% (better)
macOS BenchmarkDefer 61.200 ns/op +10.02 ns/op / +19.6% (worse)
macOS BenchmarkDirectCall 1.295 ns/op -0.166 ns/op / -11.4% (better)
macOS BenchmarkGlobalRead 1.250 ns/op -0.041 ns/op / -3.2% (better)
macOS BenchmarkGlobalWrite 1.671 ns/op +0.116 ns/op / +7.5% (worse)
macOS BenchmarkGoroutine 79933 ns/op +32789 ns/op / +69.6% (worse)
macOS BenchmarkInterfaceCall 5.055 ns/op -0.376 ns/op / -6.9% (better)
macOS BenchmarkRuntimeGetG 2.985 ns/op -1.068 ns/op / -26.4% (better)
Windows MinGW BenchmarkLookupPCRandom 13.080 ns/op -0.17 ns/op / -1.3% (better)
Windows MinGW BenchmarkMergeCompilerFlags 620.200 ns/op +7.7 ns/op / +1.3% (worse)
Windows MinGW BenchmarkMergeLinkerFlags 541.500 ns/op +12.5 ns/op / +2.4% (worse)
Windows MinGW BenchmarkChannelBuffered 30.120 ns/op -1.25 ns/op / -4.0% (better)
Windows MinGW BenchmarkChannelHandoff 911.400 ns/op -13.6 ns/op / -1.5% (better)
Windows MinGW BenchmarkDefer 56.550 ns/op -0.86 ns/op / -1.5% (better)
Windows MinGW BenchmarkDirectCall 1.549 ns/op +0.001 ns/op / +0.1% (worse)
Windows MinGW BenchmarkGlobalRead 1.547 ns/op -0.316 ns/op / -17.0% (better)
Windows MinGW BenchmarkGlobalWrite 2.471 ns/op +0.006 ns/op / +0.2% (worse)
Windows MinGW BenchmarkGoroutine 94098 ns/op -6552 ns/op / -6.5% (better)
Windows MinGW BenchmarkInterfaceCall 8.367 ns/op +0.297 ns/op / +3.7% (worse)
Windows MinGW BenchmarkRuntimeGetG 2.486 ns/op +0.309 ns/op / +14.2% (worse)
Windows MinGW 386 BenchmarkLookupPCRandom 51.980 ns/op -7.79 ns/op / -13.0% (better)
Windows MinGW 386 BenchmarkMergeCompilerFlags 574.200 ns/op -54.7 ns/op / -8.7% (better)
Windows MinGW 386 BenchmarkMergeLinkerFlags 518.100 ns/op -75.2 ns/op / -12.7% (better)
Windows MinGW 386 BenchmarkChannelBuffered 35.250 ns/op -0.85 ns/op / -2.4% (better)
Windows MinGW 386 BenchmarkChannelHandoff 3481 ns/op +37 ns/op / +1.1% (worse)
Windows MinGW 386 BenchmarkDefer 33.160 ns/op +0.14 ns/op / +0.4% (worse)
Windows MinGW 386 BenchmarkDirectCall 0.368 ns/op +0.0085 ns/op / +2.4% (worse)
Windows MinGW 386 BenchmarkGlobalRead 0.688 ns/op -0.0004 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGlobalWrite 12.690 ns/op +0.01 ns/op / +0.1% (worse)
Windows MinGW 386 BenchmarkGoroutine 178040 ns/op -3901 ns/op / -2.1% (better)
Windows MinGW 386 BenchmarkInterfaceCall 4.385 ns/op +0.244 ns/op / +5.9% (worse)
Windows MinGW 386 BenchmarkRuntimeGetG 0.959 ns/op +0.0006 ns/op / +0.1% (worse)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.040 ns/op -0.07 ns/op / -0.6% (better)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 565 ns/op -10.1 ns/op / -1.8% (better)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 538 ns/op +2.9 ns/op / +0.5% (worse)
Windows MinGW ARM64 BenchmarkChannelBuffered 37.640 ns/op -1.75 ns/op / -4.4% (better)
Windows MinGW ARM64 BenchmarkChannelHandoff 2075 ns/op -224 ns/op / -9.7% (better)
Windows MinGW ARM64 BenchmarkDefer 56.640 ns/op +0.34 ns/op / +0.6% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.590 ns/op +0.0002 ns/op / +0.03394% (worse)
Windows MinGW ARM64 BenchmarkGlobalRead 0.590 ns/op -0.0736 ns/op / -11.1% (better)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.663 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkGoroutine 60581 ns/op -1156 ns/op / -1.9% (better)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.160 ns/op +0.018 ns/op / +0.4% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.775 ns/op +0.004 ns/op / +0.2% (worse)
Windows MSVC BenchmarkLookupPCRandom 12.300 ns/op -0.04 ns/op / -0.3% (better)
Windows MSVC BenchmarkMergeCompilerFlags 538.500 ns/op -2.5 ns/op / -0.5% (better)
Windows MSVC BenchmarkMergeLinkerFlags 471.400 ns/op +7.3 ns/op / +1.6% (worse)
Windows MSVC BenchmarkChannelBuffered 29.270 ns/op +0.48 ns/op / +1.7% (worse)
Windows MSVC BenchmarkChannelHandoff 1119 ns/op -84 ns/op / -7.0% (better)
Windows MSVC BenchmarkDefer 55.110 ns/op -2.11 ns/op / -3.7% (better)
Windows MSVC BenchmarkDirectCall 1.400 ns/op -0.119 ns/op / -7.8% (better)
Windows MSVC BenchmarkGlobalRead 1.748 ns/op 0 ns/op / +0.0%
Windows MSVC BenchmarkGlobalWrite 2.794 ns/op +0.009 ns/op / +0.3% (worse)
Windows MSVC BenchmarkGoroutine 73908 ns/op +1592 ns/op / +2.2% (worse)
Windows MSVC BenchmarkInterfaceCall 9.106 ns/op +0.554 ns/op / +6.5% (worse)
Windows MSVC BenchmarkRuntimeGetG 1.810 ns/op -0.289 ns/op / -13.8% (better)
Windows MSVC 386 BenchmarkLookupPCRandom 59.930 ns/op -0.26 ns/op / -0.4% (better)
Windows MSVC 386 BenchmarkMergeCompilerFlags 644.100 ns/op -24.3 ns/op / -3.6% (better)
Windows MSVC 386 BenchmarkMergeLinkerFlags 617.700 ns/op +20 ns/op / +3.3% (worse)
Windows MSVC 386 BenchmarkChannelBuffered 45.570 ns/op 0 ns/op / +0.0%
Windows MSVC 386 BenchmarkChannelHandoff 1483 ns/op -119 ns/op / -7.4% (better)
Windows MSVC 386 BenchmarkDefer 38.250 ns/op -1.63 ns/op / -4.1% (better)
Windows MSVC 386 BenchmarkDirectCall 1.119 ns/op +0.1752 ns/op / +18.6% (worse)
Windows MSVC 386 BenchmarkGlobalRead 1.116 ns/op -0.049 ns/op / -4.2% (better)
Windows MSVC 386 BenchmarkGlobalWrite 14.780 ns/op -0.01 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGoroutine 207941 ns/op -7196 ns/op / -3.3% (better)
Windows MSVC 386 BenchmarkInterfaceCall 4.784 ns/op -0.01 ns/op / -0.2% (better)
Windows MSVC 386 BenchmarkRuntimeGetG 1.352 ns/op +0.185 ns/op / +15.9% (worse)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.160 ns/op +0.12 ns/op / +1.0% (worse)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 581.800 ns/op +10.4 ns/op / +1.8% (worse)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 541 ns/op +0.6 ns/op / +0.1% (worse)
Windows MSVC ARM64 BenchmarkChannelBuffered 37.430 ns/op -0.25 ns/op / -0.7% (better)
Windows MSVC ARM64 BenchmarkChannelHandoff 3916 ns/op -84 ns/op / -2.1% (better)
Windows MSVC ARM64 BenchmarkDefer 60.760 ns/op -1.58 ns/op / -2.5% (better)
Windows MSVC ARM64 BenchmarkDirectCall 0.590 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 BenchmarkGlobalRead 0.665 ns/op +0.0016 ns/op / +0.2% (worse)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.760 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 BenchmarkGoroutine 56868 ns/op -2279 ns/op / -3.9% (better)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.284 ns/op +0.146 ns/op / +3.5% (worse)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.813 ns/op +0.01 ns/op / +0.6% (worse)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 903.600 ns/op -16.6 ns/op / -1.8% (better)
Linux AfterFuncZeroDelivery/LLGo 36094 ns/op -2010 ns/op / -5.3% (better)
Linux CreateStop/Go 291.900 ns/op -2.4 ns/op / -0.8% (better)
Linux CreateStop/LLGo 1721 ns/op -161 ns/op / -8.6% (better)
Linux RearmStopped/Go 115.500 ns/op -0.5 ns/op / -0.4% (better)
Linux RearmStopped/LLGo 1096 ns/op -143 ns/op / -11.5% (better)
Linux ResetActive/Go 67.420 ns/op -1.7 ns/op / -2.5% (better)
Linux ResetActive/LLGo 749.700 ns/op +94.5 ns/op / +14.4% (worse)
Linux ResetHeap1024/Go 67.080 ns/op -0.27 ns/op / -0.4% (better)
Linux ResetHeap1024/LLGo 179.200 ns/op +5.3 ns/op / +3.0% (worse)
macOS AfterFuncZeroDelivery/Go 607.900 ns/op -197.5 ns/op / -24.5% (better)
macOS AfterFuncZeroDelivery/LLGo 121525 ns/op -249 ns/op / -0.2% (better)
macOS CreateStop/Go 275.100 ns/op +96 ns/op / +53.6% (worse)
macOS CreateStop/LLGo 914.300 ns/op +23.6 ns/op / +2.6% (worse)
macOS RearmStopped/Go 77.300 ns/op -3.93 ns/op / -4.8% (better)
macOS RearmStopped/LLGo 513.700 ns/op -230.4 ns/op / -31.0% (better)
macOS ResetActive/Go 76.760 ns/op +22.29 ns/op / +40.9% (worse)
macOS ResetActive/LLGo 192.800 ns/op -80.9 ns/op / -29.6% (better)
macOS ResetHeap1024/Go 57.550 ns/op -4.29 ns/op / -6.9% (better)
macOS ResetHeap1024/LLGo 103.100 ns/op -1.2 ns/op / -1.2% (better)
Windows MinGW AfterFuncZeroDelivery/Go 558.100 ns/op -0.5 ns/op / -0.1% (better)
Windows MinGW AfterFuncZeroDelivery/LLGo 195108 ns/op -1567 ns/op / -0.8% (better)
Windows MinGW CreateStop/Go 117.100 ns/op +2.7 ns/op / +2.4% (worse)
Windows MinGW CreateStop/LLGo 415.200 ns/op -14.1 ns/op / -3.3% (better)
Windows MinGW RearmStopped/Go 31.300 ns/op -0.46 ns/op / -1.4% (better)
Windows MinGW RearmStopped/LLGo 269.400 ns/op -45 ns/op / -14.3% (better)
Windows MinGW ResetActive/Go 20.070 ns/op -0.03 ns/op / -0.1% (better)
Windows MinGW ResetActive/LLGo 161.200 ns/op +0.3 ns/op / +0.2% (worse)
Windows MinGW ResetHeap1024/Go 20.530 ns/op +0.03 ns/op / +0.1% (worse)
Windows MinGW ResetHeap1024/LLGo 124.400 ns/op +0.1 ns/op / +0.1% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/Go 693.900 ns/op -0.4 ns/op / -0.1% (better)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 163395 ns/op -20562 ns/op / -11.2% (better)
Windows MinGW 386 CreateStop/Go 181.100 ns/op -14.6 ns/op / -7.5% (better)
Windows MinGW 386 CreateStop/LLGo 572 ns/op -93.3 ns/op / -14.0% (better)
Windows MinGW 386 RearmStopped/Go 64.810 ns/op +0.02 ns/op / +0.03087% (worse)
Windows MinGW 386 RearmStopped/LLGo 510.600 ns/op +88.2 ns/op / +20.9% (worse)
Windows MinGW 386 ResetActive/Go 31.060 ns/op -0.09 ns/op / -0.3% (better)
Windows MinGW 386 ResetActive/LLGo 316.900 ns/op +12.4 ns/op / +4.1% (worse)
Windows MinGW 386 ResetHeap1024/Go 31.250 ns/op -0.14 ns/op / -0.4% (better)
Windows MinGW 386 ResetHeap1024/LLGo 123.700 ns/op -13.5 ns/op / -9.8% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 661.800 ns/op -2.7 ns/op / -0.4% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 145847 ns/op -206 ns/op / -0.1% (better)
Windows MinGW ARM64 CreateStop/Go 196.400 ns/op -11.7 ns/op / -5.6% (better)
Windows MinGW ARM64 CreateStop/LLGo 357.300 ns/op -6.2 ns/op / -1.7% (better)
Windows MinGW ARM64 RearmStopped/Go 70.570 ns/op -0.08 ns/op / -0.1% (better)
Windows MinGW ARM64 RearmStopped/LLGo 253.700 ns/op +2 ns/op / +0.8% (worse)
Windows MinGW ARM64 ResetActive/Go 31.090 ns/op -0.13 ns/op / -0.4% (better)
Windows MinGW ARM64 ResetActive/LLGo 129.500 ns/op +4.5 ns/op / +3.6% (worse)
Windows MinGW ARM64 ResetHeap1024/Go 31.140 ns/op -0.1 ns/op / -0.3% (better)
Windows MinGW ARM64 ResetHeap1024/LLGo 126.800 ns/op -1 ns/op / -0.8% (better)
Windows MSVC AfterFuncZeroDelivery/Go 471.100 ns/op -41.3 ns/op / -8.1% (better)
Windows MSVC AfterFuncZeroDelivery/LLGo 143232 ns/op +277 ns/op / +0.2% (worse)
Windows MSVC CreateStop/Go 117.100 ns/op +1.4 ns/op / +1.2% (worse)
Windows MSVC CreateStop/LLGo 462.400 ns/op +11.1 ns/op / +2.5% (worse)
Windows MSVC RearmStopped/Go 31.560 ns/op +0.01 ns/op / +0.0317% (worse)
Windows MSVC RearmStopped/LLGo 280.500 ns/op +0.3 ns/op / +0.1% (worse)
Windows MSVC ResetActive/Go 19.100 ns/op +0.01 ns/op / +0.1% (worse)
Windows MSVC ResetActive/LLGo 160.100 ns/op +6.8 ns/op / +4.4% (worse)
Windows MSVC ResetHeap1024/Go 19.090 ns/op -0.08 ns/op / -0.4% (better)
Windows MSVC ResetHeap1024/LLGo 130.600 ns/op -0.3 ns/op / -0.2% (better)
Windows MSVC 386 AfterFuncZeroDelivery/Go 861.200 ns/op +0.1 ns/op / +0.01161% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 170000 ns/op -1926 ns/op / -1.1% (better)
Windows MSVC 386 CreateStop/Go 217.800 ns/op -1.3 ns/op / -0.6% (better)
Windows MSVC 386 CreateStop/LLGo 888.700 ns/op +245.8 ns/op / +38.2% (worse)
Windows MSVC 386 RearmStopped/Go 80.420 ns/op +0.23 ns/op / +0.3% (worse)
Windows MSVC 386 RearmStopped/LLGo 406.400 ns/op +12.9 ns/op / +3.3% (worse)
Windows MSVC 386 ResetActive/Go 38.690 ns/op 0 ns/op / +0.0%
Windows MSVC 386 ResetActive/LLGo 273.100 ns/op +11.5 ns/op / +4.4% (worse)
Windows MSVC 386 ResetHeap1024/Go 39 ns/op +0.02 ns/op / +0.1% (worse)
Windows MSVC 386 ResetHeap1024/LLGo 142.200 ns/op -1.3 ns/op / -0.9% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 663.700 ns/op +1.5 ns/op / +0.2% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 146318 ns/op -1746 ns/op / -1.2% (better)
Windows MSVC ARM64 CreateStop/Go 204.300 ns/op -3.7 ns/op / -1.8% (better)
Windows MSVC ARM64 CreateStop/LLGo 381.300 ns/op -12.5 ns/op / -3.2% (better)
Windows MSVC ARM64 RearmStopped/Go 70.560 ns/op -0.12 ns/op / -0.2% (better)
Windows MSVC ARM64 RearmStopped/LLGo 272 ns/op -3.4 ns/op / -1.2% (better)
Windows MSVC ARM64 ResetActive/Go 31.040 ns/op -0.16 ns/op / -0.5% (better)
Windows MSVC ARM64 ResetActive/LLGo 134.900 ns/op +1.5 ns/op / +1.1% (worse)
Windows MSVC ARM64 ResetHeap1024/Go 31.090 ns/op -0.06 ns/op / -0.2% (better)
Windows MSVC ARM64 ResetHeap1024/LLGo 137.500 ns/op +0.2 ns/op / +0.1% (worse)

Compared with f06143ba2aac measured in the same runner job.

@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

39c6afb6a9cb | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 154272 B +79 B / +0.1% (worse) 88742 B 0 B / +0.0%
cprintf/j32-goos-js/LLGo 152374 B +79 B / +0.1% (worse) 73165 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 141713 B +465 B / +0.3% (worse) 92630 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 152951 B +103 B / +0.1% (worse) 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 153663 B +103 B / +0.1% (worse) 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3204015 B -9413 B / -0.3% (better) 132442 B 0 B / +0.0%
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3182989 B -9534 B / -0.3% (better) 101635 B 0 B / +0.0%
fmtprintf/j64-emscripten-memory64/LLGo 2941982 B -8051 B / -0.3% (better) 139285 B 0 B / +0.0%
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2341037 B -6314 B / -0.3% (better) 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2337786 B -6227 B / -0.3% (better) 0 B 0 B / 0.0%
j32-emscripten/LLGo 153507 B +79 B / +0.1% (worse) 88742 B 0 B / +0.0%
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 151847 B +79 B / +0.1% (worse) 73165 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 141048 B +466 B / +0.3% (worse) 92630 B 0 B / +0.0%
reflectcall/j32-emscripten/LLGo 1543496 B -275 B / -0.01781% (better) 105908 B 0 B / +0.0%
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1545909 B -313 B / -0.02024% (better) 90331 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1428577 B +151 B / +0.01057% (worse) 111641 B 0 B / +0.0%
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1276923 B +230 B / +0.01802% (worse) 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1274566 B +265 B / +0.0208% (worse) 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 152598 B +103 B / +0.1% (worse) 0 B 0 B / 0.0%
w32-wasi/LLGo 153310 B +103 B / +0.1% (worse) 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 6.139 s +165.8 ms / +2.8% (worse)
j32-goos-js 5.835 s -3.843 ms / -0.1% (better)
j64-emscripten-memory64 5.245 s +55.05 ms / +1.1% (worse)
reflectcall/w32-wasi 25.064 s +1.379 s / +5.8% (worse)
w32-goos-wasip1 4.317 s +262.2 ms / +6.5% (worse)
w32-wasi 4.056 s +143.4 ms / +3.7% (worse)

Compared with f06143ba2aac measured in the same runner job.

@visualfc
visualfc force-pushed the fix/static-itab branch 5 times, most recently from edc884a to 3109759 Compare September 30, 2026 13:34
Known concrete-to-interface conversions emit a read-only _llgo_itab$
global (Inter, Type, Hash, Fun) with weak_odr linkage, matching
cmd/compile's go:itab symbols.

Ordinary builds use that global as the runtime vtable. The T2I site is
the only reference, so --gc-sections/-dead_strip can drop itabs that
belong to dead functions. LTO and deadcode-drop keep calling NewItab so
unused interface methods can be dropped; Fun[] would pin them. LTO still
emits the static template for de-virt and erases unused templates after
the plugin runs.

Hash is copied from the type descriptor when present. interequal
compares the (inter, _type) pair so a static itab and a dynamically
allocated itab for the same conversion compare equal.
@visualfc

visualfc commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator Author

Folded into #2704 as a single commit on latest main. This PR can be closed once #2704 is the review target.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant