Skip to content

cl: lower large constant map literals to a loop - #2706

Open
visualfc wants to merge 3 commits into
xgo-dev:mainfrom
visualfc:fix/static-maplit
Open

visualfc wants to merge 3 commits into
xgo-dev:mainfrom
visualfc:fix/static-maplit

Conversation

@visualfc

@visualfc visualfc commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Map literals with more than 25 constant keys become a counted mapassign loop over private constant key/value arrays, matching cmd/compile's maplit.

Composite values whose fields are constants or trivial iface constructors (the SSA shape from #2705) are rebuilt as LLVM constants so the value array is a ConstArray. Mixed MakeInt64/MakeBool-style boxes in one literal are allowed. Identification is by SSA (MakeMap, constant keys, complit fields), not by function name.

Depends on: #2704 and #2705. The large {string, constant.Value} literals need folded constant ifaces. Merge after #2705. This branch contains #2704–#2705 plus this commit; only the latest commit is in scope.

Test plan

  • go test ./cl -run TestCollect
  • cl/_testrt/maplit (map[string]int and mixed int/bool complits)

@fennoai fennoai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: static maplit + static itab/interface boxing

This is a well-structured change. The interface-representation split (T2I now uses static itab globals while I2I still goes through NewItab) is handled carefully, and the interequal fix correctly restores cross-representation equality. interhash hashes by tab._type, and both itab representations point _type/hash at the same ODR-merged type descriptor, so map insert/lookup stays consistent. The maplit-to-loop lowering is safe by construction (keys required constant, arrays synthesized at known length).

Findings below are mostly correctness-hardening and readability; none are blocking.

Verified sound: full 32-byte sha256 in box/itab names (no truncation collision), mapLitIndexAddr only indexes fixed-length synthesized arrays, static globals are SetGlobalConstant(true) and never mutated after Init, and the staticItab.hash computation matches abiCommonFields.

Comment thread cl/maplit.go Outdated
if classifyStaticComplit(updates, plan) {
plan.kind = mapLitStaticComplit
} else {
plan.skip = nil

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] classifyStaticComplit mutates plan on failure; plan.skip=nil is misleading

classifyStaticComplit mutates the passed plan (sets plan.structTy at the top, appends to skip) before it may return false on a later entry. After a false return, plan.structTy is left populated while plan.kind stays mapLitGeneric. The else { plan.skip = nil } here only resets one of the two mutated fields, and plan.skip is already nil on a freshly-allocated plan, so the line is both dead and misleading. This works today only because the generic branch of compileMapLitUpdate never reads plan.structTy — a fragile invariant. Suggest making classifyStaticComplit compute structTy/skip in locals and commit to plan only on success (return them, assign in the caller), removing the need for plan.skip = nil.

Comment thread cl/maplit.go
return false
}
if block != afterBlock {
return block.Index > afterBlock.Index

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] instrIsAfter treats BasicBlock.Index as execution order for all functions

instrIsAfter uses block.Index > afterBlock.Index as "executes after". Block index is the position in fn.Blocks (roughly RPO for x/tools SSA) and does not guarantee execution/dominance order across branches. collectLargeMapLits runs for every function (compile.go:759), not just the synthetic package initializer. For a large map literal inside a normal function with branching, a makeMap referrer in a higher-indexed block that is actually on a parallel/earlier path could be wrongly judged "after" the last update, permitting an unsafe delay. The sibling static-map-init pass gates on initFn.Synthetic == "package initializer". Consider restricting this optimization to synthetic init functions, or strengthening the ordering check (e.g. real dominance) and documenting the assumption.

Comment thread cl/constiface.go
if !ok || c.Value == nil {
return llssa.Expr{}, false
}
x := b.Const(c.Value, p.type_(concrete, llssa.InGo))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Constant re-typing in foldConstantMakeValue may skip Convert semantics

analyzeTrivialIfaceBox accepts a *ssa.Convert in the unwrap chain and returns mi.X.Type() as the concrete type. Here the source constant c has the parameter's type, but the value is materialized directly at concrete (= mi.X.Type()). If the chain contains a real numeric Convert (e.g. param int32 widened/narrowed to int64, or a representation-changing conversion), building b.Const(c.Value, concrete) reinterprets the untyped constant instead of applying the conversion, which can diverge from actually calling the constructor. ChangeType/ChangeInterface are representation-preserving and safe. Consider rejecting *ssa.Convert in the accepted chain (or applying the conversion to the constant). Low likelihood in practice, but worth hardening.

Comment thread ssa/interface.go
rawIntf.NumMethods() == 0 || concrete == nil {
if rawIntf.NumMethods() == 0 || concrete == nil {
return Expr{}, false
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] staticItab does method-set work before the dedup cache check

staticItab computes types.NewMethodSet(concrete) plus per-method Lookup and AssignableTo before it builds the cache key name and checks b.Pkg.VarOf(name). Since the itab is deduplicated per (interface, concrete) pair and the cache key does not need the method set, repeated boxings of the same pair recompute the full method set only to discard it on the cache hit. Hoisting the VarOf(name) early-return above the method-set work would make repeated boxings O(1).

Comment thread ssa/datastruct.go
return !v.impl.IsNil() && !v.impl.IsAConstant().IsNil()
}

var mapLitSeq atomic.Uint64

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] mapLitArray globals use a non-dedup, order-dependent counter

mapLitArray names emitted constant arrays via a process-global atomic.Uint64 (mapLitSeq), unlike the other two globals in this PR (staticItab, staticIfaceBox) which are content-addressed and thus deduplicated + deterministic. Consequences: identical constant map-literal arrays never merge (code/data-size bloat), and names depend on compilation order, hurting build reproducibility under parallel/reordered compilation. Content-hashing the array would fix both.

Comment thread runtime/internal/runtime/alg.go Outdated
if x.tab == y.tab {
return ifaceeq(x.tab, x.data, y.data)
}
// T2I uses a static itab global; I2I still goes through NewItab.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] interequal/foldConstantMakeValue comments slightly overstate the mapping

Two minor doc nits: (1) alg.go's // T2I uses a static itab global; I2I still goes through NewItab — T2I falls back to NewItab too when staticItab returns false (empty interface, unassignable, no-interface-method), so it is not a strict one-to-one mapping; the equality logic is still correct. (2) foldConstantMakeValue's doc lists Convert/ChangeType but the recognizer (analyzeTrivialIfaceBox) also accepts ChangeInterface; align the two comments.

Comment thread ssa/interface.go Outdated
}
g := b.Pkg.NewVarEx(name, prog.Pointer(typ))
g.Init(x)
if g.impl.IsNil() {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] staticIfaceBox: potential leftover global if Init fails

If g.Init(x) leaves g.impl nil, the function returns (zero, false), but the named global was already created via NewVarEx. A later call with the same name would hit the VarOf(name) cache and return that leftover, uninitialized global with true. Confirm NewVarEx does not register the var when init fails, or clean up on the failure path. Also worth a one-line note that x.impl.String() is used purely as a dedup key so correctness does not depend on LLVM's textual-print stability.

@codecov

codecov Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 84.64819% with 72 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
cl/maplit.go 78.00% 53 Missing ⚠️
ssa/datastruct.go 86.20% 8 Missing ⚠️
cl/constiface.go 87.03% 7 Missing ⚠️
ssa/interface.go 96.29% 4 Missing ⚠️

📢 Thoughts on this report? Let us know!

@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

LLGo baseline benchmarks

8d946e2e0162 | workflow run | long-term charts

Program measurements

Platform Workload File size vs base Text size vs base Build vs base Run vs base
Linux cprintf 8568 B 0 B / +0.0% 387 B 0 B / +0.0% 942.587 ms -52.15 ms / -5.2% (better) 1.251 ms -75.93 us / -5.7% (better)
Linux cprintf-lto 8408 B 0 B / +0.0% 368 B 0 B / +0.0% 960.856 ms -72 ms / -7.0% (better) 1.270 ms +3.112 us / +0.2% (worse)
Linux fmtprintf 4886536 B +6600 B / +0.1% (worse) 491158 B -10155 B / -2.0% (better) 7.058 s -312.8 ms / -4.2% (better) 3.122 ms -25.02 us / -0.8% (better)
Linux fmtprintf-lto 3591936 B -6352 B / -0.2% (better) 426655 B -11502 B / -2.6% (better) 16.349 s -351.4 ms / -2.1% (better) 2.908 ms +62.33 us / +2.2% (worse)
Linux println 675144 B +776 B / +0.1% (worse) 16649 B -206 B / -1.2% (better) 943.463 ms -69.61 ms / -6.9% (better) 1.616 ms +7.432 us / +0.5% (worse)
Linux println-lto 190048 B +56 B / +0.02947% (worse) 14089 B -184 B / -1.3% (better) 1.284 s -50.95 ms / -3.8% (better) 1.599 ms -86.54 us / -5.1% (better)
macOS cprintf 50736 B 0 B / +0.0% 4409 B 0 B / +0.0% 1.250 s +172.4 ms / +16.0% (worse) 3.708 ms -344.5 us / -8.5% (better)
macOS cprintf-lto 50496 B 0 B / +0.0% 161 B 0 B / +0.0% 1.151 s +26.9 ms / +2.4% (worse) 3.398 ms -2.09 ms / -38.1% (better)
macOS fmtprintf 1827296 B +53840 B / +3.0% (worse) 869104 B -8312 B / -0.9% (better) 5.821 s -1.228 s / -17.4% (better) 10.985 ms -1.248 ms / -10.2% (better)
macOS fmtprintf-lto 1390144 B +29216 B / +2.1% (worse) 752404 B -9512 B / -1.2% (better) 11.240 s -3.428 s / -23.4% (better) 5.439 ms -4.679 ms / -46.2% (better)
macOS println 101152 B +1808 B / +1.8% (worse) 24006 B -216 B / -0.9% (better) 2.036 s +884.5 ms / +76.8% (worse) 8.177 ms +3.925 ms / +92.3% (worse)
macOS println-lto 84032 B +368 B / +0.4% (worse) 21257 B -200 B / -0.9% (better) 2.221 s +736.6 ms / +49.6% (worse) 5.997 ms -1.358 ms / -18.5% (better)
Windows MinGW cprintf 650240 B +512 B / +0.1% (worse) 4550 B 0 B / +0.0% 1.668 s -40.43 ms / -2.4% (better) 3.544 ms +22.9 us / +0.7% (worse)
Windows MinGW cprintf-lto 43520 B 0 B / +0.0% 4486 B 0 B / +0.0% 1.730 s +38.45 ms / +2.3% (worse) 3.870 ms +377.3 us / +10.8% (worse)
Windows MinGW fmtprintf 5453312 B +22016 B / +0.4% (worse) 585734 B -14272 B / -2.4% (better) 7.240 s +55.52 ms / +0.8% (worse) 7.819 ms -366.5 us / -4.5% (better)
Windows MinGW fmtprintf-lto 4128256 B +4096 B / +0.1% (worse) 531446 B -14800 B / -2.7% (better) 14.750 s +136.4 ms / +0.9% (worse) 8.710 ms +315.5 us / +3.8% (worse)
Windows MinGW println 706560 B +1024 B / +0.1% (worse) 24998 B -192 B / -0.8% (better) 1.686 s +18.95 ms / +1.1% (worse) 7.030 ms +760.5 us / +12.1% (worse)
Windows MinGW println-lto 209408 B +512 B / +0.2% (worse) 21846 B -208 B / -0.9% (better) 2.042 s -62.31 ms / -3.0% (better) 6.348 ms -347 us / -5.2% (better)
Windows MinGW 386 cprintf 600576 B -512 B / -0.1% (better) 5326 B 0 B / +0.0% 1.660 s -20.01 ms / -1.2% (better) 5.030 ms -69.2 us / -1.4% (better)
Windows MinGW 386 cprintf-lto 103424 B 0 B / +0.0% 5094 B 0 B / +0.0% 1.700 s -31.41 ms / -1.8% (better) 5.103 ms +70.4 us / +1.4% (worse)
Windows MinGW 386 fmtprintf 4749312 B +5632 B / +0.1% (worse) 459166 B -13136 B / -2.8% (better) 7.448 s -62.78 ms / -0.8% (better) 10.587 ms -727.1 us / -6.4% (better)
Windows MinGW 386 fmtprintf-lto 4164608 B +18432 B / +0.4% (worse) 437950 B -12868 B / -2.9% (better) 14.668 s +247.8 ms / +1.7% (worse) 10.129 ms -290.6 us / -2.8% (better)
Windows MinGW 386 println 653824 B +2048 B / +0.3% (worse) 21282 B -208 B / -1.0% (better) 1.677 s -65.72 ms / -3.8% (better) 8.920 ms +174.6 us / +2.0% (worse)
Windows MinGW 386 println-lto 259072 B +512 B / +0.2% (worse) 19114 B -192 B / -1.0% (better) 1.951 s +17.53 ms / +0.9% (worse) 8.512 ms -110 us / -1.3% (better)
Windows MinGW ARM64 cprintf 660480 B 0 B / +0.0% 4408 B 0 B / +0.0% 1.905 s -40.48 ms / -2.1% (better) 6.308 ms -179.3 us / -2.8% (better)
Windows MinGW ARM64 cprintf-lto 43520 B 0 B / +0.0% 4340 B 0 B / +0.0% 1.955 s -18.1 ms / -0.9% (better) 6.526 ms -321.7 us / -4.7% (better)
Windows MinGW ARM64 fmtprintf 5350400 B +7680 B / +0.1% (worse) 497964 B -12724 B / -2.5% (better) 7.130 s +107.2 ms / +1.5% (worse) 13.581 ms +440.5 us / +3.4% (worse)
Windows MinGW ARM64 fmtprintf-lto 4321792 B +22528 B / +0.5% (worse) 464752 B -12188 B / -2.6% (better) 13.439 s -251.1 ms / -1.8% (better) 13.070 ms -736.3 us / -5.3% (better)
Windows MinGW ARM64 println 713728 B +1024 B / +0.1% (worse) 23644 B -240 B / -1.0% (better) 1.896 s -36.06 ms / -1.9% (better) 10.354 ms -1.303 ms / -11.2% (better)
Windows MinGW ARM64 println-lto 216064 B +512 B / +0.2% (worse) 21036 B -196 B / -0.9% (better) 2.158 s -43.25 ms / -2.0% (better) 11.166 ms -360.7 us / -3.1% (better)
Windows MSVC cprintf 893440 B +1536 B / +0.2% (worse) 65798 B 0 B / +0.0% 1.135 s -129.8 ms / -10.3% (better) 3.026 ms -130.5 us / -4.1% (better)
Windows MSVC cprintf-lto 289792 B 0 B / +0.0% 65734 B 0 B / +0.0% 1.179 s -2.957 ms / -0.3% (better) 3.062 ms +269.6 us / +9.7% (worse)
Windows MSVC fmtprintf 5752832 B +20480 B / +0.4% (worse) 681270 B -14272 B / -2.1% (better) 5.601 s +175.1 ms / +3.2% (worse) 8.035 ms -54.6 us / -0.7% (better)
Windows MSVC fmtprintf-lto 4451328 B +3072 B / +0.1% (worse) 631030 B -14752 B / -2.3% (better) 11.382 s +391.1 ms / +3.6% (worse) 8.706 ms +48.4 us / +0.6% (worse)
Windows MSVC println 1017344 B +512 B / +0.1% (worse) 120662 B -192 B / -0.2% (better) 1.383 s +229 ms / +19.8% (worse) 6.426 ms -801.5 us / -11.1% (better)
Windows MSVC println-lto 528384 B 0 B / +0.0% 118182 B -208 B / -0.2% (better) 1.411 s +51.26 ms / +3.8% (worse) 6.254 ms +446 us / +7.7% (worse)
Windows MSVC 386 cprintf 512512 B 0 B / +0.0% 3931 B 0 B / +0.0% 1.545 s +21.61 ms / +1.4% (worse) 5.376 ms +72.8 us / +1.4% (worse)
Windows MSVC 386 cprintf-lto 44032 B 0 B / +0.0% 3853 B 0 B / +0.0% 1.533 s -9.334 ms / -0.6% (better) 5.691 ms +205.7 us / +3.8% (worse)
Windows MSVC 386 fmtprintf 4484096 B +5632 B / +0.1% (worse) 442556 B -13136 B / -2.9% (better) 7.025 s +56.07 ms / +0.8% (worse) 10.957 ms -75.7 us / -0.7% (better)
Windows MSVC 386 fmtprintf-lto 3907584 B +10752 B / +0.3% (worse) 412955 B -13440 B / -3.2% (better) 13.324 s -506.9 ms / -3.7% (better) 10.990 ms -1.736 ms / -13.6% (better)
Windows MSVC 386 println 566784 B +512 B / +0.1% (worse) 20132 B -208 B / -1.0% (better) 1.505 s -238.8 ms / -13.7% (better) 9.159 ms -390 us / -4.1% (better)
Windows MSVC 386 println-lto 199168 B 0 B / +0.0% 18293 B -208 B / -1.1% (better) 1.748 s -297.7 us / -0.01702% (better) 9.285 ms +83 us / +0.9% (worse)
Windows MSVC ARM64 cprintf 662016 B +512 B / +0.1% (worse) 4192 B 0 B / +0.0% 1.509 s +46.07 ms / +3.1% (worse) 5.992 ms -153.3 us / -2.5% (better)
Windows MSVC ARM64 cprintf-lto 47616 B 0 B / +0.0% 4084 B 0 B / +0.0% 1.541 s +71.99 ms / +4.9% (worse) 6.104 ms -137.7 us / -2.2% (better)
Windows MSVC ARM64 fmtprintf 5346304 B +7680 B / +0.1% (worse) 497896 B -12736 B / -2.5% (better) 6.389 s +81.05 ms / +1.3% (worse) 13.913 ms +1.271 ms / +10.1% (worse)
Windows MSVC ARM64 fmtprintf-lto 4329472 B +24064 B / +0.6% (worse) 465428 B -12176 B / -2.5% (better) 12.498 s +315.7 ms / +2.6% (worse) 14.079 ms +773.5 us / +5.8% (worse)
Windows MSVC ARM64 println 715776 B +1024 B / +0.1% (worse) 23668 B -240 B / -1.0% (better) 1.559 s +91.39 ms / +6.2% (worse) 11.293 ms +534.3 us / +5.0% (worse)
Windows MSVC ARM64 println-lto 220160 B 0 B / +0.0% 21188 B -192 B / -0.9% (better) 1.755 s +74.5 ms / +4.4% (worse) 11.217 ms +671.2 us / +6.4% (worse)
Core language and compiler benchmarks
Platform Benchmark ns/op vs base
Linux BenchmarkLookupPCRandom 14.580 ns/op -0.09 ns/op / -0.6% (better)
Linux BenchmarkMergeCompilerFlags 205.600 ns/op +7.7 ns/op / +3.9% (worse)
Linux BenchmarkMergeLinkerFlags 131.800 ns/op -5 ns/op / -3.7% (better)
Linux BenchmarkChannelBuffered 55.060 ns/op -0.13 ns/op / -0.2% (better)
Linux BenchmarkChannelHandoff 12867 ns/op -59 ns/op / -0.5% (better)
Linux BenchmarkDefer 49.520 ns/op -1.18 ns/op / -2.3% (better)
Linux BenchmarkDirectCall 1.165 ns/op -0.425 ns/op / -26.7% (better)
Linux BenchmarkGlobalRead 1.169 ns/op +0.003 ns/op / +0.3% (worse)
Linux BenchmarkGlobalWrite 7.792 ns/op +0.024 ns/op / +0.3% (worse)
Linux BenchmarkGoroutine 24068 ns/op -17 ns/op / -0.1% (better)
Linux BenchmarkInterfaceCall 5.888 ns/op +0.016 ns/op / +0.3% (worse)
Linux BenchmarkRuntimeGetG 2.921 ns/op +0.558 ns/op / +23.6% (worse)
macOS BenchmarkLookupPCRandom 16.730 ns/op -1.44 ns/op / -7.9% (better)
macOS BenchmarkMergeCompilerFlags 154.800 ns/op -25.1 ns/op / -14.0% (better)
macOS BenchmarkMergeLinkerFlags 109.200 ns/op +0.5 ns/op / +0.5% (worse)
macOS BenchmarkChannelBuffered 31.450 ns/op -2.38 ns/op / -7.0% (better)
macOS BenchmarkChannelHandoff 12920 ns/op +4718 ns/op / +57.5% (worse)
macOS BenchmarkDefer 44.540 ns/op -2.27 ns/op / -4.8% (better)
macOS BenchmarkDirectCall 1.382 ns/op +0.185 ns/op / +15.5% (worse)
macOS BenchmarkGlobalRead 1.192 ns/op -0.029 ns/op / -2.4% (better)
macOS BenchmarkGlobalWrite 1.564 ns/op +0.057 ns/op / +3.8% (worse)
macOS BenchmarkGoroutine 45472 ns/op -32582 ns/op / -41.7% (better)
macOS BenchmarkInterfaceCall 6.297 ns/op +0.91 ns/op / +16.9% (worse)
macOS BenchmarkRuntimeGetG 2.590 ns/op -0.14 ns/op / -5.1% (better)
Windows MinGW BenchmarkLookupPCRandom 12.950 ns/op -0.27 ns/op / -2.0% (better)
Windows MinGW BenchmarkMergeCompilerFlags 615.500 ns/op -10.8 ns/op / -1.7% (better)
Windows MinGW BenchmarkMergeLinkerFlags 543.800 ns/op +4.3 ns/op / +0.8% (worse)
Windows MinGW BenchmarkChannelBuffered 30.740 ns/op +0.25 ns/op / +0.8% (worse)
Windows MinGW BenchmarkChannelHandoff 944.100 ns/op +11.1 ns/op / +1.2% (worse)
Windows MinGW BenchmarkDefer 62.530 ns/op +4.92 ns/op / +8.5% (worse)
Windows MinGW BenchmarkDirectCall 1.547 ns/op -0.009 ns/op / -0.6% (better)
Windows MinGW BenchmarkGlobalRead 1.859 ns/op +0.308 ns/op / +19.9% (worse)
Windows MinGW BenchmarkGlobalWrite 2.449 ns/op -0.019 ns/op / -0.8% (better)
Windows MinGW BenchmarkGoroutine 90585 ns/op -292 ns/op / -0.3% (better)
Windows MinGW BenchmarkInterfaceCall 8.370 ns/op -0.014 ns/op / -0.2% (better)
Windows MinGW BenchmarkRuntimeGetG 1.860 ns/op -0.621 ns/op / -25.0% (better)
Windows MinGW 386 BenchmarkLookupPCRandom 26.480 ns/op -0.16 ns/op / -0.6% (better)
Windows MinGW 386 BenchmarkMergeCompilerFlags 777.500 ns/op +48 ns/op / +6.6% (worse)
Windows MinGW 386 BenchmarkMergeLinkerFlags 679.500 ns/op -21.9 ns/op / -3.1% (better)
Windows MinGW 386 BenchmarkChannelBuffered 38.890 ns/op -2.34 ns/op / -5.7% (better)
Windows MinGW 386 BenchmarkChannelHandoff 959.800 ns/op +38.9 ns/op / +4.2% (worse)
Windows MinGW 386 BenchmarkDefer 43.570 ns/op -0.25 ns/op / -0.6% (better)
Windows MinGW 386 BenchmarkDirectCall 1.548 ns/op -0.001 ns/op / -0.1% (better)
Windows MinGW 386 BenchmarkGlobalRead 1.864 ns/op +0.315 ns/op / +20.3% (worse)
Windows MinGW 386 BenchmarkGlobalWrite 7.785 ns/op +0.012 ns/op / +0.2% (worse)
Windows MinGW 386 BenchmarkGoroutine 106841 ns/op +627 ns/op / +0.6% (worse)
Windows MinGW 386 BenchmarkInterfaceCall 7.998 ns/op -0.392 ns/op / -4.7% (better)
Windows MinGW 386 BenchmarkRuntimeGetG 1.925 ns/op -0.005 ns/op / -0.3% (better)
Windows MinGW ARM64 BenchmarkLookupPCRandom 12.110 ns/op +0.04 ns/op / +0.3% (worse)
Windows MinGW ARM64 BenchmarkMergeCompilerFlags 579.100 ns/op +13.9 ns/op / +2.5% (worse)
Windows MinGW ARM64 BenchmarkMergeLinkerFlags 529.200 ns/op +9 ns/op / +1.7% (worse)
Windows MinGW ARM64 BenchmarkChannelBuffered 39.080 ns/op 0 ns/op / +0.0%
Windows MinGW ARM64 BenchmarkChannelHandoff 1611 ns/op +14 ns/op / +0.9% (worse)
Windows MinGW ARM64 BenchmarkDefer 57.290 ns/op +2 ns/op / +3.6% (worse)
Windows MinGW ARM64 BenchmarkDirectCall 0.589 ns/op -0.0741 ns/op / -11.2% (better)
Windows MinGW ARM64 BenchmarkGlobalRead 0.884 ns/op +0.2206 ns/op / +33.2% (worse)
Windows MinGW ARM64 BenchmarkGlobalWrite 0.663 ns/op -0.0737 ns/op / -10.0% (better)
Windows MinGW ARM64 BenchmarkGoroutine 60735 ns/op +1407 ns/op / +2.4% (worse)
Windows MinGW ARM64 BenchmarkInterfaceCall 4.185 ns/op +0.048 ns/op / +1.2% (worse)
Windows MinGW ARM64 BenchmarkRuntimeGetG 1.800 ns/op +0.031 ns/op / +1.8% (worse)
Windows MSVC BenchmarkLookupPCRandom 9.626 ns/op +1.209 ns/op / +14.4% (worse)
Windows MSVC BenchmarkMergeCompilerFlags 461.100 ns/op +68.7 ns/op / +17.5% (worse)
Windows MSVC BenchmarkMergeLinkerFlags 414.300 ns/op +67.8 ns/op / +19.6% (worse)
Windows MSVC BenchmarkChannelBuffered 31.310 ns/op +2.62 ns/op / +9.1% (worse)
Windows MSVC BenchmarkChannelHandoff 7161 ns/op -3250 ns/op / -31.2% (better)
Windows MSVC BenchmarkDefer 35.320 ns/op -0.12 ns/op / -0.3% (better)
Windows MSVC BenchmarkDirectCall 0.343 ns/op +0.0726 ns/op / +26.9% (worse)
Windows MSVC BenchmarkGlobalRead 0.364 ns/op +0.013 ns/op / +3.7% (worse)
Windows MSVC BenchmarkGlobalWrite 7.109 ns/op +0.011 ns/op / +0.2% (worse)
Windows MSVC BenchmarkGoroutine 65557 ns/op +3767 ns/op / +6.1% (worse)
Windows MSVC BenchmarkInterfaceCall 5.341 ns/op +0.469 ns/op / +9.6% (worse)
Windows MSVC BenchmarkRuntimeGetG 0.893 ns/op -0.0963 ns/op / -9.7% (better)
Windows MSVC 386 BenchmarkLookupPCRandom 26.540 ns/op -0.05 ns/op / -0.2% (better)
Windows MSVC 386 BenchmarkMergeCompilerFlags 752.800 ns/op -23.1 ns/op / -3.0% (better)
Windows MSVC 386 BenchmarkMergeLinkerFlags 687.800 ns/op -19.4 ns/op / -2.7% (better)
Windows MSVC 386 BenchmarkChannelBuffered 39.520 ns/op +0.8 ns/op / +2.1% (worse)
Windows MSVC 386 BenchmarkChannelHandoff 867 ns/op -14.9 ns/op / -1.7% (better)
Windows MSVC 386 BenchmarkDefer 49.040 ns/op +4.16 ns/op / +9.3% (worse)
Windows MSVC 386 BenchmarkDirectCall 1.547 ns/op -0.001 ns/op / -0.1% (better)
Windows MSVC 386 BenchmarkGlobalRead 1.866 ns/op +0.316 ns/op / +20.4% (worse)
Windows MSVC 386 BenchmarkGlobalWrite 7.786 ns/op +0.014 ns/op / +0.2% (worse)
Windows MSVC 386 BenchmarkGoroutine 112781 ns/op +2884 ns/op / +2.6% (worse)
Windows MSVC 386 BenchmarkInterfaceCall 8.077 ns/op -0.283 ns/op / -3.4% (better)
Windows MSVC 386 BenchmarkRuntimeGetG 2.168 ns/op +0.001 ns/op / +0.04615% (worse)
Windows MSVC ARM64 BenchmarkLookupPCRandom 12.110 ns/op +0.03 ns/op / +0.2% (worse)
Windows MSVC ARM64 BenchmarkMergeCompilerFlags 566.500 ns/op -3.3 ns/op / -0.6% (better)
Windows MSVC ARM64 BenchmarkMergeLinkerFlags 541.300 ns/op +6.5 ns/op / +1.2% (worse)
Windows MSVC ARM64 BenchmarkChannelBuffered 37.450 ns/op -1.03 ns/op / -2.7% (better)
Windows MSVC ARM64 BenchmarkChannelHandoff 1936 ns/op +48 ns/op / +2.5% (worse)
Windows MSVC ARM64 BenchmarkDefer 62.140 ns/op +5 ns/op / +8.8% (worse)
Windows MSVC ARM64 BenchmarkDirectCall 0.589 ns/op -0.0738 ns/op / -11.1% (better)
Windows MSVC ARM64 BenchmarkGlobalRead 0.590 ns/op -0.0739 ns/op / -11.1% (better)
Windows MSVC ARM64 BenchmarkGlobalWrite 3.794 ns/op 0 ns/op / +0.0%
Windows MSVC ARM64 BenchmarkGoroutine 57460 ns/op +3908 ns/op / +7.3% (worse)
Windows MSVC ARM64 BenchmarkInterfaceCall 4.139 ns/op -0.005 ns/op / -0.1% (better)
Windows MSVC ARM64 BenchmarkRuntimeGetG 1.809 ns/op +0.005 ns/op / +0.3% (worse)
Timer runtime benchmarks
Platform Operation and runtime ns/op vs base
Linux AfterFuncZeroDelivery/Go 903.900 ns/op +3.2 ns/op / +0.4% (worse)
Linux AfterFuncZeroDelivery/LLGo 44045 ns/op +7241 ns/op / +19.7% (worse)
Linux CreateStop/Go 289.200 ns/op +0.2 ns/op / +0.1% (worse)
Linux CreateStop/LLGo 1843 ns/op +168 ns/op / +10.0% (worse)
Linux RearmStopped/Go 116 ns/op -0.1 ns/op / -0.1% (better)
Linux RearmStopped/LLGo 1197 ns/op -116 ns/op / -8.8% (better)
Linux ResetActive/Go 68.800 ns/op +0.13 ns/op / +0.2% (worse)
Linux ResetActive/LLGo 758.300 ns/op -12.9 ns/op / -1.7% (better)
Linux ResetHeap1024/Go 67.320 ns/op +0.2 ns/op / +0.3% (worse)
Linux ResetHeap1024/LLGo 177.700 ns/op 0 ns/op / +0.0%
macOS AfterFuncZeroDelivery/Go 553.100 ns/op -141.7 ns/op / -20.4% (better)
macOS AfterFuncZeroDelivery/LLGo 99882 ns/op -27842 ns/op / -21.8% (better)
macOS CreateStop/Go 189.600 ns/op -80.6 ns/op / -29.8% (better)
macOS CreateStop/LLGo 440.200 ns/op -702.8 ns/op / -61.5% (better)
macOS RearmStopped/Go 71.790 ns/op -10.48 ns/op / -12.7% (better)
macOS RearmStopped/LLGo 439.200 ns/op +8.6 ns/op / +2.0% (worse)
macOS ResetActive/Go 57.630 ns/op -4.82 ns/op / -7.7% (better)
macOS ResetActive/LLGo 242.500 ns/op -23 ns/op / -8.7% (better)
macOS ResetHeap1024/Go 50.470 ns/op -12.73 ns/op / -20.1% (better)
macOS ResetHeap1024/LLGo 86.270 ns/op -97.23 ns/op / -53.0% (better)
Windows MinGW AfterFuncZeroDelivery/Go 546.300 ns/op -11.4 ns/op / -2.0% (better)
Windows MinGW AfterFuncZeroDelivery/LLGo 182562 ns/op +349 ns/op / +0.2% (worse)
Windows MinGW CreateStop/Go 114.800 ns/op -0.6 ns/op / -0.5% (better)
Windows MinGW CreateStop/LLGo 425.400 ns/op +3.7 ns/op / +0.9% (worse)
Windows MinGW RearmStopped/Go 31.510 ns/op +0.18 ns/op / +0.6% (worse)
Windows MinGW RearmStopped/LLGo 269.100 ns/op +1.8 ns/op / +0.7% (worse)
Windows MinGW ResetActive/Go 20 ns/op -0.04 ns/op / -0.2% (better)
Windows MinGW ResetActive/LLGo 154.600 ns/op -10.4 ns/op / -6.3% (better)
Windows MinGW ResetHeap1024/Go 20.460 ns/op +0.04 ns/op / +0.2% (worse)
Windows MinGW ResetHeap1024/LLGo 126.800 ns/op +3.5 ns/op / +2.8% (worse)
Windows MinGW 386 AfterFuncZeroDelivery/Go 954.900 ns/op -2.5 ns/op / -0.3% (better)
Windows MinGW 386 AfterFuncZeroDelivery/LLGo 203194 ns/op +18 ns/op / +0.008859% (worse)
Windows MinGW 386 CreateStop/Go 190.600 ns/op +0.2 ns/op / +0.1% (worse)
Windows MinGW 386 CreateStop/LLGo 497.500 ns/op -31.4 ns/op / -5.9% (better)
Windows MinGW 386 RearmStopped/Go 63.480 ns/op -0.16 ns/op / -0.3% (better)
Windows MinGW 386 RearmStopped/LLGo 355.800 ns/op -5.2 ns/op / -1.4% (better)
Windows MinGW 386 ResetActive/Go 39.100 ns/op -0.06 ns/op / -0.2% (better)
Windows MinGW 386 ResetActive/LLGo 1015 ns/op +26 ns/op / +2.6% (worse)
Windows MinGW 386 ResetHeap1024/Go 39.370 ns/op -0.12 ns/op / -0.3% (better)
Windows MinGW 386 ResetHeap1024/LLGo 186.300 ns/op -2.7 ns/op / -1.4% (better)
Windows MinGW ARM64 AfterFuncZeroDelivery/Go 668.200 ns/op +5.4 ns/op / +0.8% (worse)
Windows MinGW ARM64 AfterFuncZeroDelivery/LLGo 146681 ns/op +2166 ns/op / +1.5% (worse)
Windows MinGW ARM64 CreateStop/Go 192.100 ns/op -3.3 ns/op / -1.7% (better)
Windows MinGW ARM64 CreateStop/LLGo 411.800 ns/op +4.2 ns/op / +1.0% (worse)
Windows MinGW ARM64 RearmStopped/Go 70.630 ns/op -0.04 ns/op / -0.1% (better)
Windows MinGW ARM64 RearmStopped/LLGo 254.600 ns/op -4.1 ns/op / -1.6% (better)
Windows MinGW ARM64 ResetActive/Go 31.150 ns/op -0.37 ns/op / -1.2% (better)
Windows MinGW ARM64 ResetActive/LLGo 134.800 ns/op -0.7 ns/op / -0.5% (better)
Windows MinGW ARM64 ResetHeap1024/Go 31.170 ns/op -0.01 ns/op / -0.03207% (better)
Windows MinGW ARM64 ResetHeap1024/LLGo 126 ns/op +1.2 ns/op / +1.0% (worse)
Windows MSVC AfterFuncZeroDelivery/Go 545.700 ns/op +9.8 ns/op / +1.8% (worse)
Windows MSVC AfterFuncZeroDelivery/LLGo 147833 ns/op +7461 ns/op / +5.3% (worse)
Windows MSVC CreateStop/Go 152.600 ns/op +17.5 ns/op / +13.0% (worse)
Windows MSVC CreateStop/LLGo 576 ns/op -368.4 ns/op / -39.0% (better)
Windows MSVC RearmStopped/Go 57.800 ns/op +5.41 ns/op / +10.3% (worse)
Windows MSVC RearmStopped/LLGo 280 ns/op -14 ns/op / -4.8% (better)
Windows MSVC ResetActive/Go 25.650 ns/op +3.14 ns/op / +13.9% (worse)
Windows MSVC ResetActive/LLGo 166.900 ns/op -42.5 ns/op / -20.3% (better)
Windows MSVC ResetHeap1024/Go 25.620 ns/op +0.06 ns/op / +0.2% (worse)
Windows MSVC ResetHeap1024/LLGo 106.900 ns/op +5.4 ns/op / +5.3% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/Go 959.600 ns/op +6.4 ns/op / +0.7% (worse)
Windows MSVC 386 AfterFuncZeroDelivery/LLGo 197241 ns/op -6567 ns/op / -3.2% (better)
Windows MSVC 386 CreateStop/Go 193.200 ns/op -7.9 ns/op / -3.9% (better)
Windows MSVC 386 CreateStop/LLGo 479.800 ns/op +5.2 ns/op / +1.1% (worse)
Windows MSVC 386 RearmStopped/Go 63.160 ns/op -0.19 ns/op / -0.3% (better)
Windows MSVC 386 RearmStopped/LLGo 318.500 ns/op +4 ns/op / +1.3% (worse)
Windows MSVC 386 ResetActive/Go 38.960 ns/op -0.05 ns/op / -0.1% (better)
Windows MSVC 386 ResetActive/LLGo 896.800 ns/op -39.7 ns/op / -4.2% (better)
Windows MSVC 386 ResetHeap1024/Go 39.340 ns/op -0.09 ns/op / -0.2% (better)
Windows MSVC 386 ResetHeap1024/LLGo 171.400 ns/op +1.5 ns/op / +0.9% (worse)
Windows MSVC ARM64 AfterFuncZeroDelivery/Go 663.500 ns/op -2.9 ns/op / -0.4% (better)
Windows MSVC ARM64 AfterFuncZeroDelivery/LLGo 144326 ns/op -14757 ns/op / -9.3% (better)
Windows MSVC ARM64 CreateStop/Go 194.600 ns/op -4.7 ns/op / -2.4% (better)
Windows MSVC ARM64 CreateStop/LLGo 446.300 ns/op +19.4 ns/op / +4.5% (worse)
Windows MSVC ARM64 RearmStopped/Go 70.530 ns/op -0.08 ns/op / -0.1% (better)
Windows MSVC ARM64 RearmStopped/LLGo 283.100 ns/op +8.9 ns/op / +3.2% (worse)
Windows MSVC ARM64 ResetActive/Go 30.990 ns/op +0.06 ns/op / +0.2% (worse)
Windows MSVC ARM64 ResetActive/LLGo 147.700 ns/op +0.9 ns/op / +0.6% (worse)
Windows MSVC ARM64 ResetHeap1024/Go 31.120 ns/op +0.17 ns/op / +0.5% (worse)
Windows MSVC ARM64 ResetHeap1024/LLGo 139.800 ns/op +2.7 ns/op / +2.0% (worse)

Compared with d98b43f90a3d measured in the same runner job.

@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

LLGo WebAssembly build benchmarks

b08ce437f781 | workflow run | long-term charts

WebAssembly output sizes
Example, profile and compiler Wasm module vs base Generated JS glue vs base
cprintf/j32-emscripten/LLGo 153510 B -683 B / -0.4% (better) 88742 B 0 B / +0.0%
cprintf/j32-goos-js/LLGo 151500 B -795 B / -0.5% (better) 73165 B 0 B / +0.0%
cprintf/j64-emscripten-memory64/LLGo 141143 B -105 B / -0.1% (better) 92630 B 0 B / +0.0%
cprintf/w32-goos-wasip1/LLGo 152668 B -180 B / -0.1% (better) 0 B 0 B / 0.0%
cprintf/w32-wasi/LLGo 153380 B -180 B / -0.1% (better) 0 B 0 B / 0.0%
fmtprintf/j32-emscripten/LLGo 3167263 B -46165 B / -1.4% (better) 132442 B 0 B / +0.0%
fmtprintf/j32-goos-js/Go 2526852 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/j32-goos-js/LLGo 3146457 B -46066 B / -1.4% (better) 101635 B 0 B / +0.0%
fmtprintf/j64-emscripten-memory64/LLGo 2908923 B -41110 B / -1.4% (better) 139285 B 0 B / +0.0%
fmtprintf/w32-goos-wasip1/Go 2500019 B 0 B / +0.0% 0 B 0 B / 0.0%
fmtprintf/w32-goos-wasip1/LLGo 2330312 B -17039 B / -0.7% (better) 0 B 0 B / 0.0%
fmtprintf/w32-wasi/LLGo 2326952 B -17061 B / -0.7% (better) 0 B 0 B / 0.0%
j32-emscripten/LLGo 152744 B -684 B / -0.4% (better) 88742 B 0 B / +0.0%
j32-goos-js/Go 1895533 B 0 B / +0.0% 0 B 0 B / 0.0%
j32-goos-js/LLGo 150972 B -796 B / -0.5% (better) 73165 B 0 B / +0.0%
j64-emscripten-memory64/LLGo 140468 B -114 B / -0.1% (better) 92629 B -1 B / -0.00108% (better)
reflectcall/j32-emscripten/LLGo 1531975 B -11796 B / -0.8% (better) 105908 B 0 B / +0.0%
reflectcall/j32-goos-js/Go 2191221 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/j32-goos-js/LLGo 1534219 B -12003 B / -0.8% (better) 90331 B 0 B / +0.0%
reflectcall/j64-emscripten-memory64/LLGo 1419174 B -9252 B / -0.6% (better) 111641 B 0 B / +0.0%
reflectcall/w32-goos-wasip1/Go 2205707 B 0 B / +0.0% 0 B 0 B / 0.0%
reflectcall/w32-goos-wasip1/LLGo 1271574 B -5119 B / -0.4% (better) 0 B 0 B / 0.0%
reflectcall/w32-wasi/LLGo 1269190 B -5111 B / -0.4% (better) 0 B 0 B / 0.0%
w32-goos-wasip1/Go 1909947 B 0 B / +0.0% 0 B 0 B / 0.0%
w32-goos-wasip1/LLGo 152315 B -180 B / -0.1% (better) 0 B 0 B / 0.0%
w32-wasi/LLGo 153027 B -180 B / -0.1% (better) 0 B 0 B / 0.0%
LLGo WebAssembly build measurements
Example and profile Build vs base
j32-emscripten 6.169 s -76.13 ms / -1.2% (better)
j32-goos-js 6.109 s +59.71 ms / +1.0% (worse)
j64-emscripten-memory64 5.416 s -70.14 ms / -1.3% (better)
reflectcall/w32-wasi 23.060 s -632.9 ms / -2.7% (better)
w32-goos-wasip1 4.246 s -68.74 ms / -1.6% (better)
w32-wasi 4.081 s -159.4 ms / -3.8% (better)

Compared with d98b43f90a3d measured in the same runner job.

@visualfc
visualfc force-pushed the fix/static-maplit branch 4 times, most recently from 1c5017c to 8d946e2 Compare October 2, 2026 02:23
Known concrete-to-interface conversions emit a read-only _llgo_itab$
global (Inter, Type, Hash, Fun) with weak_odr linkage, matching
cmd/compile's go:itab symbols.

Ordinary builds use that global as the runtime vtable. The T2I site is
the only reference, so --gc-sections/-dead_strip can drop itabs that
belong to dead functions. LTO and deadcode-drop keep calling NewItab so
unused interface methods can be dropped; Fun[] would pin them. LTO still
emits the static template for de-virt and erases unused templates after
the plugin runs.

Non-direct interface values (integers, bools, small aggregates) still
use IfaceIndir, matching Go: data is a pointer to a copy. Compile-time
constants go in a WeakODR _llgo_ifacebox$ global instead of AllocU,
like cmd/compile's static temps.

Box names hash the constant's bits, not LLVM's printed form, so
same-length strings in different packages do not collide under
Windows COMDAT. Repeated boxing of the same LLVM value reuses a
per-package pointer cache. Zero-sized constants stay AllocU so boxing
does not mutate the shared sentinel.

Together, a constant T2I is {itab, box} with no runtime call or heap
allocation. Hash is copied from the type descriptor when present.
interequal compares the (inter, _type) pair so a static itab and a
dynamically allocated itab for the same conversion compare equal.
A call whose SSA body is MakeInterface of a Convert/ChangeType of a
constant parameter is lowered at the call site to MakeInterface of
that concrete value. No function-name matching: constant.MakeInt64
and user helpers such as boxMyInt(x int64) any { return myInt(x) }
use the same path.

Depends on static itabs and iface boxes so the folded value is a
compile-time {itab, box} pair.
Map literals with more than 25 constant keys become a counted
mapassign loop over private constant key/value arrays, matching
cmd/compile's maplit. Composite values whose fields are constants or
trivial iface constructors (PR 3) are rebuilt as LLVM constants so
the value array is a ConstArray.

Runtime values stay unrolled. Spilling them into alloca [N x T] makes
LLVM default<Os> SLP scalarize the array.

Depends on trivial iface folding so {string, MakeInt64/MakeBool}
entries are compile-time structs.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant