Skip to content

[6.0] Reduce invocation overhead for intercepted bean methods - #3546

Open
AngeloRubens wants to merge 17 commits into
weld:6.0from
AngeloRubens:pr/proxy4-6.0
Open

AngeloRubens wants to merge 17 commits into
weld:6.0from
AngeloRubens:pr/proxy4-6.0

Conversation

@AngeloRubens

@AngeloRubens AngeloRubens commented Oct 5, 2026 •

Copy link
Copy Markdown

Summary

Reduces the steady-state cost of calls to methods on intercepted beans. This PR includes the prerequisite client-proxy and interception-stack changes because they have not yet been merged upstream, followed by invocation-chain caching, smaller invocation contexts, shared empty arguments, per-thread stack reuse in InvocationContext.proceed(), and cached invokers. For eligible instance methods, a LambdaMetafactory function lets HotSpot inline interceptor dispatch. Reflection remains the fallback. A regression fix checks that a cached privileged handle cannot bypass a missing or revoked setAccessible override.

The Weld 6 backport is submitted separately against 6.0.

Validation

  • Full Weld 6 CI on JDK 17/21 including unit tests, in-container tests, CDI TCK (WildFly and SE), relaxed mode, signature test and examples: Weld 6 run 37281293013.
  • Targeted access-override and cross-thread stack tests: 10 passed on each version, no failures or errors, run 37281149853.
  • The focused tests and all code were formatted by Weld's Maven formatter in CI before the full runs.

Performance

Interception scenarios derived from OWB benchmarks

The CDI benchmark scenarios and harness in this comparison follow the work of Matej Novotny, creator of OpenWebBeans and author of the Weld Core MicroBenchmarks. Results below are from the same GitHub Actions run, executed sequentially with one thread, two forks, 2 × 1 s warmup and 4 × 1 s measurement iterations. Values are JMH throughput in million benchmark operations per second (ops/µs); these are method calls in these scenarios, not HTTP requests.

Scenario Weld 6 base OWB 4.1.1 Weld 6 + PR #3546 PR vs Weld base PR vs OWB
Method intercepted 9.3 M/s 15.3 M/s 50.3 M/s 5.43× (+443%) 3.30× (+230%)
Class intercepted 8.9 M/s 16.0 M/s 49.2 M/s 5.53× (+453%) 3.07× (+207%)

In these two interception scenarios, the PR moves Weld from below OWB to above it. This does not describe every benchmark in the suite: for example, OWB remains faster in the application-scoped and request-scoped scenarios. The quick comparison has only two forks and relatively short iterations, so the exact ratios need confirmation with longer repeated runs. Full report and raw JMH results, run 37314014340.

Official Weld JMH suite

As an independent cross-check, Weld's official microbenchmark suite was run against base and PR sequentially on the same runner, with four threads, five forks, and 5 × 5 warmup/measurement iterations. It exercises its own Decorator and Interceptor workloads, which are distinct from the interception scenarios above.

Benchmark Weld 6 base Weld 6 + PR #3546 Change
Decorator 276.60 ops/s 280.91 ops/s +1.56%
Interceptor 345.16 ops/s 337.40 ops/s −2.25%

The official-suite differences are small relative to their JMH error margins and do not establish a general throughput reversal. Official JMH run 37335224493.

In the separate proxy-3-to-proxy-4 incremental comparison, Boot/shutdown was 44.95 ± 3.15 → 48.15 ± 2.76 ms. A separate ten-fork cold-start run invokes intercepted methods on a fresh container: Weld 6 was 686.08 ± 22.97 ms plain and 691.57 ± 16.62 ms with first intercepted calls. These runs show no clear startup regression. Cold-start results, run 37281596417.

The Weld 7 GC profile measured 0.0008 ± 0.0221 B/op for the intercepted-method benchmark, down from about 64 B/op in its previous proxy-3 profile. This is supporting evidence for the shared invoker implementation; it depends on JIT inlining/escape analysis and is not guaranteed for polymorphic call sites. Profile, run 37281324917.

No CDI self-invocation semantics are intentionally changed. The earlier client-proxy and stack changes are included for review because this optimization series depends on them.

Origin of the work

This work originated in the analysis of OpenWebBeans (OWB). Its approach to invocation and interception performance was the starting point for this optimization series, which was adapted to Weld internals and compatibility requirements.

@AngeloRubens
AngeloRubens requested a review from manovotn as a code owner October 5, 2026 12:42
…s when none exists

Every business method invocation on a client proxy calls
InterceptionDecorationContext.startIfNotEmpty(). When the caller is not
intercepted and no RequestScopedCache is active (e.g. Weld SE, or any
invocation outside of an HTTP/EJB request), getStack() allocated a new
Stack (+ ArrayDeque), set it into the thread-local and then immediately
removed it again in removeIfEmpty(). The net effect is nothing, but
ThreadLocal.set()/remove() plus the allocation dominate the cost of a
client proxy invocation.

Return null straight away if there is no stack on the current thread.
This is semantically identical: an absent stack is an empty stack and
the old code never left anything behind in that case (with an active
RequestScopedCache the empty stack would have been cached; it is now
created lazily by the first startIfNotOnTop()/getStack() instead).
… stack

Outside of an active RequestScopedCache each intercepted invocation (and
each non-intercepted method of an intercepted subclass) creates a Stack,
sets it into the thread-local and removes it again once the stack gets
empty. ThreadLocal.remove() clears the weak reference of the map entry
(native Reference.clear0) and expunges the entry, so the next set()
allocates a new WeakReference entry again. The profile of an intercepted
@ApplicationScoped bean shows remove() + expunge + entry allocation as
the most expensive part of the invocation.

Set the thread-local value to null instead: the entry is reused, and a
null value does not retain any Weld class, so there is no class loader
leak. Also allocate the ArrayDeque with a small initial capacity as the
stack is usually very shallow.
ContextBeanInstance.getInstance() called Container.isSet(contextId) on
each invocation to detect usage of a client proxy after shutdown. This is
a ConcurrentHashMap lookup in RegistrySingletonProvider and shows up in
the profile of @ApplicationScoped invocations.

ContextBeanInstance now keeps the Container it was created for. Container
gets a volatile cleanedUp flag set at the beginning of cleanup(); while
it is false the container is guaranteed to be registered, so the
registry lookup is only performed after the container was cleaned up.
Behaviour is unchanged (including the case where a new container was
registered under the same id afterwards).
Every invocation of an intercepted subclass method (both intercepted and
non-intercepted methods push the subclass handler) outside of a request
allocated a new Stack and ArrayDeque, set the thread-local and cleared it
again once the stack became empty. Within a request the stack was kept in
the RequestScopedCache instead.

The thread-local now holds a small Object[] holder (a JDK class) that
references the per-thread stack strongly only while it is not empty and
weakly otherwise. An idle thread never retains a Weld class, so there is
no class loader leak, and the stack is reused until the next GC clears
the weak reference. The RequestScopedCache registration and the validity
tracking of the stack are no longer needed.
…rst used

Client proxies call InterceptionDecorationContext.startIfNotEmpty() on every
invocation, i.e. a ThreadLocal lookup. As long as no thread has ever created
an interception context (no intercepted or decorated method has been
invoked yet, which is always the case in deployments without interceptors
and decorators) the lookup cannot find anything.

startIfNotEmpty() now invokes a constant MethodHandle guarded by a
SwitchPoint: it returns null without any lookup until the first per-thread
stack holder is created, when the SwitchPoint is invalidated (once, before
the holder is published to that thread). The JIT compiles the guard to
nothing, so applications that do use interception do not pay for an
additional check either. A thread can only have an active stack if it
created its holder itself, so the switch cannot be observed too late.
… check in client proxies

ContextBeanInstance keeps the bean as RIBean when possible so that the
client proxy invocation goes straight to the bean's ContextualInstanceStrategy
(for @ApplicationScoped beans a volatile read of the cached instance)
instead of going through the instanceof check in ContextualInstance.
Invalidation is unchanged: the strategy remains the single place where the
instance is cached.
Replace the ArrayDeque inside InterceptionDecorationContext.Stack with an
array and a size index that grows on demand. push/pop/peek/startIfNotOnTop
need fewer branches and no ring-buffer index arithmetic. Popped slots are
cleared so that no handler is retained, and the weakly referenced idle stack
(no class loader leak) is unchanged. Null elements are still rejected and
popping an empty Stack still throws NoSuchElementException.
Interceptor methods (@AroundInvoke etc.) and the proceed method of an
around invoke chain (WeldSubclass.method$$super() or the decorated method)
used to be invoked by Method.invoke(). They are now invoked through a
MethodHandle adapted to (Object, Object)Object or (Object, Object[])Object,
created lazily upon the first invocation and cached per method
(MethodInvoker; per interceptor metadata for interceptor methods, per
declaring class via a ClassValue for proceed methods).

The semantics of Method.invoke() are kept exactly:
- the handle is only used if the receiver is an instance of the declaring
  class and every argument matches the parameter type exactly (reference
  type or exact wrapper of the primitive type), i.e. if the invocation
  cannot fail before reaching the method; otherwise (null receiver, wrong
  arguments, widening primitive conversions, static methods...) the
  Method is invoked reflectively and throws the same exceptions as before
- any Throwable thrown by the invoked method is wrapped in an
  InvocationTargetException, so that the existing unwrapping code paths
  are unchanged
- accessibility is never changed; a handle that cannot be created (e.g.
  module readability) results in reflective invocations
…hed interception chains

InterceptorMethodHandler.invoke() called Reflections.ensureAccessible()
(AccessibleObject.canAccess()) and TargetClassInterceptorMetadata
.isInterceptorMethod() (a hash set lookup) upon every invocation, before
looking up the cached interception chain. A cached chain implies that both
checks have been performed already for the given method and proceed method
(the chain is created after them), and neither result can change, so the
chain is now looked up first and used directly.
… instance

Avoids the ConcurrentHashMap lookup in InterceptorMethodHandler for the
first intercepted method that has been invoked on a given instance (e.g. a
bean with a single intercepted method). The field is set at most once, the
chain is immutable (final fields only), other methods keep using the map.
…ocation contexts

The timer field of AbstractInvocationContext was always null and the
constructor field is only needed for around construct interception
(SimpleInvocationContext). Moving the latter to SimpleInvocationContext
reduces the size of the InvocationContext allocated upon every intercepted
invocation by 8 bytes (compressed oops), which compensates for the method
handle invoker field. No behavior change: getTimer() returned null before
as well, getConstructor() and the parameter checks are unchanged.
…out parameters

The intercepted subclass passes a shared, immutable empty array to the
method handler for methods without parameters instead of allocating a new
one upon every invocation (InvocationContext.getParameters() returns an
empty array as before; setParameters() replaces the array, it never writes
into it).
…nvoking thread

An around invoke InvocationContext now keeps the interception stack of the
thread that created it. proceed() uses it directly if it is called on the
same thread (Thread.currentThread() comparison) instead of looking up the
stack in the thread-local again. While the context references the stack,
it is guaranteed to be the stack getStack() returns on that thread (see
InterceptionDecorationContext.getStack()), so the behavior is the same.
On another thread (e.g. asynchronous proceed), the thread-local lookup is
performed as before.

The additional field does not increase the size of the context (it fits
into the padding left by the removal of the timer/constructor fields).
For methods with a non-void return type and no parameter or a single
reference type parameter (e.g. @AroundInvoke interceptor methods and
WeldSubclass.getter$$super()), MethodInvoker creates a Function/BiFunction
via LambdaMetafactory (in the package of the declaring class, requires the
package to be open to Weld) so that the JIT can inline the invocation,
which is not possible with a method handle stored in a field. Same
argument checks and exception wrapping as for the method handle; if the
lambda cannot be created, the method handle (or reflection) is used.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant