[6.0] Reduce invocation overhead for intercepted bean methods - #3546
Open
AngeloRubens wants to merge 17 commits into
Open
AngeloRubens wants to merge 17 commits into
AngeloRubens wants to merge 17 commits into
Conversation
…s when none exists Every business method invocation on a client proxy calls InterceptionDecorationContext.startIfNotEmpty(). When the caller is not intercepted and no RequestScopedCache is active (e.g. Weld SE, or any invocation outside of an HTTP/EJB request), getStack() allocated a new Stack (+ ArrayDeque), set it into the thread-local and then immediately removed it again in removeIfEmpty(). The net effect is nothing, but ThreadLocal.set()/remove() plus the allocation dominate the cost of a client proxy invocation. Return null straight away if there is no stack on the current thread. This is semantically identical: an absent stack is an empty stack and the old code never left anything behind in that case (with an active RequestScopedCache the empty stack would have been cached; it is now created lazily by the first startIfNotOnTop()/getStack() instead).
… stack Outside of an active RequestScopedCache each intercepted invocation (and each non-intercepted method of an intercepted subclass) creates a Stack, sets it into the thread-local and removes it again once the stack gets empty. ThreadLocal.remove() clears the weak reference of the map entry (native Reference.clear0) and expunges the entry, so the next set() allocates a new WeakReference entry again. The profile of an intercepted @ApplicationScoped bean shows remove() + expunge + entry allocation as the most expensive part of the invocation. Set the thread-local value to null instead: the entry is reused, and a null value does not retain any Weld class, so there is no class loader leak. Also allocate the ArrayDeque with a small initial capacity as the stack is usually very shallow.
ContextBeanInstance.getInstance() called Container.isSet(contextId) on each invocation to detect usage of a client proxy after shutdown. This is a ConcurrentHashMap lookup in RegistrySingletonProvider and shows up in the profile of @ApplicationScoped invocations. ContextBeanInstance now keeps the Container it was created for. Container gets a volatile cleanedUp flag set at the beginning of cleanup(); while it is false the container is guaranteed to be registered, so the registry lookup is only performed after the container was cleaned up. Behaviour is unchanged (including the case where a new container was registered under the same id afterwards).
Every invocation of an intercepted subclass method (both intercepted and non-intercepted methods push the subclass handler) outside of a request allocated a new Stack and ArrayDeque, set the thread-local and cleared it again once the stack became empty. Within a request the stack was kept in the RequestScopedCache instead. The thread-local now holds a small Object[] holder (a JDK class) that references the per-thread stack strongly only while it is not empty and weakly otherwise. An idle thread never retains a Weld class, so there is no class loader leak, and the stack is reused until the next GC clears the weak reference. The RequestScopedCache registration and the validity tracking of the stack are no longer needed.
…rst used Client proxies call InterceptionDecorationContext.startIfNotEmpty() on every invocation, i.e. a ThreadLocal lookup. As long as no thread has ever created an interception context (no intercepted or decorated method has been invoked yet, which is always the case in deployments without interceptors and decorators) the lookup cannot find anything. startIfNotEmpty() now invokes a constant MethodHandle guarded by a SwitchPoint: it returns null without any lookup until the first per-thread stack holder is created, when the SwitchPoint is invalidated (once, before the holder is published to that thread). The JIT compiles the guard to nothing, so applications that do use interception do not pay for an additional check either. A thread can only have an active stack if it created its holder itself, so the switch cannot be observed too late.
… check in client proxies ContextBeanInstance keeps the bean as RIBean when possible so that the client proxy invocation goes straight to the bean's ContextualInstanceStrategy (for @ApplicationScoped beans a volatile read of the cached instance) instead of going through the instanceof check in ContextualInstance. Invalidation is unchanged: the strategy remains the single place where the instance is cached.
Replace the ArrayDeque inside InterceptionDecorationContext.Stack with an array and a size index that grows on demand. push/pop/peek/startIfNotOnTop need fewer branches and no ring-buffer index arithmetic. Popped slots are cleared so that no handler is retained, and the weakly referenced idle stack (no class loader leak) is unchanged. Null elements are still rejected and popping an empty Stack still throws NoSuchElementException.
Interceptor methods (@AroundInvoke etc.) and the proceed method of an around invoke chain (WeldSubclass.method$$super() or the decorated method) used to be invoked by Method.invoke(). They are now invoked through a MethodHandle adapted to (Object, Object)Object or (Object, Object[])Object, created lazily upon the first invocation and cached per method (MethodInvoker; per interceptor metadata for interceptor methods, per declaring class via a ClassValue for proceed methods). The semantics of Method.invoke() are kept exactly: - the handle is only used if the receiver is an instance of the declaring class and every argument matches the parameter type exactly (reference type or exact wrapper of the primitive type), i.e. if the invocation cannot fail before reaching the method; otherwise (null receiver, wrong arguments, widening primitive conversions, static methods...) the Method is invoked reflectively and throws the same exceptions as before - any Throwable thrown by the invoked method is wrapped in an InvocationTargetException, so that the existing unwrapping code paths are unchanged - accessibility is never changed; a handle that cannot be created (e.g. module readability) results in reflective invocations
…hed interception chains InterceptorMethodHandler.invoke() called Reflections.ensureAccessible() (AccessibleObject.canAccess()) and TargetClassInterceptorMetadata .isInterceptorMethod() (a hash set lookup) upon every invocation, before looking up the cached interception chain. A cached chain implies that both checks have been performed already for the given method and proceed method (the chain is created after them), and neither result can change, so the chain is now looked up first and used directly.
… instance Avoids the ConcurrentHashMap lookup in InterceptorMethodHandler for the first intercepted method that has been invoked on a given instance (e.g. a bean with a single intercepted method). The field is set at most once, the chain is immutable (final fields only), other methods keep using the map.
…ocation contexts The timer field of AbstractInvocationContext was always null and the constructor field is only needed for around construct interception (SimpleInvocationContext). Moving the latter to SimpleInvocationContext reduces the size of the InvocationContext allocated upon every intercepted invocation by 8 bytes (compressed oops), which compensates for the method handle invoker field. No behavior change: getTimer() returned null before as well, getConstructor() and the parameter checks are unchanged.
…out parameters The intercepted subclass passes a shared, immutable empty array to the method handler for methods without parameters instead of allocating a new one upon every invocation (InvocationContext.getParameters() returns an empty array as before; setParameters() replaces the array, it never writes into it).
…nvoking thread An around invoke InvocationContext now keeps the interception stack of the thread that created it. proceed() uses it directly if it is called on the same thread (Thread.currentThread() comparison) instead of looking up the stack in the thread-local again. While the context references the stack, it is guaranteed to be the stack getStack() returns on that thread (see InterceptionDecorationContext.getStack()), so the behavior is the same. On another thread (e.g. asynchronous proceed), the thread-local lookup is performed as before. The additional field does not increase the size of the context (it fits into the padding left by the removal of the timer/constructor fields).
For methods with a non-void return type and no parameter or a single reference type parameter (e.g. @AroundInvoke interceptor methods and WeldSubclass.getter$$super()), MethodInvoker creates a Function/BiFunction via LambdaMetafactory (in the package of the declaring class, requires the package to be open to Weld) so that the JIT can inline the invocation, which is not possible with a method handle stored in a field. Same argument checks and exception wrapping as for the method handle; if the lambda cannot be created, the method handle (or reflection) is used.
github-actions
Bot
force-pushed
the
pr/proxy4-6.0
branch
from
October 5, 2026 12:53
e2cf6fc to
9d8a605
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Reduces the steady-state cost of calls to methods on intercepted beans. This PR includes the prerequisite client-proxy and interception-stack changes because they have not yet been merged upstream, followed by invocation-chain caching, smaller invocation contexts, shared empty arguments, per-thread stack reuse in
InvocationContext.proceed(), and cached invokers. For eligible instance methods, a LambdaMetafactory function lets HotSpot inline interceptor dispatch. Reflection remains the fallback. A regression fix checks that a cached privileged handle cannot bypass a missing or revokedsetAccessibleoverride.The Weld 6 backport is submitted separately against
6.0.Validation
Performance
Interception scenarios derived from OWB benchmarks
The CDI benchmark scenarios and harness in this comparison follow the work of Matej Novotny, creator of OpenWebBeans and author of the Weld Core MicroBenchmarks. Results below are from the same GitHub Actions run, executed sequentially with one thread, two forks, 2 × 1 s warmup and 4 × 1 s measurement iterations. Values are JMH throughput in million benchmark operations per second (
ops/µs); these are method calls in these scenarios, not HTTP requests.In these two interception scenarios, the PR moves Weld from below OWB to above it. This does not describe every benchmark in the suite: for example, OWB remains faster in the application-scoped and request-scoped scenarios. The quick comparison has only two forks and relatively short iterations, so the exact ratios need confirmation with longer repeated runs. Full report and raw JMH results, run 37314014340.
Official Weld JMH suite
As an independent cross-check, Weld's official microbenchmark suite was run against base and PR sequentially on the same runner, with four threads, five forks, and 5 × 5 warmup/measurement iterations. It exercises its own Decorator and Interceptor workloads, which are distinct from the interception scenarios above.
The official-suite differences are small relative to their JMH error margins and do not establish a general throughput reversal. Official JMH run 37335224493.
In the separate proxy-3-to-proxy-4 incremental comparison, Boot/shutdown was 44.95 ± 3.15 → 48.15 ± 2.76 ms. A separate ten-fork cold-start run invokes intercepted methods on a fresh container: Weld 6 was 686.08 ± 22.97 ms plain and 691.57 ± 16.62 ms with first intercepted calls. These runs show no clear startup regression. Cold-start results, run 37281596417.
The Weld 7 GC profile measured 0.0008 ± 0.0221 B/op for the intercepted-method benchmark, down from about 64 B/op in its previous proxy-3 profile. This is supporting evidence for the shared invoker implementation; it depends on JIT inlining/escape analysis and is not guaranteed for polymorphic call sites. Profile, run 37281324917.
No CDI self-invocation semantics are intentionally changed. The earlier client-proxy and stack changes are included for review because this optimization series depends on them.
Origin of the work
This work originated in the analysis of OpenWebBeans (OWB). Its approach to invocation and interception performance was the starting point for this optimization series, which was adapted to Weld internals and compatibility requirements.