CacheBackend uses Kubernetes label selectors to find the engine pods it injects cache wiring into. This page explains the model, the lifecycle, and the common ways it goes wrong.
Three actors participate in the binding:
- CacheBackend — the namespaced CR you create. Its
spec.engineSelector.matchLabelsis a label selector over pods in the same namespace, with the same semantics asService.spec.selector. - Engine pod — a vLLM (or other supported runtime) pod, typically owned by a user-managed Deployment. Its
template.metadata.labelsare what the selector matches against. - Mutating Pod webhook — the controller's admission webhook intercepts pod CREATE and stamps the matched engine pod with the LMCache engine wiring (env vars + CLI args). The kvevent-subscriber observation sidecar is appended in addition only when the controller is started with
--kvevent-subscriber-imageset; lifecycle step 3 below has the full conditions.
+-----------------------+
| CacheBackend (CR) |
| spec.engineSelector |
+-----------+-----------+
| label-selector match (at pod CREATE)
v
+-----------------+ pod CREATE +------------+-----------+
| Engine | -----------+----> | Mutating Pod webhook |
| Deployment | | | (matches selector; |
| template: | | | injects engine config;|
| labels: {...} | | | +subscriber sidecar*) |
+-----------------+ | +------------+-----------+
| |
| v
| +----------+--------------+
| | Engine pod |
| | env: LMCACHE_* |
| | args: --kv-... |
| | sidecar: subscriber* |
| +----------+--------------+
| |
| v subscriber publishes
| +----------+-----------+
| | lmcache-server pod |
+------> | (managed by the CR; |
| endpoint published |
| in status.endpoint) |
+----------------------+
The match is evaluated once at pod CREATE by the mutating webhook. The wiring is sticky to the life of the pod; relabeling an existing pod does not re-evaluate it. To opt a pod out regardless of label match, set inferencecache.io/skip-inject: "true" on the pod template. Skipped pods are stamped with inferencecache.io/inject-skipped: "skip-inject-annotation" and receive a SkippedByOperator Event, so an intentional opt-out is distinguishable from selector drift.
* The kvevent-subscriber sidecar is opt-in. It is appended only when the controller is started with --kvevent-subscriber-image set (empty by default) AND the matched CacheBackend has spec.observation.modelID configured; otherwise the engine is wired without it. The default install does not auto-attach the sidecar.
Native SGLang HiCache exception.
type: SGLangHiCacheis engine-local: the controller creates no backend workload or endpoint, and the webhook does not wait forstatus.endpoint. It injects the typedspec.hiCachelaunch flags directly into matching SGLang Pods. The LMCache lifecycle below remains the endpoint-bearing path.
- Apply the CacheBackend.
kubectl apply -f cachebackend.yaml. The reconciler creates the managed lmcache-server Deployment + Service and publishes the resolved address instatus.endpoint. - Deploy the engine. Apply an engine Deployment whose pod template labels include every key/value in
spec.engineSelector.matchLabels. New pods from that Deployment hit the mutating webhook at admission time. - The webhook claims matching pods (precondition: status.endpoint published). When the matched CacheBackend has
status.endpointpopulated by the time admission runs, the webhook injects LMCache env vars, the--kv-transfer-configCLI arg, and stamps TWO annotations:inferencecache.io/injected-by: <ns>/<name>(operator-readable identity, visible inkubectl describe pod) andinferencecache.io/injected-by-uid: <cache.UID>(the matched CR'smetadata.uid— an apiserver-assigned identifier that the events controller cross-checks against the live CR's UID before emitting. The check catches casual copy-paste of an injected pod's annotations into a fresh template, but it isn't a security boundary: UIDs are readable metadata, so a pod creator withgetRBAC on CacheBackends can stamp the pair correctly). When the controller is started with--kvevent-subscriber-imageset (empty by default) AND the matched CacheBackend has a model id configured, the kvevent-subscriber sidecar is also appended; otherwise the engine is wired without the sidecar. A controller watching pods then validates the UID annotation against the live CR's UID and records aNormal InjectedByCacheBackendevent on the now-persisted pod, visible inkubectl describe pod. (The event is emitted from a controller rather than the webhook because the apiserver does not assignmetadata.uiduntil after mutating admission — an event recorded from the webhook would haveinvolvedObject.uid=""and would not surface under describe. The UID annotation is what lets the controller distinguish a real webhook injection from a user-suppliedinjected-byannotation when the webhook is unreachable underfailurePolicy=Ignore.) If the CacheBackend'sstatus.endpointis empty at admission time (the operator applied the engine Deployment before the reconciler had a chance to publish it), the webhook fail-opens — the pod is admitted unwired, no annotations are stamped, no Event is emitted. Because admission is CREATE-only, this is permanent for that pod; recovery iskubectl rollout restart deploy/<engine>so new pods re-enter admission. Note thatstatus.matchedEnginePodsstill counts the unwired pod (the selector matches its labels), soMatched > 0does not guarantee the pod was injected — the per-pod annotations and Event are the authoritative wiring signals. - KV events flow (when the sidecar is configured). When the subscriber sidecar is present, it streams the engine's KV-cache events to the policy server's index, which surfaces them in
CacheBackend.status(the index-participation status fields). Without the sidecar, the binding is still observable —status.matchedEnginePodsand the per-pod annotation/Event still materialize from step 3 — but no KV events flow and the index-participation fields stay unset. - Observe and debug.
kubectl get cachebackendshows theMatchedcolumn — the snapshot count of pods whose labels currently match.kubectl describe pod <engine-pod>shows which CacheBackend (if any) claimed it.
A single CacheBackend with a matching engine Deployment. The label app: qwen-demo appears in two places — that's what binds them:
apiVersion: inferencecache.io/v1alpha1
kind: CacheBackend
metadata:
name: qwen-demo-cache # <-- CR name; deliberately distinct from the engine Deployment name (see note below)
spec:
runtime: VLLM
type: LMCache
integration:
role: ReadWrite
engineSelector:
matchLabels:
app: qwen-demo # <-- selector key/value (1 of 2; binding is by label, not by resource name)
observation:
modelID: Qwen/Qwen2.5-0.5B-Instruct
remoteStorage:
provider: LMCacheServer
ownership: Managed
lmCacheServer: {}
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: qwen-engine # <-- engine Deployment name; must differ from the CR name above (the controller reconciles the CR into an lmcache-server Deployment whose name equals the CR's name, so sharing names would collide on Create)
spec:
replicas: 1
selector:
matchLabels:
app: qwen-demo
template:
metadata:
labels:
app: qwen-demo # <-- selector key/value (2 of 2; this is what the webhook sees)
spec:
containers:
- name: vllm
# Arch-tagged: swap to `:latest-arm64` on Apple Silicon hosts.
image: vllm/vllm-openai-cpu:latest-x86_64
args: ["--model", "Qwen/Qwen2.5-0.5B-Instruct"]A copy-pasteable version of this pair ships at config/samples/cachebackend-with-engine.yaml.
| Symptom | Cause | How to detect | Fix |
|---|---|---|---|
| Engine pod runs uncached; no LMCache env on its container | Selector and pod labels don't overlap (typo, drift after a Deployment rename, etc.) | kubectl describe pod <engine-pod> shows no InjectedByCacheBackend or SkippedByOperator event and neither inferencecache.io/injected-by nor inferencecache.io/inject-skipped; kubectl get cachebackend shows Matched: 0; kubectl get cachebackend <name> -o jsonpath='{.status.engineSelectorMessage}' echoes the selector that matched no pods; kubectl describe cachebackend <name> shows a Normal EngineSelectorUnmatched Event when the CR first observes zero matches or transitions from matched to zero |
Reconcile the label sets: either fix engineSelector.matchLabels on the CR or fix the Deployment's pod template labels |
| Multiple CacheBackends overlap on the same pod | Two CacheBackends in the namespace have selectors that both match | The webhook picks the lexicographically-first match by metadata.name (sort is deterministic so the picked CR is reproducible across re-admissions) and stamps inferencecache.io/injected-by on the pod — kubectl describe pod and that annotation name the CR that actually injected. There is no admission validator for selector overlap today; multi-match is a misconfiguration the operator must avoid by hand |
Pick one CR for the pod; delete or narrow the other so each engine pod's labels match exactly one CacheBackend |
| Engine pod was labeled after creation but still uncached | Label match is evaluated once at pod CREATE; relabeling later has no effect | kubectl describe pod <engine-pod> shows no InjectedByCacheBackend event |
Delete the pod (kubectl delete pod <engine-pod>); the Deployment will recreate it and the new pod will re-enter admission |
| CacheBackend was deleted, but engine pods are still running with the old wiring | Wiring is sticky to the pod's lifetime; deleting the CR does not retract env vars from already-admitted pods | Engine logs show LMCache connect failures to a no-longer-existing Service | Rolling-restart the engine Deployment to admit fresh pods (which will match no CR and run uncached) |
| Engine pod intentionally runs uncached | The pod template has inferencecache.io/skip-inject set to a truthy value |
kubectl get pod <pod> -o jsonpath='{.metadata.annotations.inferencecache\.io/inject-skipped}' returns skip-inject-annotation, and kubectl describe pod <pod> shows a Normal SkippedByOperator Event |
No fix needed if the opt-out was intentional; remove the skip annotation from the pod template and restart if the pod should be wired |
| Pod that should be skipped still gets wiring | The inferencecache.io/skip-inject: "true" annotation was missing or set on the Deployment, not the pod template |
kubectl get pod <pod> -o yaml shows no inferencecache.io/skip-inject annotation |
Add the annotation under spec.template.metadata.annotations of the Deployment and restart |
The inferencecache.io/skip-inject annotation is the explicit escape hatch: any non-empty value other than the falsey set opts the pod out. The falsey set is what Go's strconv.ParseBool accepts as false (false/FALSE/False/f/F/0) plus the case-insensitive synonyms no/off/disable/disabled. Use the annotation for pods you've already pre-wired or that should run vanilla for a debugging experiment. A skipped pod keeps the original skip-inject annotation, gets inferencecache.io/inject-skipped: "skip-inject-annotation", and does not get inferencecache.io/injected-by.