Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -61,6 +61,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- **The Substrate floor is `1.0.3` and `substrate.images.agentgateway` the agentgateway line's `2.1.2`: the gateways' client certificates follow their rotation whatever order kubelet writes a pod's bundles in** (giantswarm/giantswarm#37915). 2.1.1 re-read a *managed* file on every fetch, but the certificate files a Substrate policy's `policies.backendTLS` names are read at parse time and were never managed, so a rotation reached `substrateEgressActorResolution` and `substrateIngress` only when another resource's event reloaded the config after the podidentity bundle was on disk — the egress pod when kubelet wrote both bundles in one sync, the router by the cadence of its servicedns rotation. 2.1.2 registers those files as managed dependencies, watched and re-read every 60 s (giantswarm/agentgateway-upstream#23). Substrate `1.0.3` (`release-1.0`) pins it; `components.substrate*.versionRange` is `>=1.0.3 <1.1.0`, the BOM pin and the derived worker image follow, and the gateway pods roll once onto the new data plane. The platform's own controller and data planes move with it: `agentgateway.controller.image.tag` and `agentgateway.proxy.image.tag` are `2.1.2` in both charts (the packaging chart's 2.7.0 carries the same), so the line is one release everywhere it runs.
- **The kagent floor is `1.0.2`: a golden boot that failed Substrate's boot bound is started over instead of failing the AgentTemplate for good** (giantswarm/kagent-upstream#72). The controller retried only the unattributed `GoldenActorCrashed`. A `GoldenActorNotReady` from Substrate's boot bound (three boots in a row whose workload missed its readiness probe, about a minute apart, which is every boot while a skill repository or the egress is unreachable) was final, so a template that booted during an egress restart stayed `ActorTemplateFailed` until someone deleted it. kagent 1.0.2 starts such a boot over within the crash retry's budget: five retries per revision, backing off from 20 s to 2 min. Meanwhile it reports `Ready=False ActorTemplateRetrying` with Substrate's message, and `ActorTemplateFailed` with the count once the budget is spent. `components.kagent.versionRange` and `components.kagent-crds.versionRange` are `>=1.0.2 <1.1.0` and the BOM pins `1.0.2`. The previous range already admitted 1.0.2; the floor keeps an installation from resolving 1.0.1.
- **The Substrate floor is `1.0.2` and `substrate.images.agentgateway` the agentgateway line's `2.1.1`: the egress gateway's client certificate follows its rotation** (giantswarm/giantswarm#37915). atenet-egress and atenet-router authenticate to ate-api with the client certificate of a PodCertificate projected volume that kubelet rotates 30 minutes before its 24-hour expiry; agentgateway 2.0.0 read that file once, so a day after pod start every actor CONNECT was answered `403 Forbidden: actor identity check denied … CertificateExpired` and every sandbox's first outbound request was reset by peer — the Swarmgeist outage of 2026-09-23. The Substrate line's `1.0.2` (its `release-1.0` branch: 1.0.1 plus the pin and `podCertificates.maxExpirationSeconds`, unset by default) runs the agentgateway line's `2.1.1`, which re-reads a managed file resource on every fetch (giantswarm/agentgateway-upstream#20). `components.substrate*.versionRange` is `>=1.0.2 <1.1.0`: `make verify-substrate-images` checks the pin against every release the range admits, and `1.0.0` and `1.0.1` name `2.0.0`; the BOM pin and the derived worker image follow, so the pool rolls once with the control plane. The `agentgateway.controller.image` and `agentgateway.proxy.image` of the platform's own data planes are giantswarm/agent-platform#647's. `make verify-worker-image` now admits a Substrate pin that leads the kagent build's stamp by a patch of the same minor (kagent 1.0.1 is stamped against Substrate 1.0.0): a patch of the line changes no runtime contract, which is what the range shape `>=X.Y.Z <X.(Y+1).0` already encodes, so a data-plane fix ships without a kagent rebuild; a stamp newer than the pin, or in another minor, still fails.
- **Substrate's OTLP export reaches the collector with a tenant: `substrate.podLabels` and the connectivity chart's Substrate egress follow `substrate.otel`** (giantswarm/giantswarm#36711). The meta chart sets `substrate.podLabels` to `observability.giantswarm.io/tenant: giantswarm`, the label otlp-gateway routes a headerless export by; the substrate chart takes the key from the line's `1.1.0` (giantswarm/substrate#55, carrying kagent-dev/substrate#48), an older release ignores it. The `substrate-ate-api-server`, `-ate-controller`, `-atelet` and `-atenet-router` Cilium policies open the namespace and port of each enabled signal's endpoint (`agent-platform.substrate.otlpEgress`); ate-api-server's fixed rule to the cluster entity on 4317 is gone, and no endpoint opens nothing. Trace sampling stays at the chart's `parentbased_traceidratio` 0.01. `make verify-substrate-otlp` asserts it. **Every installation with Substrate on gets the four changed policies; the Substrate pods roll once on the upgrade to `1.1.0`** (the pod label).
- **The agentgateway data plane exports traces: the connectivity chart renders a Gateway-scoped `-tracing` policy, its egress policy and the tenant pod label** (giantswarm/giantswarm#36711). The proxy reads `OTEL_EXPORTER_OTLP_*` only as overrides of a tracing configuration, so the env of `gateway.parameters.dataPlaneEnv` started no exporter and no data-plane span reached Tempo. New: `AgentgatewayPolicy <release>-tracing` on the data-plane Gateway with `frontend.tracing` to the endpoint and protocol the env names (`global.observability.traces.otlp` wins when its endpoint is set), at agentgateway's default sampling: a request with a sampled `traceparent` is traced, one without is not. `<release>-dataplane-otlp-egress` allows DNS plus the endpoint's namespace on its port, in both policy flavours. The exporter sends no headers, so the new `gateway.parameters.podLabels` carries `observability.giantswarm.io/tenant: giantswarm` from the meta chart, the label otlp-gateway routes by. `verify-metric-labels` now counts templates with a `frontend.metrics` section, not any `frontend` section. `make verify-dataplane-tracing` asserts all of it. **Every installation with the agentgateway data plane gets two new objects and rolls the data-plane pods once** (the pod label).
- **muster's OTLP export reaches the collector: the connectivity chart renders `-muster-otlp-egress`** (giantswarm/giantswarm#36711). muster's own network policy (muster sub-chart) allows egress to the cluster entity on 80/443 only, so with `muster.muster.observability.otel.endpoint` set every export was dropped at the SYN, and no muster span or OTLP log reached Tempo or Loki. The new policy selects muster by name and allows DNS plus the endpoint's namespace on its port (the cluster entity for a host that is not an in-cluster Service), in both policy flavours, through the same `agent-platform.otlpTarget` / `agent-platform.otlpEgressRule` helpers as `-klausgateway-otlp-egress`. It renders exactly while the endpoint is set, with `networkPolicy.enabled` and the muster component on. `make verify-muster-otlp` asserts it. **Every installation with muster on gets one new object**; no pod rolls.
- **The agentgateway data plane declares cpu and memory on both sides of its budget, so `CPU_LIMIT` is the pod's and not the node's** (giantswarm/agent-platform#629, mechanism corrected in #631). `gateway.parameters.dataPlaneResources` carried an `ephemeral-storage` request and limit only — the pair the `require-emptydir-requests-and-limits` Kyverno policy needs for the writable `/tmp` emptyDir the controller injects. The container's cpu and memory requests came from the agentgateway controller's own defaults (100m / 128Mi, merged in under the container's name), and it had no cpu or memory limit at all. The cpu limit is load-bearing beyond capping the container: the generated pod carries **`CPU_LIMIT`** as a `resourceFieldRef` on this container's own `limits.cpu` (`divisor: 1`), and a `resourceFieldRef` against an unset limit resolves to the **node's** allocatable capacity, not the pod's budget — so the data plane sized its worker threads for whatever node it landed on. On gazelle, a 3.4 CPU node gave `CPU_LIMIT` 4; it is now 2. The block declares `requests` cpu 100m / memory 128Mi (what the controller already injected, so the render of the requests is unchanged — stated here so the whole budget is in one place and a controller default change on an agentgateway bump cannot move it) and `limits` cpu 2000m / memory 512Mi, the ephemeral-storage pair unchanged. `limits.cpu` is read as whole cores rounded up, so the value IS what the proxy sees. The memory limit feeds no env var — the data plane is Rust, and the pod carries no `GOMEMLIMIT`/`GOMAXPROCS` — so it is a plain cgroup ceiling; 512Mi is roughly 15x the working set measured on gazelle (12-33 MiB across both replicas). Both charts carry the block: the meta chart forwards its whole values tree to the connectivity release and a forwarded copy of a default shadows the child's, so a connectivity-only default would never have reached an installation (the #303 shape, and #522 again). `make verify-dataplane-ha` asserts cpu, memory and ephemeral-storage on each side of the rendered container's budget — each side cut out on its own, because a `grep` window over `limits:` runs on into `requests:` and passes a container that has no limit at all — that an installation's own limits reach the container in millicores, and that the meta chart forwards the two limits; `tests/verify-mirrored-values.py` holds all six leaves equal between the two charts, so the budget cannot drift the way the egress lists did. **Every installation gets a data-plane pod roll.** The PodDisruptionBudget does not cover it — a budget governs evictions, not a Deployment's own rolling update — and the `AgentgatewayParameters` sets no `strategy`, so the roll runs at whatever the agentgateway controller defaults to across the two replicas.
Expand Down
32 changes: 32 additions & 0 deletions Makefile.custom.mk
Original file line number Diff line number Diff line change
Expand Up @@ -3529,6 +3529,38 @@ verify-muster-otlp: ## Assert muster's OTLP egress (giantswarm/giantswarm#36711)
@helm template t $(CONNECTIVITY_DIR) $(VM) --set muster.muster.observability.otel.endpoint=$(MUSTER_OTLP_EP) --set components.muster.enabled=false 2>/dev/null | grep -q 'name: $(MUSTER_OTLP_POLICY)$$' && { echo "FAIL: the OTLP policy renders with the component off"; exit 1; } || true
@echo "ok: $@"

SUBSTRATE_OTLP_EXPORTERS := substrate-ate-api-server substrate-ate-controller substrate-atelet substrate-atenet-router
SUBSTRATE_OTLP_VM := $(VM) --set components.substrate.enabled=true --set substrate.otel.endpoint=$(MUSTER_OTLP_EP)

.PHONY: verify-substrate-otlp
verify-substrate-otlp: ## Assert Substrate's OTLP export reaches a tenant (giantswarm/giantswarm#36711): the meta chart forwards substrate.otel.endpoint to the connectivity release and substrate.podLabels (the observability.giantswarm.io/tenant label otlp-gateway routes a headerless export by) to the substrate release; the connectivity chart's policies of the four exporters (ate-api-server, ate-controller, atelet, atenet-router) each open the endpoint's namespace on its port, no cluster-wide 4317 rule remains, a signal's own endpoint adds its destination and a disabled signal's does not, an endpoint that is not an in-cluster Service is the cluster entity on its port, and no endpoint opens nothing.
@echo "====> $@ ($(CHART_DIR) + $(CONNECTIVITY_DIR))"
@helm template t $(CHART_DIR) $(VM) $(SUBSTRATE_ON) >/tmp/vso-meta.out 2>&1 || { cat /tmp/vso-meta.out; exit 1; }
@$(PICK) /tmp/vso-meta.out HelmRelease substrate | grep -A1 '^ podLabels:$$' | grep -q '^ observability.giantswarm.io/tenant: giantswarm$$' || { echo "FAIL: the substrate release does not receive the tenant pod label"; exit 1; }
@$(PICK) /tmp/vso-meta.out HelmRelease agent-platform-connectivity | grep -A1 '^ otel:$$' | grep -q '^ endpoint: $(MUSTER_OTLP_EP)$$' || { echo "FAIL: the connectivity release does not receive substrate.otel.endpoint"; exit 1; }
@echo "ok: meta chart forwards the endpoint and the tenant label"
@helm template t $(CONNECTIVITY_DIR) $(SUBSTRATE_OTLP_VM) >/tmp/vso.out 2>&1 || { cat /tmp/vso.out; exit 1; }
@for p in $(SUBSTRATE_OTLP_EXPORTERS); do \
$(PICK) /tmp/vso.out CiliumNetworkPolicy $$p >/tmp/vso-pol.out || { echo "FAIL: no CiliumNetworkPolicy $$p"; exit 1; }; \
grep -B1 -A4 '^ - matchLabels:$$' /tmp/vso-pol.out | grep -A4 'io.kubernetes.pod.namespace: kube-system$$' | grep -q 'port: "4317"' || { echo "FAIL: $$p does not open kube-system:4317"; cat /tmp/vso-pol.out; exit 1; }; \
if grep -B6 'port: "4317"' /tmp/vso-pol.out | grep -q -- '- cluster$$'; then echo "FAIL: $$p still opens 4317 to the whole cluster"; exit 1; fi; \
done
@if $(PICK) /tmp/vso.out CiliumNetworkPolicy substrate-atenet-egress | grep -q 'port: "4317"'; then echo "FAIL: atenet-egress exports nothing of its own, yet opens 4317 with kagent off"; exit 1; fi
@echo "ok: the four exporters open the endpoint"
@helm template t $(CONNECTIVITY_DIR) $(SUBSTRATE_OTLP_VM) --set substrate.otel.traces.endpoint=http://tempo-gw.tracing.svc:4317 >/tmp/vso-sig.out 2>&1 || { cat /tmp/vso-sig.out; exit 1; }
@$(PICK) /tmp/vso-sig.out CiliumNetworkPolicy substrate-atelet | grep -q 'io.kubernetes.pod.namespace: tracing$$' || { echo "FAIL: a signal's own endpoint does not add its destination"; exit 1; }
@$(PICK) /tmp/vso-sig.out CiliumNetworkPolicy substrate-atelet | grep -q 'io.kubernetes.pod.namespace: kube-system$$' || { echo "FAIL: the other signals' shared endpoint is gone"; exit 1; }
@helm template t $(CONNECTIVITY_DIR) $(SUBSTRATE_OTLP_VM) --set substrate.otel.traces.endpoint=http://tempo-gw.tracing.svc:4317 --set substrate.otel.traces.enabled=false >/tmp/vso-off.out 2>&1 || { cat /tmp/vso-off.out; exit 1; }
@if $(PICK) /tmp/vso-off.out CiliumNetworkPolicy substrate-atelet | grep -q 'io.kubernetes.pod.namespace: tracing$$'; then echo "FAIL: a disabled signal's endpoint is opened"; exit 1; fi
@echo "ok: per-signal endpoints"
@helm template t $(CONNECTIVITY_DIR) $(VM) --set components.substrate.enabled=true --set substrate.otel.endpoint=https://collector.example.com >/tmp/vso-ext.out 2>&1 || { cat /tmp/vso-ext.out; exit 1; }
@$(PICK) /tmp/vso-ext.out CiliumNetworkPolicy substrate-ate-controller | grep -A4 -- '- cluster$$' | grep -q 'port: "443"' || { echo "FAIL: an external collector is not the cluster entity on 443"; exit 1; }
@helm template t $(CONNECTIVITY_DIR) $(VM) --set components.substrate.enabled=true >/tmp/vso-none.out 2>&1 || { cat /tmp/vso-none.out; exit 1; }
@for p in $(SUBSTRATE_OTLP_EXPORTERS); do \
if $(PICK) /tmp/vso-none.out CiliumNetworkPolicy $$p | grep -q 'OTLP gateway\|port: "4317"'; then echo "FAIL: $$p opens an OTLP destination with no endpoint"; exit 1; fi; \
done
@echo "ok: $@"

.PHONY: verify-kagent-storage-version
verify-kagent-storage-version: ## Assert the kagent CRDs' storage-version hooks of the 3.x → 4.x cut-over (#396): with kagent on, the backup Job (pre-install,pre-upgrade, -7: records the objects of modelconfigs/modelproviderconfigs/remotemcpservers.kagent.dev still stored at v1alpha2 into the migration ConfigMap, sets the crds policy of the HelmRelease the CRDs' Flux labels name to Skip (#416), deletes those CRDs and watches them stay absent for 60 s — a re-created one is deleted again and fails the hook naming the owner) and the restore Job (post-install,post-upgrade, 0: waits for modelconfigs.kagent.dev to serve v1alpha3, re-creates the recorded ModelConfigs no Helm release owned at kagent.dev/v1alpha3, tolerates AlreadyExists, marks restored-at) as the hook identity in the helm image, the identity at their events; with the engine off (the fleet) the same pair and nothing else; with kagent off none of it; the kagent namespace follows kagent.namespaceOverride; helm lint.
@echo "====> $@ ($(CHART_DIR))"
Expand Down
32 changes: 32 additions & 0 deletions helm/agent-platform-connectivity/templates/_helpers.tpl
Original file line number Diff line number Diff line change
Expand Up @@ -2114,6 +2114,38 @@ Usage: include "agent-platform.kagent.otlpTargets" . | fromJsonArray
{{- $targets | toJson -}}
{{- end -}}

{{/*
The cilium egress rules to the OTLP gateways Substrate's control plane exports
to (ate-api-server, ate-controller, atelet and the atenet router), one per
distinct destination of the signals that are on. A signal exports when it is
enabled (the substrate chart's default) and resolves to an endpoint: its own
substrate.otel.<signal>.endpoint, else substrate.otel.endpoint — the chart's
substrate.otel.signalEndpoint. The chart's exporters speak gRPC only. Empty
when no signal has an endpoint.
*/}}
{{- define "agent-platform.substrate.otlpEgress" -}}
{{- $otel := dig "otel" dict (.Values.substrate | default dict) -}}
{{- $seen := dict -}}
{{- $first := true -}}
{{- range $signal := list "traces" "metrics" "logs" -}}
{{- $cfg := index $otel $signal | default dict -}}
{{- if ne (toString (dig "enabled" true $cfg)) "false" -}}
{{- $endpoint := $cfg.endpoint | default $otel.endpoint | default "" | toString | trim -}}
{{- if $endpoint -}}
{{- $t := include "agent-platform.otlpTarget" (dict "endpoint" $endpoint "protocol" "grpc") | fromJson -}}
{{- $key := printf "%s:%s" $t.namespace $t.port -}}
{{- if not (hasKey $seen $key) -}}
{{- $_ := set $seen $key true -}}
{{- if not $first }}
{{ end -}}
{{- $first = false -}}
{{- include "agent-platform.otlpEgressRule" (dict "target" $t "who" "Substrate's exporters send to") -}}
{{- end -}}
{{- end -}}
{{- end -}}
{{- end -}}
{{- end -}}

{{/*
The cilium egress rules to the OTLP gateways kagent's exporters send to
(agent-platform.kagent.otlpTargets), one per destination
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,9 @@ what (the substrate chart's arguments and Service definitions):
store :9000.
ate-controller, dns, podcertificate-controller ──▶ the apiserver (and
ate-controller ──▶ ate-api :443); dns answers the actors' lookups.
ate-api-server, ate-controller, atelet, atenet-router ──▶ the OTLP gateway
substrate.otel names (agent-platform.substrate.otlpEgress); atenet-egress
exports nothing.
Prometheus scrapes the annotated pods (atelet 9090, ate-api-server 9090,
ate-controller 8080, atenet-router 9090, the agentgateway sidecars 15020) —
admitted from the cluster entity, as the platform's other metrics ports are.
Expand Down Expand Up @@ -176,19 +179,16 @@ spec:
protocol: TCP
{{- end }}
# The snapshot store: an S3 (or S3-compatible) endpoint, and a CAPA
# cluster's IRSA token exchange with STS; the OTLP gateway.
# cluster's IRSA token exchange with STS.
- toEntities:
- world
toPorts:
- ports:
- port: "443"
protocol: TCP
- toEntities:
- cluster
toPorts:
- ports:
- port: "4317"
protocol: TCP
{{- with include "agent-platform.substrate.otlpEgress" . }}
{{- . | nindent 4 }}
{{- end }}
{{- if $rustfs }}
- toEndpoints:
- matchLabels:
Expand Down Expand Up @@ -242,6 +242,9 @@ spec:
- ports:
- port: "443"
protocol: TCP
{{- with include "agent-platform.substrate.otlpEgress" . }}
{{- . | nindent 4 }}
{{- end }}
---
# atelet: the per-node agent. Its gRPC API (hostPort 8085, mTLS with the pod
# identity CA) is driven by ate-api-server; the source of hostPort traffic is
Expand Down Expand Up @@ -289,6 +292,9 @@ spec:
- ports:
- port: "443"
protocol: TCP
{{- with include "agent-platform.substrate.otlpEgress" . }}
{{- . | nindent 4 }}
{{- end }}
{{- if $rustfs }}
- toEndpoints:
- matchLabels:
Expand Down Expand Up @@ -375,6 +381,9 @@ spec:
protocol: TCP
- port: "8443"
protocol: TCP
{{- with include "agent-platform.substrate.otlpEgress" . }}
{{- . | nindent 4 }}
{{- end }}
---
# atenet-egress: the actors' egress gateway. The worker pods' tunnels come in
# on 8443; every connection an actor opens leaves from here, so this is where
Expand Down
Loading
Loading