Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 7 additions & 3 deletions assets/agw-docs/pages/agentgateway/llm/failover.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,11 @@ For {{< reuse "agw-docs/snippets/agentgateway.md" >}}, you can set up failover a
Failover in {{< reuse "agw-docs/snippets/agentgateway.md" >}} has two parts:

- **Priority groups** in the {{< reuse "agw-docs/snippets/backend.md" >}} define the failover order. Each group is a tier. Models within the same group are load balanced equally. When all models in a group are evicted, requests fail over to the next group.
- **A health policy** in an {{< reuse "agw-docs/snippets/policy.md" >}} defines what counts as an unhealthy response (such as 5xx errors or 429 rate limits) and how to evict unhealthy backends. Without a health policy, backends are not evicted and failover does not occur.
- **A health policy** in an {{< reuse "agw-docs/snippets/policy.md" >}} defines what counts as an unhealthy response (such as 5xx errors or 429 rate limits) and how to evict unhealthy backends. {{< version include-if="1.0.x,1.1.x,1.2.x,1.3.x,1.4.x,1.5.x,2.2.x" >}}Without a health policy, backends are not evicted and failover does not occur.{{< /version >}}{{< version exclude-if="1.0.x,1.1.x,1.2.x,1.3.x,1.4.x,1.5.x,2.2.x" >}}Without a health policy, an {{< reuse "agw-docs/snippets/backend.md" >}} with more than one priority group uses default eviction, as described after this list.{{< /version >}}

{{< version exclude-if="1.0.x,1.1.x,1.2.x,1.3.x,1.4.x,1.5.x,2.2.x" >}}
When no health policy targets an {{< reuse "agw-docs/snippets/backend.md" >}} that has more than one priority group, default eviction applies. A single 5xx response or connection failure evicts the backend for 3 seconds, and each repeated eviction lasts longer. To classify more responses as unhealthy, such as 429, or to tune eviction, add a health policy. A health policy replaces the default eviction instead of adding to it. If the health policy has no `eviction` block, a backend is evicted only when a retry policy's `backoff`, or a `Retry-After` header on a response that the policy classifies as unhealthy, supplies an eviction duration. To keep the default eviction settings in your health policy, set `eviction: {}`.
{{< /version >}}

This approach increases the resiliency of your network environment by ensuring that apps that call LLMs can keep working without problems, even if one model has issues.

Expand Down Expand Up @@ -234,7 +238,7 @@ For weight-based traffic distribution within a priority group (such as 80/20 spl
```


3. Create an {{< reuse "agw-docs/snippets/policy.md" >}} with a health policy that targets the {{< reuse "agw-docs/snippets/backend.md" >}}. The health policy defines which responses are considered unhealthy and how to evict backends. Without this policy, backends are not evicted and failover does not occur.
3. Create an {{< reuse "agw-docs/snippets/policy.md" >}} with a health policy that targets the {{< reuse "agw-docs/snippets/backend.md" >}}. The health policy defines which responses are considered unhealthy and how to evict backends. {{< version include-if="1.0.x,1.1.x,1.2.x,1.3.x,1.4.x,1.5.x,2.2.x" >}}Without this policy, backends are not evicted and failover does not occur.{{< /version >}}{{< version exclude-if="1.0.x,1.1.x,1.2.x,1.3.x,1.4.x,1.5.x,2.2.x" >}}Without this policy, default eviction still fails over on 5xx responses and connection failures. The policy replaces the default eviction, so each of the following examples sets its own `eviction` settings. The first example also evicts on 429 rate-limit responses.{{< /version >}}

The `unhealthyCondition` field is an optional [CEL expression](https://github.com/cel-expr/cel-spec) that classifies each response. When you set it, `true` means the response counts as unhealthy toward eviction. The `eviction` settings control how many failures and how long an unhealthy backend stays out of its priority group.

Expand Down Expand Up @@ -727,7 +731,7 @@ Retries and eviction do different jobs here, and transparent failover needs both
* The retry supplies that next attempt inside the same client request, so the client never sees the 500.

> [!IMPORTANT]
> A retry policy on its own does not fail over. Without a health policy, no backend is evicted, so every retry returns to the same highest-priority group and the client still receives the error. To fail over transparently, configure both policies.
> A retry fails over only when the failing backend is evicted first. {{< version include-if="1.0.x,1.1.x,1.2.x,1.3.x,1.4.x,1.5.x,2.2.x" >}}Without a health policy, no backend is evicted, so every retry returns to the same highest-priority group and the client still receives the error. To fail over transparently, configure both policies.{{< /version >}}{{< version exclude-if="1.0.x,1.1.x,1.2.x,1.3.x,1.4.x,1.5.x,2.2.x" >}}Without a health policy, default eviction covers 5xx responses and connection failures, so a retry policy on its own fails over on those errors. To retry on other responses, such as 429, add a health policy that classifies them as unhealthy. If that health policy has no `eviction` block, set `backoff` on the retry policy. Otherwise, the backend is not evicted, every retry returns to the same highest-priority group, and the client still receives the error.{{< /version >}}

## Cleanup

Expand Down
6 changes: 6 additions & 0 deletions assets/agw-docs/pages/agentgateway/llm/load-balancing.md
Original file line number Diff line number Diff line change
Expand Up @@ -467,6 +467,7 @@ For a complete guide on traffic splitting patterns, see [Traffic splitting]({{<

## Known limitations

{{< version include-if="1.0.x,1.1.x,1.2.x,1.3.x,1.4.x,1.5.x,2.2.x" >}}
> [!WARNING]
> **Rate-limit-based eviction only**: Provider eviction and failover currently only trigger on 429 (Too Many Requests) responses with proper rate-limit headers (`Retry-After` or `x-ratelimit-reset`). Eviction does NOT trigger on:
> - 503 Service Unavailable responses
Expand All @@ -475,6 +476,11 @@ For a complete guide on traffic splitting patterns, see [Traffic splitting]({{<
> - Other error codes (404, 500, etc.)
>
> Providers that return non-429 errors receive degraded health scores (EWMA) and lower priority within their group, but are not evicted or failed over. This means traffic may still be routed to consistently failing providers, though at reduced rates.
{{< /version >}}
{{< version exclude-if="1.0.x,1.1.x,1.2.x,1.3.x,1.4.x,1.5.x,2.2.x" >}}
> [!WARNING]
> **Eviction within a single priority group needs a health policy**: Default eviction applies only to an {{< reuse "agw-docs/snippets/backend.md" >}} with more than one priority group. When all providers share one group and no health policy targets the {{< reuse "agw-docs/snippets/backend.md" >}}, providers that return errors receive lower health scores (EWMA) and fewer requests, but are not evicted. Traffic can still reach a provider that fails consistently, at a reduced rate. To evict failing providers, add a health policy with an `eviction` block. For more information, see [Failover]({{< link-hextra path="/documentation/llm/failover/" >}}).
{{< /version >}}

## Monitoring load balancing

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -359,7 +359,7 @@ Use `virtualModel.conditional` to select a target with a CEL expression. Targets

Use `virtualModel.failover` to group targets by priority. Lower values are preferred. Targets in the same priority group are selected by a score that considers health and latency. The next group is used only when every target in the current group is degraded.

Failover depends on eviction. Configure `policies.health` on the concrete target models to define when a target is evicted. Without a health policy, targets are never evicted and failover does not occur.
Failover depends on eviction. A virtual model with more than one priority group uses default eviction: a single 5xx response or connection failure evicts a target for 3 seconds, and each repeated eviction lasts longer. To change when a target is evicted, configure `policies.health` on the concrete target models. A health policy replaces the default eviction instead of adding to it, so include an `eviction` block, as in the following example.

1. Create a model that points at an address with no backing workload, so that requests to it always fail. In a real deployment, the target would be a healthy primary provider.

Expand Down
19 changes: 9 additions & 10 deletions content/docs/standalone/main/documentation/llm/virtual-models.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,10 +30,10 @@ test:
# WHAT THIS TEST DOES NOT VALIDATE (and why):
# * That traffic is actually split 90/10 by `weight` - external dependency;
# observing the split needs many live completions against OpenAI.
# * That failover moves to a lower `priority` target after `health.eviction`
# removes the primary, and that same-priority targets are load balanced by
# health and latency - external dependency; triggering a real upstream
# failure needs live providers.
# * That failover moves to a lower `priority` target after eviction removes
# the primary, and that same-priority targets are load balanced by health
# and latency - external dependency; triggering a real upstream failure
# needs live providers.
# * That `when` expressions select a target by request header - requires
# config/traffic the page omits; the page shows no request example, and
# confirming which internal target served a response needs a live provider
Expand Down Expand Up @@ -156,7 +156,7 @@ assert_models config-weighted.yaml '["gpt-4o-public","smart"]'

### Failover routing

Use failover (also called automatic fallback) to keep serving when a primary model fails or becomes unavailable. Configure `routing.failover.targets` with `priority` on the virtual model, and configure `health.eviction` on the concrete target models so unhealthy backends can leave the active set.
Use failover (also called automatic fallback) to keep serving when a primary model fails or becomes unavailable. Configure `routing.failover.targets` with `priority` on the virtual model. When a virtual model has more than one priority group, agentgateway enables default eviction for target models that do not set their own `health` policy, so unhealthy backends can leave the active set.

Failover has two levels of grouping:

Expand All @@ -171,12 +171,11 @@ Configure health on the concrete `llm.models[]` entries that the virtual model t

| Setting | What it does |
| -- | -- |
| No `health` policy | Unhealthy responses (by default, `5xx` or connection failures) still lower the endpoint health score used for within-group load balancing. Endpoints are never evicted, so traffic never fails over to the next priority. |
| `health` without `eviction` | Same score-based weighting within a group. Eviction (and thus cross-priority failover) happens only when agentgateway can derive an eviction duration from elsewhere: `backoff` on a retry policy, or a `Retry-After` header on a 429 that is classified as unhealthy. |
| `health.eviction` | Removes an unhealthy endpoint from the active set for a backoff period. When every endpoint in a priority group is evicted, later requests use the next priority. |
| No `health` policy | The default unhealthy classifier covers `5xx` responses, non-zero gRPC statuses, and connection failures. When the virtual model has more than one priority group, these failures use default eviction: a single unhealthy response evicts the endpoint for `3s`, and each repeated eviction lasts longer. Traffic fails over to the next priority. |
| `health` without `eviction` | Setting `health` replaces the default eviction instead of adding to it. You can use `health.unhealthyExpression` to classify additional responses, such as `429`, as unhealthy. Without an `eviction` block, the endpoint is evicted only when a retry policy's `backoff`, or a `Retry-After` header on a response that is classified as unhealthy, supplies an eviction duration. To keep the default eviction settings, add `eviction: {}`. |
| `health.eviction` | Override how long an unhealthy endpoint leaves the active set, and which thresholds trigger eviction. When every endpoint in a priority group is evicted, later requests use the next priority. |

> [!WARNING]
> Setting `routing.failover` alone does **not** switch to a lower-priority target after errors. You must set `health.eviction` on the primary (and typically backup) concrete models. Without eviction, requests keep hitting the highest-priority group forever.
You do not need a `health` policy for basic failover on server errors or connection failures. To classify rate-limit responses, tune eviction timing, or change eviction thresholds, add a health policy that includes an `eviction` configuration.

Useful `health` fields:

Expand Down
Loading