SIGN IN SIGN UP

composite: fail over to the next sub-cluster when a sub-cluster has no hosts (#46308)

## Commit Message

composite: fail over to the next sub-cluster when a sub-cluster has no
hosts

The composite cluster failed a request with 503 no_healthy_upstream when
the sub-cluster mapped
to the current attempt had no host available, for example because DNS
resolution returned an empty
endpoint list or because all of its hosts had been ejected. Host
selection now advances to the
following sub-clusters in the configured list within the same attempt,
and only fails when none of
them can provide a host.

Guarded by
`envoy.reloadable_features.composite_cluster_skip_clusters_without_hosts`.

## Additional Description

### Documented behavior vs. actual behavior

The `composite_cluster`
[docs](https://www.envoyproxy.io/docs/envoy/latest/intro/arch_overview/upstream/composite_cluster)
stated:

> *"If a selected sub-cluster has no healthy hosts available, the
request will fail according to that sub-cluster's load balancing
behavior, **potentially triggering another retry attempt if
configured**."*

In practice that retry path was unreachable: `no_healthy_upstream` is a
routing-level rejection
raised before any upstream connection attempt, and no `retry_on`
condition covers it. A composite
cluster whose first sub-cluster had zero endpoints failed with `503
no_healthy_upstream` regardless
of the retry policy, so per-request failover did not work at all in that
scenario.

### Change from the previous revision of this PR

The first revision of this PR added a core `no-healthy-upstream` retry
condition
(`RETRY_ON_NO_HEALTHY_UPSTREAM`) plus a hook in
`Filter::createConnPoolOrHandleFailure()`.
Per review feedback from @wbpcode and @agrawroh, that a core feature
serving one extension is the
wrong layer, and that the composite cluster should pick a sub-cluster
that has healthy hosts itself. The previous approach has been fully
reverted and the fix now lives entirely in the extension.

Reverted back to `main`: `envoy/router/router.h`,
`source/common/http/headers.h`,
`source/common/router/retry_state_impl.{h,cc}`,
`source/common/router/router.cc`,
`test/common/router/{router_test,retry_state_impl_test}.cc`,
`test/mocks/router/mocks.{h,cc}`,
`api/envoy/config/route/v3/route_components.proto`,
`docs/root/configuration/http/http_filters/router_filter.rst`.

### Fix

`CompositeClusterLoadBalancer::chooseHost()` still maps the attempt
count to a starting cluster
index, but host selection now runs through a new
`selectHostWithFailover()` helper that walks
forward from that index:

* A sub-cluster is skipped when it is not present in the cluster
manager, or when its load
  balancer returns no host.
* Selection returns immediately when a host is selected, or when the
sub-cluster started
asynchronous host selection (`HostSelectionResponse::cancelable` set),
so an in-flight async
  selection is not abandoned or duplicated.
* When no sub-cluster can provide a host, the `details` and
`failure_status` of the last sub-cluster
tried are propagated, so the resulting local reply keeps the same
diagnostics as before.

Availability is determined by asking the sub-cluster's load balancer
rather than by inspecting its
priority set, so nested composite/aggregate sub-clusters,
`CLUSTER_PROVIDED` sub-clusters, panic
mode and override-host behavior all keep working. When a sub-cluster can
serve the attempt, its
load balancer is the only one invoked, so there are no wasted host
selections on the happy path.

### Behavior notes

* Selection remains driven by the attempt count: an attempt that failed
over to a later
sub-cluster does not shift the mapping of subsequent attempts. With
`[primary, secondary,
fallback]` and an empty `primary`, attempt 1 uses `secondary` and
attempt 2 also maps to
`secondary`. This is documented in the arch overview. If reviewers
prefer that a failover also
advance the following attempts, that can be done by recording the used
index in filter state —
  happy to add it here or as a follow-up.
* `peekAnotherHost()` and `selectExistingConnection()` keep the plain
attempt mapping. Every
in-tree `selectExistingConnection()` implementation returns `nullopt`,
so there is no practical
  mismatch, and prefetching does not need the failover.

### Minimal reproduction

A minimal Docker Compose reproduction of the original bug is available
at:

https://github.com/JonSchaeffer/playground/tree/main/envoy/composite-bug-minimal-repro

## Risk Level

Low — the change is scoped to the composite cluster extension, and it is
behind a reloadable
runtime guard.

## Testing

* Unit tests in `test/extensions/clusters/composite/cluster_test.cc`:
failover when the selected
sub-cluster has no host, failover when the selected sub-cluster is
missing from the cluster
manager, no host plus preserved failure details when no sub-cluster can
serve the request,
asynchronous host selection returned without trying other sub-clusters,
and legacy behavior with
  the runtime guard disabled.
* Integration tests in
`test/extensions/clusters/composite/cluster_integration_test.cc`:
failover
with `num_retries: 0` (proving no retry is involved), failover across
two consecutive
sub-clusters without endpoints, `503` when no sub-cluster has endpoints,
and `503` with the
runtime guard disabled. The two `no-healthy-upstream` retry tests from
the previous revision were
  removed.
* Format and spelling checks: `bazel run
//tools/code_format:check_format` and
  `bazel run //tools/spelling:check_spelling_pedantic` both pass.

## Docs Changes

* `docs/root/intro/arch_overview/upstream/composite_cluster.rst`: new
"Hosts availability" section,
updated the sub-cluster health / deterministic routing considerations,
and documented the runtime
  guard and the attempt-progression caveat.
* `api/envoy/extensions/clusters/composite/v3/cluster.proto`:
comment-only update describing the
  failover. No field or type changes.

## Release Notes


`changelogs/current/bug_fixes/composite_cluster__skip-clusters-without-hosts.rst`.

## Runtime guard


`envoy.reloadable_features.composite_cluster_skip_clusters_without_hosts`
(enabled by default).
Setting it to `false` restores the previous behavior of failing the
attempt when the sub-cluster
mapped to that attempt has no host available.

---

> **Note:** This PR was developed with AI assistance (Claude). The
author has reviewed and
> understands all changes.

---------

Signed-off-by: JonSchaeffer <jon.schaeffer@fastly.com>
J
Jonathan Schaeffer committed
7ba51266cc60cd560c37ea9cb3e2adf70676b6cf
Parent: 329dd51
Committed by GitHub <noreply@github.com> on 9/5/2026, 9:03:53 PM