07-spring-cloud

Client-Side Load Balancing with Spring Cloud LoadBalancer

What @LoadBalanced actually does, how client-side differs from server-side balancing, and writing a zone-aware or weighted strategy.

October 9, 2026
spring-cloudload-balancingLoadBalancedribbonservice-discoveryresiliencemicroservices

Two Places to Decide

Discovery returns a list of instances. Something must pick one.

Server-side load balancing puts a dedicated component in the path — NGINX, a cloud load balancer, a Kubernetes Service. The client calls one address and the balancer forwards.

Client-side load balancing gives the caller the instance list and lets it choose. No extra hop.

Server-sideClient-side
Extra network hopYesNo
Single point of failureThe balancerNone
Client complexityNoneLibrary + registry awareness
Routing logicCentralisedPer client
Language-agnosticYesNeeds a library per language
Health awarenessThe balancer probesFrom the registry

Neither is better in general. Client-side avoids a hop and a shared failure point and allows caller-specific routing; server-side keeps clients simple and works for any language.

Spring Cloud LoadBalancer

Ribbon is gone — Spring Cloud LoadBalancer replaced it, and spring-cloud-starter-loadbalancer comes in transitively with the discovery starters.

java
@Configuration
public class ClientConfig {
 
    @Bean
    @LoadBalanced
    RestClient.Builder loadBalancedRestClientBuilder() {
        return RestClient.builder();
    }
}
java
@Service
public class InventoryService {
 
    private final RestClient client;
 
    public InventoryService(@LoadBalanced RestClient.Builder builder) {
        this.client = builder.baseUrl("http://inventory").build();
    }
 
    public StockLevel stockFor(String sku) {
        return client.get().uri("/stock/{sku}", sku).retrieve().body(StockLevel.class);
    }
}

What @LoadBalanced does, concretely: it adds an interceptor that treats the host portion of the URL as a service name, asks the DiscoveryClient for instances, chooses one, and rewrites the URL to that instance's address. http://inventory is never resolved by DNS.

⚠️

You need two builder beans if some calls are load-balanced and others are to external URLs. An unannotated RestClient.Builder injected alongside an annotated one is ambiguous, and @LoadBalanced on a client calling https://api.stripe.com fails — it tries to resolve api.stripe.com as a registered service name. This is the @Qualifier disambiguation from phase 1, and the error message (No instances available for api.stripe.com) is clear once you know what it means.

Strategies

The default is round robin. Random is built in:

java
@Configuration
public class InventoryLoadBalancerConfig {
 
    @Bean
    ReactorLoadBalancer<ServiceInstance> randomLoadBalancer(
            Environment environment,
            LoadBalancerClientFactory factory) {
        String name = environment.getProperty(LoadBalancerClientFactory.PROPERTY_NAME);
        return new RandomLoadBalancer(
                factory.getLazyProvider(name, ServiceInstanceListSupplier.class), name);
    }
}
java
@LoadBalancerClient(value = "inventory", configuration = InventoryLoadBalancerConfig.class)
@Configuration
public class ClientConfig { }

Round robin is fine when instances are homogeneous and requests cost roughly the same. It is not load-aware: an instance that is slow because it's garbage collecting, or on a noisy neighbour, receives exactly the same share as a healthy one.

✅

Round robin distributes requests, not load. If one instance is degraded, round robin keeps feeding it — so a slow instance becomes a latency spike in your p99 rather than being routed around. The fix is not a cleverer balancer so much as the resilience measures in the next guide: a timeout bounds the damage, and a circuit breaker per instance stops sending to one that is failing.

Zone Awareness

In a multi-zone deployment, cross-zone traffic costs latency and often money. Prefer same-zone instances:

yaml
spring:
  cloud:
    loadbalancer:
      configurations: zone-preference
eureka:
  instance:
    metadata-map:
      zone: eu-west-1a

Instances in the caller's zone are preferred, with other zones used when none are available locally — which is the behaviour you want: cheaper and faster normally, still available if a zone fails.

A custom supplier

For weighted routing — a canary, or instances of differing capacity:

java
public class WeightedInstanceSupplier implements ServiceInstanceListSupplier {
 
    private final ServiceInstanceListSupplier delegate;
 
    @Override
    public Flux<List<ServiceInstance>> get() {
        return delegate.get().map(this::expandByWeight);
    }
 
    private List<ServiceInstance> expandByWeight(List<ServiceInstance> instances) {
        List<ServiceInstance> weighted = new ArrayList<>();
        for (ServiceInstance instance : instances) {
            int weight = Integer.parseInt(
                    instance.getMetadata().getOrDefault("weight", "1"));
            for (int i = 0; i < weight; i++) {
                weighted.add(instance);           // repeat → higher selection odds
            }
        }
        return weighted;
    }
 
    @Override
    public String getServiceId() {
        return delegate.getServiceId();
    }
}

Repeating an instance in the list makes round robin pick it proportionally more often — a simple trick that avoids writing a balancer.

Health Checks and Retries

By default, the balancer trusts the registry. Given Eureka's slow eviction, that means sending traffic to dead instances for up to two minutes. Add client-side health checking:

yaml
spring:
  cloud:
    loadbalancer:
      health-check:
        initial-delay: 0s
        interval: 10s
        path:
          default: /actuator/health/readiness
      retry:
        enabled: true
        max-retries-on-same-service-instance: 0
        max-retries-on-next-service-instance: 2
        retryable-status-codes: 502,503,504

Two settings here matter most.

max-retries-on-next-service-instance: 2 retries against a different instance. This is the most valuable resilience setting in client-side load balancing: a single dead or restarting instance becomes invisible to callers, because the retry lands elsewhere.

max-retries-on-same-service-instance: 0 — don't retry the instance that just failed. If it's down, it's down.

🚨

Load-balancer retries are only safe for idempotent requests. The default retries only GET, which is correct. Enabling retries for POST here duplicates non-idempotent operations — the same hazard as gateway retries and client retries, and now in a third place.

Count your retry layers. Gateway retry (3) × load-balancer retry (3) × a Resilience4j @Retry (3) is 27 attempts for one logical request. Each layer looks reasonable alone and together they amplify a partial outage into a self-inflicted denial of service. Retry at one layer. Decide which, and disable the others.

That multiplication problem is worth a moment, because it's a genuinely common production incident. A downstream service starts failing, every caller's retry stack multiplies out, and the downstream receives an order of magnitude more traffic exactly when it's least able to serve it — the retry storm. One retry layer plus a circuit breaker is the configuration that degrades rather than amplifies.

Check yourself

A service calls an external payment API using a @LoadBalanced RestClient.Builder with baseUrl https://api.payments.example.com. Every call fails with 'No instances available for api.payments.example.com'. Why?

When You Don't Need It

On Kubernetes, a Service already load-balances across healthy pods, with readiness probes controlling membership and kube-proxy doing the distribution. Calling http://inventory as a Service name with a plain RestClient gets you load balancing, health awareness and no library.

Client-side balancing earns its place when you need caller-specific routing logic — zone preference, weighting, a bespoke strategy — or when you're not on a platform that provides it. A service mesh (Istio, Linkerd) is the third option: client-side behaviour with language-agnostic sidecars, at the cost of operating the mesh.

The Mental Model, Restated

  1. @LoadBalanced treats the URL host as a service name, resolved via discovery.
  2. Keep separate builders for internal service names and external URLs.
  3. Client-side avoids a hop and a shared failure point; server-side keeps clients simple.
  4. Round robin distributes requests, not load — a degraded instance still gets its share.
  5. Zone preference saves latency and cross-zone cost, with fallback when a zone fails.
  6. max-retries-on-next-service-instance is the high-value setting — it hides a single dead instance.
  7. Retry at one layer only. Stacked retries multiply into a retry storm.
  8. Kubernetes Services may already be enough.

What Next

Retries, timeouts and routing around a failed instance handle transient failure. They don't help when a dependency is properly down — they make it worse, by continuing to call it. The next guide covers circuit breakers and the rest of the Resilience4j toolkit: the state machine, bulkheads, rate limiters, fallbacks, and the order to combine them in.