Circuit Breakers and Resilience4j: Failing Fast on Purpose
The state machine, why a bulkhead matters as much as the breaker, fallbacks that degrade instead of erroring, and the order to stack the decorators in.
Retrying a Dead Service Makes It Worse
A timeout bounds one slow call. A retry recovers from a blip. Neither helps when a dependency is genuinely down — then every request waits for its timeout, every retry multiplies the load on a service already failing, and your own threads and connections fill with doomed work.
A circuit breaker stops calling a dependency that is failing. Requests fail immediately — in microseconds, not seconds — which protects your resources and gives the dependency room to recover.
The counter-intuitive part: failing fast is the feature. The breaker doesn't make failed requests succeed; it makes them cheap, and stops your service dying alongside its dependency.
Setup
dependencies {
implementation 'org.springframework.cloud:spring-cloud-starter-circuitbreaker-resilience4j'
}Resilience4j replaced Hystrix, which has been in maintenance since 2018. Spring Cloud CircuitBreaker is the abstraction; Resilience4j is the implementation you'd pick.
The State Machine
- CLOSED — normal. Outcomes are recorded in a sliding window.
- OPEN — the failure rate crossed the threshold. Calls fail immediately with
CallNotPermittedException. - HALF_OPEN — after a wait, a limited number of trial calls probe whether the dependency recovered.
resilience4j:
circuitbreaker:
instances:
payments:
sliding-window-type: COUNT_BASED
sliding-window-size: 20
minimum-number-of-calls: 10
failure-rate-threshold: 50 # percent
slow-call-duration-threshold: 2s
slow-call-rate-threshold: 50
wait-duration-in-open-state: 30s
permitted-number-of-calls-in-half-open-state: 3
automatic-transition-from-open-to-half-open-enabled: true
register-health-indicator: trueThree of those settings deserve attention.
minimum-number-of-calls: 10 prevents the breaker opening on tiny samples. Without it, the first two calls failing is a 100% failure rate and the circuit opens on a coincidence.
slow-call-duration-threshold is the under-used one. A dependency that responds in 8 seconds without erroring is arguably worse than one that fails — it holds your resources while technically succeeding. Counting slow calls as failures catches the degradation that a pure error-rate threshold misses.
wait-duration-in-open-state: 30s is the recovery-attempt interval. Too short and you hammer a recovering service; too long and you stay down after it's healthy.
Not every exception should trip the breaker. A 404 from a downstream service, or a 400 caused by your own bad request, says nothing about the dependency's health — counting them opens the circuit for everyone because one caller sent malformed input.
resilience4j:
circuitbreaker:
instances:
payments:
record-exceptions:
- java.io.IOException
- java.util.concurrent.TimeoutException
- org.springframework.web.client.HttpServerErrorException
ignore-exceptions:
- com.example.shop.PaymentDeclinedException
- org.springframework.web.client.HttpClientErrorException$NotFoundA declined payment is a successful call with a business outcome. Record infrastructure failures; ignore domain results.
Using It
@Service
public class PaymentService {
@CircuitBreaker(name = "payments", fallbackMethod = "paymentUnavailable")
@Retry(name = "payments")
@TimeLimiter(name = "payments")
public PaymentResult charge(Order order) {
return paymentClient.post()
.uri("/charges")
.body(ChargeRequest.from(order))
.retrieve()
.body(PaymentResult.class);
}
private PaymentResult paymentUnavailable(Order order, Throwable cause) {
log.warn("Payment unavailable for order {}: {}", order.getId(), cause.getMessage());
return PaymentResult.deferred(order.getId());
}
}The fallback method must have the same signature plus a Throwable parameter. Resilience4j can also pick a more specific fallback by exception type, which lets you distinguish "circuit open" from "timed out".
These are annotation-based, so they are proxies — which means the self-invocation and non-public limitations from the AOP guide apply for the fifth time in this roadmap. A @CircuitBreaker method called from within the same bean has no circuit breaker, silently. If you take one thing from the repeated appearances of this caveat: when an annotation that should be doing something appears to do nothing, check whether the call came through the proxy.
Fallbacks Worth Having
A fallback that rethrows is pointless. A useful one degrades:
// 1. Serve stale cached data
private StockLevel stockFallback(String sku, Throwable cause) {
return staleCache.get(sku).orElse(StockLevel.unknown(sku));
}
// 2. A safe default
private List<Recommendation> recommendationsFallback(Long userId, Throwable cause) {
return popularItems(); // not personalised, but not empty
}
// 3. Queue for later
private PaymentResult paymentFallback(Order order, Throwable cause) {
outbox.enqueue(new PendingCharge(order.getId()));
return PaymentResult.deferred(order.getId());
}
// 4. An honest error
private Order orderFallback(Long id, Throwable cause) {
throw new ServiceUnavailableException("Order service unavailable", cause);
}The fourth is legitimate and important: not everything can be degraded. A fallback that invents a plausible-looking value where correctness matters is worse than an error — serving a stale balance or a fabricated payment confirmation is a correctness bug dressed as resilience. Map it to a 503 with Retry-After, as the exception handling guide describes, and let the caller decide.
Choose by asking: is a wrong answer worse than no answer? Recommendations, yes degrade. Account balances, no.
Bulkheads
A circuit breaker reacts after a failure rate accumulates. A bulkhead prevents one dependency consuming all your capacity in the first place — named for ship compartments that stop one breach sinking the vessel.
resilience4j:
bulkhead:
instances:
payments:
max-concurrent-calls: 20
max-wait-duration: 100ms
thread-pool-bulkhead:
instances:
reports:
max-thread-pool-size: 8
core-thread-pool-size: 4
queue-capacity: 20@Bulkhead(name = "payments", type = Bulkhead.Type.SEMAPHORE)
public PaymentResult charge(Order order) { }Two types: semaphore (caps concurrent calls, caller's thread) and thread pool (separate pool, full isolation). Semaphore is lighter and usually right.
Bulkheads matter more now that virtual threads are available. The thread pool used to impose an accidental concurrency limit on every downstream call; with virtual threads that limit is gone and thousands of concurrent calls can hit one dependency. A bulkhead is how you reinstate the limit deliberately — the explicit replacement for a constraint you used to get by accident.
Combining Them
All four decorators on one method apply in a fixed order, outermost first:
Retry → CircuitBreaker → RateLimiter → TimeLimiter → Bulkhead → your method
This ordering is deliberate and worth reasoning through. Retry is outermost, so a retry gets a fresh circuit-breaker decision — and when the circuit is open, retries fail instantly rather than waiting. TimeLimiter is inside the breaker, so a timeout is recorded as a failure and can open it.
If you want different semantics — retry inside the breaker, so one logical call's retries count as a single outcome — compose the decorators programmatically rather than with annotations.
Count your retry layers across the whole stack. A gateway Retry=3, a load-balancer max-retries-on-next-service-instance: 2, and a Resilience4j @Retry(3) multiply: one logical request becomes dozens of attempts. Each configuration looks reasonable in isolation.
During a partial outage this is the retry storm — the dependency receives an order of magnitude more traffic precisely when it can least serve it, and your own resources fill with retries. Retry at exactly one layer, and let the circuit breaker handle sustained failure.
Observability
management:
endpoints:
web:
exposure:
include: health,info,prometheus,circuitbreakers
health:
circuitbreakers:
enabled: true/actuator/circuitbreakers shows each breaker's state, and Micrometer publishes:
| Metric | Use |
|---|---|
resilience4j_circuitbreaker_state | Alert on this — a breaker stuck OPEN |
resilience4j_circuitbreaker_calls | Tagged successful/failed/not_permitted |
resilience4j_bulkhead_available_concurrent_calls | Saturation |
resilience4j_retry_calls | Whether retries are helping or just amplifying |
A breaker that has been OPEN for ten minutes is an incident, and not_permitted call volume tells you how much traffic you're shedding.
Think carefully before including circuit-breaker state in your readiness probe. A breaker open on a non-essential dependency would mark the pod unready and remove it from the load balancer — turning a partial degradation into a total outage, which is the opposite of what the breaker is for. This is the same judgement as the dependency-in-readiness question from phase 6.
Check yourself
A circuit breaker on a payment service is configured with failure-rate-threshold: 50 and records all exceptions. Customers with expired cards get PaymentDeclinedException. During a promotion, many declines occur and the circuit opens — blocking payments for everyone, including valid cards. What went wrong?
A Starting Configuration
resilience4j:
circuitbreaker:
configs:
default:
sliding-window-size: 20
minimum-number-of-calls: 10
failure-rate-threshold: 50
slow-call-duration-threshold: 2s
slow-call-rate-threshold: 50
wait-duration-in-open-state: 30s
permitted-number-of-calls-in-half-open-state: 3
register-health-indicator: true
record-exceptions:
- java.io.IOException
- java.util.concurrent.TimeoutException
- org.springframework.web.client.HttpServerErrorException
instances:
payments:
base-config: default
inventory:
base-config: default
failure-rate-threshold: 70 # more tolerant; degradation is acceptable
timelimiter:
configs:
default:
timeout-duration: 3s
bulkhead:
configs:
default:
max-concurrent-calls: 25Use configs.default with base-config so each dependency inherits sensible values and overrides only what differs — and tune per dependency, since a payment provider and a recommendations service warrant different tolerances.
The Mental Model, Restated
- Failing fast is the feature. A breaker protects you and gives the dependency room to recover.
- CLOSED → OPEN → HALF_OPEN, driven by a sliding window with a minimum call count.
- Count slow calls as failures. A slow dependency can be worse than a failing one.
- Record infrastructure failures; ignore business outcomes. A decline is not an outage.
- These are proxies — self-invocation bypasses them.
- Fallbacks should degrade, or fail honestly. A wrong answer can be worse than none.
- Bulkheads prevent saturation, and matter more with virtual threads.
- Retry at one layer only. Stacked retries become a retry storm.
- Alert on breaker state, and think twice before putting it in readiness.
What's Next
Each service now has its own configuration, duplicated across environments and deployments. The next guide covers Spring Cloud Config Server — Git-backed centralised configuration, encrypted secrets, and refresh without restart — along with an honest comparison against Kubernetes ConfigMaps and a dedicated secrets manager.