07-spring-cloud

Spring Cloud Gateway: Routing, Filters and Rate Limiting

Predicates and filters, what belongs at the edge versus in every service, distributed rate limiting with Redis, and why the gateway is reactive.

October 9, 2026
spring-cloudapi-gatewaygatewayroutingrate-limitingrediswebfluxmicroservices

One Door

Without a gateway, every client needs to know every service's address, and every service needs to implement authentication, rate limiting, CORS and request logging. Both halves of that are a problem: clients become coupled to your internal topology, and cross-cutting concerns are reimplemented N times with N subtly different behaviours.

An API gateway is a single entry point that routes to internal services and handles the concerns that belong at the edge.

Setup

groovy
dependencies {
    implementation 'org.springframework.cloud:spring-cloud-starter-gateway'
}
yaml
spring:
  cloud:
    gateway:
      routes:
        - id: orders
          uri: lb://orders                      # lb:// = via load balancer + discovery
          predicates:
            - Path=/api/orders/**
          filters:
            - RewritePath=/api/orders/(?<seg>.*), /$\{seg}
        - id: inventory
          uri: lb://inventory
          predicates:
            - Path=/api/inventory/**
            - Method=GET,POST
          filters:
            - RewritePath=/api/inventory/(?<seg>.*), /$\{seg}
            - AddRequestHeader=X-Gateway, shop-gateway
⚠️

Spring Cloud Gateway is built on WebFlux and Netty, not the servlet stack. Adding spring-boot-starter-web to a gateway application breaks it — Spring Boot configures a servlet web application and the gateway's reactive routing never engages. The gateway must be its own deployable with no servlet starter.

There is now a servlet-based variant (spring-cloud-starter-gateway-mvc) if you need the blocking stack, with a reduced filter set. For a pure proxy, reactive is the better fit: a gateway spends nearly all its time waiting on I/O with no business logic, which is precisely what the reactive model is good at.

Predicates: Which Route Matches

yaml
predicates:
  - Path=/api/orders/**
  - Method=GET,POST
  - Host=api.example.com
  - Header=X-Tenant, \d+               # name, regex
  - Query=version, 2
  - After=2026-01-01T00:00:00Z[UTC]
  - Weight=orders-group, 90            # for canary traffic splitting

Multiple predicates on one route are ANDed. Routes are evaluated in order, first match wins — the same ordering discipline as Spring Security's rule chain, and the same failure mode when a broad route shadows a specific one.

Weight is how you do a canary release at the edge:

yaml
- id: orders-v1
  uri: lb://orders
  predicates: [ Path=/api/orders/**, "Weight=orders, 90" ]
- id: orders-v2
  uri: lb://orders-v2
  predicates: [ Path=/api/orders/**, "Weight=orders, 10" ]

Filters: What Happens to the Request

yaml
filters:
  - RewritePath=/api/orders/(?<seg>.*), /$\{seg}
  - AddRequestHeader=X-Source, gateway
  - RemoveRequestHeader=X-Internal-Secret
  - AddResponseHeader=X-Served-By, gateway
  - CircuitBreaker=ordersCircuitBreaker
  - Retry=3
  - RequestRateLimiter
  - SetStatus=404

Filters can be per-route or global:

yaml
spring:
  cloud:
    gateway:
      default-filters:
        - RemoveRequestHeader=Cookie          # strip browser cookies from internal calls
        - AddRequestHeader=X-Gateway, shop

RemoveRequestHeader=Cookie as a default filter is a useful pattern: it prevents a browser cookie reaching internal services that might otherwise treat it as authentication.

A global filter in code

java
@Component
public class CorrelationIdFilter implements GlobalFilter, Ordered {
 
    private static final String HEADER = "X-Correlation-Id";
 
    @Override
    public Mono<Void> filter(ServerWebExchange exchange, GatewayFilterChain chain) {
        String correlationId = Optional
                .ofNullable(exchange.getRequest().getHeaders().getFirst(HEADER))
                .orElseGet(() -> UUID.randomUUID().toString());
 
        ServerHttpRequest mutated = exchange.getRequest().mutate()
                .header(HEADER, correlationId)
                .build();
 
        exchange.getResponse().getHeaders().add(HEADER, correlationId);
        return chain.filter(exchange.mutate().request(mutated).build());
    }
 
    @Override
    public int getOrder() {
        return Ordered.HIGHEST_PRECEDENCE;
    }
}

Every request gets a correlation ID, propagated inward and returned to the client. Note this complements rather than replaces trace context — Micrometer Tracing already propagates traceparent, and as the observability guide argued, the trace ID is the better join key. A correlation ID is useful when you need an identifier a client can quote without depending on your tracing setup.

🚨

A gateway filter must never block. filter returns a Mono and runs on a Netty event-loop thread; a JDBC call, a RestClient without reactive handling, or Thread.sleep blocks the event loop and stalls every concurrent request on that thread, not just this one. This is the single most damaging mistake in gateway code, and it looks like the gateway randomly hanging under load.

If a filter needs to call a database or another service, use WebClient and compose reactively, or Mono.fromCallable(...).subscribeOn(Schedulers.boundedElastic()) to move the blocking work off the event loop.

Rate Limiting

groovy
dependencies {
    implementation 'org.springframework.boot:spring-boot-starter-data-redis-reactive'
}
yaml
filters:
  - name: RequestRateLimiter
    args:
      redis-rate-limiter.replenishRate: 100     # sustained requests/sec
      redis-rate-limiter.burstCapacity: 200     # bucket size
      redis-rate-limiter.requestedTokens: 1
      key-resolver: "#{@userKeyResolver}"
java
@Bean
KeyResolver userKeyResolver() {
    return exchange -> exchange.getPrincipal()
            .map(Principal::getName)
            .switchIfEmpty(Mono.just(
                    Optional.ofNullable(exchange.getRequest().getRemoteAddress())
                            .map(a -> a.getAddress().getHostAddress())
                            .orElse("unknown")));
}

A token-bucket limiter backed by Redis — so the limit is shared across gateway instances. An in-memory limiter with four gateway replicas gives each client four times the intended quota, which is the kind of bug nobody notices until the limit matters.

replenishRate is the sustained rate and burstCapacity the bucket size, so a client can burst to 200 then settles at 100/sec. Burst capacity above replenish rate is what makes limits tolerable for bursty but legitimate clients.

⚠️

Rate limiting by IP is unreliable. Clients behind corporate NAT or a mobile carrier share an IP, so one limit covers thousands of users; conversely an attacker rotating IPs evades it. Prefer an authenticated identity — user ID, API key, tenant — and fall back to IP only for unauthenticated endpoints, where it's a blunt instrument rather than a precise one.

Also check that X-Forwarded-For is trustworthy: if your gateway sits behind another proxy, the remote address is that proxy's, and a client-supplied X-Forwarded-For header can be spoofed unless the trusted hop count is configured.

What Belongs at the Edge

The useful discipline is deciding this explicitly:

ConcernWhereWhy
TLS terminationGatewayOne certificate to manage
Authentication (token validation)GatewayReject invalid tokens once, early
AuthorisationServiceNeeds domain context the gateway lacks
Rate limitingGatewayMust be global across the system
CORSGatewayOne policy for all clients
Request logging / correlationGatewayUniform coverage
Routing, canary splittingGatewayClients stay decoupled from topology
Business validationServiceDomain logic
Response aggregationNeither — see below
🚨

The gateway validating a token does not mean services can skip authorisation. Two reasons. The gateway establishes who the caller is; only the service knows whether this caller may touch this order — the object-level check from the authorisation guide. And anything that can reach a service directly — another service, a misconfigured network policy, a port-forward — bypasses the gateway entirely.

Treat the gateway as defence in depth, not a perimeter. Services authenticate and authorise independently. A system whose internal services trust any request they receive is one network misconfiguration away from total compromise.

On response aggregation: a gateway that calls three services and merges results is no longer a gateway — it's a service with business logic, deployed as shared infrastructure, where a bug affects all traffic. If you need aggregation, build a backend-for-frontend service that the gateway routes to, or use GraphQL. Keep the gateway a proxy.

Resilience at the Gateway

yaml
filters:
  - name: CircuitBreaker
    args:
      name: ordersCircuitBreaker
      fallbackUri: forward:/fallback/orders
  - name: Retry
    args:
      retries: 3
      methods: GET                       # idempotent only
      backoff:
        firstBackoff: 50ms
        maxBackoff: 500ms
        factor: 2
java
@RestController
public class FallbackController {
 
    @GetMapping("/fallback/orders")
    ResponseEntity<ProblemDetail> ordersUnavailable() {
        ProblemDetail problem = ProblemDetail.forStatusAndDetail(
                HttpStatus.SERVICE_UNAVAILABLE, "Orders are temporarily unavailable");
        problem.setTitle("Service unavailable");
        return ResponseEntity.status(503)
                .header(HttpHeaders.RETRY_AFTER, "30")
                .body(problem);
    }
}

methods: GET is essential — retrying a POST at the gateway can duplicate a payment, for exactly the reasons the outbound HTTP guide set out. The gateway cannot know whether your POST is idempotent, so restrict it to methods that are by definition.

Check yourself

A gateway runs four replicas with a RequestRateLimiter configured at 100 req/s per user, using an in-memory limiter rather than Redis. What is the effective limit?

Alternatives

Spring Cloud Gateway is not the only option, and sometimes not the best:

  • Kubernetes Ingress / Gateway API — routing and TLS at the platform level, no application to operate. Often enough on its own.
  • Envoy, NGINX, Traefik — purpose-built proxies, faster and more battle-tested for pure routing.
  • A managed API gateway — AWS API Gateway, Kong, Apigee. Routing plus developer portal, key management, quota billing.

Spring Cloud Gateway's advantage is that filters are Java you can write and test — valuable when your edge logic is genuinely bespoke (a legacy authentication scheme, tenant-specific routing). If your needs are path routing, TLS and rate limiting, a platform ingress is less to run.

The Mental Model, Restated

  1. One entry point so clients don't know your topology and cross-cutting concerns live in one place.
  2. Gateway is reactive — never add starter-web, never block in a filter.
  3. Predicates AND together; routes match in order. First match wins.
  4. Rate limiting must be Redis-backed to be shared across replicas, and keyed on identity rather than IP.
  5. Authenticate at the edge, authorise in the service. Defence in depth, not a perimeter.
  6. No response aggregation in the gateway. Use a BFF.
  7. Retry only idempotent methods at the edge.
  8. A platform ingress may be sufficient — write a gateway when the filter logic is genuinely yours.

What's Next

Both the previous guides used lb:// and @LoadBalanced without explaining them. The next guide covers client-side load balancing: how an instance is chosen, how that differs from a load balancer in front of the service, and the custom strategies worth knowing.