06-production-cloud-native

Graceful Shutdown and Probes: Deployments Users Don't Notice

Why SIGTERM drops in-flight requests by default, the race between deregistration and termination, and probe wiring that makes a rollout invisible.

October 9, 2026
spring-bootgraceful-shutdownkuberneteslivenessreadinessprobeszero-downtimedevops

What Happens When a Pod Is Deleted

Kubernetes sends SIGTERM. The JVM's default response is to run shutdown hooks and exit — immediately, while requests are in flight. Those requests get a connection reset. Mid-transaction work is abandoned. A consumer that had taken a message hasn't acknowledged it.

A rolling deployment of four replicas does this four times. If your deploys produce a small error spike, this is usually why.

Two separate problems hide here, and both need fixing:

  1. In-flight requests are killed when the process exits.
  2. New requests keep arriving for a short period after termination begins, because load balancers take time to notice.

Graceful Shutdown

Spring Boot handles the first problem with two properties:

yaml
server:
  shutdown: graceful
spring:
  lifecycle:
    timeout-per-shutdown-phase: 30s

With this, on SIGTERM the web server stops accepting new connections and lets in-flight requests complete, up to the timeout. Only then does the context close.

Set the timeout above your slowest realistic request — p99.9, not the average. Too short and you're back to dropping requests; too long and deployments crawl. 30 seconds suits most APIs.

🚨

terminationGracePeriodSeconds must exceed your shutdown timeout. Kubernetes sends SIGTERM, waits that period, then sends SIGKILL — which is unblockable. With a 30-second Spring timeout and Kubernetes' 30-second default, SIGKILL can arrive exactly as Spring is finishing, defeating the entire configuration.

yaml
spec:
  terminationGracePeriodSeconds: 45     # > Spring's 30s, with headroom

Graceful shutdown covers more than HTTP. The context close also stops @Scheduled tasks, drains @Async executors (if you configured waitForTasksToCompleteOnShutdown, as the scheduling guide advised), closes the Hikari pool, and stops message listeners so they finish the message in hand before releasing it.

The Deregistration Race

The subtler problem. When a pod is deleted, two things happen concurrently:

  1. SIGTERM is sent to the container.
  2. The pod's removal propagates to Service endpoints, kube-proxy on every node, and any ingress controller.

The second takes time — typically a few seconds. So a pod that has already begun shutting down keeps receiving new requests, which it now refuses because graceful shutdown stopped accepting connections. The client sees a connection refused.

The fix is counter-intuitive: delay the start of shutdown so deregistration completes first.

yaml
lifecycle:
  preStop:
    exec:
      command: ["sh", "-c", "sleep 5"]

The preStop hook runs before SIGTERM. During those 5 seconds the pod still serves traffic normally while its removal propagates. Only then does Spring begin its graceful shutdown.

⚠️

A distroless image has no shell, so sh -c sleep 5 fails. Use httpGet against an endpoint, or a sleep binary if the image has one. This is one of the concrete costs of distroless mentioned in the previous guide — worth knowing before you discover it during a deployment.

Spring Boot can also take part explicitly, by marking itself out of service before shutting down:

yaml
management:
  endpoint:
    health:
      probes:
        enabled: true
  health:
    livenessstate:
      enabled: true
    readinessstate:
      enabled: true

The readinessState becomes REFUSING_TRAFFIC when shutdown begins, so a readiness probe fails promptly — which is the signal a load balancer actually watches.

Probes

Phase 4 established the distinction; here it is wired up.

Liveness — is the process broken beyond recovery? Failing means restart. Readiness — can it serve traffic? Failing means stop routing, no restart. Startup — has it finished starting? Suppresses the other two until it passes.

yaml
management:
  endpoint:
    health:
      probes:
        enabled: true
      group:
        liveness:
          include: livenessState
        readiness:
          include: readinessState,db
yaml
startupProbe:
  httpGet: { path: /actuator/health/liveness, port: 9001 }
  failureThreshold: 30
  periodSeconds: 5              # allows up to 150s to start
 
livenessProbe:
  httpGet: { path: /actuator/health/liveness, port: 9001 }
  periodSeconds: 10
  failureThreshold: 3
  timeoutSeconds: 3
 
readinessProbe:
  httpGet: { path: /actuator/health/readiness, port: 9001 }
  periodSeconds: 5
  failureThreshold: 2
  timeoutSeconds: 3

Three points carry the weight here.

A startupProbe replaces a generous initialDelaySeconds on liveness. A JVM can take 20–60 seconds to start; without a startup probe, liveness must be lenient enough to allow that, which makes it slow to detect a genuinely hung process later. The startup probe lets liveness be strict once running.

Liveness includes only livenessState — no database, no cache, no downstream service. A dependency check here causes the fleet-wide crash loop from phase 4.

Readiness is checked more often with a lower threshold, because removing a pod from rotation is cheap and reversible while restarting it is not.

🚨

Probing the management port (9001) rather than the application port is deliberate: it keeps Actuator off your public ingress, per the hardening guidance in phase 4. But make sure the management port is reachable from the kubelet — a NetworkPolicy that only allows ingress on 8080 silently fails every probe, and the symptom is pods that never become ready with no application error.

Custom indicators belong in readiness only

java
@Component
public class InventoryHealthIndicator implements HealthIndicator {
 
    @Override
    public Health health() {
        try {
            return Health.up().withDetail("latencyMs", client.ping().toMillis()).build();
        } catch (Exception ex) {
            return Health.down(ex).build();
        }
    }
}
yaml
management:
  endpoint:
    health:
      group:
        readiness:
          include: readinessState,db,inventory

Then ask the honest question: if the inventory service is down, should this pod stop serving entirely? If some endpoints still work, excluding it keeps them available. Marking yourself unready because a dependency is unhealthy propagates the outage outward — the opposite of what a circuit breaker does.

Health checks also run on every probe, so a slow indicator makes probes time out and your pod fail for the wrong reason. Keep them fast, and consider caching:

yaml
management:
  health:
    inventory:
      cache:
        time-to-live: 10s

Draining Non-HTTP Work

A service that consumes messages or runs jobs needs shutdown handling beyond HTTP.

Message listeners stop in the context close. A Kafka or RabbitMQ listener finishes the message it holds and doesn't take another — provided the shutdown timeout is long enough to finish. If not, the message is redelivered, which is fine because your consumers are idempotent, as the messaging guide required.

Long-running scheduled jobs are the awkward case. A 10-minute archive job cannot finish inside a 30-second window, so it is interrupted mid-run. Make such jobs resumable and chunked — commit progress as you go, so the next run continues rather than restarting — or move them out of the request-serving deployment entirely, which is also the answer to the multiple-instance problem from phase 4.

@PreDestroy for anything else:

java
@Component
public class OutboxPublisher {
 
    @PreDestroy
    void drain() {
        log.info("Draining outbox before shutdown");
        publishPending(Duration.ofSeconds(10));
    }
}

Keep it bounded — a @PreDestroy that blocks past the grace period is killed by SIGKILL anyway.

Check yourself

A service has graceful shutdown enabled with a 30s timeout and terminationGracePeriodSeconds: 45. Deployments still produce a burst of connection-refused errors at the start of each pod's termination. What is missing?

A Working Configuration

yaml
# application.yml
server:
  shutdown: graceful
  port: 8080
spring:
  lifecycle:
    timeout-per-shutdown-phase: 30s
management:
  server:
    port: 9001
  endpoint:
    health:
      probes:
        enabled: true
      show-details: when-authorized
      group:
        liveness:
          include: livenessState
        readiness:
          include: readinessState,db
yaml
# deployment.yaml
spec:
  template:
    spec:
      terminationGracePeriodSeconds: 45
      containers:
        - name: shop
          ports:
            - { name: http, containerPort: 8080 }
            - { name: management, containerPort: 9001 }
          lifecycle:
            preStop:
              exec: { command: ["sh", "-c", "sleep 5"] }
          startupProbe:
            httpGet: { path: /actuator/health/liveness, port: management }
            failureThreshold: 30
            periodSeconds: 5
          livenessProbe:
            httpGet: { path: /actuator/health/liveness, port: management }
            periodSeconds: 10
            failureThreshold: 3
          readinessProbe:
            httpGet: { path: /actuator/health/readiness, port: management }
            periodSeconds: 5
            failureThreshold: 2

Verify it rather than trusting it: run a load generator against the service and roll the deployment. A correct configuration produces zero failed requests. That test is worth automating — this configuration has several interacting parts, and a regression in any one of them is invisible until a deployment under real traffic.

The Mental Model, Restated

  1. Default SIGTERM handling drops in-flight requests. server.shutdown: graceful fixes it.
  2. terminationGracePeriodSeconds must exceed the Spring timeout, or SIGKILL cuts it short.
  3. Deregistration races with SIGTERM — a preStop sleep is what closes that window.
  4. A startupProbe lets liveness stay strict without tolerating slow JVM startup.
  5. Liveness depends on nothing external. Readiness may, deliberately.
  6. Health indicators run on every probe — keep them fast, cache if needed.
  7. Chunk and resume long jobs; they cannot finish in a grace period.
  8. Test it with a rolling deploy under load. Zero errors or it isn't working.

What's Next

Startup time has come up twice now — as the reason for a startup probe, and as a cost paid on every deployment and scale-up. The next guide covers the option that removes it almost entirely: compiling to a GraalVM native image, what the closed-world assumption demands in return, and whether the trade is worth making.