Production Hardening: A Go-Live Review
Locking down Actuator, sizing the JVM inside a container, the 12-factor checks that matter, and a consolidated checklist across all six phases.
What This Guide Is
Not new concepts — a review. Six phases have each introduced decisions that matter in production, and this guide consolidates them into something you can run through before a go-live, plus the few hardening items that haven't come up yet.
Lock Down Actuator
The most commonly misconfigured thing in a Spring Boot deployment.
management:
server:
port: 9001 # not on the public ingress
address: 0.0.0.0 # but reachable by the kubelet
endpoints:
web:
exposure:
include: health,info,prometheus
endpoint:
health:
show-details: when-authorized
probes:
enabled: true
info:
git:
mode: fullThree layers, all of them:
1. Enumerate endpoints. Never include: "*". Phase 4 lists why: env prints configuration, heapdump downloads your memory with credentials in it, loggers lets anyone flood your disk, shutdown is a one-request outage.
2. Separate port, excluded from public routing.
3. Authenticate anything beyond probes:
.requestMatchers("/actuator/health/**", "/actuator/info", "/actuator/prometheus").permitAll()
.requestMatchers("/actuator/**").hasRole("OPS")Verify this from outside your cluster rather than trusting the configuration. curl https://your-api.example.com/actuator/env should return 404 or 401. Internet-wide scans for exposed Actuator endpoints are routine and automated, and an exposed /actuator/heapdump is a credential breach, not an information leak.
Disable DevTools in production. It should never ship — on the developmentOnly configuration it is excluded from bootJar, and it disables itself when running from a fat JAR anyway — but confirm nobody has moved it to implementation. It exposes a restart endpoint and relaxes several defaults. In Spring Boot 4, DevTools is deprecated anyway.
JVM and Container Sizing
The pairing that causes most container OOM-kills:
resources:
requests: { memory: "768Mi", cpu: "500m" }
limits: { memory: "768Mi", cpu: "2000m" }JAVA_TOOL_OPTIONS=-XX:MaxRAMPercentage=70 -XX:+UseZGC -XX:+ExitOnOutOfMemoryError
Why MaxRAMPercentage and not -Xmx. The JVM needs memory beyond the heap: metaspace, thread stacks, code cache, direct buffers, GC structures. -Xmx700m in a 768Mi container is killed by the kernel — exit code 137, no stack trace, no OutOfMemoryError, because the process was killed rather than failing. 70% leaves room for the rest.
Set memory requests equal to limits. Java doesn't release heap readily, so a burstable memory allocation invites eviction. CPU can burst; memory shouldn't.
-XX:+ExitOnOutOfMemoryError. A JVM that has exhausted its heap is usually unable to serve requests but still passes a liveness probe, so it lingers in a degraded state. Exiting lets the orchestrator replace it.
Garbage collector. ZGC for low pause times (sub-millisecond, and it's the practical default on modern JDKs for latency-sensitive services); G1 remains a solid general choice; avoid SerialGC, which the JVM may select automatically in a container with less than 2 CPUs and ~1792MB — a frequent cause of "it's slow in Kubernetes but fine locally".
CPU limits and the JVM don't mix well. The JVM reads its CPU count to size GC threads, the common ForkJoinPool, and connection pools. A low limits.cpu triggers CFS throttling that can stall GC mid-collection, producing latency spikes that look like application problems. Prefer generous or absent CPU limits with accurate requests, and set -XX:ActiveProcessorCount explicitly if you must cap.
The 12-Factor Checks That Matter
Not all twelve are equally relevant to a Spring Boot service. These four are:
Config in the environment. One artefact per build, behaviour from environment variables — as phase 1 established. Secrets never in the image or the repository.
Stateless processes. No in-memory session state and no local disk you expect to persist. If you have sessions, externalise them (Spring Session + Redis); if you cache in-process, accept the per-instance staleness phase 3 described.
Disposability. Fast startup, graceful shutdown on SIGTERM, and tolerance of being killed at any moment. The previous guides cover this.
Logs as event streams. stdout, structured, collected by the platform.
The one most often violated in practice is statelessness — usually by accident, via an in-process cache or a scheduled job assuming it's the only instance.
Dependency and Supply-Chain Hygiene
Not in the roadmap spec, and worth more than most of what is:
plugins {
id 'org.owasp.dependencycheck' version '12.1.0'
id 'com.github.ben-manes.versions' version '0.52.0'
}./gradlew dependencyCheckAnalyze # known CVEs in your dependencies
./gradlew dependencyUpdates # newer versions availableStay on a supported Spring Boot line. OSS support for a minor line is roughly 12–18 months; running past it means unpatched CVEs. As phase 3's Flyway guidance put it for schemas and the Spring Boot 4 post puts it for upgrades: plan against your support window, not against the release announcement. An unsupported Boot version is a security problem, not untidiness.
Generate an SBOM. Buildpacks do it automatically; the cyclonedx-gradle-plugin otherwise, or Spring Boot 3.3+'s own springBoot { sbom() } support. When the next Log4Shell happens, the question "are we affected" needs an answer in minutes.
Pin and scan base images. Rebuild periodically even without code changes — base image CVEs accumulate independently of your application.
The Go-Live Checklist
Drawn from all six phases.
Configuration
- No secrets in the repository or image; all from the environment
- Typed
@ConfigurationPropertieswith@Validatedso misconfiguration fails at startup ddl-auto: validate, Flyway owning the schemaspring.jpa.open-in-view: false
Security
- Rules ordered specific→general, ending in
authenticated()ordenyAll() - CSRF enabled for cookie/session auth; disabled only for header-authenticated APIs
- CORS origins enumerated, never
*with credentials - Passwords with BCrypt (cost ≥ 10) or Argon2id
- Negative authorisation tests: 401 anonymous, 403 under-privileged, one user's data not visible to another
- Actuator locked down and verified from outside
Data
- Hikari
connection-timeoutlowered to a few seconds max-lifetimebelow any infrastructure idle timeout- Total pool across all instances fits
max_connections, with deployment-overlap headroom leak-detection-thresholdset- No network I/O inside
@Transactional - N+1 queries checked, with query-count assertions on key paths
Resilience
- Connect and read timeouts on every outbound client
- Retries only on idempotent operations; idempotency keys where not
- Message consumers idempotent; DLQ configured with
default-requeue-rejected: falseand monitored - Scheduled jobs: exceptions caught, concurrency-safe across instances
Operations
server.shutdown: graceful,terminationGracePeriodSecondsgreater than its timeoutpreStophook to cover the deregistration race- Liveness on
livenessStateonly; readiness chosen deliberately - Structured JSON logs to stdout with
trace.id - Prometheus histograms; alerts on error ratio and latency, not heap
- Trace ID returned as the client error reference
/actuator/infocarries git SHA and build time; images tagged immutably
Verification
- A rolling deployment under load produces zero failed requests
- Migrations tested from empty against the production database engine
- Dependency CVE scan in CI; SBOM generated
Run the deployment test before go-live and keep it in CI. It exercises graceful shutdown, the preStop hook, probe configuration, grace periods and readiness gating simultaneously — and those parts interact, so a regression in any one is invisible until a real deployment under real traffic. It is the single highest-value test in this list.
Check yourself
A service is configured with container memory limit 1Gi and JAVA_TOOL_OPTIONS=-Xmx1g. It runs fine in staging and is OOM-killed under production load with exit code 137 and no Java stack trace. Why?
What Hardening Doesn't Cover
Worth saying plainly, so the checklist isn't mistaken for completeness. It makes a single service operable. It says nothing about:
- Capacity — how many instances, and what they can actually serve. That needs load testing against production-shaped data.
- Disaster recovery — backups that have been restored in a drill, and a measured RTO/RPO.
- Cross-service failure — one service correctly hardened can still be taken down by a dependency, which is phase 7.
- Operational readiness — runbooks, on-call, and someone who knows what the alerts mean at 3am.
A green checklist is a necessary condition, not a sufficient one.
The Mental Model, Restated
- Actuator: enumerate endpoints, separate port, authenticate — and verify from outside.
MaxRAMPercentageat 70%, never-Xmxat the container limit. Requests equal limits for memory.- ZGC or G1; make sure you haven't been given SerialGC by a small container.
- Beware CPU limits — throttling stalls GC and produces latency spikes.
- Config from the environment, stateless processes, graceful disposability, logs to stdout.
- Scan dependencies, generate an SBOM, stay on a supported Boot line.
- Test a rolling deployment under load. Zero errors or the configuration is wrong.
Phase 6 in Four Sentences
A fat JAR in one COPY wastes the layer cache, so layer by change frequency — or let buildpacks do it, with a memory calculator and a non-root user included. Default SIGTERM handling drops in-flight requests, and even with graceful shutdown a preStop hook is needed because deregistration races with termination. Native images buy sub-100ms startup and much less memory for the closed-world assumption and slow builds, while virtual threads remove the thread-pool ceiling and move the bottleneck to every pool you had implicitly sized. Observability pays off only when metrics, traces and logs are joined by a trace ID — which is also the best thing to hand a client as an error reference.
What's Next
Phase 7 is the last: what changes when one service becomes several. Service discovery, an API gateway, client-side load balancing, circuit breakers that stop a dependency's failure becoming yours, centralised configuration, and tracing that spans the whole system — plus an honest look at when you don't need any of it.