Actuator and Micrometer: Making a Service Observable
The endpoints worth exposing, why liveness and readiness are different questions, custom metrics with Micrometer, and keeping label cardinality under control.
Three Questions a Running Service Must Answer
- Is it alive? Should the orchestrator restart it?
- Can it take traffic? Should the load balancer route to it?
- What is it doing? Latency, throughput, errors, resource use.
Actuator answers all three over HTTP, and Micrometer provides the metrics underneath. Both come from one dependency:
dependencies {
implementation 'org.springframework.boot:spring-boot-starter-actuator'
}This guide is the Spring half. Phase 6 covers the stack it feeds — Prometheus, Grafana, log aggregation and tracing backends.
Exposure Is Opt-In
By default only /actuator/health is exposed over HTTP. That default is correct, and the common mistake is to replace it with include: "*":
management:
endpoints:
web:
exposure:
include: health,info,metrics,prometheus # name them| Endpoint | Gives you |
|---|---|
health | Aggregate health plus per-component detail |
info | Build and git metadata |
metrics | Metric names and values |
prometheus | Everything in Prometheus scrape format |
loggers | Read and change log levels at runtime |
env | Every property and its source |
configprops | Bound @ConfigurationProperties |
threaddump | A thread dump |
heapdump | Downloads a heap dump |
conditions | The auto-configuration report from phase 1 |
shutdown | Shuts the application down (disabled by default) |
include: "*" on a publicly reachable port is a serious exposure. /actuator/env prints configuration including property names and many values; /actuator/heapdump downloads your heap — which contains credentials, tokens and personal data in plain memory; /actuator/loggers lets anyone turn on DEBUG logging and flood your disk; /actuator/shutdown is a one-request denial of service. Enumerate what you need, and treat the whole namespace as privileged.
Two protections worth applying together. A separate port, which keeps Actuator off your public ingress entirely:
management:
server:
port: 9001 # scrape and probe this; don't route it publicly
endpoints:
web:
exposure:
include: health,info,prometheus,metricsAnd authentication, for anything beyond probes:
.requestMatchers("/actuator/health/**", "/actuator/info").permitAll()
.requestMatchers("/actuator/**").hasRole("OPS")Note health/** rather than health — the liveness and readiness groups below are sub-paths.
Health: Liveness ≠ Readiness
The distinction that matters most operationally, and the one most often collapsed into a single probe.
Liveness — is the process broken beyond recovery? A failing liveness probe means restart me.
Readiness — can it serve traffic right now? A failing readiness probe means stop sending requests, without a restart.
Confusing them causes a specific and severe failure. If liveness includes a database check, a brief database outage makes every instance fail liveness, so the orchestrator restarts them all — repeatedly, while the database is still down. You have converted a recoverable dependency blip into a crash loop across the fleet.
Liveness should depend on almost nothing. Readiness may depend on dependencies.
Spring Boot has built-in support:
management:
endpoint:
health:
probes:
enabled: true
show-details: when-authorized
group:
liveness:
include: livenessState # just "is the context running"
readiness:
include: readinessState,db # add what you need to serveThat gives you /actuator/health/liveness and /actuator/health/readiness, which map directly onto Kubernetes probes.
Even readiness deserves thought. If every instance's readiness depends on the database, a database outage removes the entire service from the load balancer — clients get connection failures rather than a degraded response. For a service that genuinely cannot function without its database that may be right; for one that can serve cached or static responses, excluding db from readiness keeps part of the service useful. There is no universal answer, which is the point: decide it deliberately.
Custom health indicators
@Component
public class InventoryHealthIndicator implements HealthIndicator {
private final InventoryClient client;
@Override
public Health health() {
try {
return Health.up()
.withDetail("latencyMs", client.ping().toMillis())
.build();
} catch (Exception ex) {
return Health.down(ex).withDetail("endpoint", client.baseUrl()).build();
}
}
}Registered automatically and named from the class — this appears as inventory. Two cautions: a health check is executed on every probe, so it must be fast and must not be an expensive query; and Health.down(ex) includes the exception message, so only expose it where show-details is restricted.
Metrics
Micrometer is a facade over metrics backends — you instrument once and choose the backend by dependency, the same pattern as SLF4J for logging and Spring's cache abstraction.
Out of the box you get JVM memory and GC, CPU, thread counts, HTTP server request timings, DataSource/HikariCP pool stats, cache statistics, and RestClient call timings. A long way toward observable before writing any code.
For Prometheus:
dependencies {
runtimeOnly 'io.micrometer:micrometer-registry-prometheus'
}The four meter types
@Service
public class OrderService {
private final Counter placed;
private final Timer processing;
private final DistributionSummary values;
public OrderService(MeterRegistry registry, OrderRepository repository) {
this.placed = Counter.builder("orders.placed")
.description("Orders successfully placed")
.register(registry);
this.processing = Timer.builder("orders.processing")
.publishPercentiles(0.5, 0.95, 0.99)
.register(registry);
this.values = DistributionSummary.builder("orders.value")
.baseUnit("GBP")
.register(registry);
Gauge.builder("orders.pending", repository, r -> r.countByStatus(OrderStatus.NEW))
.register(registry);
}
public Order place(NewOrderRequest request) {
return processing.record(() -> {
Order order = doPlace(request);
placed.increment();
values.record(order.getTotal().doubleValue());
return order;
});
}
}| Type | Measures | Example |
|---|---|---|
| Counter | A monotonically increasing total | Orders placed, errors |
| Gauge | A current value that goes up and down | Queue depth, pool size |
| Timer | Duration and count together | Request latency |
| DistributionSummary | Distribution of a non-time value | Payload size, order value |
Or annotate, with @Timed (needs @EnableAspectJAutoProxy and the AOP starter — the proxy mechanism again):
@Timed(value = "orders.processing", percentiles = { 0.5, 0.95, 0.99 })
public Order place(NewOrderRequest request) { }Tag cardinality is the way to break your metrics backend. Every distinct combination of tag values is a separate time series. registry.counter("orders.placed", "customerId", id) with 100,000 customers creates 100,000 series — and Prometheus will struggle or fall over. Tags must be bounded and low-cardinality: status, region, endpoint template, error type. Never a user ID, order ID, email, raw URL or free-text input. This is the same discipline as keeping URIs as templates in the outbound HTTP guide.
Percentiles need one caveat: publishPercentiles computes them per instance, and percentiles cannot be averaged across instances — the mean of eight p99s is not the fleet p99. For correct aggregation, publish a histogram instead and let Prometheus compute quantiles:
Timer.builder("orders.processing")
.publishPercentileHistogram() // buckets, aggregatable across instances
.register(registry);Common tags
@Bean
MeterRegistryCustomizer<MeterRegistry> commonTags(
@Value("${spring.application.name}") String app,
@Value("${app.environment}") String env) {
return registry -> registry.config().commonTags("application", app, "environment", env);
}Applied to every metric, so dashboards can filter by service and environment without each instrumentation site remembering.
Check yourself
A Kubernetes deployment uses the same /actuator/health path for both liveness and readiness probes, and health includes a database check. The database becomes briefly unavailable. What happens?
Tracing
In a multi-service system, metrics tell you that latency is high and tracing tells you where. Micrometer Tracing (which replaced Spring Cloud Sleuth) propagates a trace context across service boundaries:
dependencies {
implementation 'io.micrometer:micrometer-tracing-bridge-otel'
runtimeOnly 'io.opentelemetry:opentelemetry-exporter-otlp'
}management:
tracing:
sampling:
probability: 0.1 # 10% in production
otlp:
tracing:
endpoint: http://collector:4318/v1/tracesIncoming requests, outbound RestClient calls and database queries are instrumented automatically. The sampling rate is the knob that matters: 100% is usually unaffordable in volume and cost, and 1–10% is typical. Tail-based sampling — keeping all traces that contain an error — is better when your collector supports it, because the traces you most want are the rare ones.
Put the trace ID in your log output and the three pillars connect: a metric shows a latency spike, a trace shows which call caused it, and the trace ID finds every log line for that request across every service.
logging:
pattern:
level: "%5p [${spring.application.name:},%X{traceId:-},%X{spanId:-}]"This also closes the loop on the error-reference pattern from the exception handling guide — use the trace ID as the reference you return to clients and you can jump straight from a support ticket to the full request history.
Build Information in /info
springBoot {
buildInfo()
}management:
info:
git:
mode: full
env:
enabled: true/actuator/info then reports the version and commit actually running — which is the first thing you want during an incident and surprisingly hard to establish otherwise.
The Mental Model, Restated
- Expose endpoints by name, never
"*".env,heapdumpandloggersare privileged. - Use a separate management port and authenticate everything but the probes.
- Liveness ≠ readiness. Liveness depends on almost nothing; a dependency check there causes crash loops.
- Counter, Gauge, Timer, DistributionSummary — pick by what you're measuring.
- Keep tag cardinality bounded. No IDs, emails or free text as tags.
- Publish histograms, not per-instance percentiles, so quantiles aggregate correctly.
- Sample traces and put the trace ID in your logs.
What's Next
@Timed needed the AOP starter, and @Transactional, @Cacheable and @PreAuthorize have all now appeared with the same self-invocation caveat. The next guide covers the mechanism itself: pointcuts, advice types, what Spring AOP can and cannot intercept, and when writing your own aspect is the right call rather than a clever way to hide behaviour.