04-advanced-spring

Actuator and Micrometer: Making a Service Observable

The endpoints worth exposing, why liveness and readiness are different questions, custom metrics with Micrometer, and keeping label cardinality under control.

October 9, 2026
spring-bootactuatormicrometerobservabilitymetricsprometheushealth-checkstracing

Three Questions a Running Service Must Answer

  • Is it alive? Should the orchestrator restart it?
  • Can it take traffic? Should the load balancer route to it?
  • What is it doing? Latency, throughput, errors, resource use.

Actuator answers all three over HTTP, and Micrometer provides the metrics underneath. Both come from one dependency:

groovy
dependencies {
    implementation 'org.springframework.boot:spring-boot-starter-actuator'
}

This guide is the Spring half. Phase 6 covers the stack it feeds — Prometheus, Grafana, log aggregation and tracing backends.

Exposure Is Opt-In

By default only /actuator/health is exposed over HTTP. That default is correct, and the common mistake is to replace it with include: "*":

yaml
management:
  endpoints:
    web:
      exposure:
        include: health,info,metrics,prometheus     # name them
EndpointGives you
healthAggregate health plus per-component detail
infoBuild and git metadata
metricsMetric names and values
prometheusEverything in Prometheus scrape format
loggersRead and change log levels at runtime
envEvery property and its source
configpropsBound @ConfigurationProperties
threaddumpA thread dump
heapdumpDownloads a heap dump
conditionsThe auto-configuration report from phase 1
shutdownShuts the application down (disabled by default)
🚨

include: "*" on a publicly reachable port is a serious exposure. /actuator/env prints configuration including property names and many values; /actuator/heapdump downloads your heap — which contains credentials, tokens and personal data in plain memory; /actuator/loggers lets anyone turn on DEBUG logging and flood your disk; /actuator/shutdown is a one-request denial of service. Enumerate what you need, and treat the whole namespace as privileged.

Two protections worth applying together. A separate port, which keeps Actuator off your public ingress entirely:

yaml
management:
  server:
    port: 9001          # scrape and probe this; don't route it publicly
  endpoints:
    web:
      exposure:
        include: health,info,prometheus,metrics

And authentication, for anything beyond probes:

java
.requestMatchers("/actuator/health/**", "/actuator/info").permitAll()
.requestMatchers("/actuator/**").hasRole("OPS")

Note health/** rather than health — the liveness and readiness groups below are sub-paths.

Health: Liveness ≠ Readiness

The distinction that matters most operationally, and the one most often collapsed into a single probe.

Liveness — is the process broken beyond recovery? A failing liveness probe means restart me.

Readiness — can it serve traffic right now? A failing readiness probe means stop sending requests, without a restart.

Confusing them causes a specific and severe failure. If liveness includes a database check, a brief database outage makes every instance fail liveness, so the orchestrator restarts them all — repeatedly, while the database is still down. You have converted a recoverable dependency blip into a crash loop across the fleet.

Liveness should depend on almost nothing. Readiness may depend on dependencies.

Spring Boot has built-in support:

yaml
management:
  endpoint:
    health:
      probes:
        enabled: true
      show-details: when-authorized
      group:
        liveness:
          include: livenessState          # just "is the context running"
        readiness:
          include: readinessState,db      # add what you need to serve

That gives you /actuator/health/liveness and /actuator/health/readiness, which map directly onto Kubernetes probes.

⚠️

Even readiness deserves thought. If every instance's readiness depends on the database, a database outage removes the entire service from the load balancer — clients get connection failures rather than a degraded response. For a service that genuinely cannot function without its database that may be right; for one that can serve cached or static responses, excluding db from readiness keeps part of the service useful. There is no universal answer, which is the point: decide it deliberately.

Custom health indicators

java
@Component
public class InventoryHealthIndicator implements HealthIndicator {
 
    private final InventoryClient client;
 
    @Override
    public Health health() {
        try {
            return Health.up()
                    .withDetail("latencyMs", client.ping().toMillis())
                    .build();
        } catch (Exception ex) {
            return Health.down(ex).withDetail("endpoint", client.baseUrl()).build();
        }
    }
}

Registered automatically and named from the class — this appears as inventory. Two cautions: a health check is executed on every probe, so it must be fast and must not be an expensive query; and Health.down(ex) includes the exception message, so only expose it where show-details is restricted.

Metrics

Micrometer is a facade over metrics backends — you instrument once and choose the backend by dependency, the same pattern as SLF4J for logging and Spring's cache abstraction.

Out of the box you get JVM memory and GC, CPU, thread counts, HTTP server request timings, DataSource/HikariCP pool stats, cache statistics, and RestClient call timings. A long way toward observable before writing any code.

For Prometheus:

groovy
dependencies {
    runtimeOnly 'io.micrometer:micrometer-registry-prometheus'
}

The four meter types

java
@Service
public class OrderService {
 
    private final Counter placed;
    private final Timer processing;
    private final DistributionSummary values;
 
    public OrderService(MeterRegistry registry, OrderRepository repository) {
        this.placed = Counter.builder("orders.placed")
                .description("Orders successfully placed")
                .register(registry);
        this.processing = Timer.builder("orders.processing")
                .publishPercentiles(0.5, 0.95, 0.99)
                .register(registry);
        this.values = DistributionSummary.builder("orders.value")
                .baseUnit("GBP")
                .register(registry);
        Gauge.builder("orders.pending", repository, r -> r.countByStatus(OrderStatus.NEW))
                .register(registry);
    }
 
    public Order place(NewOrderRequest request) {
        return processing.record(() -> {
            Order order = doPlace(request);
            placed.increment();
            values.record(order.getTotal().doubleValue());
            return order;
        });
    }
}
TypeMeasuresExample
CounterA monotonically increasing totalOrders placed, errors
GaugeA current value that goes up and downQueue depth, pool size
TimerDuration and count togetherRequest latency
DistributionSummaryDistribution of a non-time valuePayload size, order value

Or annotate, with @Timed (needs @EnableAspectJAutoProxy and the AOP starter — the proxy mechanism again):

java
@Timed(value = "orders.processing", percentiles = { 0.5, 0.95, 0.99 })
public Order place(NewOrderRequest request) { }
🚨

Tag cardinality is the way to break your metrics backend. Every distinct combination of tag values is a separate time series. registry.counter("orders.placed", "customerId", id) with 100,000 customers creates 100,000 series — and Prometheus will struggle or fall over. Tags must be bounded and low-cardinality: status, region, endpoint template, error type. Never a user ID, order ID, email, raw URL or free-text input. This is the same discipline as keeping URIs as templates in the outbound HTTP guide.

Percentiles need one caveat: publishPercentiles computes them per instance, and percentiles cannot be averaged across instances — the mean of eight p99s is not the fleet p99. For correct aggregation, publish a histogram instead and let Prometheus compute quantiles:

java
Timer.builder("orders.processing")
     .publishPercentileHistogram()      // buckets, aggregatable across instances
     .register(registry);

Common tags

java
@Bean
MeterRegistryCustomizer<MeterRegistry> commonTags(
        @Value("${spring.application.name}") String app,
        @Value("${app.environment}") String env) {
    return registry -> registry.config().commonTags("application", app, "environment", env);
}

Applied to every metric, so dashboards can filter by service and environment without each instrumentation site remembering.

Check yourself

A Kubernetes deployment uses the same /actuator/health path for both liveness and readiness probes, and health includes a database check. The database becomes briefly unavailable. What happens?

Tracing

In a multi-service system, metrics tell you that latency is high and tracing tells you where. Micrometer Tracing (which replaced Spring Cloud Sleuth) propagates a trace context across service boundaries:

groovy
dependencies {
    implementation 'io.micrometer:micrometer-tracing-bridge-otel'
    runtimeOnly 'io.opentelemetry:opentelemetry-exporter-otlp'
}
yaml
management:
  tracing:
    sampling:
      probability: 0.1          # 10% in production
  otlp:
    tracing:
      endpoint: http://collector:4318/v1/traces

Incoming requests, outbound RestClient calls and database queries are instrumented automatically. The sampling rate is the knob that matters: 100% is usually unaffordable in volume and cost, and 1–10% is typical. Tail-based sampling — keeping all traces that contain an error — is better when your collector supports it, because the traces you most want are the rare ones.

✅

Put the trace ID in your log output and the three pillars connect: a metric shows a latency spike, a trace shows which call caused it, and the trace ID finds every log line for that request across every service.

yaml
logging:
  pattern:
    level: "%5p [${spring.application.name:},%X{traceId:-},%X{spanId:-}]"

This also closes the loop on the error-reference pattern from the exception handling guide — use the trace ID as the reference you return to clients and you can jump straight from a support ticket to the full request history.

Build Information in /info

groovy
springBoot {
    buildInfo()
}
yaml
management:
  info:
    git:
      mode: full
    env:
      enabled: true

/actuator/info then reports the version and commit actually running — which is the first thing you want during an incident and surprisingly hard to establish otherwise.

The Mental Model, Restated

  1. Expose endpoints by name, never "*". env, heapdump and loggers are privileged.
  2. Use a separate management port and authenticate everything but the probes.
  3. Liveness ≠ readiness. Liveness depends on almost nothing; a dependency check there causes crash loops.
  4. Counter, Gauge, Timer, DistributionSummary — pick by what you're measuring.
  5. Keep tag cardinality bounded. No IDs, emails or free text as tags.
  6. Publish histograms, not per-instance percentiles, so quantiles aggregate correctly.
  7. Sample traces and put the trace ID in your logs.

What's Next

@Timed needed the AOP starter, and @Transactional, @Cacheable and @PreAuthorize have all now appeared with the same self-invocation caveat. The next guide covers the mechanism itself: pointcuts, advice types, what Spring AOP can and cannot intercept, and when writing your own aspect is the right call rather than a clever way to hide behaviour.