06-production-cloud-native

The Observability Stack: Metrics, Logs and Traces Together

Prometheus and Grafana for metrics, structured JSON logs that are actually searchable, and the trace ID that joins all three into one debugging workflow.

October 9, 2026
spring-bootobservabilityprometheusgrafanaloggingelktracingopentelemetryslo

Three Signals, One Question

Phase 4 covered the Spring side — Actuator endpoints, Micrometer meters, health groups. This guide assembles the stack around it and, more importantly, makes the three signals work together.

Each answers a different question:

SignalAnswersGood atBad at
MetricsIs something wrong?Cheap, aggregated, alertableExplaining a specific request
TracesWhere is the time going?Latency breakdown across servicesAggregate trends
LogsWhat exactly happened?Full detail for one eventVolume and cost

The workflow that matters: a metric alerts, a trace localises, logs explain. That only works if the three are linked — and the link is the trace ID.

Metrics: Prometheus and Grafana

groovy
dependencies {
    runtimeOnly 'io.micrometer:micrometer-registry-prometheus'
}
yaml
management:
  server:
    port: 9001
  endpoints:
    web:
      exposure:
        include: health,info,prometheus
  metrics:
    distribution:
      percentiles-histogram:
        http.server.requests: true
      slo:
        http.server.requests: 50ms,100ms,200ms,500ms,1s

Prometheus scrapes /actuator/prometheus. In Kubernetes, annotate the pod or use a ServiceMonitor:

yaml
metadata:
  annotations:
    prometheus.io/scrape: "true"
    prometheus.io/port: "9001"
    prometheus.io/path: "/actuator/prometheus"

Note percentiles-histogram: true rather than publishPercentiles. As phase 4 explained, per-instance percentiles cannot be averaged across instances — the mean of eight p99s is not the fleet p99. A histogram ships buckets, and Prometheus computes the quantile correctly:

promql
histogram_quantile(0.99,
  sum(rate(http_server_requests_seconds_bucket[5m])) by (le, uri))

Queries worth having

promql
# Request rate by endpoint
sum(rate(http_server_requests_seconds_count[5m])) by (uri)
 
# Error ratio — the key SLO signal
sum(rate(http_server_requests_seconds_count{status=~"5.."}[5m]))
  / sum(rate(http_server_requests_seconds_count[5m]))
 
# Connection pool saturation
hikaricp_connections_active / hikaricp_connections_max
 
# Threads waiting for a connection — the leading indicator
hikaricp_connections_pending
 
# GC pause time proportion
sum(rate(jvm_gc_pause_seconds_sum[5m])) by (instance)
 
# Heap utilisation after collection
jvm_memory_used_bytes{area="heap"} / jvm_memory_max_bytes{area="heap"}
✅

Alert on symptoms users feel, not on causes. "Error ratio above 1% for 5 minutes" and "p99 above 500ms for 10 minutes" are worth waking someone for. "Heap above 80%" usually isn't — a JVM running near its heap limit is often working exactly as configured. Cause metrics belong on dashboards for diagnosis; alerts should map to user impact, or they train people to ignore the pager.

A dashboard that earns its place

Four rows, in this order:

  1. Request rate, error ratio, p50/p95/p99 — the service's user-visible behaviour.
  2. Dependencies — hikaricp_connections_pending, outbound client latency, cache hit ratio.
  3. JVM — heap after GC, GC pause proportion, thread count.
  4. Business — orders placed, payments failed, queue depth.

That fourth row is the one teams skip and miss most. A deployment that breaks checkout while keeping HTTP 200s is invisible in rows 1–3 and obvious in row 4.

Logs: Make Them Structured

A plain-text log line is readable by a human and expensive for a machine. grep across twelve services does not scale, and parsing free text is fragile.

Spring Boot 3.4+ has built-in structured logging — no custom Logback XML:

yaml
logging:
  structured:
    format:
      console: ecs        # or logstash, gelf
  level:
    root: INFO
    com.example.shop: INFO
json
{
  "@timestamp": "2026-10-09T09:14:22.116Z",
  "log.level": "ERROR",
  "service.name": "shop",
  "trace.id": "4f2a9c1b8e7d6a5f",
  "span.id": "a1b2c3d4",
  "message": "Payment authorisation failed",
  "error.type": "com.example.shop.PaymentDeclinedException"
}

Now a log store can index and filter on fields. The crucial field is trace.id — the join key between logs and traces.

Log to stdout

yaml
logging:
  file:
    name:              # leave unset — don't write files in a container

In a container, log to stdout and let the platform collect it. Writing to a file means log rotation, a volume, disk-full risk, and logs lost when the pod is deleted. The collection agent (Fluent Bit, Promtail, Vector) reads the container's stdout.

Add context, not noise

java
@Component
public class TenantLoggingFilter extends OncePerRequestFilter {
 
    @Override
    protected void doFilterInternal(HttpServletRequest request,
                                    HttpServletResponse response,
                                    FilterChain chain) throws ServletException, IOException {
        String tenant = request.getHeader("X-Tenant");
        try {
            if (tenant != null) {
                MDC.put("tenant", tenant);
            }
            chain.doFilter(request, response);
        } finally {
            MDC.clear();                           // always — threads are reused
        }
    }
}

That finally is mandatory on platform threads: a pooled thread retaining MDC state attributes the next request's logs to the previous tenant. With virtual threads the risk disappears, but write the finally regardless.

🚨

Logs are a data-protection surface. Never log passwords, tokens, full card numbers, or personal data beyond what you can justify. The common accidents: logging a whole request body on a validation failure (which includes the password on a registration request), logging an exception whose message embeds a token, and DEBUG on org.springframework.security in production. Logs are widely readable, aggregated into systems with long retention, and often outside your primary data-protection review.

Levels, used properly

LevelUseVolume
ERRORNeeds human attention — 5xx, failed jobsLow
WARNUnexpected but handled — retry, fallbackLow
INFOSignificant business eventsModerate
DEBUGDiagnostic detailOff in production
TRACEVery fine detailNever in production

The rule from the exception handling guide holds: 4xx is information, 5xx is an incident. Logging every 404 at ERROR makes the level meaningless — and a log-based alert on error rate then fires on normal traffic.

Change levels at runtime without a redeploy, via Actuator:

bash
curl -X POST http://localhost:9001/actuator/loggers/com.example.shop \
  -H 'Content-Type: application/json' -d '{"configuredLevel":"DEBUG"}'

Invaluable during an incident — and the reason the loggers endpoint must be authenticated.

Traces

groovy
dependencies {
    implementation 'io.micrometer:micrometer-tracing-bridge-otel'
    runtimeOnly 'io.opentelemetry:opentelemetry-exporter-otlp'
}
yaml
management:
  tracing:
    sampling:
      probability: 0.1
  otlp:
    tracing:
      endpoint: http://otel-collector:4318/v1/traces

Instrumented automatically: incoming requests, RestClient/WebClient calls, JDBC queries, Kafka and Rabbit listeners, scheduled tasks. Context propagates via W3C traceparent headers, so a trace spans services without application code.

Prefer OTLP over Zipkin or Jaeger-native exporters. OpenTelemetry is the vendor-neutral standard, and an OTel Collector can fan out to Tempo, Jaeger, Honeycomb or Datadog without changing your application.

Sampling

100% tracing is usually unaffordable in both volume and cost. The options:

  • Head sampling (probability: 0.1) — decide at the first span. Simple, and loses 90% of errors.
  • Tail sampling — the Collector decides after seeing the whole trace, so it can keep all traces containing an error or exceeding a latency threshold, plus a small random sample of healthy ones.

Tail sampling is strictly better when available, because the traces you most want are exactly the rare ones head sampling discards.

Custom spans

java
@Observed(name = "orders.place", contextualName = "place-order")
public Order place(NewOrderRequest request) { }

Or manually, with a tag:

java
Span span = tracer.nextSpan().name("inventory-reservation").start();
try (var ignored = tracer.withSpan(span)) {
    span.tag("sku", request.sku());
    return inventory.reserve(request);
} finally {
    span.end();
}

Span tags tolerate higher cardinality than metric labels — a trace is one record, not a time series — so a SKU or order ID is acceptable here and never acceptable as a metric tag.

Joining the Three

The payoff. With trace.id in your logs and a tracing backend that indexes it:

In Grafana, with Loki and Tempo configured as linked data sources, each step is a click. Without the trace ID in logs, each step is a guess based on timestamps.

✅

Return the trace ID to clients as the error reference from the exception handling guide, rather than a fresh UUID:

java
problem.setProperty("reference", tracer.currentSpan().context().traceId());

A support ticket then carries a key that resolves directly to the full cross-service history of that request. This is the single highest-leverage line in this guide.

A Stack That Works

NeedOption
MetricsPrometheus (or Mimir / VictoriaMetrics at scale)
LogsLoki (cheap) or Elasticsearch (powerful search)
TracesTempo (cheap) or Jaeger
DashboardsGrafana
CollectionOpenTelemetry Collector

Loki vs Elasticsearch is the main decision. Loki indexes only labels and stores log content compressed, making it far cheaper — good when you know which service and time window you want. Elasticsearch indexes everything, making full-text search across all logs fast and the storage bill large. For a team already filtering by service and trace ID, Loki is usually the better trade.

The roadmap mentions ELK, and Logstash's heavyweight JVM agent is now usually replaced by Fluent Bit or Vector — both much lighter for the same job.

SLOs, Briefly

Metrics without a target are just graphs. An SLO states the target:

99.5% of requests complete successfully in under 300ms, measured over 30 days.

That gives an error budget — 0.5%, about 3.6 hours a month. Which makes the engineering conversation concrete: budget remaining means you can ship riskier changes; budget exhausted means stability comes first.

promql
# Fraction of requests inside the 300ms objective
sum(rate(http_server_requests_seconds_bucket{le="0.3",status!~"5.."}[30d]))
  / sum(rate(http_server_requests_seconds_count[30d]))

Those slo buckets configured at the top of this guide are what make that query possible — the bucket boundary must exist to be queried.

Check yourself

An alert fires for elevated p99 latency. The team has Prometheus, Grafana and Tempo, and plain-text logs without trace IDs. Why is diagnosis slow?

The Mental Model, Restated

  1. Metrics alert, traces localise, logs explain — and only if linked by trace ID.
  2. Publish histograms, not per-instance percentiles, so quantiles aggregate.
  3. Alert on user-visible symptoms, not on causes like heap usage.
  4. Include a business-metrics row on your dashboard.
  5. Structured JSON logs to stdout, never files in a container.
  6. Clear the MDC in a finally, and never log secrets or personal data.
  7. Use OTLP and prefer tail sampling, which keeps the error traces.
  8. Return the trace ID as the client's error reference.
  9. An SLO turns metrics into an error budget, which turns reliability into a decision.

What's Next

The final guide in this phase is a hardening pass: locking down Actuator properly, JVM and container sizing, the 12-factor checklist, and a go-live review drawing together the decisions made across all six phases.