The Observability Stack: Metrics, Logs and Traces Together
Prometheus and Grafana for metrics, structured JSON logs that are actually searchable, and the trace ID that joins all three into one debugging workflow.
Three Signals, One Question
Phase 4 covered the Spring side — Actuator endpoints, Micrometer meters, health groups. This guide assembles the stack around it and, more importantly, makes the three signals work together.
Each answers a different question:
| Signal | Answers | Good at | Bad at |
|---|---|---|---|
| Metrics | Is something wrong? | Cheap, aggregated, alertable | Explaining a specific request |
| Traces | Where is the time going? | Latency breakdown across services | Aggregate trends |
| Logs | What exactly happened? | Full detail for one event | Volume and cost |
The workflow that matters: a metric alerts, a trace localises, logs explain. That only works if the three are linked — and the link is the trace ID.
Metrics: Prometheus and Grafana
dependencies {
runtimeOnly 'io.micrometer:micrometer-registry-prometheus'
}management:
server:
port: 9001
endpoints:
web:
exposure:
include: health,info,prometheus
metrics:
distribution:
percentiles-histogram:
http.server.requests: true
slo:
http.server.requests: 50ms,100ms,200ms,500ms,1sPrometheus scrapes /actuator/prometheus. In Kubernetes, annotate the pod or use a ServiceMonitor:
metadata:
annotations:
prometheus.io/scrape: "true"
prometheus.io/port: "9001"
prometheus.io/path: "/actuator/prometheus"Note percentiles-histogram: true rather than publishPercentiles. As phase 4 explained, per-instance percentiles cannot be averaged across instances — the mean of eight p99s is not the fleet p99. A histogram ships buckets, and Prometheus computes the quantile correctly:
histogram_quantile(0.99,
sum(rate(http_server_requests_seconds_bucket[5m])) by (le, uri))Queries worth having
# Request rate by endpoint
sum(rate(http_server_requests_seconds_count[5m])) by (uri)
# Error ratio — the key SLO signal
sum(rate(http_server_requests_seconds_count{status=~"5.."}[5m]))
/ sum(rate(http_server_requests_seconds_count[5m]))
# Connection pool saturation
hikaricp_connections_active / hikaricp_connections_max
# Threads waiting for a connection — the leading indicator
hikaricp_connections_pending
# GC pause time proportion
sum(rate(jvm_gc_pause_seconds_sum[5m])) by (instance)
# Heap utilisation after collection
jvm_memory_used_bytes{area="heap"} / jvm_memory_max_bytes{area="heap"}Alert on symptoms users feel, not on causes. "Error ratio above 1% for 5 minutes" and "p99 above 500ms for 10 minutes" are worth waking someone for. "Heap above 80%" usually isn't — a JVM running near its heap limit is often working exactly as configured. Cause metrics belong on dashboards for diagnosis; alerts should map to user impact, or they train people to ignore the pager.
A dashboard that earns its place
Four rows, in this order:
- Request rate, error ratio, p50/p95/p99 — the service's user-visible behaviour.
- Dependencies —
hikaricp_connections_pending, outbound client latency, cache hit ratio. - JVM — heap after GC, GC pause proportion, thread count.
- Business — orders placed, payments failed, queue depth.
That fourth row is the one teams skip and miss most. A deployment that breaks checkout while keeping HTTP 200s is invisible in rows 1–3 and obvious in row 4.
Logs: Make Them Structured
A plain-text log line is readable by a human and expensive for a machine. grep across twelve services does not scale, and parsing free text is fragile.
Spring Boot 3.4+ has built-in structured logging — no custom Logback XML:
logging:
structured:
format:
console: ecs # or logstash, gelf
level:
root: INFO
com.example.shop: INFO{
"@timestamp": "2026-10-09T09:14:22.116Z",
"log.level": "ERROR",
"service.name": "shop",
"trace.id": "4f2a9c1b8e7d6a5f",
"span.id": "a1b2c3d4",
"message": "Payment authorisation failed",
"error.type": "com.example.shop.PaymentDeclinedException"
}Now a log store can index and filter on fields. The crucial field is trace.id — the join key between logs and traces.
Log to stdout
logging:
file:
name: # leave unset — don't write files in a containerIn a container, log to stdout and let the platform collect it. Writing to a file means log rotation, a volume, disk-full risk, and logs lost when the pod is deleted. The collection agent (Fluent Bit, Promtail, Vector) reads the container's stdout.
Add context, not noise
@Component
public class TenantLoggingFilter extends OncePerRequestFilter {
@Override
protected void doFilterInternal(HttpServletRequest request,
HttpServletResponse response,
FilterChain chain) throws ServletException, IOException {
String tenant = request.getHeader("X-Tenant");
try {
if (tenant != null) {
MDC.put("tenant", tenant);
}
chain.doFilter(request, response);
} finally {
MDC.clear(); // always — threads are reused
}
}
}That finally is mandatory on platform threads: a pooled thread retaining MDC state attributes the next request's logs to the previous tenant. With virtual threads the risk disappears, but write the finally regardless.
Logs are a data-protection surface. Never log passwords, tokens, full card numbers, or personal data beyond what you can justify. The common accidents: logging a whole request body on a validation failure (which includes the password on a registration request), logging an exception whose message embeds a token, and DEBUG on org.springframework.security in production. Logs are widely readable, aggregated into systems with long retention, and often outside your primary data-protection review.
Levels, used properly
| Level | Use | Volume |
|---|---|---|
ERROR | Needs human attention — 5xx, failed jobs | Low |
WARN | Unexpected but handled — retry, fallback | Low |
INFO | Significant business events | Moderate |
DEBUG | Diagnostic detail | Off in production |
TRACE | Very fine detail | Never in production |
The rule from the exception handling guide holds: 4xx is information, 5xx is an incident. Logging every 404 at ERROR makes the level meaningless — and a log-based alert on error rate then fires on normal traffic.
Change levels at runtime without a redeploy, via Actuator:
curl -X POST http://localhost:9001/actuator/loggers/com.example.shop \
-H 'Content-Type: application/json' -d '{"configuredLevel":"DEBUG"}'Invaluable during an incident — and the reason the loggers endpoint must be authenticated.
Traces
dependencies {
implementation 'io.micrometer:micrometer-tracing-bridge-otel'
runtimeOnly 'io.opentelemetry:opentelemetry-exporter-otlp'
}management:
tracing:
sampling:
probability: 0.1
otlp:
tracing:
endpoint: http://otel-collector:4318/v1/tracesInstrumented automatically: incoming requests, RestClient/WebClient calls, JDBC queries, Kafka and Rabbit listeners, scheduled tasks. Context propagates via W3C traceparent headers, so a trace spans services without application code.
Prefer OTLP over Zipkin or Jaeger-native exporters. OpenTelemetry is the vendor-neutral standard, and an OTel Collector can fan out to Tempo, Jaeger, Honeycomb or Datadog without changing your application.
Sampling
100% tracing is usually unaffordable in both volume and cost. The options:
- Head sampling (
probability: 0.1) — decide at the first span. Simple, and loses 90% of errors. - Tail sampling — the Collector decides after seeing the whole trace, so it can keep all traces containing an error or exceeding a latency threshold, plus a small random sample of healthy ones.
Tail sampling is strictly better when available, because the traces you most want are exactly the rare ones head sampling discards.
Custom spans
@Observed(name = "orders.place", contextualName = "place-order")
public Order place(NewOrderRequest request) { }Or manually, with a tag:
Span span = tracer.nextSpan().name("inventory-reservation").start();
try (var ignored = tracer.withSpan(span)) {
span.tag("sku", request.sku());
return inventory.reserve(request);
} finally {
span.end();
}Span tags tolerate higher cardinality than metric labels — a trace is one record, not a time series — so a SKU or order ID is acceptable here and never acceptable as a metric tag.
Joining the Three
The payoff. With trace.id in your logs and a tracing backend that indexes it:
In Grafana, with Loki and Tempo configured as linked data sources, each step is a click. Without the trace ID in logs, each step is a guess based on timestamps.
Return the trace ID to clients as the error reference from the exception handling guide, rather than a fresh UUID:
problem.setProperty("reference", tracer.currentSpan().context().traceId());A support ticket then carries a key that resolves directly to the full cross-service history of that request. This is the single highest-leverage line in this guide.
A Stack That Works
| Need | Option |
|---|---|
| Metrics | Prometheus (or Mimir / VictoriaMetrics at scale) |
| Logs | Loki (cheap) or Elasticsearch (powerful search) |
| Traces | Tempo (cheap) or Jaeger |
| Dashboards | Grafana |
| Collection | OpenTelemetry Collector |
Loki vs Elasticsearch is the main decision. Loki indexes only labels and stores log content compressed, making it far cheaper — good when you know which service and time window you want. Elasticsearch indexes everything, making full-text search across all logs fast and the storage bill large. For a team already filtering by service and trace ID, Loki is usually the better trade.
The roadmap mentions ELK, and Logstash's heavyweight JVM agent is now usually replaced by Fluent Bit or Vector — both much lighter for the same job.
SLOs, Briefly
Metrics without a target are just graphs. An SLO states the target:
99.5% of requests complete successfully in under 300ms, measured over 30 days.
That gives an error budget — 0.5%, about 3.6 hours a month. Which makes the engineering conversation concrete: budget remaining means you can ship riskier changes; budget exhausted means stability comes first.
# Fraction of requests inside the 300ms objective
sum(rate(http_server_requests_seconds_bucket{le="0.3",status!~"5.."}[30d]))
/ sum(rate(http_server_requests_seconds_count[30d]))Those slo buckets configured at the top of this guide are what make that query possible — the bucket boundary must exist to be queried.
Check yourself
An alert fires for elevated p99 latency. The team has Prometheus, Grafana and Tempo, and plain-text logs without trace IDs. Why is diagnosis slow?
The Mental Model, Restated
- Metrics alert, traces localise, logs explain — and only if linked by trace ID.
- Publish histograms, not per-instance percentiles, so quantiles aggregate.
- Alert on user-visible symptoms, not on causes like heap usage.
- Include a business-metrics row on your dashboard.
- Structured JSON logs to stdout, never files in a container.
- Clear the MDC in a
finally, and never log secrets or personal data. - Use OTLP and prefer tail sampling, which keeps the error traces.
- Return the trace ID as the client's error reference.
- An SLO turns metrics into an error budget, which turns reliability into a decision.
What's Next
The final guide in this phase is a hardening pass: locking down Actuator properly, JVM and container sizing, the 12-factor checklist, and a go-live review drawing together the decisions made across all six phases.