Distributed Tracing End to End
How trace context crosses service boundaries, sampling strategies that keep the traces you need, and correlating traces with logs and metrics.
The Question a Single Service Cannot Answer
A checkout request is slow. It touched the gateway, the orders service, inventory, payments and a notification publisher. Each service's own metrics say its p99 is fine.
Metrics are aggregates — they tell you a population is slow, never which call in which request. Logs are per-service and, without a shared key, uncorrelatable across six of them. Tracing is the signal that follows one request through the whole system.
A trace is one request's journey. A span is one operation within it, with a parent, a duration and tags. The tree shows where the time went — and in that example, immediately.
Setup
dependencies {
implementation 'org.springframework.boot:spring-boot-starter-actuator'
implementation 'io.micrometer:micrometer-tracing-bridge-otel'
runtimeOnly 'io.opentelemetry:opentelemetry-exporter-otlp'
}spring:
application:
name: orders # the service name in every trace
management:
tracing:
sampling:
probability: 0.1
otlp:
tracing:
endpoint: http://otel-collector:4318/v1/traces
logging:
pattern:
level: "%5p [${spring.application.name:},%X{traceId:-},%X{spanId:-}]"Micrometer Tracing (which replaced Spring Cloud Sleuth) instruments automatically: incoming HTTP requests, RestClient and WebClient calls, JDBC queries, Kafka and RabbitMQ listeners, scheduled tasks, and @Async methods.
Prefer the OTLP exporter over Zipkin- or Jaeger-native ones. OpenTelemetry is the vendor-neutral standard, and a Collector can fan out to Tempo, Jaeger, Zipkin or a commercial backend without any application change.
How Context Crosses a Boundary
The mechanism is a header. The outbound interceptor adds it; the inbound filter reads it.
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
││ └─ trace id (32 hex) ────────┘ └─ span id ───┘ └ flags
└─ version
That's the W3C Trace Context standard, and the default in current Micrometer Tracing. Older systems use B3 (X-B3-TraceId, X-B3-SpanId, X-B3-Sampled), which Spring still supports:
management:
tracing:
propagation:
type: w3c # or b3, or b3_multiAll services in one system must agree on the propagation format. A service emitting W3C headers calling one that reads only B3 produces a broken trace — the downstream starts a new trace rather than continuing yours, and you get two disconnected fragments instead of one tree. Nothing errors; the traces just don't join, which is easy to misread as missing instrumentation.
Spring can accept multiple formats on the way in while emitting one. When migrating a mixed estate, that's the setting to use.
Where propagation breaks
Automatic instrumentation covers the common paths. It misses:
Manually built clients. A RestClient constructed with RestClient.create() instead of the injected builder has no interceptor — the point the outbound HTTP guide made about always injecting the auto-configured builder.
Raw threads. Context lives in a ThreadLocal, so work handed to an unwrapped executor loses it — the same propagation problem as SecurityContext in phase 4:
@Bean
ThreadPoolTaskExecutor taskExecutor(ObservationRegistry registry) {
var executor = new ThreadPoolTaskExecutor();
executor.setTaskDecorator(new ContextPropagatingTaskDecorator());
executor.initialize();
return executor;
}Messaging, sometimes. Kafka and Rabbit listeners are instrumented, but a custom envelope that doesn't carry trace headers breaks the chain between producer and consumer. Asynchronous flows need the context in the message.
Sampling
Tracing every request is usually unaffordable in volume, storage and backend cost.
Head sampling decides at the first span:
management:
tracing:
sampling:
probability: 0.1 # 10%The decision propagates in the traceparent flags, so every service in a trace makes the same choice — you never get half a trace. That's the main virtue, and the main flaw: 90% of your errors are discarded too.
Tail sampling decides in the Collector, after the whole trace is visible:
processors:
tail_sampling:
decision_wait: 10s
policies:
- name: errors
type: status_code
status_code: { status_codes: [ERROR] }
- name: slow
type: latency
latency: { threshold_ms: 1000 }
- name: sample-the-rest
type: probabilistic
probabilistic: { sampling_percentage: 5 }Keep every errored or slow trace, plus 5% of healthy ones. This is strictly better, because the traces worth having are exactly the rare ones head sampling throws away. The cost is that the Collector must buffer spans until it decides.
Set probability: 1.0 in development and use tail sampling in production. The common mistake is 10% head sampling in production and then being unable to investigate a reported incident — a 90% chance the trace for the request the customer complained about was never recorded.
Custom Spans and Tags
Automatic spans cover I/O. Add spans for meaningful internal work:
@Observed(name = "orders.pricing", contextualName = "calculate-pricing")
public Price calculate(Order order) { }Or manually, with the tags that make a span diagnostic:
public PaymentResult charge(Order order) {
Span span = tracer.nextSpan().name("payment-authorisation").start();
try (var ignored = tracer.withSpan(span)) {
span.tag("order.id", order.getId().toString());
span.tag("payment.provider", "stripe");
span.tag("amount.currency", order.getCurrency());
return gateway.authorise(order);
} catch (Exception ex) {
span.error(ex);
throw ex;
} finally {
span.end();
}
}Span tags tolerate high cardinality. A span is one record, not a time series, so an order ID here is fine — and genuinely useful, because you can search traces by it. The same value as a metric tag would create a time series per order and break your metrics backend, which is the rule from phase 4.
This asymmetry is worth holding onto: high-cardinality identifiers belong in traces and logs, never in metric labels.
Span tags are stored and widely readable, so the same data-protection rules as logs apply: no tokens, passwords, full card numbers, or personal data beyond what you can justify. A span tagged with a customer's email is a personal-data record in your tracing backend, usually with retention nobody has reviewed.
Correlating the Three Signals
The real payoff, and the thing to configure even if you do nothing else.
With traceId in the log pattern and structured JSON logs from the observability guide, each signal reaches the next in one step:
And close the loop with the client, by returning the trace ID as the error reference from the exception handling guide:
@ExceptionHandler(Exception.class)
ProblemDetail onUnexpected(Exception ex) {
String traceId = Optional.ofNullable(tracer.currentSpan())
.map(s -> s.context().traceId())
.orElse("unavailable");
log.error("Unhandled exception [trace={}]", traceId, ex);
ProblemDetail problem = ProblemDetail.forStatusAndDetail(
HttpStatus.INTERNAL_SERVER_ERROR,
"An unexpected error occurred. Quote reference " + traceId + " to support.");
problem.setProperty("reference", traceId);
return problem;
}A support ticket now carries a key that resolves to the complete cross-service history of that one request. This is the highest-leverage line in this phase — it turns "a customer says checkout failed yesterday" from an investigation into a lookup.
Cost and Overhead
Instrumentation overhead is small — a few percent at most, mostly in span creation and export, and the exporter batches asynchronously off the request path.
The real cost is storage and backend. Spans are numerous: one request across six services with database calls can be 30 spans. At 10,000 requests per second, unsampled, that's 300,000 spans per second. Which is why sampling is not optional at scale, and why tail sampling — keeping the valuable 5% rather than a random 10% — is the better economics as well as the better signal.
Grafana Tempo is designed for this: it indexes only trace IDs and stores spans in object storage, making retention cheap when your access pattern is "look up this trace ID" — which, with the correlation above, it is.
Check yourself
A trace for a checkout request shows the gateway and orders service, then stops — the payments service appears as a separate, unconnected trace. Both services have tracing configured and export to the same backend. What is the most likely cause?
The Mental Model, Restated
- Tracing follows one request across services — the signal metrics and per-service logs cannot provide.
- Context travels in a header. W3C
traceparentby default; all services must agree on the format. - Propagation breaks on hand-built clients and raw threads. Inject the builder; decorate executors.
- Head sampling is consistent but discards errors; tail sampling keeps them. Prefer tail.
- High-cardinality tags belong in spans, never in metric labels.
- No secrets or personal data in span tags — same rules as logs.
traceIdin logs is what joins the three signals.- Return the trace ID as the client's error reference.
Phase 7 in Four Sentences
Service discovery replaces addresses with names, but eviction is slow enough that callers still need timeouts and retries — and on Kubernetes the platform may already provide all of it. A gateway centralises TLS, authentication, rate limiting and routing, provided you never block its event loop and never let services trust it as a perimeter. Circuit breakers make failure cheap rather than making it succeed, with the standing caveats that business outcomes are not infrastructure failures and retries must happen at exactly one layer. Centralised configuration solves duplication across services without being a secrets manager, and distributed tracing is what makes the whole thing diagnosable — especially once the trace ID reaches your logs and your clients.
The Roadmap, Finished
Seven phases, and a few ideas that kept recurring:
The container owning instantiation is what makes Spring work. @Transactional, @Cacheable, @PreAuthorize, @Timed, @Async, @CircuitBreaker — all proxies, all silently bypassed by a self-call. That one mechanism explains more Spring behaviour than any other.
Defaults are chosen for getting started, not for production. Eager @ManyToOne, ordinal enums, a 30-second connection timeout, a single-thread scheduler, an unbounded cache, infinite HTTP timeouts, open-in-view, -Xmx at the container limit. Each is one line to fix and each has caused real outages.
Make the implicit explicit. A thread pool was an undeclared concurrency limit; virtual threads remove it, so declare a bulkhead. A fetch type was an undeclared query plan; declare a fetch join. A retry at three layers is an undeclared multiplication; pick one.
Boundaries are where the work is. DTOs at the API edge, validation at the boundary and invariants in the domain, transactions in the service layer, authentication at the edge and authorisation in the service. Most of the bugs in this roadmap came from a responsibility sitting one layer away from where it belonged.
Observability is a design concern, not an afterthought. The trace ID in a log line, the error reference handed to a client, a bounded metric label, an alert on a symptom — each is a small decision made while writing the code, and together they decide whether an incident takes ten minutes or a day.
From here: the system design track for the architecture these services sit inside, Docker and Kubernetes for the platform beneath them, and the backend engineer roadmap for the surrounding discipline.