06-production-cloud-native

Virtual Threads: One Property, Several Consequences

What changes when Tomcat stops using platform threads, why pinning matters, and every pool you must re-examine before enabling it.

October 9, 2026
spring-bootvirtual-threadsloomjava-21concurrencyperformancetomcat

The Constraint Being Removed

On the servlet stack, each request holds one thread for its entire duration — including every millisecond blocked on a database query or an outbound HTTP call. An OS thread costs around 1 MB of stack and a context switch to schedule, so Tomcat caps its pool (200 by default).

For an I/O-bound service that ceiling is almost pure waste. A request spending 95% of its time waiting holds a thread that could serve 20 others. At 200 threads you are limited not by CPU or by the database, but by an accounting artefact.

Virtual threads are JVM-managed threads, cheap enough to create by the million. When one blocks, the JVM unmounts it from its carrier OS thread and runs something else. Blocking stops being expensive.

yaml
spring:
  threads:
    virtual:
      enabled: true

That's the whole configuration. On Java 21+, Spring Boot switches Tomcat's request handling, the @Async executor and scheduled task execution to virtual threads.

What Actually Changes

Concurrency is no longer bounded by a thread pool. Thousands of requests can be in flight.

Blocking code stays blocking, and that's the point. This is the key difference from WebFlux: you keep writing ordinary sequential code with JdbcClient and RestClient, no Mono, no Flux, no reactive operators. Virtual threads deliver scalability without the programming-model change.

Thread-per-request remains the mental model. Stack traces are readable, debuggers work, ThreadLocal works, and your profiler shows what you expect.

💡

This is why virtual threads are the pragmatic answer to most "should we rewrite this reactive?" questions, as the Spring MVC guide noted. WebFlux remains right for streaming, server-sent events and very high connection counts with minimal per-connection work. For an ordinary blocking service, virtual threads remove the scalability argument at a fraction of the complexity.

The Bottleneck Moves

Here is the part that catches teams out. Virtual threads don't make your system faster; they remove one limit and expose the next.

That 200-thread pool was doing unacknowledged work: it was a concurrency limit on everything downstream. At most 200 requests could be querying your database or calling a payment provider. Remove it, and 5,000 requests arrive at a connection pool of 10.

🚨

Enabling virtual threads without revisiting your limits can make things worse. Three specific risks:

The connection pool becomes the queue. Thousands of virtual threads wait on 10 connections, and with Hikari's 30-second default connection-timeout they wait a long time while holding request state. Lower that timeout so you shed load fast rather than queueing deeply.

You become the load spike. A downstream service sized for your 200 concurrent calls now receives thousands — you have moved your outage to your dependency.

Memory, not threads, becomes the limit. Virtual threads are cheap; the request objects, buffers and entities they hold are not.

So enabling virtual threads is a prompt to reinstate the limits deliberately, where they belong:

java
// an explicit concurrency limit on a downstream call
private final Semaphore inventoryCalls = new Semaphore(50);
 
public StockLevel stockFor(String sku) {
    inventoryCalls.acquire();
    try {
        return client.get().uri("/stock/{sku}", sku).retrieve().body(StockLevel.class);
    } finally {
        inventoryCalls.release();
    }
}

Spring Framework 7's @ConcurrencyLimit does this declaratively, and a Resilience4j bulkhead — covered in phase 7 — is the fuller version. The principle: a limit that was implicit must become explicit.

Pinning

A virtual thread normally unmounts when it blocks. In two situations it cannot, and it pins its carrier OS thread instead:

  1. Inside a synchronized block or method
  2. Inside a native frame (JNI)

A pinned virtual thread blocking holds an OS thread, and with a default carrier pool sized to your core count, a handful of pinned blocking threads can stall the scheduler.

java
// pins the carrier if the call blocks
public synchronized Token refresh() {
    return authClient.fetchToken();          // network call inside synchronized
}
 
// does not pin
private final ReentrantLock lock = new ReentrantLock();
 
public Token refresh() {
    lock.lock();
    try {
        return authClient.fetchToken();
    } finally {
        lock.unlock();
    }
}

ReentrantLock is virtual-thread aware: a thread blocking on it unmounts properly.

⚠️

Java 24 (JEP 491) largely fixed pinning on synchronized, so on a current JDK this is much less of a concern than when virtual threads shipped. Two caveats remain. Native frames still pin. And third-party libraries — older JDBC drivers especially — may hold synchronized around blocking I/O in ways that behave differently across JDK versions.

The practical advice: on Java 21–23, replace synchronized-around-blocking-I/O with ReentrantLock. On 24+, it matters much less, but prefer ReentrantLock in new code anyway for its clearer semantics.

Detect pinning while testing:

bash
java -Djdk.tracePinnedThreads=full -jar app.jar

This prints a stack trace whenever a virtual thread pins while blocking — the fastest way to find the offending library. It's a diagnostic flag, not a production one.

ThreadLocal Still Works, With a Caveat

ThreadLocal works on virtual threads, which is why SecurityContextHolder and MDC logging context continue to function.

The caveat is scale: a ThreadLocal holding something sizeable, multiplied by 10,000 virtual threads, is 10,000 copies. A pattern that was fine with 200 threads can become a memory problem.

There's an upside too. The dangerous case from the filter chain guide — a pooled thread retaining state across requests — doesn't apply, since each request gets a fresh virtual thread that is then discarded. Context leakage between users through thread reuse stops being possible.

Scoped values (ScopedValue, finalised in Java 25) are the forward-looking replacement: immutable, explicitly bounded in scope, and cheaper across many threads.

Where Virtual Threads Don't Help

CPU-bound work. Virtual threads solve blocking. A request doing heavy computation needs a core, and you cannot have more useful parallelism than you have cores. Running image processing on 10,000 virtual threads just thrashes.

Already-reactive applications. WebFlux is non-blocking already; there's nothing to unblock.

Work bounded by something else. If your database maxes out at 200 concurrent queries, allowing 10,000 concurrent requests just moves the queue.

✅

Keep a bounded platform-thread executor for CPU-bound work, and use virtual threads for I/O:

java
@Bean("cpuBound")
ExecutorService cpuBoundExecutor() {
    return Executors.newFixedThreadPool(Runtime.getRuntime().availableProcessors());
}

Virtual threads are not a universal replacement for thread pools — they are the right tool for waiting, and a fixed pool is still the right tool for computing.

Check yourself

A service enabling spring.threads.virtual.enabled=true sees p99 latency get worse under load, and connection-pool timeout errors appear. What happened?

An Adoption Checklist

  1. Confirm Java 21+ (24+ preferred, for the pinning fix).
  2. Audit synchronized around blocking I/O and convert to ReentrantLock.
  3. Run with -Djdk.tracePinnedThreads=full under load in staging.
  4. Revisit every pool. Connection pool size and connection-timeout, @Async executors, HTTP client maxConnPerRoute.
  5. Add explicit concurrency limits to downstream calls that were implicitly limited.
  6. Check your downstream dependencies' capacity. You are about to send them more.
  7. Load test before and after. Compare p99 and error rate, not just throughput.
  8. Watch memory, which is now the binding constraint.

Step 4 is the one that gets skipped, and the one that causes the regression in the quiz above.

The Mental Model, Restated

  1. One property, and Tomcat, @Async and scheduling switch over.
  2. Blocking code stays blocking — that's the appeal versus reactive.
  3. The bottleneck moves. The thread pool was an implicit limit on everything downstream.
  4. Reinstate limits explicitly — semaphore, @ConcurrencyLimit, or a bulkhead.
  5. Pinning on synchronized matters on Java 21–23, largely fixed in 24+. Prefer ReentrantLock.
  6. ThreadLocal works, and per-request virtual threads remove the cross-request leakage risk.
  7. No help for CPU-bound work. Keep a fixed pool for that.
  8. Load test before and after. Throughput alone will mislead you.

What's Next

Virtual threads make it possible to run far more concurrent work, which makes it far more important to know what the system is actually doing. The next guide assembles the full observability stack around the Micrometer and Actuator foundations from phase 4: Prometheus and Grafana for metrics, structured JSON logs shipped to a searchable store, and traces correlated with both.