You sized a thread pool, tuned it for peak traffic, then watched latency climb when every request blocked on the database. Adding more platform threads helped until memory and context switching did not.
Virtual threads are the language’s answer for that shape of concurrency: keep writing simple blocking code, and let the JVM multiplex many tasks onto a small set of OS threads.
A virtual thread is a lightweight thread scheduled by the JVM, not by the operating system. When it blocks on I/O, the JVM can unmount it from its carrier and run something else. Reach for virtual threads when work is mostly waiting — HTTP calls, JDBC, file I/O — not when every task burns CPU without pause.
The problem virtual threads solve
Classic server code often looks like “one request, one thread.” That model is easy to reason about, but each platform thread costs roughly a megabyte of stack plus OS bookkeeping. A fixed pool of a few hundred threads becomes the ceiling for concurrent blocked work.
Before virtual threads, teams either:
- Tuned larger pools and hoped memory held, or
- Rewrote I/O in async / reactive style so a few threads could juggle many connections.
Virtual threads restore the simple model at scale: start a thread per task, write blocking calls, and let the runtime handle the multiplexing.
// Platform-thread pool: concurrency capped by pool size
try (var executor = Executors.newFixedThreadPool(200)) {
for (int i = 0; i < 10_000; i++) {
executor.submit(() -> fetchOrder(i)); // blocks a scarce OS thread
}
}
// Virtual threads: one task per virtual thread, cheap to create
try (var executor = Executors.newVirtualThreadPerTaskExecutor()) {
for (int i = 0; i < 10_000; i++) {
executor.submit(() -> fetchOrder(i)); // blocks without holding an OS thread
}
}
Same blocking fetchOrder. Different cost when thousands wait on the network at once.
When virtual threads shipped
| Release | Status | Spec |
|---|---|---|
| Java 19 | First preview | JEP 425 |
| Java 20 | Second preview | JEP 436 |
| Java 21 LTS | Standard feature | JEP 444 |
| Java 24 | Synchronize without pinning (preview / incremental) | JEP 491 |
| Java 25 LTS | Pinning fix + matured Loom story for production | JEP 491 and related LTS packaging |
Use Java 21+ for virtual threads with no preview flag. Prefer Java 25 LTS when you can: carrier pinning inside synchronized is far less painful than on early 21–23 deployments.
Mental model
Think in two layers:
| Kind | Who schedules it | Cost | Best for |
|---|---|---|---|
| Platform thread | OS | Expensive (stack + kernel) | CPU-bound work, small pools, native code boundaries |
| Virtual thread | JVM on a pool of carriers | Cheap (many thousands–millions) | Blocking I/O, request-per-task servers |
- A virtual thread runs on a carrier (a platform thread) while it is executing Java code.
- On a blocking JDK call the JVM can unmount the virtual thread and reuse the carrier.
- When the I/O completes, the virtual thread is mounted again (possibly on a different carrier).
You do not manage carriers yourself. You start virtual threads (or submit to a virtual-thread-per-task executor) and write ordinary sequential code.
Basics in code
Start one virtual thread:
Thread t = Thread.startVirtualThread(() -> {
System.out.println("running on " + Thread.currentThread());
});
t.join();
Or use the builder when you need a name or to inspect properties:
Thread worker = Thread.ofVirtual()
.name("order-fetch-", 0)
.start(() -> fetchOrder(42));
System.out.println(worker.isVirtual()); // true
worker.join();
For fan-out work, prefer an executor that creates a new virtual thread per task and shuts down cleanly with try-with-resources:
List<Callable<Order>> tasks = orderIds.stream()
.map(id -> (Callable<Order>) () -> fetchOrder(id))
.toList();
try (var executor = Executors.newVirtualThreadPerTaskExecutor()) {
List<Future<Order>> futures = executor.invokeAll(tasks);
for (Future<Order> future : futures) {
process(future.get());
}
}
invokeAll still waits for every task. The win is that thousands of those tasks can sit blocked on I/O without needing thousands of OS threads.
Use cases
1. Throughput for blocking HTTP and JDBC
Servlet / Spring MVC style “thread per request” becomes viable again at high concurrency. Each request can call other services and the database with blocking APIs; virtual threads absorb the wait time.
public OrderDetail loadOrderDetail(String orderId) {
// Each call may block — fine on a virtual thread
Order order = orderClient.get(orderId);
Customer customer = customerClient.get(order.customerId());
List<LineItem> items = inventoryClient.listItems(orderId);
return new OrderDetail(order, customer, items);
}
Run loadOrderDetail on a virtual thread (or behind a virtual-thread executor / server config) and you keep readable sequential code instead of a reactive chain for the same I/O.
2. Fan-out / fan-in without reactive ceremony
Fetch many independent resources in parallel, then combine:
public record Prices(BigDecimal spot, BigDecimal forward, BigDecimal vol) {}
Prices fetchPrices(String symbol) throws Exception {
try (var executor = Executors.newVirtualThreadPerTaskExecutor()) {
Future<BigDecimal> spot = executor.submit(() -> spotService.get(symbol));
Future<BigDecimal> forward = executor.submit(() -> forwardService.get(symbol));
Future<BigDecimal> vol = executor.submit(() -> volService.get(symbol));
return new Prices(spot.get(), forward.get(), vol.get());
}
}
For richer cancellation and “fail one, cancel siblings” semantics, look next at structured concurrency (StructuredTaskScope) on modern JDKs — a natural follow-up to this post.
3. Replacing oversized application pools
If a service exists only to absorb blocking calls (newFixedThreadPool(400) around HTTP clients), swap in newVirtualThreadPerTaskExecutor() and size downstream limits instead: connection pools, rate limits, bulkheads. Virtual threads move the bottleneck to where it belongs — shared resources — not artificial thread caps.
4. Spring Boot teaser
Spring Boot 3.2+ can run Tomcat/Jetty request handling on virtual threads:
spring:
threads:
virtual:
enabled: true
Keep this post language-first: turn the flag on only after you understand pinning, pool limits, and observability. The Boot-side deep dive is Spring Boot Virtual Threads.
When they win vs when they don’t
| Workload | Prefer |
|---|---|
| Many concurrent blocking I/O calls | Virtual threads |
| Short CPU bursts between waits | Virtual threads still fine |
| Tight CPU-bound loops (crypto, codecs, number crunching) | Platform threads / sized ForkJoinPool |
| Heavy use of native code that blocks carriers | Measure carefully; may need isolation |
Virtual threads do not add CPU cores. If eight cores are saturated with pure computation, a million virtual threads will not finish faster — they will only queue more work.
Gotchas
Pinning (and what Java 25 improved)
A virtual thread is pinned to its carrier when the JVM cannot unmount it during blocking — historically common inside synchronized blocks that perform blocking I/O, and with some native frames.
On Java 21–23, prefer ReentrantLock (or shrink synchronized regions) when a critical section might block. On Java 24/25, JEP 491 largely removes pinning for synchronized, which is why 25 LTS is the comfortable production target for Loom-heavy services.
The lock-choice companion is Locks After Loom: when synchronized still wins, when ReentrantLock still wins, and what still pins after JEP 491.
Keep the JDK Flight Recorder pinning events in your toolbox when you adopt virtual threads on older 21.x lines.
ThreadLocal sprawl
ThreadLocal on millions of short-lived virtual threads can allocate a lot of garbage and surprise caches that assumed a small thread pool. Prefer explicit arguments, scoped values (finalized on Java 25), or request context objects passed down the call stack.
Still bound shared resources
Virtual threads make it easy to open 50,000 concurrent DB calls. Your connection pool of 20 will not like that. Cap concurrency at the pool, semaphore, or resilience layer — not by going back to a tiny thread pool as the only brake.
Mixing with old thread-pool assumptions
Code that caches expensive objects in a static “one per worker thread” map, or that relies on thread interruption quirks of a specific pool, needs a second look. Treat “thread identity is scarce and long-lived” as an outdated assumption.
Cheat sheet
Thread.startVirtualThread(runnable)
Thread.ofVirtual().name(...).start(runnable)
Executors.newVirtualThreadPerTaskExecutor()
Java 21+ (JEP 444); prefer Java 25 LTS for pinning fixes
Good: blocking HTTP, JDBC, fan-out I/O, request-per-task servers
Avoid as a CPU booster: pure compute still needs cores
Watch: ThreadLocal, connection pools, synchronized+block on older JDKs
Spring Boot: spring.threads.virtual.enabled=true — see Boot virtual threads post
Do:
- Write straightforward blocking code on virtual threads for I/O-bound services.
- Bound connection pools and downstream rate limits explicitly.
- Prefer Java 25 LTS when virtual threads are central to your runtime model.
Don’t:
- Expect virtual threads to speed up CPU-bound work.
- Unleash unbounded fan-out against a tiny JDBC pool.
- Stuff large ThreadLocal state into short-lived virtual threads.
Pros and cons
Pros
- Cheap concurrency for blocking I/O (HTTP, JDBC, fan-out) without a reactive rewrite
- Simple “one task, one thread” model via
Thread.startVirtualThread/Executors.newVirtualThreadPerTaskExecutor() - Fits request-per-task servers; Spring Boot can enable with
spring.threads.virtual.enabled - Stronger production story on Java 25 LTS (JEP 491 pinning improvements)
Cons
- No win for CPU-bound work — virtual threads do not add cores
- Easy to overwhelm shared resources (connection pools, rate limits) without explicit bounds
ThreadLocalcost and sprawl on short-lived virtual threads- Pinning /
synchronized+ blocking still a concern on older Java 21–23 lines
Wrap-up
Virtual threads make the “one task, one thread” style affordable again for I/O-bound Java. They became a standard feature in Java 21 (JEP 444) and are a stronger default on Java 25 LTS, where synchronized pinning is far less of a migration tax.
Start with newVirtualThreadPerTaskExecutor() around blocking fan-out. Keep CPU-heavy work on ordinary pools. Size databases and HTTP clients on purpose. When you want structured cancellation and safer context propagation, step up to structured concurrency (still preview) and scoped values — the next chapter in the same Loom story.