For fifteen years the industry told you that blocking a thread was a sin. Then Java shipped virtual threads, and thread-per-request came back into fashion. Meanwhile Go had been doing cheap goroutines since 2012, and Kotlin had been selling suspending functions as the only correct answer since 2018. In 2026 all three approaches are mature, load-tested and deeply opinionated — and picking the wrong one for your service shows up as tail latency you cannot explain.

Here is the honest comparison: what each model actually costs, where each one breaks, and which one you should reach for.

TL;DR — Quick Verdict

  • If your codebase is synchronous Java and you simply want more throughput without rewriting to callbacks or reactive streams, use virtual threads. It is the smallest possible change for the largest possible win.
  • If you are building a new network service from scratch and want the simplest mental model with the best observability, use Go goroutines. Channels plus context remain the least surprising concurrency toolkit in production.
  • If you already live on the JVM with Kotlin, or you need composable cancellation and streaming as first-class citizens, use Kotlin coroutines. suspend is a language feature, not a library trick, and it composes better than either alternative.

Everything below is the detail behind those three sentences.

Side-by-Side Comparison

Java Virtual ThreadsGo GoroutinesKotlin Coroutines
RuntimeJVM (JDK 21+, permanent in JEP 444)Go runtime since 1.0 (2012)Kotlin + kotlinx.coroutines 1.11.0
Unit of concurrencyVirtual thread (JVM-scheduled)Goroutine (runtime-scheduled)Coroutine (compiler-transformed state machine)
Cost per unit~hundreds of bytes to a few KB of heap~2 KB initial stackContinuation object per suspension point
Blocking I/OYes — blocking is the pointYes — blocks the goroutine, not the OS threadYes, if you stay on the right dispatcher
CancellationThread interruption / StructuredTaskScopecontext.ContextStructured by default; Job.cancel()
Structured concurrencyStructuredTaskScope (preview track)errgroup.Group, WaitGroupBuilt into coroutineScope { }
Syntax costNone — same Thread APIgo keywordsuspend + compiler plugin
Streaming primitivesBlockingQueue, Flow via Reactive StreamsChannels, selectFlow, Channel
DebuggabilityExcellent — standard thread dumps, JFRGood — goroutine dumps, pprof, traceFair — coroutine dumps exist but stack traces differ
Function coloringNone (no async keyword)NoneYes — suspend infects callers
Best-known failure modePinning (fixed for synchronized in JDK 24)Goroutine leaksScope leaks and blocked dispatcher threads

Decision Matrix

Your situationPickWhy
Spring Boot app on JDK 21+Virtual threadsOne property flip: spring.threads.virtual.enabled=true
Thousands of concurrent HTTP calls from one serviceVirtual threadsScales like async without rewriting business logic
New CLI, agent, or proxy daemonGoroutinesZero ceremony, tiny stacks, excellent runtime tooling
Heavy CPU work in JVM serviceNeitherUse a bounded platform-thread pool; cheap threads do not create CPU
Android / multiplatform shared logicKotlin coroutinesCancellation and lifecycle integration are the whole point
Complex pipelines with backpressureKotlin coroutinesFlow operators compose; virtual threads have no equivalent
You love try/finally and hate callbacksVirtual threads or goroutinesBoth keep straight-line code straight

Java Virtual Threads — Blocking Is Allowed Again

Virtual threads are a JVM-managed unit of concurrency scheduled by the JDK onto a small pool of carrier (platform) threads. The API surface is deliberately boring — Thread.ofVirtual() and Executors.newVirtualThreadPerTaskExecutor():

1
2
3
4
5
6
7
8
9
try (var executor = Executors.newVirtualThreadPerTaskExecutor()) {
    List<Future<String>> futures = urls.stream()
        .map(url -> executor.submit(() -> fetch(url)))
        .toList();

    for (var f : futures) {
        System.out.println(f.get());   // plain blocking get()
    }
}

That code spawns one virtual thread per URL with no reactive plumbing. Ten thousand concurrent requests is unremarkable; a million virtual threads on 8 GB of heap is achievable because each costs a few hundred bytes rather than a 1 MB stack reserved upfront.

Two details matter in production. First, pinning: a virtual thread that blocks inside a synchronized block or a native frame used to pin its carrier thread, which silently destroys throughput. JDK 24’s JEP 491 fixed the synchronized case, so the remaining risk is native code and Object.wait-heavy legacy paths. Turn on JFR’s jdk.VirtualThreadPinned event instead of guessing.

Second, structured concurrency is the idiomatic way to fan out and fail fast:

1
2
3
4
5
6
7
try (var scope = new StructuredTaskScope.ShutdownOnFailure()) {
    Subtask<User> user    = scope.fork(() -> fetchUser(id));
    Subtask<Orders> order = scope.fork(() -> fetchOrders(id));

    scope.join().throwIfFailed();
    return new Profile(user.get(), order.get());
}

If either subtask fails, the other is cancelled automatically and the scope releases every thread it created. That is the same guarantee Kotlin has had from day one, arriving in Java as an opt-in API on the preview track rather than a language feature.

Where it wins: migrating existing blocking code at nearly zero rewrite cost. Where it loses: no streaming-backpressure story (you reach for Flow/Reactive Streams anyway), no language-level cancellation, and CPU-bound work gains nothing.

Go Goroutines — The Boring Baseline That Everyone Copies

Go’s model is a function call prefixed with go and a runtime scheduler that multiplexes thousands of goroutines onto GOMAXPROCS OS threads (default: logical CPU count since Go 1.5). Goroutines start with a ~2 KB stack that grows in place, far cheaper than an OS thread but not as cheap as a virtual thread’s heap object.

Communication is explicit through channels, and cancellation is explicit through context:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
func gather(ctx context.Context, ids []string) ([]Item, error) {
    g, ctx := errgroup.WithContext(ctx)
    results := make([]Item, len(ids))

    for i, id := range ids {
        i, id := i, id
        g.Go(func() error {
            item, err := fetch(ctx, id)
            if err != nil {
                return fmt.Errorf("fetch %s: %w", id, err)
            }
            results[i] = item
            return nil
        })
    }
    return results, g.Wait()   // first error cancels the rest
}

That pattern — errgroup.WithContext, per-index capture, g.Wait() — is the Go equivalent of StructuredTaskScope. The difference is that it is a library convention rather than something the runtime enforces; nothing stops a developer from launching a bare go func() with no context and leaking it forever.

Go’s runtime earned its reputation through iteration: asynchronous preemption arrived in Go 1.14, eliminating the “tight loop starves the scheduler” class of bugs, and the runtime ships pprof, execution tracing, and goroutine dumps (SIGQUIT) that make production debugging genuinely tractable.

Where it wins: the simplest correct model, best-in-class diagnostics, and stacks that stay small even under abusive concurrency. Where it loses: no scheduler-level work stealing between blocking syscalls the way the JVM does with carrier threads, unbuffered channel deadlocks are still a rite of passage, and goroutine leaks have no automatic safety net.

Kotlin Coroutines — Concurrency as a Language Feature

Kotlin’s approach is the most invasive and the most powerful: suspend functions are compiled into state machines, so a “suspension” is a return, not a blocked thread. kotlinx.coroutines 1.11.0 is the runtime, and structured concurrency is not an API you opt into — it is how scopes work by construction:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
suspend fun loadProfile(id: String): Profile = coroutineScope {
    val user   = async { fetchUser(id) }
    val orders = async { fetchOrders(id) }

    Profile(user.await(), orders.await())   // failure cancels the sibling
}

// Bridging JVM virtual threads when an API is blocking-only:
val vThreads = Executors.newVirtualThreadPerTaskExecutor().asCoroutineDispatcher()

val html = withContext(vThreads) { blockingHttpClient.fetch(url) }

Two properties make this hard to beat. Cancellation is cooperative and free — job.cancel() propagates down the tree, so a client disconnecting cancels the database query. And Flow gives you a composable abstraction for streams with operators like debounce, buffer, flatMapLatest and retryWhen, which the other two models simply do not have.

The costs are real. Function coloring is the biggest: the moment one function is suspend, every caller must be too, and the boundary between blocking and suspending code becomes an architectural decision. Every blocking call must be pushed onto Dispatchers.IO (or the virtual-thread dispatcher above), because blocking on the wrong dispatcher starves a shared pool. Diagnosing a suspended stack trace is also an acquired taste — you get coroutine names rather than a clean call stack.

Where it wins: Android, multiplatform, and any pipeline that needs cancellation plus backpressure plus retries in a few lines. Where it loses: incremental adoption — you do not bolt suspend onto an existing JVM service casually, you design for it.

Pitfalls and Migration Traps

  • Virtual threads are not faster CPU. They accelerate waiting. If your bottleneck is JSON serialization or encryption, a thousand virtual threads just contend harder for the same cores.
  • Semaphore instead of thread-pool limits. The classic migration bug: a service that limited concurrency by sizing a fixed thread pool now has unbounded virtual threads and hammers a downstream database. Replace the pool with an explicit Semaphore — typically around 2× your database connection count.
  • Watch for pinning regressions after library upgrades. JFR’s pinned-thread event is the only reliable detector. Re-check after major framework upgrades; a single blocking native call in a hot path is enough to erase the gains.
  • Always pass a context in Go. A go func() without ctx is a future leak. Lint for it; contextcheck and revive catch most cases.
  • Never block on a coroutine dispatcher. This includes Thread.sleep, JDBC, and legacy synchronized file I/O. Wrap them in withContext(Dispatchers.IO) or the virtual-thread dispatcher, or you will see mysterious latency spikes whenever the pool saturates.
  • Structured concurrency changes error semantics. With coroutineScope, StructuredTaskScope or errgroup, one failing child cancels siblings. Code that relied on “collect every result regardless of failure” needs a supervisor-style scope (supervisorScope) or an explicit non-cancelling policy.
  • Keep an escape hatch for CPU work. All three models shine at I/O concurrency. For parallel number-crunching, use bounded pools or work-stealing parallelism (ForkJoinPool on the JVM, worker pools in Go) and leave the lightweight units for waiting.

If you want to go deeper on the primitives underneath these runtimes, our comparison of coroutine libraries such as libaco, greenlet and libco explains what a continuation actually costs, the async I/O runtime comparison covering libuv, Boost.Asio and Tokio shows how event loops are structured below the language level, and for the opposite style of concurrency our actor model frameworks round-up covers message-passing systems in depth.

Frequently Asked Questions

Are Java virtual threads faster than goroutines?

Neither is universally faster — they optimize the same thing (cheap concurrency for I/O-bound work) with different constant factors. Goroutines start around 2 KB and grow; virtual threads are heap-allocated continuation objects whose footprint is smaller per idle task. Under identical workloads Go’s scheduler normally has lower overhead because it also owns the network poller, while virtual threads win whenever the surrounding ecosystem is already JVM code. Benchmark your service, not a microbenchmark.

Can I use virtual threads in Spring Boot today?

Yes. On JDK 21 or newer, setting spring.threads.virtual.enabled=true switches Tomcat’s request handling to virtual threads. The main follow-up work is auditing connection pools and any synchronized-heavy libraries, since concurrency limits that used to be implicit in the thread pool must become explicit semaphores.

Do Kotlin coroutines work with virtual threads?

They complement each other. kotlinx.coroutines 1.11.0 can run on a dispatcher backed by Executors.newVirtualThreadPerTaskExecutor(), which is the standard bridge for blocking third-party clients. You keep coroutine cancellation and Flow composition while the actual blocking happens on a virtual thread instead of consuming a shared pool thread.

Why does everyone complain about “function coloring” in Kotlin?

Because a suspend function can only be called from another suspending context. That propagates upward through your call graph: repositories, services and controllers all become suspend. Java virtual threads and Go goroutines avoid coloring by letting blocking code stay blocking — which is exactly why they are easier to retrofit into existing codebases.

Which model should a new backend service use in 2026?

If you are choosing the language at the same time, pick Go for maximum operational simplicity, or Kotlin on the JVM if you need streaming composition and multiplatform reach. If the codebase already exists in Java, adopt virtual threads and structured concurrency incrementally — that migration is measured in days, not quarters.


💰 想测试你的市场判断力?我用 Polymarket 做预测市场交易——这是全球最大的预测市场平台,从大选结果到技术监管时间线,什么都可以押注。和赌博不同,这是真正的信息市场:你懂的信息越多,胜率越高。我靠预测技术相关事件的走向已经赚了不少。用我的邀请链接注册:Polymarket.com