Language-Level Concurrency and Multithreading Questions
Per-language and per-runtime concurrency: the threading and async APIs each language and platform provides (goroutines, channels, worker pools and pipelines in Go, threads, executors, ThreadPoolExecutor tuning and CompletableFuture in Java, Kotlin coroutines, dispatchers, Flow and structured concurrency, Swift GCD with queues and QoS, DispatchGroup, OperationQueue, actors and async/await, Android Handler/Looper, HandlerThread and thread pools, Flutter isolates, C++ std::thread and atomics, Python threads versus multiprocessing versus asyncio tasks, gather and graceful shutdown), each language's memory model and visibility guarantees (Java happens-before and volatile, C++11 memory orders such as relaxed, acquire and release, Objective-C and Swift atomic versus nonatomic), mobile main-thread rules and JNI thread attachment, cancellation and shutdown idioms, and the idioms for coordinating shared state safely in that language, including thread-safe caches, singletons and bounded queues. Also covers reproducing, testing and diagnosing races and deadlocks in a specific language or app, and migrating callback, GCD or thread-pool code to structured concurrency. Covers choosing and using a language's concurrency primitives correctly. Boundary: general synchronization theory, deadlock and lock-free algorithm internals, OS scheduling, database isolation levels, and callback or event-loop architecture are covered elsewhere.
Describe Kotlin coroutines' Dispatchers.IO, Dispatchers.Main, and Dispatchers.Default. For an app that downloads images and decodes them, specify which dispatcher you would use for network download, for image decoding, and for updating the UI, and justify each choice based on blocking vs CPU-bound characteristics.
Sample Answer
Direct answer
A coroutine is a lightweight unit of work that can pause without holding a thread and resume later; a function marked suspend is one that may pause this way, and while it is paused its thread is free to run other work. A dispatcher decides which thread (or threads) a coroutine runs on. For an app that downloads and decodes images: run the network download on Dispatchers.IO (blocking work that mostly waits), the image decoding on Dispatchers.Default (CPU-bound work), and the UI update on Dispatchers.Main (the single UI thread). Switch between them with withContext(...) (launch and async start new coroutines inside a scope, the object that owns them and cancels them together, such as viewModelScope on Android, which is cancelled when the screen's ViewModel is cleared; withContext instead moves the current coroutine to another dispatcher for one block), which suspends the caller, runs the block on the new dispatcher, and returns the result. The rule behind all three choices: put work that waits on IO, work that computes on Default, and work that touches UI objects on Main.
The three dispatchers
| Dispatcher | Backing threads | Meant for | Wrong use |
|---|---|---|---|
Dispatchers.Default | Shared pool; the maximum number of threads equals the number of CPU cores, with a minimum of two | CPU-bound work: decoding, parsing, sorting, hashing. It is also what launch and async use when neither you nor the enclosing scope supplies a dispatcher (a scope such as viewModelScope supplies Dispatchers.Main.immediate, so launches inside it start on the main thread) | Blocking calls: a thread sleeping on a socket is a core's worth of capacity wasted |
Dispatchers.IO | Shared pool that creates threads on demand; defaults to a limit of 64 threads or the number of cores, whichever is larger | Blocking I/O: network, files, databases | CPU-heavy loops: they would compete for the same threads as the I/O waits |
Dispatchers.Main | One thread, the app's UI thread (on Android it comes from the kotlinx-coroutines-android artifact) | Updating views and UI state; short work only | Anything slow: it blocks drawing and input |
Dispatchers.IO and Dispatchers.Default share threads. The kotlinx.coroutines documentation says that withContext(Dispatchers.IO) while already on Default typically does not cause a real thread switch, because the implementation tries to keep execution on the same thread. So the cost of switching between them is small, and the two are separated by their parallelism limits, not by separate thread pools.
Per-stage choice and justification
- Network download:
Dispatchers.IO. The download is a blocking call. While it waits, the thread does no computation, so you want many of them parked at once.IOallows up to 64 (or the core count if larger), far more thanDefaultcould ever give you. - Image decoding:
Dispatchers.Default. Decoding keeps a core busy the whole time. More threads than cores adds only context switching, so the pool sized to the number of cores is the right fit. Running decoding onIOwould let 64 decodes fight over a handful of cores. - Updating the UI:
Dispatchers.Main. Views may only be touched from the UI thread, and the code there should be tiny (set the bitmap), otherwise the screen stutters.
suspend fun loadImage(url: String): Bitmap {
val bytes = withContext(Dispatchers.IO) { download(url) } // waits on the network
return withContext(Dispatchers.Default) { decode(bytes) } // computes
}
// in the UI layer, on Dispatchers.Main (for example viewModelScope.launch):
imageView.setImageBitmap(loadImage(url)) // runs back on Main
In numbers: on a machine with 14 cores, Default runs at most 14 threads and IO at most 64 (the larger of 64 and the core count); on a 2-core phone, Default has 2 and IO still has 64. Android's guidance is that suspend functions should be main-safe, meaning safe to call from the main thread: the function that does blocking work is responsible for moving itself off the main thread with withContext, so callers never need to know. It also says not to hardcode dispatchers, but to inject them (as constructor parameters with defaults), so tests can substitute a test dispatcher.
Runnable check on the JVM
The program below (Kotlin 2.0.21, kotlinx-coroutines 1.9.0, run on JDK 21) cannot use the real Dispatchers.Main, which needs an Android or Swing artifact, so a single thread named main-standin plays the UI thread. It runs the three stages in order, then tests a claim from the table: how many blocking tasks each dispatcher can hold at the same time. It starts cores + 1 tasks that each wait until all have started; that can only succeed if the dispatcher can run cores + 1 threads at once. Step by step: runBlocking is the entry point that blocks main until the coroutines inside finish (a program needs one to start coroutines from plain code). allBlockedTogether creates a CountDownLatch set to n, a counter that threads wait on until it has been counted down n times. Each of n coroutines, launched on the dispatcher under test, counts the latch down once and then blocks on it (up to 2 seconds) like a thread stuck on a network read. The caller waits up to 1 second for the count to reach zero: it reaches zero only if all n coroutines are running at the same moment. Then cancelAndJoin() cancels each coroutine and waits for it to end. Two small helpers appear in the stage lines: Thread.currentThread().name.substringBefore('-', "?") keeps the text of the thread name before its first dash (DefaultDispatcher-worker-3 becomes DefaultDispatcher, and "?" is the fallback when there is no dash), and the Thread.sleep(50) stands in for a blocking network wait.
import kotlinx.coroutines.*
import java.util.concurrent.CountDownLatch
import java.util.concurrent.Executors
import java.util.concurrent.TimeUnit
fun decode(bytes: ByteArray): Long { // CPU-bound stand-in for image decoding
var h = 0L
repeat(2_000_000) { i -> h = h * 31 + bytes[i % bytes.size] }
return h
}
fun main() = runBlocking {
// Stand-in for Dispatchers.Main (only exists with an Android or Swing artifact): one named thread.
val mainExecutor = Executors.newSingleThreadExecutor { Thread(it, "main-standin") }
val mainThread = mainExecutor.asCoroutineDispatcher()
val stages = mutableListOf<String>()
val bytes = withContext(Dispatchers.IO) { // blocking download
stages += "download on ${Thread.currentThread().name.substringBefore('-', "?")}-pool"
Thread.sleep(50) // blocking call: parks a thread, not CPU
ByteArray(1024) { it.toByte() }
}
val bitmap = withContext(Dispatchers.Default) { // CPU-bound decode
stages += "decode on ${Thread.currentThread().name.substringBefore('-', "?")}-pool"
decode(bytes)
}
withContext(mainThread) { // UI update
stages += "update on ${Thread.currentThread().name}"
}
println(stages)
println("decoded value is deterministic: $bitmap")
// How many blocked threads can each dispatcher hold at once?
val n = Runtime.getRuntime().availableProcessors() + 1
suspend fun allBlockedTogether(d: CoroutineDispatcher): Boolean {
val latch = CountDownLatch(n)
val jobs = List(n) { launch(d) { latch.countDown(); latch.await(2, TimeUnit.SECONDS) } }
val ok = latch.await(1, TimeUnit.SECONDS) // true only if all n were running at once
jobs.forEach { it.cancelAndJoin() }
return ok
}
println("cores+1 = $n blocking tasks at once on Dispatchers.Default: ${allBlockedTogether(Dispatchers.Default)}")
println("cores+1 = $n blocking tasks at once on Dispatchers.IO: ${allBlockedTogether(Dispatchers.IO)}")
mainExecutor.shutdown()
}
[download on DefaultDispatcher-pool, decode on DefaultDispatcher-pool, update on main-standin]
decoded value is deterministic: -3386939593423404480
cores+1 = 15 blocking tasks at once on Dispatchers.Default: false
cores+1 = 15 blocking tasks at once on Dispatchers.IO: true
The output was identical in repeated runs (the container had 14 CPUs, so cores + 1 is 15; on another machine the number differs). Default prints false on any machine with two or more cores, because its limit is the core count and cores + 1 is one more than that. On a one-core machine its limit is the minimum of two, which equals cores + 1, so that line prints true there. IO prints true as long as cores + 1 is at most 64. The stage line shows download and decode both on DefaultDispatcher worker threads (the shared pool: IO and Default threads carry the same name prefix) and the UI step on the dedicated thread. The last two lines are the real difference: Default could not run 15 blocked tasks at once because its limit is the core count, while IO could.
Trade-offs and pitfalls
- Blocking inside
Default. A blocking call onDefaultholds one of only a few threads; enough of them stall all CPU work in the app. The experiment above is that starvation in miniature. withContext(Dispatchers.Main)from an already-main coroutine does no useful work; do not scatter it. Put the switch at the boundary of the function that needs it.- Many concurrent blocking operations for one resource:
Dispatchers.IO.limitedParallelism(n)gives a view with its own limit for that resource (for example 4 connections to one database), and the documentation says such views are not restricted by the 64 limit ofIOitself. - Cancellation.
withContextcreates a child scope: if the calling coroutine is cancelled (for example the screen closes) before the block starts, the block finishes immediately with aCancellationException(the exception Kotlin uses to unwind a cancelled coroutine; it is treated as normal, not as a failure). A blocking call that never checks for cancellation, such as oneThread.sleepor a blocking read, keeps running until it returns, so prefer cancellable APIs. - Do not decode on
Main"because it is only a thumbnail". Many small stalls add up to dropped frames. - Many images at once. Launch one coroutine per image; the dispatchers bound the real parallelism, so 200 queued downloads do not create 200 threads.
Compare Kotlin's synchronized, volatile, AtomicInteger/AtomicReference with coroutine patterns for sharing mutable state. For an Android app, when would each mechanism be appropriate? Discuss visibility and atomicity guarantees under the JVM memory model and expected performance characteristics.
Sample Answer
Direct answer
Every mechanism answers two separate questions: visibility (will another thread see this write?) and atomicity (can two steps be interleaved by another thread?). @Volatile gives visibility and ordering for single reads and writes but no atomicity for read-modify-write. synchronized and AtomicInteger/AtomicReference give both. Mutex gives both for coroutines and may be held across a suspension point (a call to a suspend function, where a coroutine can pause and later resume, possibly on a different thread), which synchronized may not. Confining the state to one coroutine (single owner, others send messages) avoids the question; this is called confinement. Default on Android: confine UI state to the main thread and expose it as a StateFlow; use an atomic for an independent counter or flag; use Mutex when a critical section must call a suspend function; use synchronized for short critical sections in plain non-suspending code.
What each guarantees under the JVM memory model
The Java memory model (which Kotlin on the JVM and on Android inherits) only guarantees that one thread sees another's write if a happens-before edge connects them: a write to a @Volatile field before a later read of it, releasing a monitor (the lock behind synchronized) before the next thread acquires the same monitor, or Thread.start and join. A happens-before edge is a rule that says everything one thread did before the edge is visible to the other thread after it.
The reason a rule is needed: hardware caches are kept coherent, but a core's registers and store buffer are private, and the JIT compiler (the part of the JVM that turns hot code into machine code at run time) and the CPU are free to keep a value in a register or reorder accesses instead of re-reading memory. A thread that spins on a plain flag may therefore keep looking at its own old value for as long as nothing forces a fresh read.
Read the last column as: may I hold this while the coroutine pauses at a suspend call such as delay(10) or a network call? A paused coroutine can resume on another thread, and a lock belongs to a thread, so a thread-owned lock cannot safely be held across the pause.
| Mechanism | Visibility | Atomic read-modify-write | Blocks a thread? | Allowed across a suspension point? |
|---|---|---|---|---|
plain var | no guarantee | no | no | n/a |
@Volatile var | yes (volatile write then read) | no (n++ is read, add, write) | no | yes, but protects nothing compound |
synchronized / ReentrantLock | yes (unlock then lock) | yes, for everything done under the lock | a contending thread waits | no: the compiler rejects a suspension point inside synchronized and inside ReentrantLock.withLock {} (the kotlin.concurrent helper, not Mutex.withLock); a bare lock() / unlock() pair compiles but is still unsafe, because a lock must be released by the thread that took it |
AtomicInteger / AtomicReference | yes | yes, one operation at a time (incrementAndGet, compareAndSet, updateAndGet) | no (retry loop) | yes |
Mutex (kotlinx.coroutines) | yes | yes, for everything done under the lock | suspends the coroutine, not the thread | yes |
| confinement to one coroutine / thread | by message passing or thread ownership | n/a, only one writer | no | yes |
Demonstrations that ran
1. Lost updates. A coroutine is a lightweight task that can pause and resume; Dispatchers.Default is the shared pool of threads that runs them. Eight coroutines on Dispatchers.Default each add 1 to a shared counter 250,000 times (expected total 2,000,000), ten rounds per mechanism, including a single owner coroutine that receives increments through a channel. Compile and run with Kotlin 2.1.20 and kotlinx-coroutines 1.10.1 on JDK 21:
kotlinc -cp kotlinx-coroutines-core-jvm-1.10.1.jar SharedStateDemo.kt -d out then
java -cp out:kotlinx-coroutines-core-jvm-1.10.1.jar:<kotlin-stdlib.jar> SharedStateDemoKt.
import kotlinx.coroutines.*
import kotlinx.coroutines.channels.Channel
import kotlinx.coroutines.sync.Mutex
import kotlinx.coroutines.sync.withLock
import java.util.concurrent.atomic.AtomicInteger
const val WORKERS = 8
const val PER_WORKER = 250_000
const val EXPECTED = WORKERS * PER_WORKER
class PlainCounter { var n = 0; fun inc() { n++ } }
class VolatileCounter { @Volatile var n = 0; fun inc() { n++ } }
class SyncCounter { private var n = 0; fun inc() = synchronized(this) { n++ }; fun get() = synchronized(this) { n } }
class AtomicCounter { val n = AtomicInteger(); fun inc() { n.incrementAndGet() } }
class MutexCounter { private var n = 0; private val m = Mutex()
suspend fun inc() = m.withLock { n++ }; suspend fun get() = m.withLock { n } }
// Run `body` from WORKERS coroutines on Dispatchers.Default (a shared pool of threads).
suspend fun hammer(body: suspend () -> Unit) = coroutineScope {
repeat(WORKERS) { launch(Dispatchers.Default) { repeat(PER_WORKER) { body() } } }
}
fun main() = runBlocking {
val results = linkedMapOf<String, IntArray>()
fun record(name: String, got: Int) { results.getOrPut(name) { IntArray(2) }.also { if (got != EXPECTED) it[0]++; it[1]++ } }
repeat(10) {
val plain = PlainCounter(); hammer { plain.inc() }; record("plain var", plain.n)
val vol = VolatileCounter(); hammer { vol.inc() }; record("@Volatile var", vol.n)
val syn = SyncCounter(); hammer { syn.inc() }; record("synchronized", syn.get())
val atom = AtomicCounter(); hammer { atom.inc() }; record("AtomicInteger", atom.n.get())
val mx = MutexCounter(); hammer { mx.inc() }; record("Mutex", mx.get())
// Confinement: one coroutine owns the state; others send it messages.
var owned = 0
val inbox = Channel<Int>(Channel.UNLIMITED)
val owner = launch(Dispatchers.Default.limitedParallelism(1)) { for (d in inbox) owned += d }
hammer { inbox.send(1) }
inbox.close(); owner.join()
record("confined owner", owned)
}
for ((name, r) in results) println("%-15s wrong in %d of %d rounds".format(name, r[0], r[1]))
}
How to read the code: hammer(body) launches eight coroutines and each runs body 250,000 times; coroutineScope makes hammer return only after all eight finish. record keeps a two-slot tally per mechanism (IntArray(2): slot 0 counts wrong rounds, slot 1 counts rounds) in a map created on first use (getOrPut), and also just runs the tally update on that array. The confinement case starts one owner coroutine on a dispatcher limited to one thread (limitedParallelism(1), meaning at most one coroutine of that view runs at a time). It reads increments from a Channel, a queue that coroutines send messages into; Channel.UNLIMITED means a send never waits for room. Only the owner touches owned, so no lock is needed. The compile commands are needed because the program uses the external coroutines library, so the compiler and the run both need its jar on the classpath alongside the Kotlin standard library.
plain var wrong in <N> of 10 rounds
@Volatile var wrong in <N> of 10 rounds
synchronized wrong in 0 of 10 rounds
AtomicInteger wrong in 0 of 10 rounds
Mutex wrong in 0 of 10 rounds
confined owner wrong in 0 of 10 rounds
The four correct mechanisms (synchronized, AtomicInteger, Mutex and the confined owner) print 0 on every run and every CPU count. For plain var and @Volatile var, <N> is the racy part: with two or more CPUs available it is usually most or all of the ten rounds and varies between runs, and the result to read is that they are wrong at all. Restricted to a single CPU, both can pass every round, because the threads of Dispatchers.Default (which has at least two, even on a one-core machine) are rarely preempted between the read and the write of n++. A clean run on one CPU therefore proves nothing, and the failure needs real parallelism to show reliably. With several CPUs the failures illustrate the table's point that volatile does not make n++ atomic. Against the Android default in the direct answer: the atomic row and confinement (with StateFlow) cover most app code; @Volatile, Mutex and synchronized are for the specific situations listed under "When to use each on Android".
2. Visibility. A worker spins on a flag and the main thread sets it after 300 ms. Run with kotlinc StopFlagDemo.kt -d out2 and java -cp out2:<kotlin-stdlib.jar> StopFlagDemoKt:
import kotlin.concurrent.thread
class PlainFlag { var stop = false }
class VolatileFlag { @Volatile var stop = false }
fun main() {
val plain = PlainFlag()
val t1 = thread(isDaemon = true) { var spins = 0L; while (!plain.stop) { spins++ } }
Thread.sleep(300); plain.stop = true
t1.join(2000)
println("plain flag: worker still running after 2 s = ${t1.isAlive}")
val vol = VolatileFlag()
val t2 = thread(isDaemon = true) { var spins = 0L; while (!vol.stop) { spins++ } }
Thread.sleep(300); vol.stop = true
t2.join(2000)
println("@Volatile flag: worker still running after 2 s = ${t2.isAlive}")
}
plain flag: worker still running after 2 s = true
@Volatile flag: worker still running after 2 s = false
The loop while (!plain.stop) { spins++ } has nothing in it that forces a fresh read, so the worker keeps using the value it first saw. The plain flag's worker usually never notices the write on a HotSpot JVM (the JIT compiler may legally hoist the read out of the loop, meaning read stop once before the loop and reuse it, since nothing orders the two threads), while the volatile flag stopped at once. The memory model permits the plain-flag outcome but does not force it, so another JVM or flags may behave differently, which is why the rule is "use @Volatile or a lock", not "test and see".
3. The compiler guards the suspension rule. This file:
import kotlinx.coroutines.delay
val lock = Any()
var shared = 0
suspend fun update() {
synchronized(lock) {
shared++
delay(10)
}
}
fails to compile with error: the 'delay' suspension point is inside a critical section. A coroutine that suspends inside synchronized could resume on a different thread, and monitors belong to threads, so the language forbids it. Mutex.withLock has no such restriction because a suspended coroutine just releases its thread.
When to use each on Android
- UI state: keep it on the main thread (
Dispatchers.Main) and publish it withStateFlow.StateFlowis a kotlinx.coroutines holder for the current value that collectors observe.MutableStateFlow.update { ... }is an atomic compare-and-set loop (read the current value, compute the new one, store it only if nobody changed it meanwhile, otherwise repeat); the lambda may run more than once, so it must have no side effects. - Independent counter, flag or latest-value holder touched from several threads:
AtomicInteger,AtomicBoolean,AtomicReference. For a "stop" flag where only one thread writes,@Volatileis enough. - A compound invariant over several fields in non-suspending code (a small cache plus its size):
synchronizedorReentrantLock, with every access under the same lock, kept short and never calling out to unknown code. - A critical section that must call
suspendcode (read the token, refresh it once if expired, store it):Mutex. Hold it for as short a time as possible, and never across user interaction. - Mutable state with a natural owner (a repository cache, a download queue): confine it to one coroutine, or to a dispatcher created with
limitedParallelism(1)(a view of a thread pool that runs at most one task at a time), and let other code send requests. This removes the visibility and atomicity questions at the cost of a message hop. Share one such dispatcher instance for the state; two separate single-thread views do not exclude each other.
Performance characteristics (reasoning, not measurement)
- Uncontended
synchronizedand atomics are cheap; the cost appears under contention. A contended lock can block and wake threads, while a contended atomic retries its compare-and-set loop on the same cache line (the small block of memory, typically 64 bytes, that cores pass between each other), and aMutexsuspends and resumes coroutines. - Many threads hammering one hot counter favour a striped structure such as
java.util.concurrent.atomic.LongAdder, which spreads updates over cells and trades exact atomic reads for throughput. - Message-passing confinement adds an allocation and a dispatch per operation, so it suits coarse-grained operations, not per-pixel counters.
- On a phone with a handful of cores the dominant cost is usually holding any lock on the main thread: a lock that a background thread holds for a slow operation freezes the UI. Keep main-thread critical sections to memory-only work.
Pitfalls
@Volatileon a field that is incremented or checked-then-set.- A
synchronizedblock around a call to an unknown callback (deadlock risk). - Using a
Mutexthat is not reentrant from code that already holds it: the secondlock()suspends forever.Mutexis not reentrant. - Mixing mechanisms on the same variable (some writers under a lock, others bare).
Compare the trade-offs of using a serial DispatchQueue, NSLock, and the Swift actor model to protect shared mutable state. For each approach describe performance characteristics, fairness, deadlock risk, composability, and Objective-C interoperability considerations and give an example scenario where it would be the preferred solution.
Sample Answer
Recommendation first
For new Swift code, default to an actor for state that is touched from async code. Use an NSLock (or OSAllocatedUnfairLock) for a tiny, hot critical section that synchronous code or Objective-C must call. Use a serial DispatchQueue when you need strict first-in-first-out ordering of operations, fire-and-forget writes, or you are inside an existing GCD (Grand Central Dispatch, Apple's queue-based concurrency library) code base. Terms: OSAllocatedUnfairLock is a lightweight lock type from Apple's os module. An actor is a Swift reference type whose mutable state can only be touched by one caller at a time, enforced by the compiler; a serial queue runs submitted blocks one after another; NSLock is a mutual-exclusion lock. The comparison uses four words precisely. Fairness: whether waiting callers are served in the order they arrived. Composability: whether two individually correct operations can be combined into a correct larger one. Isolation: the compiler rule that only code running on the actor may touch its state. Reentrancy: while an actor method is suspended at an await, other calls to the same actor may run (shown in the example below).
Comparison
Serial DispatchQueue | NSLock | actor | |
|---|---|---|---|
| Performance | Each operation is a block dispatched to a queue. async returns quickly but costs an allocation and a later hop; sync makes the caller wait. | Lock and unlock are a few instructions when uncontended; the caller does the work on its own thread. | A call from outside is an await: the task may suspend and resume on the actor's executor (the serial context that runs the actor's code). Cheap for coarse operations, wasteful for one-integer updates. |
| Fairness | FIFO: blocks start in the order they were submitted. | No ordering promise; a thread that releases and relocks can beat waiters. | Not guaranteed FIFO. SE-0306 (the Swift Evolution proposal that introduced actors) says tasks awaiting an actor are not guaranteed to run in the order they awaited it, and that the runtime aims to avoid priority inversions, using techniques such as priority escalation. |
| Deadlock risk | sync onto the queue you are already running on traps. A sync cycle between two queues deadlocks. | Not reentrant: locking twice on one thread deadlocks (NSRecursiveLock exists). Two locks taken in different orders deadlock. | No lock-order deadlock inside one actor, and no self-deadlock, but it has the reentrancy problem below. Because an actor gives up its isolation while suspended at an await, two actors that await each other do not deadlock the way two locks do (the two-actor run below completes); a stall needs code that blocks a thread inside an actor, such as DispatchSemaphore.wait(), or an await on something that never completes. |
| Composability | Blocks submitted from many places stay ordered; a queue can be told to run its blocks on another queue, so several queues can share one underlying serial queue. Hard to know from a call site that a queue is involved. | Poor: you cannot combine two locked operations into one without exposing the lock, and the compiler will not check you hold it. | Best compiler support: isolation is part of the type, and misuse is a compile error in Swift 6 mode. But multi-step logic must tolerate interleaving at each await. |
| Objective-C interop | Fully usable. | Fully usable. | Members can be exposed to Objective-C only if they are async or not isolated to the actor (SE-0306). A synchronous actor method such as func balance() -> Int cannot be marked @objc, so Objective-C code cannot call it; if Objective-C must read the state synchronously, that points to a lock or a queue. |
| Preferred for | Ordered event processing, a log or analytics writer, legacy GCD APIs that already take a queue. | A hot counter, a small cache, state read from synchronous or Objective-C code. | A repository, session store or download manager used from async code. |
Evidence from runs
The programs below were compiled in a swift:6.0 container (Swift 6.0.3, aarch64 Linux) with swiftc -O -swift-version 6.
Cost of one operation. This program updates a counter 1,000,000 times from one thread (no contention) through each mechanism and prints the average time per update:
import Foundation
final class LockCounter: @unchecked Sendable { // @unchecked Sendable: the lock makes it safe, the compiler cannot check that
private let lock = NSLock()
private var n = 0
func inc() { lock.lock(); n += 1; lock.unlock() }
}
final class QueueCounter: @unchecked Sendable {
private let q = DispatchQueue(label: "counter")
private var n = 0
func incSync() { q.sync { n += 1 } }
func incAsync() { q.async { self.n += 1 } }
func drain() -> Int { q.sync { n } } // waits until every queued block has run
}
actor ActorCounter {
private var n = 0
func inc() { n += 1 }
}
func ns(_ body: () -> Void) -> Double {
let t0 = DispatchTime.now().uptimeNanoseconds
body()
return Double(DispatchTime.now().uptimeNanoseconds - t0)
}
@main
struct Bench {
static func main() async {
let N = 1_000_000
let lc = LockCounter()
let a = ns { for _ in 0..<N { lc.inc() } }
let qc = QueueCounter()
let b = ns { for _ in 0..<N { qc.incSync() } }
let qa = QueueCounter()
let c = ns { for _ in 0..<N { qa.incAsync() }; _ = qa.drain() }
let ac = ActorCounter()
let t0 = DispatchTime.now().uptimeNanoseconds
for _ in 0..<N { await ac.inc() }
let d = Double(DispatchTime.now().uptimeNanoseconds - t0)
print("NSLock lock/unlock : \(Int(a / Double(N))) ns per update")
print("serial queue, sync : \(Int(b / Double(N))) ns per update")
print("serial queue, async+drain : \(Int(c / Double(N))) ns per update")
print("actor, await from outside : \(Int(d / Double(N))) ns per update")
}
}
Compiled with swiftc -parse-as-library -O -swift-version 6, one run printed:
NSLock lock/unlock : 3 ns per update
serial queue, sync : 34 ns per update
serial queue, async+drain : 261 ns per update
actor, await from outside : 16988 ns per update
The figures move from run to run and with the number of CPUs given to the container, but their sizes stay apart: the lock at a few nanoseconds, the sync queue at a few tens of nanoseconds, the async block at roughly a hundred to a few hundred nanoseconds (it grows with more CPUs), and the actor call at about a microsecond with one CPU and several to over ten microseconds with eight. The ordering is lock cheapest, then a sync queue hop, then an async block, then an await on an actor, which here includes switching the calling task onto the actor's executor on every call. Treat that ordering as what this setup shows, not a law, because it depends on the platform, the runtime version and the workload. These are illustrative figures for a Linux aarch64 container with Swift 6.0.3, not an iPhone: Apple's runtime and scheduler differ, so measure on a target device with Instruments or XCTest measure before a number drives a decision. What carries over is the shape: a lock is a few instructions, a queue adds an allocation or a wait, and an actor await adds a task switch, which is why an actor suits coarse operations and not one-integer updates in a hot loop.
Actor reentrancy. An actor method that checks, then awaits, then acts is not atomic, because other calls run during the suspension:
import Foundation
// Actor reentrancy: state can change across an await inside an actor method.
actor Account {
var balance = 100
func withdraw(_ amount: Int, label: String) async -> Bool {
guard balance >= amount else { return false } // check
try? await Task.sleep(nanoseconds: 10_000_000) // suspension point: other calls run here
balance -= amount // act, on a stale check
print("\(label) withdrew \(amount), balance now \(balance)")
return true
}
}
let acct = Account()
async let a = acct.withdraw(80, label: "T1")
async let b = acct.withdraw(80, label: "T2")
_ = await (a, b)
print("final balance:", await acct.balance)
Run output:
T1 withdrew 80, balance now 20
T2 withdrew 80, balance now -60
final balance: -60
Timeline: T1 enters withdraw, sees balance = 100 >= 80, and suspends at the await (the 10 ms sleep). The actor is free, so T2 enters, also sees balance = 100 >= 80, and suspends. T1 usually resumes first, after its 10 ms, and subtracts: balance 20. T2 resumes and subtracts from the stale check: balance -60. Which call resumes first is not guaranteed (the two printed lines can swap, with the same final result), but the final balance of -60 does not depend on that order, because both calls already passed the check. async let just starts each call as a child task running concurrently, and await (a, b) collects both. The fix: do the check and the mutation with no await between them, or reserve the funds before suspending.
Actors awaiting each other. Two actors that call each other do not deadlock, because an actor is not held while it is suspended at an await:
import Foundation
actor Ping {
var peer: Pong?
func link(_ p: Pong) { peer = p }
func hit(_ depth: Int) async -> Int {
if depth == 0 { return 0 }
return 1 + (await peer!.hit(depth - 1))
}
}
actor Pong {
var peer: Ping?
func link(_ p: Ping) { peer = p }
func hit(_ depth: Int) async -> Int {
if depth == 0 { return 0 }
return 1 + (await peer!.hit(depth - 1))
}
}
let a = Ping(), b = Pong()
await a.link(b); await b.link(a)
print("levels completed:", await a.hit(10))
Compiled with swiftc -O -swift-version 6 in the same container, it prints levels completed: 10 on every run, including with the container limited to one CPU.
Serial queue self-deadlock.
import Foundation
let q = DispatchQueue(label: "state")
q.sync {
print("outer block running")
q.sync { print("inner block") } // sync onto the queue we are already on
}
print("unreachable")
Run in the same container, the process dies with Signal 5 ... Program crashed: System trap ... __DISPATCH_WAIT_FOR_QUEUE__ in libdispatch.so: libdispatch detects the synchronous call onto its own queue and traps rather than hanging.
Which one to pick, and what flips it
- Call sites are
async, state is a model with several fields: actor. It flips to a lock if profiling shows theawaithop in a hot loop, or if a synchronous API (a delegate, aUITableViewDataSourcemethod) must read the state immediately. - Several writers produce events that must be applied in order: serial queue.
- One shared counter or flag read on the main thread without
await: a lock, with the invariant documented next to it.
Design a thread-safe caching layer on Android using Kotlin coroutines and Flow that exposes a Flow of cached values, supports refresh on demand, avoids blocking the main thread, and safely invalidates stale entries. Describe APIs, coroutine scopes, dispatcher choices, and how you handle concurrent refreshes for the same key.
Sample Answer
Direct answer
Build a KeyedCache<K, V> that keeps one MutableStateFlow<Entry?> per key, where an Entry holds the value, the time it was loaded and an invalidated flag. Four rules make it safe and non-blocking:
- Reads never block.
observe(key)returns aFlow<V>backed by theStateFlow; it emits the cached value immediately (even if stale) and the refreshed value when it lands. Collectors run on whatever dispatcher the UI chooses (Main on Android). - Loads run off the main thread. The loader is wrapped in
withContext(ioDispatcher), with the dispatcher injected so tests can replace it. - One load per key at a time.
refresh(key)looks up an in-flightDeferred<V>for that key under aMutexand joins it if present; otherwise it starts one. Concurrent refreshes for the same key (rotation, two screens, a pull-to-refresh during a TTL refresh) share one network call. - The cache owns its loads. Loads are started with
asyncin a cache-owned scope with aSupervisorJob, not in the caller's scope. If the screen that asked leaves, only that caller stops waiting; the load finishes and the next screen gets the value. A failing load fails its awaiters but not the cache.
Terms: a dispatcher decides which thread or thread pool a coroutine runs on: Dispatchers.Main is the UI thread, Dispatchers.IO is a pool sized for blocking work such as network and disk, and Dispatchers.Default is a pool sized to the CPU count for computation. A CoroutineScope owns the coroutines launched in it and cancels them when the scope is cancelled; a SupervisorJob in the scope makes a failing child coroutine not cancel its siblings or the scope. withContext(d) runs a block on dispatcher d and suspends the caller until it returns. A Flow is Kotlin's asynchronous stream of values; a StateFlow is a Flow that always holds a current value and replays it to new collectors; a Mutex is a coroutine-friendly lock whose withLock suspends rather than blocks a thread; a Deferred is a coroutine's future result (scope.async { ... } starts one and await() waits for it); stale means older than the time-to-live (TTL) or explicitly invalidated.
API
| Call | Behaviour |
|---|---|
observe(key): Flow<V> | Emits cached value (stale allowed), triggers a refresh if missing, expired or invalidated, emits each new distinct value. |
suspend refresh(key): V | Loads now, coalesced with any load already running for that key. Throws the loader's exception to every waiter. |
suspend invalidate(key) | Marks the entry stale; keeps serving the old value; reloads at once only if someone is watching. |
| constructor | scope (cache-owned, SupervisorJob), ioDispatcher, ttlNanos, maxEntries, injectable clock, loader: suspend (K) -> V. |
How refresh works for two concurrent callers
refresh is the core of the design; the rest is support. Suppose callers P and Q both call refresh("user:1") at nearly the same moment:
- P enters
lock.withLock. Q suspends, waiting for the lock, without blocking its thread. - P looks in
inFlight["user:1"], finds nothing, startsscope.async { ... }(the load, running on the IO dispatcher viawithContext), stores thatDeferredininFlight, and leaves the lock. - Q gets the lock, finds P's
DeferredininFlight, and takes it instead of starting another. - Both call
load.await(). When the loader returns, theasyncblock writes the newEntryinto theStateFlow(which wakes every collector), then removes the key frominFlightin itsfinally, and both callers receive the same value. - A later
refreshfindsinFlightempty and starts a new load.
The Mutex guards only the check-then-insert on the map, so it is held for microseconds; the slow network work happens outside it.
On the other helpers used in the code: emitAll(flow) forwards every value of another flow from inside a flow { } builder; distinctUntilChanged() drops a value equal to the previous one; watchers is a per-key count of active observe collectors, kept under the lock; it decides which keys reload on invalidate and which are protected from eviction. NonCancellable is a context that lets the cleanup in a finally block run to completion even though the collector is being cancelled.
Complete code with checks
import kotlinx.coroutines.*
import kotlinx.coroutines.flow.*
import kotlinx.coroutines.sync.Mutex
import kotlinx.coroutines.sync.withLock
import java.util.concurrent.atomic.AtomicInteger
class KeyedCache<K : Any, V : Any>(
private val scope: CoroutineScope, // app or repository scope with a SupervisorJob
private val ioDispatcher: CoroutineDispatcher, // injected so tests can replace it
private val ttlNanos: Long,
private val maxEntries: Int = 256,
private val nanoTime: () -> Long = System::nanoTime,
private val loader: suspend (K) -> V,
) {
private data class Entry<V>(val value: V, val loadedAt: Long, val invalidated: Boolean = false)
private class Slot<V> {
val state = MutableStateFlow<Entry<V>?>(null)
var watchers = 0 // guarded by lock
}
private val lock = Mutex()
private val slots = LinkedHashMap<K, Slot<V>>() // guarded by lock
private val inFlight = HashMap<K, Deferred<V>>() // guarded by lock
private suspend fun slotFor(key: K, watch: Boolean = false): Slot<V> = lock.withLock {
val slot = slots.getOrPut(key) { Slot() }
if (watch) slot.watchers++ // counted under the lock, before any eviction can run
evictIdleLocked(keep = key)
slot
}
private suspend fun unwatch(slot: Slot<V>) = withContext(NonCancellable) { lock.withLock { slot.watchers-- } }
private fun evictIdleLocked(keep: K) {
if (slots.size <= maxEntries) return
val it = slots.entries.iterator()
while (it.hasNext() && slots.size > maxEntries) {
val (k, slot) = it.next()
if (k != keep && slot.watchers == 0 && k !in inFlight) it.remove() // never evict a watched key
}
}
private fun isStale(e: Entry<V>?) = e == null || e.invalidated || nanoTime() - e.loadedAt > ttlNanos
/** Cold flow of values. Emits the cached value at once (even if stale), refreshes when stale. */
fun observe(key: K): Flow<V> = flow {
val slot = slotFor(key, watch = true)
try {
if (isStale(slot.state.value)) scope.launch { runCatching { refresh(key) } } // owned by the cache, not the collector
emitAll(slot.state.filterNotNull().map { it.value }.distinctUntilChanged())
} finally {
unwatch(slot)
}
}
/** One load per key at a time: concurrent callers share the same Deferred. */
suspend fun refresh(key: K): V {
val slot = slotFor(key)
val load = lock.withLock {
inFlight[key] ?: scope.async {
try {
val value = withContext(ioDispatcher) { loader(key) }
slot.state.value = Entry(value, nanoTime())
value
} finally {
lock.withLock { inFlight.remove(key) }
}
}.also { inFlight[key] = it }
}
return load.await() // a cancelled caller stops waiting; the shared load keeps running
}
/** Mark stale. Watchers keep seeing the old value until the reload lands. */
suspend fun invalidate(key: K) {
val slot = slotFor(key)
slot.state.update { it?.copy(invalidated = true) }
if (lock.withLock { slot.watchers } > 0) scope.launch { runCatching { refresh(key) } }
}
}
// ------------------------------- checks -------------------------------
@OptIn(DelicateCoroutinesApi::class, ExperimentalCoroutinesApi::class)
fun main() = runBlocking {
val mainSim = newSingleThreadContext("main-sim")
val scope = CoroutineScope(SupervisorJob() + Dispatchers.Default)
val calls = AtomicInteger()
var failNext = false
var clock = 0L
val cache = KeyedCache<String, String>(scope, Dispatchers.IO, ttlNanos = 1_000, nanoTime = { clock }) { key ->
val n = calls.incrementAndGet()
Thread.sleep(200) // blocking work, runs on IO
if (failNext) { failNext = false; throw IllegalStateException("backend down") }
"$key-v$n"
}
// 1. 20 concurrent refreshes of one key cause one load
val results = (1..20).map { async { cache.refresh("user:1") } }.awaitAll()
println("1. loads=${calls.get()}, distinct results=${results.toSet()}")
// 2. a collector on the main thread sees the cached value, then the refreshed one after invalidate
val seen = mutableListOf<String>()
val watcher = launch(mainSim) { cache.observe("user:1").collect { seen += it } }
delay(100)
cache.invalidate("user:1")
delay(500)
println("2. watcher saw $seen, loads=${calls.get()}")
// 3. TTL: advancing the clock past the TTL makes a new observer trigger a reload
clock = 5_000
val third = withTimeout(2000) { cache.observe("user:1").take(2).toList() }
println("3. stale-then-fresh: $third, loads=${calls.get()}")
// 4. failure: refresh throws, last good value is still served, next refresh retries
failNext = true
val err = runCatching { cache.refresh("user:1") }.exceptionOrNull()
val stillServed = cache.observe("user:1").first()
println("4. refresh failed with '${err?.message}', observe still serves '$stillServed'")
println(" retry gives '${cache.refresh("user:1")}', loads=${calls.get()}")
// 5. a collector that leaves does not cancel the shared load
val before = calls.get()
val impatient = launch { cache.refresh("user:2") }
delay(50); impatient.cancel()
delay(400)
println("5. after caller cancelled, user:2 loaded=${cache.observe("user:2").first()}, loads=${calls.get() - before}")
// 6. the main thread keeps ticking while a 200 ms blocking load runs on IO
var last = System.nanoTime(); var worstGapMs = 0L; var ticks = 0
val ticker = launch(mainSim) { while (true) { delay(10); val now = System.nanoTime(); ticks++; worstGapMs = maxOf(worstGapMs, (now - last) / 1_000_000); last = now } }
cache.refresh("user:3")
ticker.cancel(); watcher.cancel()
println("6. ticks during a 200 ms load >= 10: ${ticks >= 10}, worst gap < 100 ms: ${worstGapMs < 100}")
// 7. a size limit never evicts a watched key: four watched keys with maxEntries = 2 all get a value
val small = KeyedCache<Int, String>(scope, Dispatchers.IO, ttlNanos = 1_000_000_000L, maxEntries = 2) { k -> delay(20); "v$k" }
val got = java.util.concurrent.ConcurrentHashMap<Int, String>()
val watchers = (0..3).map { k -> launch { small.observe(k).collect { got[k] = it } } }
delay(500)
println("7. watched keys with a value: ${got.toSortedMap()}")
watchers.forEach { it.cancel() }
scope.cancel(); mainSim.close()
}
Compiled with Kotlin 2.0.21 against kotlinx-coroutines-core 1.9.0 and run on Temurin 21 in a Linux aarch64 container (--cpus=4). A single thread named main-sim stands in for Android's main thread. Re-runs at 1, 2, 4 and 8 CPUs printed the same output:
1. loads=1, distinct results=[user:1-v1]
2. watcher saw [user:1-v1, user:1-v2], loads=2
3. stale-then-fresh: [user:1-v2, user:1-v3], loads=3
4. refresh failed with 'backend down', observe still serves 'user:1-v3'
retry gives 'user:1-v5', loads=5
5. after caller cancelled, user:2 loaded=user:2-v6, loads=1
6. ticks during a 200 ms load >= 10: true, worst gap < 100 ms: true
7. watched keys with a value: {0=v0, 1=v1, 2=v2, 3=v3}
Reading the output: check 1 launches 20 concurrent refreshes and the loader runs once. Check 2 shows a watcher receiving the cached value and then the new one after invalidate. Check 3 advances the injected clock past the TTL; a new observer gets the stale value first, then the fresh one. In check 4 the failed load is the fourth call (so v4 was never produced), the old value is still served, and the retry is the fifth call, hence v5. Check 5 cancels the caller 50 ms into a 200 ms load, and the cache still finishes it: the next read returns user:2-v6 and exactly one load happened. Check 6 prints two booleans, not raw times: the 10 ms ticker on the stand-in main thread ran at least 10 times while a 200 ms blocking load ran (a blocked main thread would barely tick, so this count is what makes the check able to fail), and the worst gap between ticks stayed under 100 ms. Check 7 builds a second cache with maxEntries = 2 and watches four keys at once: all four keys receive their value, because a key with a watcher is never evicted, so the limit is soft while keys are watched.
Choices and what flips them
- Core versus extras. The four rules at the top, the
Mutex-guardedinFlightmap and the cache-owned scope are what the design needs. The items below are refinements you add when a requirement calls for them. - Dispatchers. Collectors on Main, loader on IO (blocking HTTP or disk), no work on Default unless the loader parses large payloads. Keep the
Mutexsections short and free of suspension calls other thanwithLockitself; the code does map lookups and nothing else inside the lock. - Stale policy. Stale-while-revalidate: show old data immediately, replace it when fresh data arrives. If showing old data is unacceptable (balances, permissions), have the UI call
refreshand render a loading state instead of usingobserve's first emission. - TTL is checked on use, not by a timer. Expiry is evaluated when a collector starts and when
refreshorinvalidateis called. A collector that stays subscribed for an hour will not see an automatic reload; if that is needed, add a periodicinvalidate, or collect withrepeatOnLifecycleso re-entering the screen restartsobserveand re-checks staleness. - Same value, no emission.
distinctUntilChangedmeans a refresh that returns an equal value does not wake the UI. Remove it if the UI must react to every refresh (a "last updated" label). - Memory. Entries accumulate per key.
maxEntriesis a soft limit: it removes the oldest-inserted keys that have no watchers and no load in flight, and never the key being requested, so a watched key is never evicted from under its collector. The watcher count is raised under the lock before any eviction can run; a check on theStateFlow'ssubscriptionCountwould treat a key whose collector has not subscribed yet as idle, evict it, and leave that collector waiting on a flow nobody writes to. With more watched keys thanmaxEntriesthe cache simply holds more than the limit until some are released (check 7). Because the order is insertion order, this is not a true LRU; switch the map to an access-ordered one if recency matters. - Process death and persistence. (Process death is Android killing the app's process in the background, which wipes everything in memory.) This cache is in memory. For data that should survive restarts, make the loader read through a Room database (or disk cache) and treat this class as the in-memory layer in front of it.
- Android lifecycle. Collect with
repeatOnLifecycle(STARTED)orcollectAsStateWithLifecycle()so a stopped screen stops collecting.repeatOnLifecycle(STARTED)runs its block whenever the screen is at least started, cancels it when the screen stops, and restarts it on return; the Android documentation describescollectAsStateWithLifecycleas starting collection when the lifecycle isSTARTEDand stopping it whenSTOPPED. The cache's own scope should live as long as the process or the repository singleton, and be cancelled in tests. - Per-key coalescing vs one lock. A single global lock around loads would serialise unrelated keys; the per-key
Deferredmap lets different keys load in parallel and only same-key requests merge.
Known limits of this version: invalidate while a load is already in flight keeps the in-flight result (it may predate the invalidation); if that matters, store a version number in Entry and reload when the arriving result is older than the latest invalidation.
Implement a thread-safe lazily initialized singleton in both Kotlin and Swift. Explain the memory visibility and double-initialization problems each platform's mechanism solves, and why the pattern you chose is safe on that platform.
Sample Answer
Direct answer
Use the language's own lazy-initialisation guarantee instead of writing the locking yourself.
- Kotlin:
object AppConfig { ... }(an object declaration) or aby lazy { ... }property. The Kotlin docs state that "the initialization of an object declaration is thread-safe and done on first access", andlazydefaults toLazyThreadSafetyMode.SYNCHRONIZED, which "uses a lock to ensure that only a single thread can initialize a Lazy instance, and ensures that initialized value is visible by all threads". - Swift:
static let shared = Config(). The Swift book says "Stored type properties are lazily initialized on their first access. They're guaranteed to be initialized only once, even when accessed by multiple threads simultaneously, and they don't need to be marked with the lazy modifier."
Two problems are being solved. Double initialisation is two threads both seeing "not created yet" and both constructing an instance, so callers end up holding different "singletons". Memory visibility is a thread seeing the reference to the instance but not yet seeing the writes that filled in its fields (the compiler or CPU may reorder writes, and without a synchronisation edge another thread is allowed to observe them out of order). The language mechanisms below fix both with one lock or one run-once primitive, which is why they are safe. "Publishing" an object means making its reference reachable by other threads; safe publication means they see it fully built. The pattern to memorise is object in Kotlin and static let in Swift; the other forms below show what those two replace and when you would still need them.
Kotlin (JVM and Android)
import java.util.concurrent.CountDownLatch
import java.util.concurrent.atomic.AtomicInteger
import kotlin.concurrent.thread
val created = AtomicInteger()
class Config {
val values: Map<String, String>
init {
created.incrementAndGet()
Thread.sleep(2) // widen the window for a race
values = mapOf("url" to "https://api.example.com", "retries" to "3")
}
}
// 1. object declaration: the JVM initialises the class once, under a class-init lock.
object AppConfig {
val config = Config()
}
// 2. lazy with the default SYNCHRONIZED mode.
object LazyHolder {
val config: Config by lazy { Config() }
}
// 3. Same as 2 but with no synchronisation: the race, shown on purpose.
object UnsafeHolder {
val config: Config by lazy(LazyThreadSafetyMode.NONE) { Config() }
}
// 4. Hand-written double-checked locking: needs @Volatile.
class DclHolder private constructor() { // never instantiated; it only hosts the companion
companion object { // a companion object is Kotlin's home for static-like members
@Volatile private var instance: Config? = null
fun get(): Config =
instance ?: synchronized(this) {
instance ?: Config().also { instance = it }
}
}
}
fun hammer(name: String, fetch: () -> Config) {
created.set(0)
val start = CountDownLatch(1)
val seen = java.util.Collections.synchronizedSet(HashSet<Int>())
val threads = List(16) {
thread { start.await(); seen.add(System.identityHashCode(fetch())) }
}
start.countDown()
threads.forEach { it.join() }
println("$name: initialisations=${created.get()}, distinct instances seen=${seen.size}")
}
fun main() {
hammer("object ") { AppConfig.config }
hammer("lazy (default) ") { LazyHolder.config }
hammer("lazy(NONE) ") { UnsafeHolder.config }
hammer("double-checked ") { DclHolder.get() }
}
Compiled with Kotlin 2.0.21 (kotlinc -include-runtime) and run on Temurin 21 in a Linux aarch64 container, a fresh JVM per run, 16 threads released together by a latch. One sample run printed:
object : initialisations=1, distinct instances seen=1
lazy (default) : initialisations=1, distinct instances seen=1
lazy(NONE) : initialisations=16, distinct instances seen=16
double-checked : initialisations=1, distinct instances seen=1
System.identityHashCode(x) returns a number derived from the object's identity (its address-like identity, not its contents), so two different Config objects normally give two different numbers, and the set of numbers counts how many distinct instances the 16 threads received. The three safe variants show exactly 1 initialisation every time, as their guarantees require. The lazy(NONE) line usually showed all 16 threads constructing their own instance; the 2 ms sleep in the constructor keeps the race window open, so treat the exact count as timing-dependent (16 is the worst case, not the normal result of an unsafe singleton: with the sleep removed, the same program usually showed 1 initialisation and occasionally 2, and in principle the unsynchronised lazy can also fail with a NullPointerException, because it clears its initialiser after the first computation while another thread may still be about to call it). The distinct-instance figure is built from identity hash codes, which are not guaranteed unique, so a collision can make it read lower than the number of objects actually created (so the unsafe row can in principle print a distinct count below its initialisation count); the initialisation count is the exact figure. Why each row behaves as it does:
object: on the JVM the object compiles to a class whose single instance is created in the class's static initialiser. The JVM runs a class's static initialiser once, under an internal lock for that class (the class-init lock), and makes other threads wait for it, and the completed initialisation is visible to them. There is no code for you to get wrong. The cost is that it is created at first use of the class, with no way to pass constructor arguments.by lazy(defaultSYNCHRONIZED): use it for a lazily computed property inside a class, or when you want the creation deferred but not tied to class loading. The docs describe its lock as a platform- and implementation-specific detail, so do not rely on which lock object it is.lazy(LazyThreadSafetyMode.NONE): the docs say that with no synchronisation, if accessed from several threads "its behavior is unspecified". It is correct only when the property is confined to one thread (for example only touched from the Android main thread). Thelazy(NONE)row above is the failure you get otherwise.PUBLICATIONis a third mode: the initialiser may run several times, but only one result is kept and published, which is fine when construction is cheap and has no side effects.- Double-checked locking is the hand-written pattern: a cheap unlocked read first, a lock only when the instance is still null, and a second null check inside the lock.
@Volatileon the field is required. On the JVM, a volatile write followed by a volatile read creates a happens-before edge (a guaranteed ordering) that makes the constructor's field writes visible to any thread that reads the reference. Without@Volatile, the Java memory model allows a thread to see a non-null reference to a half-constructed object. That failure is rare and hardware-dependent, and the demo above cannot show it:Config.valuesis aval, which the Kotlin compiler compiles to afinalfield, and the Java memory model's initialization safety for final fields (JLS 17.5) lets a reader that sees the reference see that field, and the map it points to, correctly initialised even when the reference was published through a data race. The hazard is real for a class with ordinary non-final fields, and that is the case@Volatileprotects. Preferlazyunless you need the extra control.
Swift
import Foundation
// @unchecked Sendable tells the compiler "this class is safe to share between threads; trust me" and switches off its check. Here that is justified because every access takes the lock.
final class Counter: @unchecked Sendable {
private let lock = NSLock()
private var n = 0
func bump() { lock.lock(); n += 1; lock.unlock() }
var value: Int { lock.lock(); defer { lock.unlock() }; return n }
}
let created = Counter()
final class Config: @unchecked Sendable {
let values: [String: String]
init() {
created.bump()
Thread.sleep(forTimeInterval: 0.002) // widen the window for a race
values = ["url": "https://api.example.com", "retries": "3"]
}
}
// 1. Idiomatic: a stored type property is lazily initialised once, thread-safely.
final class SafeStore: @unchecked Sendable {
static let shared = Config()
}
// 2. Hand-rolled check-then-set with no synchronisation: the race, shown on purpose.
final class UnsafeStore: @unchecked Sendable {
nonisolated(unsafe) private static var instance: Config? // opts this variable out of Swift 6's data-race checking
static func get() -> Config {
if let existing = instance { return existing }
let fresh = Config()
instance = fresh
return fresh
}
}
// 3. Hand-rolled but locked, when the instance must be created with arguments.
final class LockedStore: @unchecked Sendable {
private static let lock = NSLock()
nonisolated(unsafe) private static var instance: Config?
static func get() -> Config {
lock.lock(); defer { lock.unlock() }
if let existing = instance { return existing }
let fresh = Config()
instance = fresh
return fresh
}
}
// ObjectIdentifier is Swift's identity token for a class instance: equal only for the very same object.
final class IdentitySet: @unchecked Sendable {
private let lock = NSLock()
private var ids = Set<ObjectIdentifier>()
func add(_ c: Config) { lock.lock(); ids.insert(ObjectIdentifier(c)); lock.unlock() }
var count: Int { lock.lock(); defer { lock.unlock() }; return ids.count }
}
func hammer(_ name: String, _ fetch: @escaping @Sendable () -> Config) {
let before = created.value
let group = DispatchGroup()
let seen = IdentitySet()
for _ in 0..<16 {
group.enter()
Thread.detachNewThread {
let c = fetch()
seen.add(c)
group.leave()
}
}
group.wait()
print("\(name): initialisations=\(created.value - before), distinct instances seen=\(seen.count)")
}
hammer("static let ") { SafeStore.shared }
hammer("unsynchronised ") { UnsafeStore.get() }
hammer("locked ") { LockedStore.get() }
Compiled with swiftc -swift-version 6 in a swift:6.0 container (Swift 6.0.3, Linux aarch64), one sample run printed:
static let : initialisations=1, distinct instances seen=1
unsynchronised : initialisations=16, distinct instances seen=16
locked : initialisations=1, distinct instances seen=1
The unsynchronised row usually showed all 16 threads constructing an instance, again because the constructor sleeps 2 ms; the exact count depends on timing, and 16 is the worst case rather than the normal outcome. That row is also undefined behaviour: several threads assign a strong reference to the same class property at once, which races on reference counts, and the program can crash outright (a segmentation fault or a heap-corruption abort) instead of printing anything, so a crash is a legitimate outcome of an unsynchronised singleton, not only a wrong count. The other two rows contain no such race. The hammer function starts 16 threads, each fetches the singleton and records its identity, and the program compares how many times the constructor ran with how many distinct objects came back.
static let sharedis the idiomatic singleton. The run-once and the publication guarantee both come from the language, so there is nothing to lock. A globallethas the same behaviour. Two caveats: the guarantee covers the initialisation only, so ifConfigis a class with mutable properties, reads and writes after creation still need their own synchronisation (an actor, a lock, or immutableletfields as in the sample); andstatic lettakes no arguments, so a singleton that needs configuration at startup needs the locked form.- Do not write
lazy varfor a shared singleton. The Swift book notes that if alazyproperty "is accessed by multiple threads simultaneously and the property hasn't yet been initialized, there's no guarantee that the property will be initialized only once." That is the reasonlazy varis the wrong tool here andstatic letis the right one. - Hand-rolled check-then-set (the
UnsafeStorerow) is the bug the question is about: two threads can both pass thenilcheck. Under Swift 6 it needsnonisolated(unsafe)(available from Swift 5.10) just to compile: that keyword switches off the compiler's data-race checking for that one variable, so having to write it is the compiler telling you the access is unprotected.@unchecked Sendableon the classes is the same kind of promise at class level ("safe to share across threads, do not check"), made true inCounterandIdentitySetby taking a lock on every access. - Locked form (
LockedStore): take the lock around the whole check-and-create. This is correct and is the fallback when the instance needs arguments at first use. In a Swift 6 code base anactoris the alternative, at the price ofawaiton every access.
Choosing
Kotlin: default to object; use by lazy inside a class when creation should be deferred or depends on a constructor argument; reach for double-checked locking only with @Volatile and a reason. Swift: default to static let; use a lock or an actor only when you need arguments or later mutation. Either way, a singleton that holds mutable state is a shared mutable object, and thread-safe creation does not make its later use thread-safe.
Unlock Full Question Bank
Get access to all 38 Language-Level Concurrency and Multithreading interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.