Races and atomicity — every question, written out
What an increment actually is, why a question goes stale between asking and acting, and what "atomic" does and does not promise.
Why is `count++` not safe to run on two threads at once?
Explain it plainly
Because it is not one operation. It compiles to a load, an add and a store, and another thread can run in between. If both threads load 5 before either stores, both store 6, and one increment is lost.
It is three operations — load, add, store — and a thread switch may land between any two of them. Both threads can load the same value, add one to the same value, and store the same result, so two increments produce one.
See every interleaving — Twenty interleavings; eighteen of them lose an update.
Two threads each run one `count++` as load-add-store. How many of the twenty possible interleavings produce 2?
Trace prediction
Two — only the orderings where one thread finishes entirely before the other starts.
Correctness requires the second thread to load after the first has stored, which only happens when the two sequences do not overlap at all. There are exactly two such orderings out of twenty.
See every interleaving — Click the failing outcome to filter to the eighteen.
If ninety percent of orderings lose an update, why does the bug seem rare in practice?
Code diagnosis
A thread usually gets through all three operations without being interrupted.
The set of possible orderings is not the distribution of observed ones. Uninterrupted execution is common, so the overlap window is rarely hit — until the machine is busy, the core count rises, or a preemption lands there.
Does making the counter `volatile` fix the lost update?
Comparison
No — it makes each read and write visible, and leaves the gap between them.
`volatile` guarantees that each individual read and write is published and correctly ordered. The increment is still load, add, store, and another thread can still land in the middle.
See every interleaving — Visibility is a separate guarantee from atomicity.
What happens to the twenty interleavings when the increment is made atomic?
Invariant identification
There are only two left, because there is no longer a point inside the operation to interleave at.
Atomicity means the operation has no observable inside. The scheduler chooses between operations, so an operation that cannot be split has no position for another thread to occupy.
See every interleaving — Two schedules, one outcome.
A thread atomically decrements `available`, then atomically increments `reserved`. Is that safe?
Edge case reasoning
No — another thread can observe the moment in between, where the item is in neither.
Atomicity is per operation, not per intention. Anything that requires several values to change together needs a lock, a transaction, or a redesign that packs them into one value.
Two threads run `if (!map.has(k)) map.set(k, expensive())`. What goes wrong?
Code diagnosis
Both pass the check and both run `expensive()`, each believing it created the entry.
The check was true when it was asked and describes a world that no longer exists by the time it is acted on. In real code the branch sends the email or charges the card, and one of those happens twice.
See every interleaving — 210 of 252 interleavings do the work twice.
A lock is added around the `map.set` call and the duplicate work continues. Why?
Code diagnosis
The critical section begins after the check, so the window it needed to cover is outside it.
The race is between the check and the act, so a critical section that starts after the check protects nothing that was ever at risk. The lock must be acquired before the check and held past the act.
See every interleaving — A real lock, correctly used, fixing nothing.
What is the fix for check-then-act that does not involve a lock?
Trade-off & selection
Ask and act in one indivisible operation — `putIfAbsent`, `SETNX`, `ON CONFLICT`, compare-and-swap.
Every API named putIfAbsent, SETNX or INSERT ... ON CONFLICT exists because check-then-act cannot be made safe by being careful. The condition and the action commit together or not at all.
When is it reasonable to leave a check-then-act race in place?
Trade-off & selection
When the action is idempotent and doing it twice costs only duplicated work.
Creating a directory, computing a cached value, setting a flag to true: repeating these costs work and nothing else. Knowing when the duplicate is harmless is as valuable as knowing how to prevent it.
What is word tearing, and when can it happen?
Edge case reasoning
A value larger than the machine word is written in parts, so a reader can see a combination that was never written.
A 64-bit value on a 32-bit machine is two stores, so a concurrent reader can observe the new high half with the old low half — a value neither thread ever computed. This is why the JVM guarantees atomicity for `volatile long` specifically.
What is the difference between atomicity and visibility?
Comparison
Atomicity is about being observed half-done; visibility is about being seen at all.
An operation can be atomic and invisible — indivisible but sitting unpublished in a store buffer — or visible and non-atomic, like a published load followed by a published store. Locks and sequentially consistent atomics happen to provide both.
See every interleaving — A write that has happened and is not visible.
A concurrency fix passes a load test that runs the code a million times. What has that established?
Code diagnosis
That the orderings the scheduler happened to produce were fine — which does not include the ones it did not produce.
A test explores the interleavings that occurred, and cannot tell you which it missed. That is the argument for enumeration: a fix verified against the whole space is a proof, and a passing test is evidence.
Sixteen threads increment one atomic counter and throughput is poor. What is happening, and what is the fix?
Trade-off & selection
The cache line is contended; use per-thread counters summed on read.
Every increment needs the line exclusively, so it ping-pongs between cores at roughly 80ns per transfer. Striping — `LongAdder` and friends — keeps a cell per thread and sums on read, trading exact-at-every-instant for far better throughput.
See every interleaving — The same coherence traffic, arriving deliberately.
What does a race detector like `-race` or ThreadSanitizer actually tell you?
Trade-off & selection
That two unsynchronised accesses to the same location occurred in this run, one of them a write.
Race detectors track happens-before at runtime and flag unsynchronised access pairs they actually observe. They find real bugs that produced correct output, which is their great value — and they cannot prove absence.
Define a data race precisely.
Explain it plainly
Two accesses to the same location from different threads, at least one of which is a write, with no synchronisation ordering them. Concurrent reads are not a race, and two writes ordered by a lock are not a race — the missing happens-before edge is what makes it one.
Two threads access the same memory location, at least one access is a write, and there is no happens-before relationship ordering them. All three clauses matter: concurrent reads are fine, and synchronised writes are fine.