Visibility and store buffers
advanced · commonly asked
Two threads write, then read each other. Under sequential consistency both reading zero is impossible; on real hardware it happens — and no interleaving explains it.
The problem it solves
This is the page where most readers discover their mental model was wrong.
Up to here, “concurrency bug” has meant “the operations happened in a bad order”. Every failure in this section so far has been explicable by pointing at an interleaving: here is the ordering, here is where it goes wrong. That model is intuitive, it is how everyone reasons about threads, and it is not sufficient.
Two threads. Each sets its own flag to 1, then reads the other’s. Reason about it: for both to read 0, both reads would have to happen before both writes — and each thread writes before it reads, so that ordering does not exist. Both reading zero is impossible.
On real hardware, both read zero. There is no interleaving of this program that produces that result, and it happens anyway.
The mechanism
A write is not visible when it happens. It is visible when it is published, and those are different moments.
When a core executes a store, the value goes into that core’s store buffer — a small queue in front of the cache. The core continues immediately rather than waiting for the cache line to be acquired, which is a large performance win, because acquiring a line exclusively can take a hundred nanoseconds. The write drains to cache later, and until it does, no other core can see it.
Meanwhile the same core’s loads may proceed. So thread 1’s write to x can be sitting in its store buffer, invisible to everyone, while thread 1 goes ahead and reads y. Thread 2 does the mirror image. Both reads see the old value, both writes are still in transit, and both threads come away with zero.
This is x86’s memory model — TSO, total store order — and it is one of the strongest models in common use. ARM, POWER and RISC-V permit considerably more, including loads being reordered with each other. The compiler adds its own layer, freely reordering anything it can prove single-threaded-equivalent, which is everything unless you tell it otherwise.
The essential shift: an interleaving model assumes a single global order of operations that every thread agrees on. Real hardware provides no such thing. Each core has its own view, and they converge — eventually, and not on any schedule you can reason about locally.
What the enumeration shows
The same program runs under two memory models on this page, and the store-buffer drain is a scheduled action — because when a buffered write becomes visible is precisely what the model leaves open.
Under sequential consistency: 70 schedules, and both-zero never occurs. Exhaustively. That is the intuitive model, confirmed.
Under TSO: 8,316 schedules, of which 2,900 produce r1 == 0 and r2 == 0. Thirty-five percent. An outcome that no ordering of the program’s operations can explain, reached by more than a third of the executions.
Being able to put those two numbers side by side is the point of this page. The extra outcomes are not caused by a subtler interleaving that the first model missed. They are outside the model entirely.
Then switch to memory-barrier: the fence drains the buffer before the read, and both-zero becomes unreachable again.
The numbers worth carrying
- Sequential consistency: 0 of 70 interleavings produce both-zero. TSO: 2,900 of 8,316 — about 35%.
- A store buffer is roughly 20–60 entries on modern x86. The window is short in time and enormous in instructions.
- Cost of a fence: an
mfenceor alock-prefixed instruction is roughly 20–100 cycles, because it waits for the buffer to drain. This is why fences are not free and why relaxed atomics exist. - What
volatile(Java/C#) oratomicwith sequential consistency actually buys: every write is published and every read sees the latest publication. Not atomicity — see atomic operations — publication.
Where it breaks down
The compiler reorders too, and it is not modelled here. This page models hardware reordering; a compiler can hoist a load out of a loop, sink a store past a branch, or eliminate a “redundant” read entirely. The classic symptom is a spin loop on a non-volatile flag that never terminates because the compiler cached the flag in a register. Both layers must be constrained, which is what a language-level volatile or atomic does — it is a barrier to the compiler and the hardware.
x86 is unusually strong, which makes testing misleading. TSO forbids most reorderings; ARM permits far more. Code that is racy but appears correct on an x86 laptop can fail on an ARM server or an Apple Silicon machine, and this is a genuinely common way for latent bugs to surface years later.
Data races are undefined behaviour in C++ and Rust, not merely unpredictable. The compiler is entitled to assume they do not occur and to optimise on that basis, so the result can be worse than any interleaving or any reordering would suggest.
Happens-before is the real model. Rather than reasoning about buffers, the language memory models define a partial order: if A happens-before B, then B sees A’s writes. A mutex release happens-before a subsequent acquire; a volatile write happens-before a subsequent volatile read of the same variable; a thread start happens-before everything in the thread. Reasoning in those terms is both easier and more portable than reasoning about hardware.
What people get wrong
“volatile makes it atomic.” It makes it visible. volatile int count; count++; is still three operations and still races. The single most common wrong answer about volatile, and it is exactly backwards.
“If I lock, I do not need to think about visibility.” Correct, and worth understanding rather than assuming: acquiring a lock establishes happens-before with the previous release, so a lock provides visibility as well as exclusion. That is why locked code needs no fences.
“I tested it and it never happened.” On x86, with this compiler, at this optimisation level, on this workload. Change any of them.
“Adding a println fixed it.” It probably contains a synchronised block, which added the barrier you needed. The bug is intact.
“Reordering is a compiler bug.” It is a documented licence. Both the compiler and the CPU are permitted to do it, and the specification says so.
In production
The practical rule is to work in terms of happens-before and never in terms of buffers: locks, atomics with default (sequentially consistent) ordering, and language-level volatile all establish it. If two threads touch the same data and at least one writes, they need one of those between them. That single sentence covers virtually all application code correctly.
Relaxed orderings — memory_order_relaxed, acquire, release — buy real performance and require you to reason about which specific guarantee you are giving up. They belong in libraries and hot paths written by people who will reason about them carefully, not in application logic.
Where this bites in practice: the double-checked locking idiom was broken in Java until the 2004 memory model precisely because the partially-constructed object could be published before its fields were — a visibility bug hiding inside a correct-looking atomicity fix. The modern fix is volatile on the instance field, and knowing why is exactly this page.
The follow-up questions
“Two threads write their own flag and read the other’s. Can both read zero?” — Under sequential consistency no; on real hardware yes. The whole page in one exchange.
“What does volatile guarantee?” — Visibility and ordering, not atomicity. Then say what it does not fix: count++.
“Why does locked code not need fences?” — Acquire/release establishes happens-before. This shows the model is understood rather than the rules memorised.
“Your code works on x86 and fails on ARM. Why?” — Weaker memory model, more reordering permitted. A latent data race that x86’s TSO was hiding.
In an interview
What `volatile` actually guarantees, and why it is about publication rather than atomicity. Most candidates have this exactly backwards.
- memory model
- volatile
- store buffer
- happens-before
Run these next
- Memory barriersA barrier does not make anything atomic. It forces publication, which is a different guarantee and the one the reordering was violating.
- The lost update`count++` is three operations, not one. Only two of its twenty interleavings produce 2; the failure is the common case, and an atomic increment removes the orderings rather than making them rarer.
- Atomic operationsAtomicity is not speed and not a lock. It is indivisibility: no interleaving can slot between the read and the write, so the orderings that lost updates cease to exist.