Memory barriers
advanced · occasionally asked
The same program with a fence between the write and the read. The outcome that no ordering could explain becomes unreachable again.
The problem it solves
Visibility establishes that a write is not visible when it happens but when it is published, and that the gap admits outcomes no interleaving explains. A memory barrier is the instruction that closes the gap: it forces publication, and it forbids the reordering that made the gap exploitable.
The important thing to get right about a barrier is what it is not. It does not make anything atomic. It does not exclude anybody from anything. It does not lock. It constrains ordering and visibility — a different guarantee from the ones the rest of this section has been about, and the one that the reordering was violating.
Almost every synchronisation primitive you use is built on top of one. Understanding what the barrier does is understanding why a mutex gives you visibility as well as exclusion, and why volatile fixes a spin loop but not a counter.
The mechanism
A barrier is a point that certain operations may not cross. The useful varieties:
- Release (a store barrier): everything written before it is visible to any thread that subsequently performs a matching acquire. Nothing before it may be moved after it.
- Acquire (a load barrier): everything read after it sees the writes published by a matching release. Nothing after it may be moved before it.
- Full barrier: both, plus draining the store buffer. On x86 this is
mfence, or anylock-prefixed instruction.
The pairing is essential and frequently missed. A release with no matching acquire guarantees nothing — publication requires somebody to be listening in the right way. Acquire/release always come in pairs, which is precisely why a mutex works: the unlock is a release, the next lock is an acquire, and together they establish that everything the first thread did inside its critical section is visible to the second.
That is the mechanism behind a sentence people repeat without unpacking: a lock gives you visibility as well as mutual exclusion. It gives you visibility because it contains these two barriers.
What the enumeration shows
The same Dekker-style program as the visibility page, with a fence between each thread’s write and its read.
Without the fence, under TSO: 2,900 of 8,316 schedules produce both threads reading zero. With the fence: 0 of 7,788, exhaustively. The buffered write is drained before the read is allowed to proceed, so the window the outcome depended on no longer exists.
Step through and watch the fence operation: the state panel’s buffered row for that thread empties, and the value moves into shared memory where the other thread can see it. That row — a write that has happened and is not visible — is the state most people’s mental model has no slot for, and watching it drain is the clearest way to see what a barrier does.
Note that the schedule count barely falls. The fence did not remove many orderings; it removed the ones where a write stayed invisible across a read. Compare with mutex, where a lock removes orderings wholesale — barriers and locks solve different problems and it shows in the numbers.
The numbers worth carrying
- Fenced: 0 of 7,788 produce the anomalous outcome. Unfenced: 2,900 of 8,316.
- A full barrier costs roughly 20–100 cycles, since it waits for the store buffer to drain. Acquire and release barriers are much cheaper on x86 — often free, because TSO already forbids the reorderings they prevent, so they compile to nothing but a compiler barrier.
- That last point explains a lot of confusion: on x86, correct acquire/release code and incorrect unsynchronised code frequently generate the same instructions. The bug is invisible until you run it on ARM.
- Sequentially consistent atomics — the default in C++ and what Java
volatilegives — cost a full barrier on the store side. Usingrelease/acquirewhere they suffice is a real optimisation on weakly ordered hardware.
Where it breaks down
Barriers order; they do not exclude. Two threads can both execute a fenced read-modify-write and still lose an update, because the fence says nothing about atomicity. If you need both, you need an atomic operation, which contains a barrier.
A barrier with no pair is decoration. Release without acquire, or acquire without release, guarantees nothing. Placing “a fence for safety” without knowing which operation it pairs with is a common cargo-cult and produces code that is slower and equally wrong.
The compiler needs constraining separately. A hardware fence does not stop the compiler from having already hoisted your load out of the loop. Language-level constructs — volatile, atomic, std::atomic_thread_fence — constrain both; a raw assembly fence constrains one.
Almost nobody should write these directly. Hand-placed barriers are how lock-free libraries are built, and the failure mode is a bug that appears on one architecture under one compiler at one optimisation level. Use a mutex or a sequentially consistent atomic unless you have measured that you cannot afford to.
What people get wrong
“A barrier makes the operation atomic.” It orders and publishes. Atomicity is a separate guarantee.
“I will add a fence to be safe.” Which fence, paired with what, on which side? An unpaired barrier does not make code safe and does make it slower.
“volatile in C is like volatile in Java.” It is not, and this is a genuinely dangerous confusion. C’s volatile prevents the compiler from optimising away accesses — it is for memory-mapped hardware registers and signal handlers. It provides no ordering against other threads and no barriers. Java’s and C#’s volatile are full acquire/release semantics. Same keyword, different guarantee.
“My code works, so the barriers are right.” On x86 acquire/release are frequently free, so correct and incorrect code emit identical instructions. Working is not evidence.
In production
The layers you should reach for, in order of preference: a mutex (contains both barriers, and is right for almost everything); a sequentially consistent atomic (AtomicInteger, std::atomic<T> with defaults, volatile in Java or C#) when one variable must be published; acquire/release atomics when profiling says the full barrier costs you; and explicit fences essentially never, outside a library.
The one pattern worth knowing by name is release-store / acquire-load publication: build an object fully, then publish its pointer with a release store; readers acquire-load the pointer and are guaranteed to see the fully constructed object. This is what makes safe publication safe, and its absence is what made double-checked locking broken in Java before 2004 — the reference could become visible before the fields it pointed at, so another thread could observe a partially-constructed object through a non-null pointer.
Java’s final fields deserve a mention as the pleasant special case: fields assigned in a constructor and never reassigned are guaranteed visible to any thread that sees the object, with no synchronisation at all. That is a freeze barrier at the end of the constructor, and it is one more reason immutability removes work rather than adding it.
The follow-up questions
“What does a memory barrier do?” — Constrains ordering and forces publication. Not atomicity — say that explicitly, because the interviewer is usually checking for it.
“Why does a mutex give you visibility?” — Unlock is a release, lock is an acquire, and the pair establishes happens-before. This is the answer that shows the model rather than the rule.
“What is the difference between volatile in C and in Java?” — Compiler-only versus full acquire/release. Getting this right marks someone who has worked across both.
“Why was double-checked locking broken?” — The reference could be published before the object’s fields. The fix is volatile on the field, and the reason is a release-store.
In an interview
The mechanism under every synchronisation primitive you use, and the reason a correct-looking double-checked lock was broken for a decade.
- fences
- happens-before
- volatile
- reordering
Run these next
- Visibility and store buffersA write is not visible when it happens, it is visible when it is published. This is where the interleaving model runs out and a memory model is required.
- MutexesA lock removes orderings from the space. The critical section must span the whole read-modify-write — guarding only the write leaves the window exactly where it was.
- Atomic operationsAtomicity is not speed and not a lock. It is indivisibility: no interleaving can slot between the read and the write, so the orderings that lost updates cease to exist.