The memory model — every question, written out
Where the interleaving model runs out: writes that have happened and are not visible, reordering, and what volatile actually guarantees.
Each thread sets its own flag to 1, then reads the other’s. Can both read 0?
Trace prediction
Yes on real hardware, and no under sequential consistency.
Each write can sit in its own core’s store buffer, invisible to the other core, while that core proceeds to its read. No interleaving of the program produces this, which is exactly why a memory model is needed.
See every interleaving — 2,900 of 8,316 TSO schedules; 0 of 70 under sequential consistency.
What is a store buffer and why does it exist?
Explain it plainly
When a core executes a store, the value goes into a small queue rather than straight to cache, and the core continues. That avoids stalling on the coherence traffic needed to take the line exclusively. The cost is that the write is invisible to other cores until it drains, which is why two threads can each write and then read stale values.
A queue in front of the cache holding writes that have executed but not yet been published. It exists so a core need not stall for the tens of nanoseconds it takes to acquire a cache line exclusively.
What does inserting a full barrier between the write and the read accomplish?
Invariant identification
It drains the buffered write, so the read cannot proceed until it is visible.
The anomalous outcome depended on a write being invisible when the read happened. Publishing it first removes that possibility, and the both-zero result becomes unreachable again.
See every interleaving — 0 of 7,788 schedules produce it once fenced.
What does Java’s `volatile` guarantee?
Comparison
Visibility and ordering for each access — not atomicity of a read-modify-write.
A volatile write is published and a volatile read sees the latest publication, with ordering constrained around both. The gap between a read and a subsequent write is untouched.
How does C’s `volatile` differ from Java’s?
Comparison
C’s constrains only the compiler and provides no cross-thread ordering; Java’s is a full acquire/release.
C’s `volatile` exists for memory-mapped registers and signal handlers: it stops the compiler eliding accesses and emits no barriers. Using it for thread synchronisation in C or C++ is a bug.
What is the happens-before relationship, and why reason in terms of it?
Explain it plainly
It is a partial order the language defines. A mutex release happens-before the next acquire; a volatile write happens-before a subsequent read of it; starting a thread happens-before everything in it. If two accesses are not ordered by it and one is a write, you have a data race. Reasoning this way is portable — store buffers and reordering are implementation details underneath it.
A partial order over operations: if A happens-before B, then B is guaranteed to see A’s writes. It is the portable way to reason, because it is what the language specifies rather than what any particular chip does.
Why does code inside a mutex need no explicit fences?
Invariant identification
Unlock is a release and the next lock is an acquire, so the pair establishes happens-before.
A lock provides visibility as well as mutual exclusion, and this is the mechanism. Everything written before an unlock is visible to whoever acquires next.
Code works on an x86 laptop and fails on an ARM server. What is the likely cause?
Code diagnosis
A latent data race that x86’s stronger memory model was hiding.
x86’s TSO forbids most reorderings, so incorrect code often behaves. ARM permits loads to be reordered with each other and much else besides, and the latent bug surfaces.
A thread spins on a non-volatile boolean flag and never sees it change. What happened?
Code diagnosis
The compiler hoisted the read out of the loop, so it tests a value it read once.
Absent synchronisation the compiler may assume nothing else changes the flag and read it once into a register. This is why the fix is a language-level construct — it constrains both compiler and hardware.
Why was double-checked locking broken in Java before 2004?
Code diagnosis
The reference could become visible before the fields of the object it points at.
A thread could observe a non-null reference to an object whose constructor had not finished publishing its fields. The modern fix is `volatile` on the field, which makes the publication a release store.
See every interleaving — Release-store publication is what makes it safe.
What is safe publication?
Invariant identification
Making an object visible so that any thread seeing the reference also sees its fully initialised state.
The mechanisms are a release store (volatile or atomic), publication from inside a lock, static initialisation, or final fields — Java guarantees that anything assigned in a constructor to a final field is visible to any thread that sees the object.
What does sequential consistency claim that real hardware does not provide?
Comparison
That there is one global order of operations which every thread agrees on.
Sequential consistency is the assumption that every execution corresponds to some interleaving of the program’s operations. Store buffers break it: each core has its own view, and they converge on no schedule you can reason about locally.
In C++ and Rust, what is the status of a data race?
Edge case reasoning
Undefined behaviour — the compiler may assume it does not happen and optimise accordingly.
Because the compiler may optimise on the assumption of no races, the outcome can be worse than any interleaving or reordering would suggest — including code paths being removed entirely. Rust prevents them at compile time instead.
Why do acquire and release barriers come in pairs?
Trade-off & selection
Publication needs somebody listening: a release only guarantees anything against a matching acquire.
The guarantee is relational: everything before the release is visible to whoever performs the matching acquire. An unpaired barrier is slower code with no guarantee attached, which is a common cargo-cult.
For ordinary application code, what should you actually reach for?
Trade-off & selection
A lock, or a sequentially consistent atomic — both establish happens-before without further thought.
If two threads touch the same data and at least one writes, put a lock or a default-ordering atomic between them. That single rule covers essentially all application code correctly.
A race disappears when a logging call is added. What has happened?
Code diagnosis
The logging call contains synchronisation, which supplied the barrier the code was missing.
Most loggers synchronise internally, so the call establishes happens-before as a side effect. The bug is intact and is now hidden behind a statement nobody will suspect.