Skip to main content
PRISM

Concurrency against parallelism

foundational · asked in almost every interview

The distinction everybody quotes and few hold. Watch a single core finish eight waiting tasks in a fraction of the time, with no parallelism involved.

Computing the curve…

The problem it solves

The distinction everybody quotes and few can use.

Concurrency is a property of the program: several things are in progress at once. Parallelism is a property of the machine: several things are executing at the same instant. One is a structure you write; the other is hardware you buy. A single-core machine can run a highly concurrent program and cannot run anything in parallel. A perfectly parallel workload — eight independent computations on eight cores — may involve no concurrency in the interesting sense at all, because nothing is ever interleaved.

The reason this matters practically rather than pedantically: they solve different problems, and reaching for the wrong one is a common and expensive mistake. If your service spends most of its time waiting on other services, more cores buy you almost nothing and concurrency buys you nearly everything. If it spends its time computing, the reverse.

Rob Pike’s formulation is the one to remember: concurrency is about dealing with lots of things at once; parallelism is about doing lots of things at once.

The mechanism

The difference comes down to what a task spends its time on.

A task that is waiting — for a database, an HTTP response, a disk — is holding no CPU. Interleaving lets another task use the CPU during that wait, so n tasks that each wait 80% of the time can be served by one core with the waits overlapping. Nothing executes simultaneously; the idleness is what overlaps. This is what an event loop does, what green threads do, and why a single-threaded server can hold tens of thousands of connections.

A task that is computing holds the CPU for its whole duration. Interleaving does not help: total CPU time is unchanged, and you have added context switches. The only way to finish sooner is more cores, and then Amdahl’s law sets the ceiling.

So the question that decides your architecture is not “how many cores” but what fraction of the work is waiting. That fraction is measurable, and it tells you which of the two you need.

What the enumeration shows

This is a model page: a curve rather than an interleaving space, because the ordering is not what varies.

Eight tasks, each 80% waiting. Doing them one at a time takes 8 units. Interleaving them on a single core takes 2.4 — and the core count on that line never changes, because concurrency does not need one. Drag the waiting share and watch that line move on its own.

The parallel line divides only the computing portion. At 80% waiting there is not much computing to divide, so extra cores add little on top of what interleaving already achieved. Drag the waiting share down towards zero and the two lines separate sharply: with nothing to wait for, interleaving does nothing and cores do everything.

That crossover is the whole page. The right answer depends on a property of your workload, and the chart makes the dependency visible rather than something to argue about.

The numbers worth carrying

  • Eight tasks at 80% waiting, on one core: 8 units becomes 2.4. No parallelism involved.
  • At 10% waiting, interleaving on one core saves almost nothing and eight cores give close to eight times.
  • The dividing question: is this I/O-bound or CPU-bound? I/O-bound wants concurrency — async, an event loop, more connections. CPU-bound wants parallelism — more cores, worker threads, and an eye on Amdahl.
  • A context switch costs roughly 1–10µs between OS threads and tens of nanoseconds between green threads or coroutines. That gap is why an event loop can hold 100,000 connections and 100,000 OS threads would collapse under scheduling overhead alone.

Where it breaks down

Most real workloads are both. A request parses JSON (CPU), queries a database (I/O), renders a template (CPU), writes a response (I/O). The useful figure is the ratio, and the usual architecture is concurrency for the waiting with a small pool of workers for the computing — which is exactly what Node with worker_threads or Python with asyncio plus a process pool looks like.

Concurrency has a cost even when nothing runs simultaneously. Interleaving means state can change between your suspension points, which is how a logical race appears in single-threaded JavaScript across an await. Concurrency without parallelism removes data races and not logical ones.

Parallelism needs the work to be divisible. Eight cores on a strictly sequential algorithm give one core’s throughput and seven idle. Amdahl’s ceiling is a property of the algorithm, not of the runtime.

Hyper-threading is neither, quite. Two hardware threads on one core share execution units, so they overlap stalls — a form of concurrency implemented in silicon. Typical gain is 20–30%, not 100%, which is why core counts and thread counts are different numbers and both appear on a spec sheet.

What people get wrong

“async makes it parallel.” It makes it concurrent. One thread, interleaved at await points. This is the most common confusion in the whole vocabulary.

“More cores will fix our latency.” Only if you are CPU-bound. An I/O-bound service on more cores has more idle cores.

“Single-threaded means slow.” For I/O-bound work a single well-designed loop routinely outperforms a thread-per-request server, because it avoids the memory and scheduling cost of thousands of stacks.

“Threads give you both.” They give you both and charge for both: preemption means you get parallelism on multiple cores and data races everywhere, so you pay in locks. The event loop model trades away parallelism to avoid the locks; the actor model keeps parallelism and avoids sharing instead — see actors.

In production

Diagnose before choosing. If CPU utilisation is low while latency is high, you are waiting and you want concurrency. If CPU is pinned, you are computing and you want parallelism — or a better algorithm, which is usually cheaper than either.

The architectures that follow: async I/O with a small thread count for I/O-bound services (Node, asyncio, Netty, Tokio); a worker pool sized near the core count for CPU-bound work; and for the common mixed case, both — an event loop for the waiting and an explicit pool for the computing, with the boundary between them made obvious so nobody accidentally does 200ms of CPU work on the loop and blocks everything, which is the failure on blocking the event loop.

The sizing rule worth remembering: a CPU-bound pool wants roughly one thread per core, because more only adds context switches. An I/O-bound pool wants cores × (1 + waitTime/computeTime), which for 80% waiting is five times the core count — and if that number is large, the answer is probably async rather than a bigger pool.

The follow-up questions

“What is the difference between concurrency and parallelism?” — Structure versus execution, and give the one-core example: eight waiting tasks finishing three times sooner with no second core involved.

“Your service is slow and CPU is at 15%. What do you do?” — You are I/O-bound. More cores will not help; more concurrency will.

“How many threads for a CPU-bound pool?” — About one per core. For I/O-bound, cores × (1 + wait/compute) — and consider async instead.

“Can you have parallelism without concurrency?” — Yes: eight independent computations on eight cores, never interleaved. A good sign the distinction is understood rather than recited.

In an interview

Asked constantly, answered with a slogan. Being able to say which one an I/O-bound service actually needs is the useful half.

  • concurrency
  • parallelism
  • I/O bound
  • CPU bound

Run these next

The rest of async and event loops