Blocking the event loop
foundational · asked in almost every interview
One synchronous call, and every pending request queues behind it. No preemption exists to take the CPU back.
The problem it solves
One function does two hundred milliseconds of synchronous work. During those two hundred milliseconds your server accepts no requests, fires no timers, resolves no promises, and responds to nothing.
Not because it is busy in the way a loaded system is busy. Because there is no mechanism by which anything else could run. The event loop runs tasks to completion, and completion is defined by your function returning. There is no preemption, no time slice, no scheduler waiting to take the CPU back. Every ready callback sits in a queue watching.
This is the single most useful thing a JavaScript engineer can internalise, and it is the direct cause of most “our Node service got slow and we cannot find it in the profiler” incidents — because the slow thing is not slow, it is frequent, and the cost lands on everybody else.
The mechanism
Run-to-completion is what removes data races from JavaScript. Your function cannot be interrupted, so no other code can observe your half-finished state, so you need no locks. That is a real and valuable guarantee.
It is the same guarantee as the problem. “Cannot be interrupted” and “cannot be preempted when it runs long” are one property viewed from two sides. A language that removes races by refusing to interleave has no way to interleave when you would like it to.
The consequence is that latency for everyone is bounded by the longest synchronous operation anyone performs. A thread pool degrades gracefully under a slow task — other threads keep serving. An event loop does not degrade; it stops. And because the work is queued rather than dropped, the effect appears as a latency spike affecting every concurrent request equally, which looks nothing like “one endpoint is slow” in a dashboard.
What the enumeration shows
Two requests arrive and are queued. Then a synchronous call runs, modelled as six operations that nothing can slot between.
One schedule — as always here — and the interesting number is stalled == 2: both requests were ready and neither could run. Step through and watch the blocker’s six operations execute consecutively with the two requests sitting in the macrotask queue the entire time. There is no interleaving to look for, which is precisely the finding.
Compare this against any threaded page in this section. On the lost update, the scheduler could switch between operations, and that was the source of the bug. Here it cannot, and that is the source of this one. The same property is protecting you there and hurting you here.
The numbers worth carrying
Rules of thumb for how long a synchronous call blocks everything:
- JSON.parse of 1MB: roughly 10ms.
- bcrypt at cost 12: roughly 250ms. One login blocks every other request for a quarter of a second.
- A synchronous file read: 0.1–10ms, and unbounded on a network filesystem.
- Sorting 100,000 objects: roughly 50ms.
- A regex with catastrophic backtracking: unbounded, and the basis of ReDoS attacks — a denial of service that needs one request.
The number to aim for is under 10ms per task. Above that, the tail latency of everything else is your function’s duration. And the arithmetic that makes it urgent: at 1,000 requests a second, a 100ms block means 100 requests queue behind it, each of which now waits an average of 50ms it did not need to.
Where it breaks down
Not everything can be made asynchronous. CPU-bound work is CPU-bound. Making it async changes nothing — an async function with no await runs synchronously and blocks exactly as much. The fix is to move it off the loop, not to decorate it.
Chunking helps and is fiddly. Splitting work into pieces separated by setImmediate lets the loop breathe between them. It works, and it makes the operation slower overall and the code harder to follow, and it does nothing for a single indivisible operation like a large JSON.parse.
Microtasks do not yield. await Promise.resolve() does not give the loop a chance to run a pending timer or I/O callback, because microtasks drain completely before any macrotask is taken. Yielding to the loop means setImmediate or setTimeout(…, 0), and this catches people who believe any await is a yield point.
Worker threads have a transfer cost. Moving work to a worker means serialising the input and the output, unless you use SharedArrayBuffer or transferables. For small tasks the transfer can cost more than the work.
What people get wrong
“Making it async makes it non-blocking.” async marks a function as returning a promise. It does not move any CPU work anywhere. A synchronous loop inside an async function blocks the loop for its full duration.
“Node is fast so this does not matter.” Node is fast at I/O concurrency, which is what the loop is for. It is exactly as fast as any other single thread at computation, and it has one.
“We will scale horizontally.” Every instance has the same problem, and each instance still stalls for the duration of each blocking call. You have multiplied the number of loops, not made any of them preemptible.
“The profiler will show it.” A sampling profiler shows which function uses CPU, not that its cost lands on unrelated requests. The signature to look for is event loop lag, and it needs to be measured deliberately.
In production
Measure loop lag explicitly. perf_hooks.monitorEventLoopDelay() in Node gives a histogram; the crude version is a setInterval that records how late it actually fired. p99 lag above tens of milliseconds means something is blocking, and that metric will find it when request-latency dashboards will not, because it isolates the loop from the work.
The escalation, in order:
- Find it. Loop lag plus a CPU profile. The blocking call is usually crypto, serialisation, compression, or a regex.
- Make it async if an async version exists.
fs.promisesoverfs.readFileSync,crypto.scryptoverscryptSync, streaming parsers over whole-buffer ones. - Move it off the loop.
worker_threadsfor CPU-bound work, or a separate service for anything substantial. - Chunk it if it must stay and can be split.
Two specific offenders worth naming: synchronous crypto (bcryptSync, pbkdf2Sync) in an authentication path, which is both blocking and on the critical path of every login; and regexes on user input, where catastrophic backtracking turns one crafted request into an indefinite outage. Both have async or safe alternatives, and both appear constantly.
The follow-up questions
“What happens when one request does 500ms of CPU work?” — Every other request waits 500ms. Say “no preemption”, because that is the mechanism.
“Does async fix it?” — No. async is about the return type; the CPU work is unchanged.
“How do you detect it?” — Event loop lag, measured deliberately. Request latency alone will not localise it.
“Where do you put the work instead?” — Worker threads, or a separate service. And say what the transfer costs, because for small tasks it can exceed the work.
In an interview
The single most useful thing a JavaScript engineer can internalise, and the cause of most "our Node service got slow" incidents.
- event loop
- blocking
- run-to-completion
- latency
Run these next
- The event loopThe loop finishes what it started, drains every microtask, then takes one macrotask. That discipline has no choices in it — which is why the outcome space has exactly one member.
- Condition variablesA woken thread does not resume holding the lock — it queues for it, and the condition it was promised can be false again by the time it gets in. `while` asks at the only moment the answer counts.
- DeadlockA cycle in the wait-for graph is the deadlock. Lock ordering removes it by construction, and it is only a policy when it is a total order over every lock.