Skip to main content
PRISM

Promise.all against sequential await

foundational · commonly asked

The same two requests, sent two ways. In one there is a moment when both are outstanding; in the other there never is, however fast the machine.

Enumerating the interleavings…

The problem it solves

Two independent requests, each taking 100 milliseconds. Written one way the operation takes 200ms; written another it takes 100ms.

const a = await fetchA();
const b = await fetchB();        // 200ms

const [a, b] = await Promise.all([fetchA(), fetchB()]);   // 100ms

Everybody knows this and most people state the reason wrongly. It is not that the second version is “faster” or that promises “run in parallel” — there is one thread and no parallelism anywhere. It is that in the first version the second request has not been sent while the first is in flight. await suspends the function that would have sent it. There is no moment at which both are outstanding, and that is a fact about the program rather than about how fast anything is.

The second difference is the one almost nobody mentions, and it is the one that causes incidents: the two forms behave differently when something fails.

The mechanism

await is a suspension point. The function stops, its continuation is registered as a microtask, and the event loop moves on. When the awaited promise settles, the continuation is queued and eventually resumes.

In the sequential version, fetchB() is a statement that has not executed. It sits after a suspension point, so it runs only after A has resolved. The waiting cannot overlap because the second wait has not begun.

In Promise.all, both calls are evaluated immediately — fetchA() and fetchB() are invoked before Promise.all is even called, since they are arguments — so both requests are in flight before anything is awaited. The single await then suspends until both settle. The total is the maximum of the two, not the sum.

That framing matters because it tells you when the trick applies: only when the operations are independent. If B’s input is A’s output, no amount of Promise.all will help, and the sequential form is not a mistake.

What the enumeration shows

Both programs have exactly one schedule — the event loop admits no choice — and the interesting variable is overlapped, recording whether both requests were ever outstanding simultaneously.

Promise.all: overlapped == 1. Sequential await: overlapped == 0, in the only ordering that exists. Not “usually not”; there is no interleaving in which they overlap, because the second request is not reachable until the first has returned.

Step through the sequential version and watch it: request A is sent, the task yields at the await, the loop has nothing else to run, A resumes, A completes, and only then is B queued. The yield is visible in the timeline, and so is the fact that nothing filled it.

Note that sequential await is not marked as an incorrect outcome. It is slower, not broken — putting a warning on working code would cost this section its credibility on the programs that are genuinely wrong.

The numbers worth carrying

  • Sequential: sum of the latencies. Parallel: maximum. For n independent calls of equal duration, that is an n-fold difference.
  • Ten independent 50ms calls: 500ms sequentially, 50ms with Promise.all. This is the most common single latency win available in async code, and it is usually sitting in a loop.
  • await inside a for loop is the canonical form of the bug. for (const id of ids) { await fetch(id) } is sequential; await Promise.all(ids.map(fetch)) is not.
  • Unbounded fan-out is its own hazard: Promise.all over 10,000 items opens 10,000 requests at once. Use a concurrency limiter — p-limit, a semaphore, batching — because your dependency has a utilisation curve too.

Where it breaks down

Failure semantics differ, and this is where the incidents are. Promise.all rejects as soon as any input rejects — and does not cancel the others. They keep running, and if one of them rejects afterwards you have an unhandled rejection from a promise nobody is awaiting any more. Sequential await stops at the first failure, so later work never starts at all. Neither is right; they are different, and the choice is usually made without noticing there was one.

Promise.allSettled is often what you meant. It waits for all of them and reports each outcome separately, which is what you want when partial success is useful — three of four services answered, render what you have.

Cancellation is not built in. Neither form cancels anything. If you need the other requests to stop when one fails, you need AbortController threaded through explicitly, and this is exactly the gap structured concurrency closes.

Overlap needs independence. If B needs A’s result, the sequential form is correct and the only option.

What people get wrong

“Promise.all runs them in parallel.” Concurrently. One thread. The waiting overlaps; no computation happens simultaneously. See concurrency against parallelism.

“Promise.all starts the requests.” The requests started when the functions were called, which is before Promise.all receives them. This matters: const pa = fetchA(); const pb = fetchB(); await pa; await pb; also overlaps, because both were invoked before either was awaited. Sequential await is slow because of where the calls are, not where the awaits are.

“Promise.all cancels the rest on failure.” It does not. They run to completion and their results are discarded.

“Awaiting in a loop is fine, it is the same work.” It is the same work done end to end instead of overlapped. This is the most common performance bug in async code by a wide margin.

In production

The four combinators are worth knowing precisely: all (all succeed or reject on the first failure), allSettled (wait for all, report each), race (first to settle, success or failure), any (first to succeed, reject only if all fail). Timeouts are usually race; redundant backends are usually any; partial rendering is usually allSettled.

Two practices worth adopting. Bound the fan-out — an unbounded Promise.all over a user-supplied list is a load amplifier pointed at your own dependencies. And attach a handler to every promise you create, even ones you abandon, because an unhandled rejection terminates the process in modern Node.

The general habit: when you see await inside a loop, ask whether the iterations are independent. If they are, this is free latency. If they are not, leave it alone — and if they are independent but numerous, bound the concurrency rather than removing it.

The follow-up questions

“Why is Promise.all faster?” — Both requests are in flight before either is awaited. Say “the waiting overlaps”, not “they run in parallel”.

“What happens if one rejects?” — Promise.all rejects immediately and the others keep running uncancelled. This is the half that separates a real answer from a memorised one.

“How would you limit it to five at a time?” — A semaphore or a limiter. And say why you would: your dependency has a capacity.

“Is await in a loop always wrong?” — No. It is wrong when the iterations are independent. Rate-limited or dependent sequences are correct as written.

In an interview

Asked constantly, answered with "it is faster", and rarely with the reason. The failure semantics differ too, and almost nobody mentions it.

  • Promise.all
  • await
  • concurrency
  • latency

Run these next

The rest of async and event loops