Structured concurrency
intermediate · occasionally asked
A function starts background work and returns without awaiting it. The caller believes the operation finished; the work is still going.
The problem it solves
A function starts some background work and returns without waiting for it.
The caller believes the operation is finished — it returned, after all. The work is still going: writing to state the caller has moved past, holding a connection nobody is tracking, and reporting its errors to a handler that has already gone. If the process shuts down, the work is killed mid-flight. If it fails, nobody hears. If it succeeds, it succeeds into a world that stopped caring.
Fire-and-forget is quiet in every one of those cases, which is why it survives. There is no error, no failing test, no log line. The bug is a lifetime mismatch, and lifetime is not something most languages let you see.
Structured concurrency is the rule that fixes it: a task may not outlive the scope that started it. Whatever you start inside a block finishes, or is cancelled, before that block returns. It makes lifetime a property of the shape of the code rather than of everybody remembering to await.
The mechanism
The idea is borrowed from something we already take for granted. A function’s local variables cannot outlive the function. Nobody has to remember to clean them up, because the language’s structure makes it impossible for them to escape. Structured concurrency applies the same discipline to concurrent tasks: they are scoped, and the scope does not close while a child is alive.
Concretely, a scope does three things. It waits for every child before returning. It propagates cancellation downward, so cancelling the scope cancels everything inside it. And it collects errors upward, so a child’s failure surfaces at the scope rather than vanishing.
That third property is what makes the pattern more than tidiness. In fire-and-forget code an exception in a background task has nowhere to go — the stack it belonged to is gone. In a scope, the child’s failure is the parent’s failure, which means normal error handling works again.
The name comes from Martin Sústrik and the argument from Nathaniel Smith’s “Notes on structured concurrency”, which frames go-style task spawning as the goto of concurrency: a jump with no return path, defeating the reasoning that block structure gives you.
What the enumeration shows
Two versions of the same program, each with exactly one schedule because the event loop admits no choice.
Fire-and-forget: the parent queues the child and returns, and the child then runs — with orphaned == 1 recording that it did its work after its caller had finished. The ordering is unambiguous and it is the wrong one.
With an await inside the scope: the ordering inverts. The child runs, does its work, and only then does the parent return. orphaned == 0, and childRan == 1 confirming the work was not merely skipped.
Both are single-schedule programs, so this is not a race that sometimes bites. In the fire-and-forget version the child runs after the parent returns in the only ordering that exists.
The numbers worth carrying
- Fire-and-forget: the child runs after the parent returns, in 100% of executions. This is not a race; it is the design.
- The failure is silent in three distinct ways: no error (the exception has no stack to propagate to), no timing signal (the caller returned promptly), and no shutdown safety (the work is killed if the process exits).
- Cost of the fix: the parent now takes as long as its slowest child, which is the honest number and is usually what the caller assumed anyway.
Where it breaks down
Sometimes you genuinely want a task to outlive its caller. A background refresh, a metrics flush, a cache warm. Structured concurrency does not forbid this — it requires the task to be owned by a longer-lived scope rather than by nobody. The application has a top-level scope; that is where such work belongs, and it is then subject to shutdown like everything else.
Cancellation needs cooperation. A scope can only cancel what checks for cancellation. A child doing uninterruptible CPU work, or blocked on a syscall that ignores the signal, will not stop because you asked. This is why cancellation is a request in every runtime that has it, and why the shutdown path needs a timeout and a decision about what to do when it expires.
Fan-out with early failure is subtle. If one child fails and the scope cancels the rest, the cancelled children may be mid-write. Scoped cancellation gives you a place to handle that; it does not make partial work atomic.
JavaScript has no built-in scope. AbortController gives you the signal and nothing propagates it for you. This is the one major ecosystem where the pattern has to be assembled by hand, which is why it is the one where orphaned tasks are most common.
What people get wrong
“It returned, so it is done.” The most expensive assumption on this page. A function that starts work and returns has told you nothing about that work.
“The unhandled rejection warning will catch it.” It fires at process level, long after the fact, with no context about which caller abandoned what. And a task that succeeds silently produces no warning at all while still being wrong.
“void somePromise() documents the intent.” It documents that you meant to ignore it. It does not make ignoring it safe.
“Cancellation kills the task.” It asks. What actually happens depends on whether the task is looking.
In production
The pattern by ecosystem: Kotlin coroutines have coroutineScope and structured concurrency is the default — a scope will not complete while children run. Python has asyncio.TaskGroup (3.11) and Trio’s nurseries, which is where the idea entered mainstream practice. Java has StructuredTaskScope (JEP 453, preview). Go has errgroup.Group, which waits and collects the first error and is the closest the language comes; a bare go f() is precisely the unstructured form. Swift has task groups and async let. C# has no scope construct but CancellationToken conventions approximate one.
In JavaScript, the assembly is manual and worth doing: keep the promises you create, await Promise.all on them before the enclosing operation returns, and thread an AbortSignal through so cancellation can reach them. Most server frameworks give you a per-request lifetime; making background work belong to it is the practical version of this page.
The rule to carry: every task has an owner, and the owner waits. If nothing is waiting for it, that is a decision — and it should be made explicitly, at a scope that lives long enough to be responsible for it.
The follow-up questions
“What is wrong with starting a task and not awaiting it?” — Its errors go nowhere, its lifetime is unbounded, and shutdown kills it mid-flight. Name all three; each is independently sufficient.
“When is fire-and-forget acceptable?” — When the task is owned by a longer-lived scope that will wait for it at shutdown. Never when it is owned by nobody.
“How does cancellation work?” — Cooperatively. The child must check. Say what happens when it does not.
“Your framework has no task scope. What do you do?” — Keep the promises, await them at the boundary, thread a signal through. The manual version, stated as a discipline.
In an interview
The reason nurseries, task scopes and errgroups exist, and a good probe for whether someone has debugged a leak they could not see.
- structured concurrency
- cancellation
- task scope
- orphan
Run these next
- Promise.all against sequential awaitSequential await does not merely take longer — the second request has not been sent, because await suspended the function that would have sent it. The saving is overlap, not speed.
- The event loopThe loop finishes what it started, drains every microtask, then takes one macrotask. That discipline has no choices in it — which is why the outcome space has exactly one member.
- Thread pool deadlockNo lock is involved and neither task is wrong on its own. The bug is the shape — a dependency from a resource back onto itself — so the fix is a second pool, not a bigger number.