A health-check command probes twenty hosts; a deploy command uploads forty files; a sync command fetches a hundred pages. Doing these concurrently with asyncio is easy to start — asyncio.gather or a handful of create_task calls — and surprisingly easy to get wrong at the edges. What happens to the other nineteen probes when one raises? Who notices an exception in a task nobody awaited? What does the user see when three things fail at once? asyncio.TaskGroup (Python 3.11+) answers these with structured concurrency: tasks live inside an async with block, the block does not exit until all of them have finished, a failure cancels the siblings, and every failure is reported together in an ExceptionGroup. This guide uses TaskGroups for the two modes a CLI usually needs — fail fast (make-style, stop at the first error) and keep going (make -k, run everything and report all failures) — and adds concurrency limits, timeouts, except* handling and tests. It belongs to the concurrency and async topic.
Prerequisites
- Python 3.11 or newer for
TaskGroup,asyncio.timeoutandexcept*. - Running async code in Typer and Click for calling async code from commands.
What structure buys
With unstructured tasks, a task can outlive the function that created it, and its exception surfaces — if at all — as a “Task exception was never retrieved” warning at interpreter exit. With gather, the first exception is returned to the caller while the other tasks keep running in the background unless you cancel them yourself. A TaskGroup removes both problems by construction: the async with block is a boundary that no task crosses. When any task fails with an ordinary exception, the group cancels the others, waits for them to finish cancelling, and raises an ExceptionGroup containing every failure. Ctrl+C and outer cancellation propagate into the group the same way.
The recipe
# src/mytool/checks.py
from __future__ import annotations
import asyncio
from collections.abc import Awaitable, Callable, Iterable
from dataclasses import dataclass
class CheckFailed(Exception):
"""An expected, reportable failure of one check."""
@dataclass
class Outcome:
name: str
ok: bool
detail: str
async def run_fail_fast(names: Iterable[str], check: Callable[[str], Awaitable[str]],
*, limit: int = 8, timeout: float = 30) -> dict[str, str]:
"""All-or-nothing: the first failure cancels every other check."""
gate = asyncio.Semaphore(limit)
results: dict[str, str] = {}
async def one(name: str) -> None:
async with gate:
results[name] = await check(name)
async with asyncio.timeout(timeout):
async with asyncio.TaskGroup() as tg:
for name in names:
tg.create_task(one(name), name=f"check:{name}")
return results
async def run_keep_going(names: Iterable[str], check: Callable[[str], Awaitable[str]],
*, limit: int = 8, per_check_timeout: float = 10) -> list[Outcome]:
"""Best effort: every check runs; expected failures become outcomes, bugs still propagate."""
gate = asyncio.Semaphore(limit)
async def one(name: str) -> Outcome:
async with gate:
try:
async with asyncio.timeout(per_check_timeout):
return Outcome(name, True, await check(name))
except CheckFailed as exc:
return Outcome(name, False, str(exc))
except TimeoutError:
return Outcome(name, False, f"timed out after {per_check_timeout:g}s")
async with asyncio.TaskGroup() as tg:
tasks = [tg.create_task(one(name)) for name in names]
return [task.result() for task in tasks]
def describe(group: BaseExceptionGroup) -> list[str]:
"""Flatten a (possibly nested) exception group into one line per leaf exception."""
lines = []
for exc in group.exceptions:
if isinstance(exc, BaseExceptionGroup):
lines.extend(describe(exc))
else:
lines.append(f"{type(exc).__name__}: {exc}")
return lines
Fail fast
run_fail_fast is the TaskGroup in its natural form. Each task acquires a semaphore before working, which caps concurrency at limit without changing the structure — twenty hosts, at most eight probes in flight. The whole group sits inside asyncio.timeout, so a hung probe cannot hold the command forever: when the deadline passes, the group is cancelled and the timeout surfaces as TimeoutError. When a check raises CheckFailed, its siblings are cancelled immediately, which is what the user wants from a deploy precondition check: no point testing the cache once the database is known to be down.
Keep going
run_keep_going changes one thing: each task catches the expected failure types itself and turns them into an Outcome. Expected failures therefore never reach the group, so nothing is cancelled and every check runs to completion. Each check also gets its own timeout, so one slow host becomes a “timed out” outcome instead of a cancelled run. Results are read from the tasks in creation order, so the report lists hosts in the order the user gave them, not the order they finished.
Unexpected exceptions — a KeyError from a bug — are deliberately not caught. They still propagate, cancel the siblings and fail the command, because continuing past a bug produces results nobody should trust. The distinction between expected and unexpected errors is the same one made in designing an exception hierarchy for a CLI.
Reporting exception groups
# src/mytool/cli.py
import asyncio
from typing import Annotated
import typer
from mytool.checks import CheckFailed, describe, run_fail_fast, run_keep_going
app = typer.Typer()
HOSTS = ["api", "db", "cache", "search"]
async def probe(name: str) -> str:
await asyncio.sleep(0.01)
if name == "search":
raise CheckFailed("connection refused")
return "ok"
@app.callback()
def main() -> None:
"""Health tool."""
@app.command()
def check(keep_going: Annotated[bool, typer.Option("--keep-going", "-k")] = False) -> None:
"""Check every host; stop at the first failure unless --keep-going."""
if keep_going:
outcomes = asyncio.run(run_keep_going(HOSTS, probe))
for o in outcomes:
typer.echo(f"{o.name:<8} {'ok' if o.ok else 'FAIL'} {o.detail if not o.ok else ''}".rstrip())
failed = sum(not o.ok for o in outcomes)
if failed:
typer.echo(f"{failed} of {len(outcomes)} checks failed", err=True)
raise typer.Exit(1)
return
try:
results = asyncio.run(run_fail_fast(HOSTS, probe))
except* CheckFailed as group:
for line in describe(group):
typer.echo(f"error: {line}", err=True)
raise typer.Exit(1)
except* TimeoutError:
typer.echo("error: checks did not finish in time", err=True)
raise typer.Exit(1)
typer.echo(f"all {len(results)} checks passed")
except* matches exception types inside a group: the CheckFailed branch receives a sub-group containing only those, and the TimeoutError branch handles the timeout (except* also matches a bare exception that is not in a group). Anything not matched — a bug — propagates as before and reaches the crash handler. describe flattens nested groups into one line per failure, because users want “error: CheckFailed: connection refused”, not a nested traceback tree. When several checks fail at nearly the same moment, a fail-fast group can contain more than one failure; the loop reports them all.
Choosing the mode
Offer both and pick a default by asking what a partial run is worth. Preconditions, deploy steps and anything that changes state default to fail fast: once something is wrong, continuing wastes time or makes things worse. Read-only surveys — health checks, linting many files, downloading a batch where each file is independent — default to keep going, or at least offer -k/--keep-going, the flag make users know. Exit non-zero in both modes when anything failed, so scripts notice; print the summary (“1 of 4 checks failed”) on stderr.
UX considerations
- Bound concurrency. A semaphore around the work keeps the CLI from opening hundreds of connections; make the limit a
--jobsoption, as in rate limiting concurrent requests in CLIs. - Report in input order. Users scan results against the list they supplied; completion order is noise.
- Show progress for long groups. A live count of finished, failed and running tasks is covered in showing progress for concurrent tasks.
- Clean up in
finally. Cancelled tasks receiveCancelledErrorat their currentawait; partial files and open connections are closed infinallyblocks, as in cancelling async tasks on Ctrl+C. - Never swallow
CancelledError. Anexcept Exceptiondoes not catch it (it derives fromBaseException), but a bareexcept:does and breaks cancellation for the whole group.
Testing the behaviour
Fake checks that sleep, fail or raise bugs on demand make every property testable in milliseconds — sibling cancellation, the concurrency cap, outcome order, per-check and overall timeouts, and the CLI’s messages:
# tests/test_checks.py
import asyncio
import pytest
from typer.testing import CliRunner
from mytool.checks import CheckFailed, describe, run_fail_fast, run_keep_going
from mytool.cli import app
def make_check(fail=(), slow=(), bug=()):
started, cancelled = [], []
async def check(name):
started.append(name)
try:
await asyncio.sleep(1 if name in slow else 0.01)
except asyncio.CancelledError:
cancelled.append(name)
raise
if name in bug:
raise KeyError(name)
if name in fail:
raise CheckFailed(f"{name} is down")
return "ok"
return check, started, cancelled
def test_fail_fast_cancels_siblings():
check, started, cancelled = make_check(fail={"b"}, slow={"a", "c"})
with pytest.raises(ExceptionGroup) as info:
asyncio.run(run_fail_fast(["a", "b", "c"], check))
assert describe(info.value) == ["CheckFailed: b is down"]
assert sorted(cancelled) == ["a", "c"]
def test_concurrency_limit_is_respected():
running = peak = 0
async def check(name):
nonlocal running, peak
running += 1
peak = max(peak, running)
await asyncio.sleep(0.01)
running -= 1
return "ok"
asyncio.run(run_fail_fast([str(n) for n in range(20)], check, limit=3))
assert peak == 3
def test_keep_going_reports_every_outcome_in_order():
check, *_ = make_check(fail={"b"}, slow={"c"})
outcomes = asyncio.run(run_keep_going(["a", "b", "c"], check, per_check_timeout=0.2))
assert [(o.name, o.ok) for o in outcomes] == [("a", True), ("b", False), ("c", False)]
assert outcomes[2].detail == "timed out after 0.2s"
def test_keep_going_still_surfaces_bugs():
check, *_ = make_check(bug={"b"})
with pytest.raises(ExceptionGroup) as info:
asyncio.run(run_keep_going(["a", "b"], check))
assert info.group_contains(KeyError)
def test_overall_timeout():
check, *_ = make_check(slow={"a"})
with pytest.raises(TimeoutError):
asyncio.run(run_fail_fast(["a"], check, timeout=0.05))
def test_cli_modes():
runner = CliRunner()
fast = runner.invoke(app, ["check"])
assert fast.exit_code == 1 and "error: CheckFailed: connection refused" in fast.stderr
full = runner.invoke(app, ["check", "-k"])
assert full.exit_code == 1 and "search FAIL connection refused" in full.stdout
assert "1 of 4 checks failed" in full.stderr
The cancellation test records which checks saw CancelledError, proving that the fail-fast group really stopped the slow siblings rather than waiting for them. ExceptionInfo.group_contains (pytest 8+) checks for an exception type anywhere inside a group without depending on nesting. Every test uses asyncio.run, so no plugin is needed; pytest-asyncio or anyio’s pytest plugin are fine alternatives if your suite already uses them.
Conclusion
asyncio.TaskGroup gives concurrent CLI work a clear shape: tasks start and finish inside one block, failures cancel siblings and arrive together in an ExceptionGroup. Use the group as-is for fail-fast commands, catch expected failures inside each task for keep-going mode while letting bugs propagate, cap concurrency with a semaphore, put deadlines on the group or each task with asyncio.timeout, handle groups with except* and flatten them into one line per failure, report in input order, and test cancellation and limits with fake checks.
Frequently asked questions
Should I still use asyncio.gather?
gather(..., return_exceptions=True) is a compact way to collect results and exceptions together. Without that flag, though, gather raises the first exception while the remaining tasks keep running unsupervised, and you have to cancel them yourself. A TaskGroup handles that automatically, which makes it the safer default for new code.
Do TaskGroups work with threads or processes?
Not directly; they manage asyncio tasks. Blocking work can join a group through asyncio.to_thread(fn), which runs in a thread but is awaited like any coroutine. For thread-only code, see parallelising CLI work with thread pools.
What about Trio and AnyIO?
Trio introduced nurseries, the idea TaskGroup adopts, and AnyIO offers a create_task_group() with the same semantics on top of asyncio or Trio. The recipe translates almost line for line.
How do I add tasks while the group is running?
Pass the group to a task and call tg.create_task from inside it — a crawler that discovers new pages, for example. The group waits for those tasks too.
Can I limit how long cancellation takes?
Cancellation waits for each task’s finally blocks. Keep cleanup short and bounded, and use asyncio.timeout around slow cleanup steps so a stuck task cannot block exit indefinitely.