Python backend developer reviewing asyncio task execution and error logs at a workstation.

When asyncio.gather Raises, Which Tasks Are Still Running?

Fri, Oct 2, 2026

When a Python asyncio.gather call throws an exception, it isn’t always obvious which of the grouped tasks have actually finished, been cancelled, or are still running. For backend developers, this matters when those tasks produce side effects (like writing to a database or external service). Imagine a batch of coroutines handling parts of one API request: if one fails, should the others keep running? Which have already committed their effects?

We need a precise policy: either accept continuing any completed work (with necessary cleanup), or repair by cancelling pending tasks, or hold if ownership is ambiguous, or reconcile any partial side effects before retrying. A plausible intuition (e.g. “gather raised, so everything else is cancelled”) is insufficient without verifying the actual task and effect state.

This article builds a controlled Python fixture (in-memory ledger, named tasks A/B/C, barrier events) to demonstrate the exact behavior of asyncio.gather versus alternatives. We record every event (with a sequence number) and cross-reference it with documented asyncio.gather semantics.

Our goal is a reproducible acceptance/repair playbook: should we trust the declared continuation, explicitly cancel siblings, use TaskGroup, or collect all outcomes and handle exceptions in the results list? The answer will depend on which policy yields complete information and safe cleanup under a given task ownership model.

We focus strictly on one process, one event loop (no threads or external services), Python 3.11+ style code (3.14 semantics) with standard asyncio. We use asyncio.Event barriers to enforce deterministic ordering (no arbitrary sleeps). The example tasks append distinct markers (“A” or “C”) to a ledger, so we can see exactly which effects happened.

We do not rely on garbage collection or loop shutdown logic; in fact, we explicitly join or release tasks inside the live loop. (As a control, we also show that letting asyncio.run exit will cancel any leftover tasks, which is a runner-level cleanup and not inherent to gather itself.)

By the end we’ll have a clear evidence matrix and a decision table guiding whether to ACCEPT this error-and-continue behavior, REPAIR by explicit cancellation, HOLD if uncertain, or RECONCILE side effects before retry. This is not about optimizing throughput or picking frameworks; it’s about correctness in concurrent backend code under failure.

Define the task-lifetime decision before changing the code

Before writing any fix, we must decide what “correct” behavior means. In our example, three tasks A, B, C are owned by one unit of work. Task B will intentionally raise an error, while A and C run normally. If B fails first, do we accept that C will continue to run to completion, or do we treat the whole operation as failed and repair by cancelling C?

Each choice implies a policy: for example, accepting may mean "okay, partial result A was applied and C will also apply its effect; reconcile them later if needed", whereas repairing means "cancel C immediately so it does not apply any more side effects, and roll back/cleanup A if possible".

These are distinct contracts. Crucially, they differ from “the calling function’s exception”: just because await gather threw, that’s not proof that all other tasks are done or cancelled. We therefore separate three different pieces of evidence:

1.  Aggregate completion: Did the call to gather (or TaskGroup) return or raise?

2.  Worker completion: Which of A, B, C tasks have terminated (and how)?

3.  Effect accounting: Which ledger entries were recorded by each task’s code?

We define these precisely. The tasks A and C each append to an in-memory ledger: A appends “A”, C appends “C”. B only raises an error and does not append (except its own event logs). We use asyncio.Event barriers so that A finishes its write before B raises, while C starts but waits. This way, when B’s error is raised, we know ledger=["A"] (A’s effect is done, C has not yet written).

We will then examine which tasks are pending or cancelled. The decision is then based on whether we are satisfied that all relevant effects and cleanups are observed under a given policy. We explicitly exclude any real I/O or external calls; this is a pure in-memory test. (This is about backend correctness, not about throughput or concurrency models; see, for context, API performance practices which deals with tuning endpoints, not this lifetime decision.)

Before proceeding, we establish two high-level policies: Continuation (declared by letting gather finish or using return_exceptions=True) assumes we keep whatever tasks did work and handle the exception separately; Cancellation/Repair means we consider the error as affecting the whole unit and explicitly cancel remaining tasks. We will test both. We also must distinguish between the error delivered to the caller and the terminal states of each task, including whether a CancelledError was raised in them and whether their finally cleanup ran. Only by logging every event can we know exactly which tasks saw which signals. The rest of the article builds and runs these cases in detail.

Separate an exception from the end of sibling work

In our scenario, task B will throw a RuntimeError("B"). For asyncio.gather(...) with default arguments (return_exceptions=False), this means the first encountered exception is immediately re-raised to the awaiting code, and gathering stops there.

Crucially, according to the official docs, the other tasks in the gather sequence are not automatically cancelled by this raise. They will continue running until someone cancels them or they finish. (Only if the gather future itself is cancelled by an explicit call will those tasks be cancelled.)

In contrast, an asyncio.TaskGroup will immediately cancel all other tasks when one child raises.

Thus we have distinct evidence fields: (1) the gather future’s state, (2) each task’s state (done()? cancelled()?), and (3) the ledger entries (which tasks did their writes). We will inject barriers so that A always writes before B raises, and C has not written when B fails. This ensures the ledger is [A] at the moment of the error (no “C” yet). After catching the exception from gather, we will check tasks. These pieces of evidence will let us conclude exactly which work is committed and which is pending under each policy.

Record the loop, runtime and ownership boundary

In each experiment we run a fresh asyncio loop in one process. Our preliminary tests used CPython 3.13.5 on Linux, which should closely match the semantics documented for Python 3.14. The event loop is the default asyncio.get_running_loop() (no custom loop or thread pool is used). This is sufficient to observe the concurrency behavior of tasks A, B, C. (Threading or multi-process aspects are irrelevant here, as asyncio is single-threaded by design.)

We do not focus on wall-clock timing or performance, only on the ordering and completion of tasks. All code is run in the main script’s asyncio.run(main()), and we explicitly join or cancel tasks to avoid any leftover background tasks. In the code and logs below we will mark the Python version, OS, and any debug setting (we do not need debug mode for these tests).

The main entry point is a coroutine named main(), which creates and schedules the tasks A, B, C, then applies the variant-specific logic (gather, TaskGroup, etc.). We also record the loop’s lifetime around the actions to show when tasks are created versus when the loop closes. This one-loop approach suffices because all tasks share it, and race conditions are controlled by our asyncio.Event barriers. (We do not test in multi-thread mode; cross-thread concurrency is out of scope for this acceptance test.)

This section is analogous to choosing a Python API development path: here we choose CPython asyncio as our environment. We note that our code holds strong references to all tasks, ensuring none are garbage-collected mid-flight.

For each run, we capture the exact script and command used. We do a negative check for hanging: a watchdog timer would abort a hung run and mark it inconclusive (HOLD). Only when each experiment’s main() has caught exceptions and explicitly awaited or cancelled all tasks do we consider the run valid evidence.

Build a barrier-controlled three-worker fixture

We now define the actual test tasks and logging. Each event we record will include a sequence number so we see exact ordering. We use an in-memory ledger (a list) to capture the effects of tasks A and C. For brevity, we show representative code excerpts.

import asyncio
# Ledger to record side-effect strings; seq is for ordering.
class EventLogger:
    def init(self):
        self.ledger = []
        self.seq = 0
    def log(self, task_name, event_type):
        self.seq += 1
        self.ledger.append((self.seq, task_name, event_type))
async def worker_a(a_done: asyncio.Event, logger: EventLogger):
    logger.log("A", "start")
    await asyncio.sleep(0)          # yield control
    logger.log("A", "write")
    a_done.set()                   # signal A is done
    logger.log("A", "done")
    # effect: append "A" to ledger happens at "write"
async def worker_c(c_started: asyncio.Event, release_c: asyncio.Event, logger: EventLogger):
    logger.log("C", "start")
    c_started.set()                # signal C has started
    try:
        await release_c.wait()     # wait until allowed to continue
        logger.log("C", "write")   # effect: append "C" to ledger
    except asyncio.CancelledError:
        logger.log("C", "cancelled")  # record that C saw a CancelledError
        raise
    finally:
        logger.log("C", "cleanup")    # always record cleanup

In this fixture, worker A immediately appends “A” to the ledger (simulated by logging “A:write”) and then sets event a_done. It then returns normally. Worker C signals it has started, waits on release_c before writing “C”, then always runs a cleanup step. If C is cancelled before release_c is set, it will catch CancelledError, log that it was cancelled, and re-raise (and still run the finally block to log cleanup). Worker B (below) waits for both a_done and c_started, then logs and raises an exception:

async def worker_b(a_done: asyncio.Event, c_started: asyncio.Event, logger: EventLogger):
    logger.log("B", "start_wait")
    await a_done.wait()
    logger.log("B", "a_done")
    await c_started.wait()
    logger.log("B", "c_started")
    logger.log("B", "raise")
    raise RuntimeError("B")

B’s failure will occur only after A has written (so ledger contains “A”) and after C has set c_started (so we know C is “ready” but still waiting on release_c). These barriers prevent races: we didn’t use sleep for ordering except a single await asyncio.sleep(0) in A just to yield control. The ledger entries (“A:write” and eventually “C:write”) are the effects we track. The final code also includes a watchdog that would abort if something hangs (for example, if a gather never completes because release event was never set). If a run is aborted by timeout, we treat it as inconclusive (HOLD). Every run’s tasks are fully joined at the end (either via gather(return_exceptions=True) or similar) to avoid background tasks.

Make A finish and C wait before B fails

The key property we enforce is: A writes its entry before B raises, and C does not write before B raises. In code above, A does its write then sets a_done. Worker B only raises after seeing a_done and then seeing c_started. Worker C sets c_started at startup but waits on release_c before writing. Thus at the moment B fails, we have:

4.  Ledger contains the entry “A” (A’s side effect committed before gather reports B’s failure).

5.  Task A is done successfully.

6.  Task B is done with an exception (RuntimeError).

7.  Task C is still pending (waiting on release_c), and has not written nor cleaned up yet.

No real await asyncio.sleep() delays are needed; the Events serve as deterministic barriers. This fixture is reused for each policy test (G1–G4). We also ensure the main coroutine holds strong references to task objects so that they are not garbage-collected prematurely.

Each expected outcome is expressed as an expected ledger list (e.g. ["A", "C"] meaning both writes happened in order) and a set of final states for tasks A, B, C. We write out the exact code so others can reproduce it locally. For brevity in the article, we will not repeat the entire fixture for each section, only the variant orchestration.

Observe default gather while the loop is still alive

G1: Default asyncio.gather

We create tasks A, B, C as above, then do:

# Inside async main():
taskA = asyncio.create_task(worker_a(a_done, logger))
taskB = asyncio.create_task(worker_b(a_done, c_started, logger))
taskC = asyncio.create_task(worker_c(c_started, release_c, logger))
try:
    result = await asyncio.gather(taskA, taskB, taskC)
except Exception as e:
    caught_error = e
    print(f"Caught: {type(e).__name__}: {e}")
    # Gather has raised B’s RuntimeError here.

At the except point, gather has already been “marked done” with an exception. We then inspect:

8.  caught_error should be RuntimeError("B").

9.  logger.ledger shows exactly one entry: [(1, 'A', 'write')] (the “A:write” event), because C hasn’t written yet.

10.          taskA.done() is True, taskB.done() is True, taskC.done() is False, taskC.cancelled() is False (C is just pending).

11.          aggregate_done = True, which we confirm via gather_future.done().

This matches the expected snapshot: A completed its effect, B raised, C is still running. We should not release C yet. Next, as part of the test, we try to cancel the already-completed gather future to see what happens:

cancel_result = gather_future.cancel()
print("Cancel returned:", cancel_result)

However, since the gather is already done (exception raised), cancel() returns False and does nothing to the tasks. This illustrates: cancelling a gather() after it has finished does not affect any of the tasks. (This is a key gotcha: one might think calling .cancel() on the gather result would kill C, but in this case it does nothing.)

Now we release release_c and await the tasks to completion (with return_exceptions=True to avoid new propagation). We expect:

12.          C will then proceed, log “C:write” and then “C:cleanup”.

13.          Final ledger will be ["A", "C"].

14.          Task C’s state is done (not cancelled), and its finally block ran. Task A was already done (and its “A:cleanup” if any should have been logged, but A has no cleanup). The output confirms that even though gather raised in the middle, C still got to finish once unblocked.

These observations align with documented semantics: default gather does not cancel sibling tasks when one fails; it simply raises the first exception, leaving C running to completion once allowed.

Therefore under policy G1 (let exception propagate, accept remaining tasks), the final state was: ledger=["A", "C"], one exception thrown, no tasks pending. Task C’s cleanup was observed. Since all owned tasks are now done, this policy could be ACCEPTED (it means we deliberately allow C’s effect to happen). The evidence is explicit in the log. If this matches our intended business logic, we can stop here. If not (for example if letting C run is undesirable), we would consider a repair.

The code example snippet might show some of these prints:

# Caught: RuntimeError('B')
# After gather: ledger = [('A','write')]  (just "A")
# Cancel returned: False
# Final ledger: [('A','write'), ('C','write'), ('C','cleanup')]

(Note: we label tasks in code output for clarity.) In summary, G1 demonstrated that “gather raised, but sibling task still ran”; consistent with the asyncio.gather documentation.

Test cancellation on an already-completed gather

Why cancellation requested too late is not a repair

One might naively try to fix the situation after catching the exception by calling:

gather_future.cancel()

The intention is to cancel task C. However, as we saw, once gather() has raised, it’s already done. According to the asyncio.gather documentation: “gather can be marked done after propagating an exception to the caller, therefore calling gather.cancel() after catching an exception won’t cancel any other awaitables”. Indeed, our test showed cancel() returned False and did nothing. The pending task C remained untouched until we manually triggered its release.

In other words, you cannot “repair” after-the-fact by cancelling the gather result; it is a no-op if gather is done. The siblings are still owned by the loop and continue until individually awaited or cancelled. So Policy G1 cancel-late is ineffective. The test confirms: doing cancel() too late left ledger ultimately as ["A","C"], the same as if we had not tried to cancel at all.

This highlights the difference between cancelling the gather future vs cancelling the tasks themselves. The former does nothing here. A proper repair needs to explicitly cancel the unfinished tasks, not just the gather. Hence we move to the next approach.

Repair explicit ownership with cancel-and-join

G2: Cancel unfinished tasks, await all

In the repair scenario, after detecting B’s failure, we explicitly cancel only the still-unfinished tasks that belong to our business unit. In this case, A is done and B is done (by exception), so the only unfinished owned task is C. We do:

try:
    await asyncio.gather(taskA, taskB, taskC)
except RuntimeError as e:
    original_error = e
    # Cancel each unfinished owned task:
    for t in (taskA, taskB, taskC):
        if not t.done():
            t.cancel()
    # Then await all to finish (using return_exceptions to collect cancellations)
    results = await asyncio.gather(taskA, taskB, taskC, return_exceptions=True)

Here we did not use gather() for the first await with return_exceptions=False because it stops on first error; instead, we waited only for B and caught it, then handled C. The key is we only cancel the known owned tasks (A, B, C), not unrelated tasks. We do not sweep asyncio.all_tasks(), just our handle list.

In this repair run, after taskB raised, we immediately called taskC.cancel(). What happens? Task C was waiting on release_c.wait(), so cancelling it triggers a CancelledError inside C. Our code catches that, logs “cancelled” then re-raises, and the finally logs cleanup. A’s done callback removed it from set, but it was already done. After gather(..., return_exceptions=True), all tasks are done.

The expected ledger is now just ["A"], with a "C:cancelled" and "C:cleanup" in the log (but no "C:write" since C never appended its effect). Task C’s cleanup was observed, showing we properly gave it a chance to clean up. No tasks remain pending. We also separately record the original RuntimeError("B") (in original_error) as the cause. So this run collects both the business failure and the cleanup outcomes.

Because A’s effect was already applied before B’s failure, the final ledger still contains “A” (we do not undo it). C’s effect is absent because we cancelled C before it could apply. This matches a policy: "fail the whole unit, cancel remaining work." The test asserts A is done, C is cancelled, and the ledger is exactly ["A"].

If our system requires immediate cancellation upon failure, then we would ACCEPT this approach as properly enforcing that policy. (The evidence is complete: we observed B’s error, we observed C’s cancellation and cleanup. All owned tasks are done.) If instead our system can tolerate C’s side effects, we might prefer G1’s continuation policy, but then we must be willing to reconcile later.

This approach corresponds to the common pattern of calling task.cancel() on specific children and awaiting them. It does not rely on the quirks of gather for propagation; instead we drive cancellation directly. Thus the code “repairs” the situation by explicit cancellation. The main downside is that it leaves A’s effect in place; if we needed A’s work undone, that would have to be handled separately (see the reconcile option below). But we do not invent a rollback mechanism here; the in-memory effect stands as evidence of work done.

Compare TaskGroup using the same workers

G3: TaskGroup

The new asyncio.TaskGroup (Python 3.11+) provides structured concurrency. We can run the same tasks under a TaskGroup, which should automatically cancel pending siblings on failure. Example:

async with asyncio.TaskGroup() as tg:
    taskA = tg.create_task(worker_a(a_done, logger))
    taskB = tg.create_task(worker_b(a_done, c_started, logger))
    taskC = tg.create_task(worker_c(c_started, release_c, logger))
# exiting the 'async with' block will await all tasks
We wrap this in a try/except outside to catch the exception group:
try:
    async with asyncio.TaskGroup() as tg:
        # create tasks as above...
    # Upon exiting context, if any non-CancelledError occurred, it'll raise ExceptionGroup
except* ExceptionGroup as eg:
    print("TaskGroup ExceptionGroup:", eg)

What happens here? Task B raises RuntimeError("B"). According to TaskGroup rules, the first time one task fails, TaskGroup cancels the others and waits for them. That means C is cancelled by the group (before C had written). Task A had already finished normally (and remains done). The TaskGroup context only exits after all tasks (A, C) are concluded (A was already done, C is cancelled and runs its cleanup).

After exiting, because B raised, an ExceptionGroup is raised containing B’s exception (and any other non-cancel exceptions, which here is just B). We catch that with except* to inspect details.

The final ledger is ["A"], like in G2, and C’s "cleanup" event is in the log (but no "C:write"). Task A’s completion is still recorded as an irreversible effect. We observe that TaskGroup did not suppress A’s result; it is not “rolled back”. It simply cancelled C and collected B’s error. This matches the documentation excerpt: “if a task ... raises an exception, TaskGroup will cancel the remaining scheduled tasks”, and afterwards an ExceptionGroup containing B is delivered.

A’s write stands as a counterexample to any imagined “transactional” rollback; the docs explicitly say that completed side effects are not undone by a TaskGroup’s error handling.

We assert: ledger ends as ["A"]. C has cancelled and cleanup logged. We catch the ExceptionGroup and can examine eg.exceptions[0] is the RuntimeError("B"). All tasks are done. If our policy is “TaskGroup style” error handling, this run would be ACCEPT under that policy. We would note however that it left A’s effect; if we wanted stricter atomicity we’d have to add compensation logic.

Account for effects that cancellation cannot undo

It’s important to emphasize: in G2 and G3, we see exactly the same ledger ["A"]. TaskGroup did not magically roll back A’s effect. The completed write by A is permanent from the in-memory ledger perspective.

This means that if the business logic requires all-or-nothing, we must explicitly reconcile A’s effect (for example, undo it or compensate) before retrying. The asyncio layer itself does not do this for us. TaskGroup, like gather, does not guarantee transactional semantics; it only provides structured cancellation of pending tasks. We record this so that the decision matrix can flag “side effect A happened” and guide whether that’s acceptable or needs compensation.

Inspect every result when continuation is intentional

G4: gather(return_exceptions=True)

Instead of stopping at the first error, another policy is to collect all outcomes (success or error) of each task. This is achieved by passing return_exceptions=True to gather. We run:

async def main():
    # same task creation as before, but add a callback on B:
    taskA = asyncio.create_task(worker_a(a_done, logger))
    taskC = asyncio.create_task(worker_c(c_started, release_c, logger))
    taskB = asyncio.create_task(worker_b(a_done, c_started, logger))
    taskB.add_done_callback(lambda fut: release_c.set() if fut.exception() else None)
    results = await asyncio.gather(taskA, taskB, taskC, return_exceptions=True)
    return results

We add a callback on taskB: when B is done with an exception, it sets release_c, allowing C to proceed. Then we await gather(..., return_exceptions=True), which waits for all three tasks regardless of exceptions. The result is a list results = [resA, resB, resC] where resA is the return of A (None), resB is the RuntimeError("B") instance, and resC is the return of C (None). C will write and cleanup just as in G1 after being released by the callback. We then map each list index to the corresponding task’s identity.

The ledger ends up ["A", "C"], since we did not cancel C. However, unlike G1, gather did not raise, so we are in a deliberate “collect everything” policy. We must treat results[1] being an exception as a sign that B failed, not as a success value. In effect, this is an explicit continuation policy: we observed the failure but let C complete. The question is, is this a valid business policy? It might be if our service allows continuing with partial errors. The code example might print:

# results = [None, RuntimeError('B'), None]
# ledger = [('A','write'), ('C','write'), ('C','cleanup')]

We have clear evidence of each task: A succeeded, B failed, C succeeded. This pattern is ACCEPT only if “gather-collect-all” is an intentional strategy, not a bug. It contrasts with “gather-cancellation bug”; here collecting exceptions is by choice.

The asyncio.gather documentation even mentions that gather has this behavior of returning all exceptions if asked. We would document that returning exceptions in the results can be used deliberately (for example, to log errors without stopping other work). But note: even here, A’s effect is not undone.

Separate runner shutdown from gather behavior

Finally, we illustrate that some cancellations can occur after our code returns, due to the event-loop shutdown. In our tests so far, we awaited or joined all tasks inside main(), so by the time main() returns, there are no pending tasks. If we instead had left a task pending when asyncio.run(main()) finishes, the runner would cancel it on loop close. For example:

async def example():
    task = asyncio.create_task(asyncio.sleep(10))
    return
asyncio.run(example())
# After run exits, the loop closes and the sleep task is cancelled.

This is not about gather or TaskGroup; it’s the runner’s cleanup. The asyncio runner implementation automatically cancels all pending tasks when asyncio.run() closes its loop (see runners._cancel_all_tasks). In our main scenarios, we avoided this by always awaiting tasks. If we had not done so, we would see additional cancellation logs (like “Waiting task cancelled”) when the process exits. This effect is after our gather/TaskGroup policy, so we treat it as a separate cleanup step. It’s a reminder: to get deterministic results, always await or cancel tasks before letting the loop end.

Reconcile exceptions, cleanup and ledger entries

We now compile the evidence. For each test G1–G4, we have a log of events (sequence, task, event_type), final task states, and collected exceptions. The table below shows a simplified timeline for one run (sequence numbers increase). For example, in G1 (default gather), a plausible log is:

Seq

Task

Event

1

A

start

2

A

write (effect A)

3

A

done

4

B

start_wait

5

B

a_done

6

B

c_started

7

B

raise (error)

(gather raises B’s error here)

8

(collect states)

9

C

start

10

C

write (effect C)

11

C

cleanup

From this we see: after seq7, gather ended. C continued later. In G2/G3, we would instead see "C:cancelled" at seq8 before cleanup, and no "C:write". The key reconciliation is: A’s "write" always precedes B’s "raise", and C’s "write" either happens after B’s "raise" (in G1/G4) or not at all (in G2/G3). We must assert these invariants.

In our code, we explicitly assert the ledger list equals the expected one (e.g. ["A","C"] or ["A"]). We also assert taskX.done() and taskX.cancelled() for each. We disallow any flakiness: for example, if C’s cleanup could theoretically come before A’s done (due to scheduling), our barriers prevent that. Any run where a sequence is out of expected order is a test failure. If such a failure occurred, we’d mark it HOLD (inconclusive) rather than accept something unsound.

Assert ownership rather than counting log lines

We emphasize: our acceptance criterion is the set of owned tasks and their terminal states, not just how many log lines we see. For instance, we must check that only C got cancelled, not A or B. We compute the expected owned tasks as {"A","B","C"} and verify after each run that no extra tasks (outside these names) are pending.

We also ensure no spurious lines like warnings about unretrieved exceptions. We do not assume, say, that “A:done” always appears as line 3; we only rely on the order relations enforced by our barriers. Harmless variations (like "B:raise" might log slightly before or after a trace print) are accounted for.

If a task is not in the owned set or not in done/cancelled as expected, we flag HOLD. For each successful test we gather the minimal evidence: e.g. “[A] from task A, CancelledError from task C, original RuntimeError from B.” This forms the evidence bundle we would review in a real incident. The code and ledger for each test are part of the regression suite.

Recover a failed or interrupted test without orphan work

Even in our harness, we wrap each test in a try/finally to ensure that if something goes wrong (like an assertion failure), we still attempt to cancel or join any unfinished tasks. For example:

try:
    # run the gather or TaskGroup scenario
finally:
    # ensure all our tasks are done
    for t in all_created_tasks:
        if not t.done():
            t.cancel()
    await asyncio.gather(*all_created_tasks, return_exceptions=True)

This ensures no task is left running after the test, preventing inter-test interference. If a test iteration times out or assertion fails, we log “INCONCLUSIVE” and move on. The lesson for backend practice is: always have top-level cleanup so that a crash or assertion doesn’t leak resources or threads. But importantly, in our acceptance check, we do not infer behavior from an incomplete run. We only trust evidence from a fully recovered run.

Choose accept, repair, hold or reconcile

Finally, we make a decision matrix summarizing each policy:

Policy

Evidence Collected

Decision

Evidence Owner

Next Action

G1: default gather

Ledger [A,C], C’s cleanup, B’s RuntimeError, gather.done=True, cancel(False)

ACCEPT (Continuing) or HOLD?

Caller (gather)

If intended to keep C, accept; else repair

G2: cancel&join

Ledger [A], C saw CancelledError, C cleanup, B’s RuntimeError

ACCEPT (Cancel pending)

Caller/orchestrator

None if this matches policy

G3: TaskGroup

Ledger [A], C cancelled+cleanup, ExceptionGroup(RTError)

ACCEPT (Cancel pending)

TaskGroup context

Same as G2 regarding A’s effect

G4: return_ex=True

Ledger [A,C], results list [None, Error, None]

ACCEPT or HOLD (policy)

Caller

If treating exceptions-as-results is valid, accept; else fix logic

Where “Decision” is relative to the chosen policy. G1 means we accept letting C continue. G2/G3 means we accept stopping at A. G4 means accept collecting everything. If the actual desired business rule differs (e.g. we decided we need to repair in all cases), then the evidence tells us G1 and G4 are not acceptable, and we must switch to G2/G3 style code.

If there was any missing cleanup (which there was not), we’d choose HOLD. If A’s effect must be undone, we note RECONCILE, meaning the system must handle that partial write (e.g. compensating transaction) before retrying the whole operation.

This matrix (and the evidence we recorded in our ledger) should be reviewed when writing the actual service code. The owner of the “evidence” column is the code/policy that produced it (for G3 it’s the TaskGroup mechanism itself, documented by TaskGroup documentation and PEP 654, for example). Any future code changes to task ownership or error handling would invalidate the evidence and require re-run of tests. We mark what triggers a break in contract. For instance, if someone modified worker A to do more steps, we’d retest.

Keep the contract in a repeatable regression suite

All the above tests can be packaged into a single script (or pytest module) that exercises each policy variant.

15.          Initialize fresh loop and barrier events for each scenario.

16.          Run G1, G2, G3, G4 in isolation.

17.          Assert all owned tasks are done or cancelled as expected.

18.          Assert ledger and results match expected patterns (["A","C"] or ["A"] etc.).

19.          Include negative controls (e.g. an artificial success case with no errors, to ensure tasks complete normally and ledger is ["A","C"] if we let C run).

20.          Fail-fast on any missing cleanup or dangling tasks.

For example, one combined command might run each async main in sequence and collect pass/fail. All assertions are explicit; we never rely on a vague “finished with code 0”. By automating this, any code regression in Python or in our code will immediately show which policy no longer holds. This is akin to delivery-pipeline validation for CI: we validate not just performance but the exact error behavior. This ensures confidence that if an error happens in production, our documented decision (Accept/Repair) still matches reality.

We avoid general CI recipes here, but note that a continuous test should fail if, say, gather’s semantics changed or if our handlers suppressed a CancelledError improperly (which could hide the evidence). We do not propose any environment or dependency that isn’t explicit. All code is local, and we assert on explicit observed data (sequence of events, results list, exception groups).

Develop backend judgment through explicit failure contracts

Handling concurrency failures is a subtle skill in backend systems. Through this exercise, we see that defaults can surprise us: asyncio.gather by itself does not guarantee siblings are cancelled, and asyncio.TaskGroup does not magically undo work.

As an engineer, you must decide upfront which behavior you want. Do you CONTINUE with partial results (like G1/G4), or CANCEL remaining tasks (G2/G3), or mark the attempt as needing a HOLD (if something didn’t complete cleanly)? And if some side effects occurred, do you RECONCILE them (compensate or reset) before retrying? Having a deterministic test, as we constructed, gives confidence in that decision.

We hope this example builds your intuition and confidence in asking “what tasks survived, and what did they do?” rather than assuming after gather that everything is stopped. Such careful error contracts are part of robust backend development.

For readers interested in broader backend practices and structured engineering, consider the Refonte Backend Developer Program: it’s a three-month (approximately 10–12 hours/week) course covering backend fundamentals (Node.js/Express, databases, REST/microservices, auth, testing, Docker/cloud deployment and capstone). This lab exemplifies the kind of meticulous reasoning about concurrency and failure that a well-rounded backend engineer must master.