DevOps engineer reviewing Docker Compose health checks and container restart logs at a workstation.

Your Compose Dependency Is Unhealthy. Which Service Will Restart?

Wed, Oct 7, 2026

Introduction

Consider a deployed Compose project where all containers started and the provider was healthy, yet a later failure in one service caught the operator by surprise: neither the unhealthy provider nor its dependent containers restarted automatically as expected. In this scenario, service provider lost its health state (failing its healthcheck) without exiting. Both its linked and comparison consumers continued running and kept making service requests, so no container exited. What, then, triggers restarts after startup?

This article answers that six-case question in a controlled experiment: two consumers use the same Python-based request loop, both depend on provider healthy at startup, but only one has depends_on.restart: true. We keep the model fixed and observable, and define the expected outcomes for five changes: (B) setting the provider unhealthy, (C) restoring its health, (D) a provider exit, plus (E) an explicit docker compose restart, and (F) the same restart with --no-deps.

We trace each container’s boot UUID and inspect their state. We separate health-only failures (provider unhealthy but running) from real exits and manual commands.

The expected outcomes remain labeled “expected” here until a real Docker host run provides empirical data. Below we use the official Compose docs for dependency and restart semantics, the Docker Engine docs for restart policies and healthchecks, and distinguish Compose’s explicit restart propagation from the container runtime’s restart policy. (See also the Refonte blog post on trace Docker Compose environment values for how we freeze and record our runtime model.)

Separate the four signals in the incident record

In diagnosing our scenario, we track four signals: whether a container is Running or Exited, whether it is Healthy or Unhealthy (per its healthcheck), the container’s boot generation (a UUID printed at startup), and whether it is responding successfully to application requests. None of these signals alone imply the others. A container can remain Running but become Unhealthy (health checks failing on a still-running process). A container may exit due to error and be restarted by the Engine, yielding a new boot ID. Conversely, an explicit Compose restart command will generate a new boot ID even without a container error. Successful HTTP requests depend on both provider readiness and the container’s running state.

In our test, we define provider and two equivalent consumers. Both consumers are bound to wait for provider health at startup (condition: service_healthy), but one consumer’s depends_on block includes restart: true (propagating explicit restarts) and the other has restart: false.

The provider-level restart field (such as on-failure:2) controls automatic restarts by the Docker Engine upon container exit. The dependency-level restart: field under depends_on only applies to explicit docker compose operations and excludes runtime-driven restarts.

We will test only the scenario where all services started and the provider was healthy (Compose waited for healthchecks) and only then the various failure triggers. At that point all four signals (running, health state, boot ID, request success) should align in the baseline. We observe how each case (B–F) deviates from that baseline. Each case’s decision outcome will be either Accept (behavior matches the documented contract), Repair (unexpected behavior that requires policy change), or Hold (insufficient evidence).

Freeze the Compose model and runtime boundary

We perform all tests in a clean private working directory on a disposable Linux Docker host. First, we record our environment: check that docker context is correct, and record Docker Engine and Compose CLI versions. We use a specific Python 3.12 runtime image by digest to avoid changes: for example, pulling python:3.12-slim and inspecting yields python:3.12-slim@sha256:<digest> (store that as $PYTHON_RUNTIME). We set a unique RUN_ID for this session. We save the supplied lab.py (below) and a compose.yaml in the project directory. We initialize an empty control/ directory for runtime signals, and prepare fresh evidence/ files.

Before testing, we export the environment variables, capture docker version, docker compose version, and use docker compose config --format json to save the final model. We do not alter the Compose model during the test: no changes in restart policy or dependency structure. (The Refonte article on Compose env interpolation shows how we fix and record environment and files for reproducibility.) This ensures that each test A–F runs under the same configuration, isolating the effect of each trigger.

Verify supported features before interpreting an outcome

We require Docker Compose 2.17.0 or newer for dependency-level restart: true. Confirm the installed docker compose supports --wait and --wait-timeout for startup. Also ensure docker compose restart --no-deps is available.

We resolve a real Docker image digest for the Python runtime: for example, docker pull python:3.12-slim then docker image inspect --format '{{index .RepoDigests 0}}' python:3.12-slim yields something like python:3.12-slim@sha256:<digest>. Set PYTHON_RUNTIME=<that digest>. This digest and the platform (e.g. linux/amd64) are our approved base image reference. We will not rebuild or retag images mid-test. All containers will use this resolved image reference. We then proceed to start the baseline scenario.

Build a provider and consumers with traceable identities

We implement a simple HTTP provider and two identical HTTP consumers in pure Python (standard library only). All containers generate and log a random boot UUID on startup. The provider serves a /health endpoint that returns ready/unready based on its startup delay and an external control/fail file. It also serves /work requests that echo the run ID, provider boot ID, and per-request sequence.

Each consumer loops sending requests to the provider’s /work endpoint and logs results, without exiting on errors. The lab.py mark command can set the provider’s failure marker (create /control/fail), clear that failure, or request a provider crash (exit the provider). Health and request responses include the provider’s current boot ID.

The provider’s healthcheck command (lab.py health) exits 0 only if /control/fail is absent and ≥2 seconds have passed since start. Thus the provider will take 2 seconds to become healthy initially, and subsequent fail markers will make it return HTTP 503 until clear. The actual restart-on-error policy is on-failure:2 in the Compose file. Neither consumer automatically stops on an unhealthy provider. The full lab.py is given below:

import datetime, json, os, signal, sys, threading, time, uuid
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from pathlib import Path
from urllib.error import HTTPError
from urllib.parse import parse_qs, urlencode, urlsplit
from urllib.request import ProxyHandler, Request, build_opener

RUN = os.environ["RUN_ID"]
CONTROL = Path(os.environ.get("LAB_CONTROL", "/control"))
PORT = int(os.environ.get("LAB_PORT", "18081"))
BOOT = uuid.uuid4().hex
START = time.monotonic()
PRINT_LOCK = threading.Lock()
OPENER = build_opener(ProxyHandler({}))

def emit(event, fields):
    with PRINT_LOCK:
        print(json.dumps(dict(event=event, run=RUN,
              utc=datetime.datetime.now(datetime.timezone.utc).isoformat(),
              monotonic=time.monotonic(), fields)), flush=True)

def phase():
    return json.loads((CONTROL / "phase.json").read_text())["phase"]

def ready():
    return time.monotonic() - START >= 2 and not (CONTROL / "fail").exists()

def atomic(path, text):
    temporary = path.with_name(path.name + "." + uuid.uuid4().hex + ".tmp")
    with temporary.open("x") as handle:
        handle.write(text)
    temporary.replace(path)

def mark(case, action):
    if action not in {"none", "fail", "clear", "crash"}:
        raise ValueError("unknown control action")
    token = case + ":" + uuid.uuid4().hex
    atomic(CONTROL / "phase.json", json.dumps(dict(phase=token)))
    if action == "fail":
        atomic(CONTROL / "fail", token)
    elif action == "clear":
        (CONTROL / "fail").unlink(missing_ok=True)
    elif action == "crash":
        atomic(CONTROL / "crash.once", token)
    emit("operator_mark", phase=token, action=action)

def read_http(url):
    request = Request(url, headers={"Connection": "close"})
    try:
        response = OPENER.open(request, timeout=0.75)
    except HTTPError as error:
        response = error
    with response:
        data = response.read(8193)
        if len(data) > 8192:
            raise ValueError("response too large")
        return response.getcode(), json.loads(data)

def provider():
    class Handler(BaseHTTPRequestHandler):
        def log_message(self, args): pass
        def do_GET(self):
            route = urlsplit(self.path)
            ok = ready()
            status = 200 if ok else 503
            payload = dict(run=RUN, provider_boot=BOOT, ready=ok)
            if route.path == "/work":
                query = parse_qs(route.query, strict_parsing=True)
                expected = {"run", "consumer", "consumer_boot", "seq", "nonce", "phase"}
                if set(query) != expected or any(len(v) != 1 for v in query.values()):
                    raise ValueError("invalid request fields")
                supplied = {k: v[0] for k, v in query.items()}
                if supplied["run"] != RUN:
                    raise ValueError("wrong run")
                payload.update(supplied)
                emit("provider_request", status=status, provider_boot=BOOT,
                     *{k: v for k, v in supplied.items() if k != "run"})
            elif route.path != "/health":
                self.send_error(404)
                return
            body = json.dumps(payload).encode()
            self.send_response(status)
            self.send_header("Content-Type", "application/json")
            self.send_header("Content-Length", str(len(body)))
            self.send_header("Connection", "close")
            self.end_headers()
            self.wfile.write(body)

    def crash_monitor():
        marker = CONTROL / "crash.once"
        while True:
            if marker.exists() and time.monotonic() - START >= 12:
                token = marker.read_text()
                marker.unlink()
                emit("deliberate_exit", provider_boot=BOOT, phase=token,
                     uptime=time.monotonic() - START, exit_code=17)
                os._exit(17)
            time.sleep(0.1)

    emit("boot", role="provider", boot=BOOT)
    threading.Thread(target=crash_monitor, daemon=True).start()
    ThreadingHTTPServer(("0.0.0.0", PORT), Handler).serve_forever()

def consumer(name):
    emit("boot", role=name, boot=BOOT)
    base = os.environ.get("PROVIDER_URL", "http://provider:18081")
    seq = 0
    while True:
        request = dict(run=RUN, consumer=name, consumer_boot=BOOT,
                       seq=str(seq), nonce=uuid.uuid4().hex, phase=phase())
        fields = {k: v for k, v in request.items() if k != "run"}
        try:
            status, result = read_http(base + "/work?" + urlencode(request))
            if any(result.get(k) != v for k, v in request.items()):
                raise ValueError("request identity mismatch")
            uuid.UUID(result["provider_boot"])
            if status not in (200, 503) or result.get("ready") is not (status == 200):
                raise ValueError("status/readiness mismatch")
            emit("consumer_result", status=status,
                 provider_boot=result["provider_boot"], fields)
        except Exception as error:
            emit("consumer_error", error=type(error).__name__, fields)
        seq += 1
        time.sleep(0.5)

def stop(signum, frame):
    emit("operator_stop", boot=BOOT, signal=signum)
    raise SystemExit(0)

if name == "__main__":
    mode, args = sys.argv[1:]
    signal.signal(signal.SIGTERM, stop)
    if mode == "provider":
        provider()
    elif mode == "consumer":
        consumer(args)
    elif mode == "mark":
        mark(*args)
    elif mode == "health":
        try:
            status, payload = read_http(f"http://127.0.0.1:{PORT}/health")
            raise SystemExit(0 if status == 200 and payload["run"] == RUN
                             and payload["ready"] is True else 1)
        except Exception:
            raise SystemExit(1)
    else:
        raise SystemExit("mode: provider, consumer NAME, mark CASE ACTION, health")

The compose.yaml uses this Python code for all services. Each service uses the pinned Python image, mounts lab.py and the control/ directory for inter-process signals, and shares the same RUN_ID. The provider’s service has restart: "on-failure:2" and a Compose healthcheck invoking lab.py health with a 1-second interval, 3 retries, and a 5-second start period. It depends on nothing.

The linked and comparison services (the consumers) have no restart policy (no) and depends_on entries on provider with condition: service_healthy. Crucially, only linked has restart: true under its depends_on (propagate explicit Compose restarts), while comparison has restart: false. Here is the complete Compose model:

x-base: &base
  image: ${PYTHON_RUNTIME:?Set a verified Python runtime digest}
  environment:
    RUN_ID: ${RUN_ID:?Set a unique run identifier}
  volumes:
    - ./lab.py:/lab.py:ro
    - ./control:/control:ro

services:
  provider:
    <<: base
    command: ["python", "/lab.py", "provider"]
    volumes:
      - ./lab.py:/lab.py:ro
      - ./control:/control
    restart: "on-failure:2"
    healthcheck:
      test: ["CMD", "python", "/lab.py", "health"]
      interval: 1s
      timeout: 1s
      retries: 3
      start_period: 5s
  linked:
    <<: base
    command: ["python", "/lab.py", "consumer", "linked"]
    restart: "no"
    depends_on:
      provider:
        condition: service_healthy
        restart: true
  comparison:
    <<: *base
    command: ["python", "/lab.py", "consumer", "comparison"]
    restart: "no"
    depends_on:
      provider:
        condition: service_healthy
        restart: false

We use a 2-second provider startup delay (encoded in lab.py) and the 5s healthcheck start-period so that both consumers only begin sending requests after the provider is healthy. Because neither consumer has its own restart policy, any container restart comes either from the provider’s on-failure:2 or from Compose’s explicit commands. Note that service_healthy only gates the initial startup of a consumer; after startup it does not continuously restart consumers when provider health changes.

Capture a healthy baseline and every process identity

With the above model and environment variables set, we bring up the Compose project and record the baseline. In a fresh directory with control/ and evidence/ empty, we run:

export RUN_ID="rl-compose-$(python3 -c 'import uuid; print(uuid.uuid4().hex[:12])')"
export PROJECT="$RUN_ID"
export LAB_CONTROL="$PWD/control"
python3 lab.py mark A none > evidence/A.operator.jsonl
docker version > evidence/docker-version.txt
docker compose version > evidence/compose-version.txt
docker context show > evidence/docker-context.txt
docker compose -p "$PROJECT" -f compose.yaml config --format json > evidence/model.json
docker events --filter type=container --filter "label=com.docker.compose.project=$PROJECT" --format '{{json .}}' > evidence/events.jsonl &
EVENTS_PID=$!
docker compose -p "$PROJECT" -f compose.yaml up --wait --wait-timeout 45

We expect all services to start. We capture each container’s ID (docker compose ps --all --quiet SERVICE) and run docker inspect and docker logs into evidence/ files (provider, linked, comparison). We confirm the log outputs include "boot" events with roles provider, linked, comparison and their boot UUIDs P0, L0, C0.

Both consumers should immediately report valid 200 responses from the provider (their consumer_result with "status":200 and matching provider_boot P0). We verify the health of provider as healthy via docker inspect.

At this point, all four signals match: all containers Running, provider Healthy, unique boot IDs, and successful requests. These are our Accept conditions for case A: the baseline is as expected.

Turn the provider unhealthy while it stays running

Next, we simulate case B (Unhealthy while running) by asking the provider to fail its healthcheck without exiting. We run:

python3 lab.py mark B fail > evidence/B.operator.jsonl

This creates control/fail, so new healthchecks will fail. We wait for the actual unhealthy state. The deliberate_exit event is expected only in the crash case; here, we observe the health state change to unhealthy. We watch the docker events stream for health changes, or check via docker inspect.

Meanwhile both consumers continue running and attempting requests. We expect the provider’s container to stay Running (since it did not exit) but its health to flip to “unhealthy”. The consumers will start receiving HTTP 503 responses (with "ready":false) from the provider while it’s unhealthy.

Crucially, since no container exited or was explicitly restarted, we expect no restart events and no change in any boot UUID. The expected snapshot after B is:

Case

Trigger

Provider boot

Linked boot

Comparison boot

Useful-work evidence

B

Unhealthy while running

P0

L0

C0

Consumers report HTTP 503; all containers remain running

Indeed, an unhealthy status by itself triggers no restarts. The provider’s engine policy (on-failure) does not apply because the container did not exit. The dependency-level restart:true does not apply because no explicit docker compose restart was issued.

Verify via docker inspect that State.Running=true for all containers and the Health.Status of provider is "unhealthy", while the boot IDs are still P0, L0, C0.

If we observe any container restart or exit, that would indicate an unexpected controller interfering (which would be a Repair). This fixture does not include an external restart controller. Accept case B when the evidence shows that the provider and consumers do not restart, consistent with Compose/Engine semantics.

Bound the observation instead of assuming an exact timer

We avoid assuming exactly how long the healthcheck transition takes. Instead, we wait until at least several consecutive failed probes (at least 3 given retries:3) and at least 3 consumer_result records with HTTP 503 from each consumer, but impose a 45-second timeout.

We examine the provider’s Health.Log entries (should show successive failures) and corresponding health_status: unhealthy events in docker events. If the unhealthy state and HTTP 503 evidence are recorded within 45 seconds, with unchanged boot UUIDs and all containers still running, we can confirm B for that observation window. If something unexpected happened (e.g. missing events due to the 256-event limit), we would mark inconclusive. The fixture introduces no external auto-restart controller, so B is expected to pass.

Recover provider health without restarting a service

Now we test case C (Health condition restored). We clear the failure and wait for recovery:

python3 lab.py mark C clear > evidence/C.operator.jsonl

This deletes control/fail. We then wait for the provider’s health to return to healthy (which it will after one successful probe). During this time, the consumers should resume receiving HTTP 200 responses.

Because the provider was never restarted in case B, and the container is still alive, we expect that after C all three containers retain their original boot UUIDs (P0, L0, C0). In other words, Compose or Docker do not automatically restart containers when an unhealthy health check becomes healthy.

We verify docker inspect shows provider health is healthy and no container has restarted (no changed IDs, no incremented RestartCount). The expected row for C is:

Case

Trigger

Provider boot

Linked boot

Comparison boot

Useful-work evidence

C

Health condition restored

P0

L0

C0

Consumers report HTTP 200; same UUIDs

Accept this case when the provider recovers health in place and both consumers and the provider continue running with the same process generations. (Health transitions do not themselves trigger a restart.) We independently verify each consumer still gets a valid 200 response, and all logged nonces and phase tokens align, confirming continuity. If instead the provider had restarted, that would conflict with documented behavior and require investigation (Repair). The expected recovery here requires no restart.

Let the container runtime recover one real exit

Next is case D (One application exit 17). We simulate a genuine provider failure. The provider is coded to watch for control/crash.once and exit code 17 after 12 seconds of uptime. We mark:

python3 lab.py mark D crash > evidence/D.operator.jsonl

When provider uptime ≥12s, it will log "deliberate_exit" and call os._exit(17). The provider container will exit 17, and Docker Engine, seeing restart: on-failure:2, will restart it (since exit code 17 is treated as failure).

The key question: do consumers restart? The comparison consumer has no dependency restart configured. The linked consumer does have depends_on.provider.restart: true, but note that dependency-level restart only applies to explicit Compose operations, not to this engine-driven restart of provider. Since neither consumer container actually exited, their own engine policy no won’t restart them, and Compose did not issue a restart command for them either.

We expect: provider gets a new boot (P1), but both consumers remain at L0 and C0. After the provider restarts, eventually both consumers see HTTP 200 responses from the new provider (boot P1). The expected row is:

Case

Trigger

Provider boot

Linked boot

Comparison boot

Useful-work evidence

D

One application exit 17

P1

L0

C0

Provider restarted, clients see HTTP 200 from P1

Use docker inspect to confirm the provider’s new StartedAt, and use its boot log to confirm the new boot UUID. Verify that both consumer boot UUIDs remain unchanged. Also check for Engine events: a die and start event for provider. Both consumers should remain running (State.Running) throughout. This matches the Engine’s restart policy, and Compose’s dependency restart is not invoked because it was not an explicit docker compose event.

Accept case D when the Engine recovers the provider, both consumer process generations remain unchanged, and both consumers receive current successful responses from P1. (If either consumer had restarted here, it would have been unexpected. We explicitly did not use docker stop or compose restart on them.)

Preserve the distinction between failure and an operator stop

Verify that this is a real application exit: do not use docker stop or kill the container. The provider exit must come from its own code. If the restart does not occur, check the effective policy and retry count. The expected result is that the on-failure restart policy brings the provider back up and it resumes serving. This shows the difference: an exit triggered an engine restart, unlike an unhealthy signal.

Compare explicit Compose restart with propagation disabled

Finally we test E and F with docker compose restart. First:

python3 lab.py mark E none > evidence/E.operator.jsonl
docker compose -p "$PROJECT" -f compose.yaml restart provider

This is an explicit Compose restart of the provider service. By Compose semantics, with linked.depends_on.restart: true, the linked service should also restart. The comparison service (with restart: false) should not. We record events and inspect again.

We expect: provider gets a new boot (P2), linked gets a new boot (L1), and comparison stays at C0. The provider should eventually become healthy, while both consumers should be running. We then verify both consumers receive valid 200 responses from the new provider (boot P2), confirming the system recovered.

The explicit restart does not re-run the healthcheck gate for starting linked; it simply restarts containers. If either consumer did not receive a fresh HTTP 200, that would indicate a startup timing issue, but we independently check after they’re all running.

The expected distinction is that linked restarts (its boot changes) and comparison does not. We capture all logs and IDs. The expected row is:

Case

Trigger

Provider boot

Linked boot

Comparison boot

Useful-work evidence

E

Compose restart provider

P2

L1

C0

Explicit restart propagated to linked, both get current 200

If instead linked had stayed at L0, that would violate the documented restart: true behavior. But here we expect an accept: the documentation explicitly says linked should restart when provider is explicitly restarted.

Change only the command’s propagation option

Now case F (Compose restart --no-deps):

python3 lab.py mark F none > evidence/F.operator.jsonl
docker compose -p "$PROJECT" -f compose.yaml restart --no-deps provider

This command restarts the provider but does not restart dependent services. We expect: provider gets a new boot (P3), but neither consumer restarts (they remain at L1 and C0). Both consumers will eventually get healthy replies from provider P3. The expected row:

Case

Trigger

Provider boot

Linked boot

Comparison boot

Useful-work evidence

F

Compose restart --no-deps provider

P3

L1

C0

Explicit no-deps: consumers unchanged

We then verify manually that both consumers (L1 and C0) still get valid 200 work responses from the new P3. Because we did not change the Compose file itself, we did not test “adopting a changed config”.

The key point is that the CLI’s --no-deps flag omits the linked restart. If we had run without --no-deps, case E provides the comparison with propagation. If the evidence matches, this confirms that dependency-level restart is only triggered by Compose when allowed, not by default for all restarts.

Reconcile control actions, boot identities and useful work

The following matrix gives the expected outcomes for all six cases:

Case

Trigger

Provider boot

Linked boot

Comparison boot

Useful-work evidence

A

Healthy startup

P0

L0

C0

Both consumers immediately receive current valid 200

B

Unhealthy while running

P0

L0

C0

Consumers see current 503; no restart (all running)

C

Health condition restored

P0

L0

C0

Consumers return to current 200 without any restart

D

One application exit 17

P1

L0

C0

Engine restarted provider (P1), consumers still L0/C0, see 200 from P1

E

Compose restart provider

P2

L1

C0

Explicit restart propagation: linked restarted (L1), comparison unchanged, both see 200

F

Compose restart --no-deps

P3

L1

C0

No propagation: only provider (P3) changed, consumers unchanged

Throughout, we match each request’s seq, nonce, and phase fields between consumer sends and provider responses to ensure we’re looking at the same interaction epoch. A container’s RestartCount from inspect can help correlate whether an exit happened, but we rely primarily on the emitted boot UUIDs (P0, P1, P2, etc.) to identify a new process generation.

Note that a new boot ID could come from either a restart or (unlikely here) ephemeral container replacement; however, in these tests it should match one of our triggers. We also capture docker events to catch the actual die and start events. Missing an event in the 256-event history doesn’t prove absence of restart; that is why we capture events live.

The expected logs show no restart for B or C, exactly one restart for D (provider), restarts of both provider and linked for E, and one restart for F (provider). Correlate events and container inspection with the boot identities to verify each response to its trigger.

If any container changes boot out of line with this table, mark that case for review. In particular, if B had somehow restarted the provider or either consumer, it would violate “unhealthy does nothing.” If E had failed to restart linked, it would violate the documented restart:true behavior.

Classify the outcome before changing a policy

Use this decision matrix only after recording the required evidence:

Condition

Evidence required for acceptance

Decision

Owner

B: Provider health-only failure

Provider and consumers remain running (P0, L0, C0); unhealthy health status; consumers receive 503s

Accept

Application Team (monitor ready)

C: Health restored, no restart

Boots unchanged (P0, L0, C0); all get HTTP 200 again

Accept

Application Team (depend on readiness)

D: Provider crash (exit code 17)

Engine restarts provider (P1) as on-failure policy calls for; consumers still L0/C0

Accept

Container Owner (restart policy)

E: Compose restart with deps

Linked restarts (L0→L1) through dependency restart propagation (depends_on.restart:true); comparison unchanged

Accept

Compose Operator (explicit commands)

F: Compose restart with --no-deps

Only provider restarts (P2→P3); consumers unchanged (L1, C0)

Accept

Compose Operator (explicit commands)

If any boot changed unexpectedly

Mismatch in expected boots (e.g. B restarts something)

Hold/Repair: investigate external factors or config

OPS/DevSecOps

Accept only outcomes supported by the documented contract and the recorded evidence. For example, evidence for B and C should confirm that health-state changes alone do not restart containers. D should show the Engine policy working as documented, separate from Compose dependencies. E and F should confirm that Compose’s restart:true only applies on explicit actions and that --no-deps works as documented (restart only the named service).

If any case deviates, investigate factors such as an external watchdog, a misapplied restart policy, or a Compose version mismatch (e.g. nested restart unsupported). This outcome table defines the expected results to verify on the recorded Docker host. This exercise tests the existing recovery mechanisms; it does not define a general “auto-heal” strategy.

We reference general cloud best practices to clarify scope: container healthchecks and orchestration can aid recovery, but adding depends_on.restart: true does not itself become a continuous heal loop. Health can recover in place, as expected in case C. Container restarts in this experiment arise from explicit restart commands or the configured response to process exits.

The Refonte article on broader cloud reliability practices notes that orchestration can help in recovery planning, but in our precise Docker Compose runtime, only the documented mechanisms are under test. We do not treat this lab as implying new resilience features in Compose beyond what this fixture tests.

Restore the clean fixture and repeat the comparison

After collecting evidence for A–F, we clean up without disturbing other Docker state. We clear the failure marker and confirm the provider is healthy and both consumers are running (P3, L1, C0, with both consumers receiving HTTP 200). Then we bring the project down (docker compose -p "$PROJECT" -f compose.yaml down). We stop the docker events logger using the stored PID, and preserve all evidence/ files. We do not prune the system or change unrelated resources.

All the snapshots, logs, and inspection data for cases A–F are retained for review. If one wanted to replay these tests (say to verify a fix), one would create a new project name and new control folder, and rerun the sequence to reproduce A–F afresh. This ensures each run is isolated and comparable; we do not rely on residual state or forget to reset the control/ signals.

Assign a recovery owner for each failure trigger

In practice, different teams take ownership of different aspects. The application owner should define what “ready” means (here the provider’s /health endpoint) and what to do on errors (e.g. client retry logic). The container/runtime owner configures the container’s restart policy. The Compose operator owns explicit maintenance actions (docker compose restart) and whether to use --no-deps. The site reliability reviewer audits the evidence (logs, boot IDs) against the expected behavior.

For instance, if case B had seen an unexpected restart, the container owner should check if an unintended policy or supervisor was active. If E had failed to restart the linked service, the Compose operator would verify their Compose version and syntax.

By mapping each trigger (health fail, health clear, exit, manual restart, no-deps restart) to who is responsible, we ensure clear accountability. (This aligns with dependency wave planning in cloud disaster recovery, where each service owner understands their recovery step.)

This exercise practices observation; it does not prescribe an automatic fix for every service. Always confirm the actual Docker versions, Compose schema, and healthcheck semantics when applying any policy.

Practice reviewed container operations with Refonte Learning

This controlled validation exercise is an example of the advanced container orchestration topics covered in the DevOps Engineering program. Refonte’s curriculum includes Docker/Kubernetes and monitoring. This exercise uses fault injection to practice those topics. (See the DevOps Engineering program page for details.)

We encourage readers to explore how Compose healthchecks and restarts work, and to practice running similar experiments. While such exercises are part of a learning journey, this lab is not presented as a published course module. It illustrates the critical thinking applied by DevOps professionals in incident reviews.