An exit-0 curl GET does not guarantee the checked service returned the expected API response. For example, a readiness probe might return an HTML login page or another service’s data and still exit 0. Our acceptance criterion is a narrowly defined response contract: an HTTP GET to http://BASE/ok must end with status 200 OK, no redirects, content type application/json, and a JSON body exactly matching {"ready":true,"service":"refonte-lab","revision":"r1"}. Anything else (a redirect, a login page, a wrong payload or an empty reply) fails the contract. This testing is scoped to a controlled loopback fixture; it only probes that one endpoint's contract, not full production health. We will build a disposable local server with valid and invalid endpoints, run curl commands (with and without following redirects), and record all evidence. Finally we decide: ACCEPT the probe contract if and only if every predicate matches, or HOLD/REPAIR otherwise. This approach mirrors how an API check tool (e.g. API testing with Postman) starts by explicitly defining the expected response before running the test.
Define the response your health check is allowed to accept
First, codify the health-check contract precisely, separate from generic HTTP success. Our API’s only accepted response is from GET BASE/ok returning exactly:
Status: 200 OK (final HTTP status must be 200)
Redirects: exactly 0 redirects (the final URL must be exactly BASE/ok)
Content-Type: application/json (media type essence matches JSON)
JSON schema: object with keys ready, service, revision
ready must be a JSON true boolean
service must equal "refonte-lab"
revision must equal "r1"
All fields are required with those exact types and values. For example, a 200-status HTML login page or an error XML are outside this contract, even though status 200 typically means success in HTTP semantics. Likewise, a redirect to a login URL is not acceptable, even if it eventually yields status 200. We are deliberately strict: no additional fields or types are permitted. (Different APIs might relax or extend this schema, but such changes must be in the contract and regression tests, not ad hoc during an acceptance check.)
This contract must be documented before executing the probe. For instance, if a configured probe URL returns {"ready":false}, the API owner might accept that meaning “still initializing” and treat it as healthy-enough, but in our contract that fails the exact-true test. Similarly, if a redirect is actually intended (e.g. a service moved endpoint), that must be explicitly allowed by updating the contract. We treat the contract as the final oracle, analogous to how an API test suite (as in professional API developer tools) would define expected response exactly.
Below are some examples of invalid "false positive" responses for this contract:
Login form (HTML): status 200 but Content-Type text/html or HTML body. This is not the JSON we expect.
Wrong JSON object: status 200 and JSON format, but e.g. "service":"wrong-service". This indicates the request was handled by the wrong service.
Empty response: status 204 No Content or 200 with empty body. We require our readiness fields, so empty is not acceptable.
Redirect: status 302 to a login or other URL. Even if curl exits 0 by default on a redirect, the actual API data never arrived at our caller. A redirect means the contract was not met.
Error status: any 4xx/5xx status (e.g. 503) should fail. (We’ll use --fail to catch those.)
Our acceptance logic will label only the expected JSON at /ok as passing, and treat all above examples as failures, just as a thorough API test would.
Freeze the client invocation and disposable environment
Next, capture and fix the entire client context so the probe is repeatable. This includes tool versions, paths, and environment settings. In a script we would detect something like:
$ which curl && curl --version/usr/bin/curl
curl 8.10.1 (x86_64-pc-linux-gnu) libcurl/8.10.1 OpenSSL/3.0.9 zlib/1.3Protocols: http https
Features: IPv6 HTTPS-proxy ...Here we show curl 8.10.1 (the prechecked baseline) running on Linux; OpenSSL is the TLS backend. We also note the Python runtime, e.g. Python 3.13.5. We record these for documentation but proceed assuming this lab environment. Any user should verify their curl --version and runtime versions to note differences. (We assume curl 8.10.1 here, but our playbook should test for option support in case older or newer releases are used.)
We then launch a disposable HTTP server bound to 127.0.0.1 on an ephemeral port. In Python, for example, we might do:
from http.server import HTTPServer, BaseHTTPRequestHandler# (Define a handler that implements our nine paths, see next section.)server = HTTPServer(("127.0.0.1", 0), TestHandler)port = server.server_address[1]print("Server listening on port", port)server_thread = threading.Thread(target=server.serve_forever, daemon=True)server_thread.start()Each test case (endpoint) will use its own output directory to keep evidence separate. For each curl invocation we include a fixed common argument list, for example:
curl --disable --noproxy '*' \ --silent --show-error \ --connect-timeout 2 --max-time 3 \ --max-redirs 2 ...Here, we explain each part:
--disable (or -q) as first option: ensures no default curlrc is read. Otherwise a user’s config file could inject headers or proxies unnoticed.
--noproxy '*': explicitly turn off any proxying for localhost traffic. (This lab is plaintext internal traffic; in production you’d remove this or configure corporate proxy separately.)
--silent --show-error: only output on errors. (This hides progress but still reports curl errors.)
--connect-timeout 2 and --max-time 3: bounds on DNS/connect and total time for this lab. These are arbitrary short limits chosen so the tests don't hang. (They are not production recommendations.)
--max-redirs 2: limit redirect hops to avoid infinite loops. Our loop case will test the two-hop limit, for example.
We will call curl with --dump-header, --output and --write-out to gather metadata. For instance:
curl --dump-header headers.txt --output body.bin \ --write-out "%{http_code} %{url_effective} %{num_redirects} %{content_type}" \ http://127.0.0.1:$port/okThis would run curl on /ok and write to files, while printing the final HTTP code, the effective URL, the number of redirects followed, and the media-type. In a Python script we would capture subprocess.returncode immediately after the call (to get curl’s exit code) and separately parse the output of --write-out. We keep exit code, HTTP code, URL, redirects count, and content type as distinct fields. (If curl times out or has a network error, we record that as an “inconclusive transport outcome,” not as an HTTP status code.)
Do not merge fields. In particular, do not let an unspecified exit code or timeout be quietly treated as a successful 0. We keep them separate, and we never reuse a stale body file from a prior failed transfer; each invocation starts fresh.
Neutralize hidden configuration before measuring
As noted, --disable must be the first option (or equivalently the -q short form at the very start) to prevent any default config (.curlrc) from auto-loading. By default, curl will look for a config file in home or /etc even if you also specify --config. The manual explicitly says: if -q/--disable is first, then no config is read. We also use --noproxy '*' to avoid any environment variable or default proxy interfering (though for localhost it usually shouldn’t). These steps ensure our lab runs are fully under our control. In a real production probe, you might omit disabling config if a policy requires certain headers/certs, but then you must document that choice explicitly.
Build endpoints that return convincing but wrong responses
We implement a local HTTP server fixture with nine routes, each simulating a specific case. Below is a conceptual handler (in Python) with all behaviors. We ensure correct Content-Length and suppress logs for brevity:
class TestHandler(BaseHTTPRequestHandler): def do_GET(self): path = self.path if path == "/ok": body = b'{"ready":true,"service":"refonte-lab","revision":"r1"}' self.send_response(200) self.send_header("Content-Type", "application/json") self.send_header("Content-Length", str(len(body))) self.end_headers() self.wfile.write(body) elif path == "/redirect-good": self.send_response(302) self.send_header("Location", "/ok") self.send_header("Content-Length", "0") self.end_headers() elif path == "/redirect-login": self.send_response(302) self.send_header("Location", "/login") self.send_header("Content-Length", "0") self.end_headers() elif path == "/login": body = b"<html><body>Login Page - unauthorized</body></html>" self.send_response(200) self.send_header("Content-Type", "text/html") self.send_header("Content-Length", str(len(body))) self.end_headers() self.wfile.write(body) elif path == "/wrong-json": body = b'{"ready":true,"service":"wrong-service","revision":"r1"}' self.send_response(200) self.send_header("Content-Type", "application/json") self.send_header("Content-Length", str(len(body))) self.end_headers() self.wfile.write(body) elif path == "/bad-json": bad = b'{"ready":true,"service": refonte-lab,"revision":"r1"' # malformed (unquoted value) self.send_response(200) self.send_header("Content-Type", "application/json") self.send_header("Content-Length", str(len(bad))) self.end_headers() self.wfile.write(bad) elif path == "/unhealthy": body = b'{"ready":false,"service":"refonte-lab","revision":"r1"}' self.send_response(503) self.send_header("Content-Type", "application/json") self.send_header("Content-Length", str(len(body))) self.end_headers() self.wfile.write(body) elif path == "/empty": self.send_response(204) self.send_header("Content-Length", "0") self.end_headers() elif path == "/loop": self.send_response(302) self.send_header("Location", "/loop") self.send_header("Content-Length", "0") self.end_headers() else: self.send_response(404) self.send_header("Content-Length", "0") self.end_headers() def log_message(self, format, *args): return # suppress loggingThe table below summarizes what each endpoint returns:
Path | HTTP Code | Location Header | Content-Type | Body (text) |
/ok | 200 | None | application/json | {"ready":true,"service":"refonte-lab","revision":"r1"} |
/redirect-good | 302 | Location: /ok | (none) | no body |
/redirect-login | 302 | Location: /login | (none) | no body |
/login | 200 | None | text/html | <html>...Login Page - unauthorized...</html> |
/wrong-json | 200 | None | application/json | {"ready":true,"service":"wrong-service","revision":"r1"} |
/bad-json | 200 | None | application/json | {"ready":true,"service": refonte-lab,"revision":"r1" (malformed JSON) |
/unhealthy | 503 | None | application/json | {"ready":false,"service":"refonte-lab","revision":"r1"} |
/empty | 204 | None | (none) | no body |
/loop | 302 | Location: /loop | (none) | no body |
Each “bad” endpoint is intended to look plausible but violate one contract rule: wrong service name, HTML content, an HTTP error, an endless redirect, or malformed JSON. We will test each of these to ensure our checker rejects them.
Run --fail without following redirects
We now invoke curl for each path with no --location (so redirects are not followed). We use --fail (-f) so that HTTP 4xx/5xx trigger a nonzero exit code. Here are the observed outcomes (assuming our lab versions):
/ok: curl -f exits 0, HTTP status 200, effective URL .../ok, num_redirects=0. Body is valid JSON.
/redirect-good: exits 0 (because 302 is not >=400). HTTP status 302, effective URL .../redirect-good, num_redirects=0. Headers include Location: /ok. Body is empty (Content-Length 0).
/redirect-login: exits 0, HTTP status 302, effective URL .../redirect-login, no body, with Location: /login.
/login: exits 0, HTTP status 200, effective URL .../login, Content-Type: text/html. Body is the HTML login page.
/wrong-json: exits 0, HTTP status 200, effective URL .../wrong-json, Content-Type: application/json. Body is JSON but with "service":"wrong-service".
/bad-json: exits 0, HTTP status 200, effective URL .../bad-json, Content-Type: application/json. Body is bytes that look like JSON but are malformed. Attempting to parse it will fail.
/unhealthy: exits 22 (curl’s code for HTTP error). No output body (by design, --fail suppresses error bodies). We know HTTP status was 503 from headers we captured.
/empty: exits 0, HTTP status 204, effective URL .../empty, no content type or body.
/loop: after 2 redirects limit, exits 47 (curl’s code for too many redirects). No final body.
Notice in all no-redirect cases, exit code 0 often happens even when the data is wrong. For example, /login exit 0 with HTML body, and /wrong-json exit 0. We cannot treat exit-0 as acceptance. Instead, we must look at the first response and policy: for /redirect-good, the probe ended at a 302 to /ok. It did not actually get /ok. Under our no-follow policy, that result is not acceptable. Merely seeing exit 0 on /redirect-good should still be considered a failed contract because the response was a redirect, not the expected JSON endpoint.
Read a successful 302 without calling it readiness
It can be tempting to see the 302 and code 0 from /redirect-good and think “everything’s fine” since there was no error. But recall our rule: zero redirects were allowed. Seeing a 302 means our policy was violated. In other words, the process and HTTP status came back “successful” (no error) because curl did exactly what it was told: not to follow, so it said “Done with 302.” But the probe contract requires 200 from /ok. In this case we treat /redirect-good’s result as a fail for the contract. We will not blindly accept exit 0 on any endpoint; we accept it only if all response fields match exactly.
Follow redirects and inspect the terminal response
Now we repeat all cases with --location (-L) added. This tells curl to follow the Location: headers up to --max-redirs 2. The observed final outcomes:
/ok: same as before: exit 0, status 200, URL .../ok, JSON body.
/redirect-good: now curl follows /redirect-good -> /ok, ending with exit 0, final status 200, effective URL .../ok, num_redirects=1. Body is valid JSON {"ready":true,...}. (Note: even though this eventually got the correct data, the request was redirected. In our default policy, we’re not allowing any redirects, so this still fails unless we explicitly allow this path.)
/redirect-login: follows to /login, ending exit 0, final status 200, effective URL .../login, num_redirects=1. Body is HTML login page. Again exit 0 and 200, but it’s not the JSON we want (it’s the login page).
All other endpoints (/login, /wrong-json, etc.) behave as before, since they either didn’t redirect or curl handled the redirect similarly.
/unhealthy still exits 22 (no change; curl still treats 503 as fail).
/loop now tries to follow itself: it reaches the max of 2 redirects and exits 47 (same as in no-follow).
Thus even with -L, the only case yielding the JSON contract is /ok. The /redirect-good path ended with the right data but under our “no-redirects” contract it’s effectively not allowed. We leave the default policy untouched and treat it as a separate compatibility policy below. The key is that curl’s success doesn’t equal contract success. Curl simply followed instructions: if instructed to follow, it did. But acceptance is defined by our policy, not by whether curl returned 0.
Keep exit codes and HTTP status in separate evidence fields
We will record three independent outcomes for each test: (1) the process exit code, (2) the final HTTP response status, and (3) the semantic contract result (pass/fail). Under no circumstances should we conflate these. For example, in /unhealthy, the exit code 22 tells us curl saw an HTTP error, which we interpret as HTTP status 503. The semantic verdict is "failed contract because status≠200". By contrast, for /login, exit code 0 and status 200 but content was HTML; the semantic verdict is still "failed contract due to wrong type".
We capture these fields distinctly. In a CSV or log, we might have columns like exit_code, http_code, content_type, body_digest, json_parse_ok, service_ok, revision_ok, ready_ok, and finally contract_pass. If any piece is missing or error, we mark it. For instance, if curl timed out (exit nonzero) we would mark HTTP code unknown and decision “inconclusive” or “transport failure.” If JSON parsing fails, we note that separately. Each test case (endpoint & mode) yields one row with all data. This separation prevents confusion: e.g., an exit code of 22 doesn’t directly become an “HTTP 22” (that means nothing in HTTP).
Preserve three independent outcomes
Case | Mode | Exit Code | HTTP Code | URL Effective | # Redirects | Content-Type | JSON parse ok? | Ready=true? | Service match? | Final decision |
/ok | --fail | 0 | 200 | /ok | 0 | application/json | yes | yes | yes | ACCEPT |
/ok | --fail --location | 0 | 200 | /ok | 0 | application/json | yes | yes | yes | ACCEPT |
/redirect-good | --fail | 0 | 302 | /redirect-good | 0 | (none) | N/A | N/A | N/A | HOLD (redir) |
/redirect-good | --fail --location | 0 | 200 | /ok | 1 | application/json | yes | yes | yes | HOLD (redir) |
/redirect-login | --fail | 0 | 302 | /redirect-login | 0 | (none) | N/A | N/A | N/A | HOLD (redir) |
/redirect-login | --fail --location | 0 | 200 | /login | 1 | text/html | no (HTML) | no | no | HOLD (HTML) |
/login | --fail | 0 | 200 | /login | 0 | text/html | no (HTML) | no | no | HOLD (HTML) |
/login | --fail --location | 0 | 200 | /login | 0 | text/html | no (HTML) | no | no | HOLD (HTML) |
/wrong-json | --fail | 0 | 200 | /wrong-json | 0 | application/json | yes | yes | no | HOLD (data) |
/wrong-json | --fail --location | 0 | 200 | /wrong-json | 0 | application/json | yes | yes | no | HOLD (data) |
/bad-json | --fail | 0 | 200 | /bad-json | 0 | application/json | no | N/A | N/A | HOLD (parse) |
/bad-json | --fail --location | 0 | 200 | /bad-json | 0 | application/json | no | N/A | N/A | HOLD (parse) |
/unhealthy | --fail | 22 | 503 | /unhealthy | 0 | application/json | yes | no | yes | HOLD (status) |
/unhealthy | --fail --location | 22 | 503 | /unhealthy | 0 | application/json | yes | no | yes | HOLD (status) |
/empty | --fail | 0 | 204 | /empty | 0 | (none) | N/A | N/A | N/A | HOLD (no-content) |
/empty | --fail --location | 0 | 204 | /empty | 0 | (none) | N/A | N/A | N/A | HOLD (no-content) |
/loop | --fail | 47 | (none) | (n/a) | max (2) | (none) | N/A | N/A | N/A | HOLD (loop) |
/loop | --fail --location | 47 | (none) | (n/a) | max (2) | (none) | N/A | N/A | N/A | HOLD (loop) |
In the decision column, ACCEPT means the contract fully satisfied, while HOLD indicates failure (with reason in parentheses). Note that the “redirect-good” case both with and without -L is held because our current contract forbids redirects. The “login”, “wrong-json”, “bad-json”, “unhealthy”, “empty”, and “loop” cases all fail at least one predicate (status, content type, JSON parsing, or redirect) and are correctly held. This table explicitly shows each case and which simple gates would have incorrectly passed them. For example, an “exit-only” gate would have accepted /login (exit 0) and /wrong-json; a “status 200” gate would have accepted /login and /wrong-json; only the full JSON contract gate accepts just /ok.
Reject a result table that silently drops bad cases
It’s crucial we maintain one row per test, never filtering out errors. The matrix above includes the loop and empty-case rows rather than dropping them. A polished checklist must account for every case ID, so that no failure scenario is overlooked. Any missing row or skipped column could hide a failure, which we must avoid.
Validate JSON identity instead of accepting any 200
We now implement the strict checker in Python (or similar) to enforce our JSON contract on curl’s outputs. Pseudocode:
# Suppose we have variables from curl's write-out:http_code = int(output_http_code)url_eff = output_url_effectivenum_redir = int(output_num_redirects)content_type = output_content_type # e.g. 'application/json'body = open("body.bin","rb").read()decision = True # will mark False on any violation# Criterion 1: exit code and HTTP statusif exit_code != 0: decision = Falseif http_code != 200: decision = False# Criterion 2: final URL and redirect count (no redirects)if url_eff != f"http://127.0.0.1:{port}/ok": decision = Falseif num_redir != 0: decision = False# Criterion 3: content type essenceif not content_type or not content_type.lower().startswith("application/json"): decision = False# Criterion 4: parse JSON body exactlytry: data = json.loads(body.decode('utf-8'))except Exception: decision = Falseelse: if not isinstance(data, dict): decision = False # check required fields if data.get("service") != "refonte-lab": decision = False if data.get("revision") != "r1": decision = False # for readiness, ensure it is True boolean, not truthy like 1 or "true" if data.get("ready") is not True: decision = FalseWe explicitly fail if any condition is unmet. Notably, the check data.get("ready") is not True ensures only the real JSON boolean true passes; a numeric or string true would fail.
If any check fails, we mark the record as failed contract. We output a machine-readable verdict (e.g. accept = 1 or 0). We do not throw exceptions or halt the process; we simply log the reason (e.g. “service mismatch” or “malformed JSON”).
After parsing, we consider the consistency with configuration. For instance, finding an unexpected field would also be cause for failure, but since we use equality, extra fields don’t automatically fail (unless we explicitly enforce "no extra fields"). In a tighter schema, we could compare sets of keys, but for this lab we assume extra fields are not present.
Importantly, even if a test fails, we keep the JSON content, headers, and raw body as diagnostics. For example, if /unhealthy yields 503 {"ready":false,...}, we might save that body and error code for logs, but our final verdict remains “fail”.
By the end, only the /ok cases (both modes) yield all criteria true. Every negative control yields at least one false criterion: wrong URL, wrong status, wrong type, parse error, or wrong field value.
Declare any permitted redirect as a different policy
Our default contract prohibits any redirects. If an organization decides to allow exactly the one good redirect, that must be a separate “compatibility” policy. For example, a policy that permits the known /redirect-good -> /ok path might relax “0 redirects” to “up to 1 redirect ending at /ok”. That would require an explicit change: adding logic like
# Compatibility policy example:if exit_code == 0 and url_eff == f"http://...:{port}/ok" and num_redir == 1: # allow one known good redirect pass # continue checking JSON, etc.However, this is separate from the default gate. We keep the default check as num_redir==0; the alternate check is a different branch. Importantly, running curl with -L and then checking url_effective does not prevent the request from going out. The redirect was already followed and credentials (if any) were already sent. Thus, allowing it in the check is purely retrospective. For security, one must consider that allowing redirects might expose the client to unintended URIs (SSRF concerns). In our fixture, no credentials or secrets are used, so it is safe to illustrate. But real readiness checks must consider redirect risks separately.
Do not mistake a terminal-URL check for a sending policy
Even if our final JSON check sees that the end URL is /ok, curling with -L means we did send a request to the intermediate location. That request cannot be undone. The check simply records the final URL. If credentials or sensitive headers were involved, one could have inadvertently sent them to the redirect target. Therefore, the "permitted redirect" policy is only a post-facto evaluation: it does not change that the redirect actually happened. As a principle, do not rely on the final URL after a redirect as proof that the redirect was safe or intended; it only tells you what happened, not what should have happened.
Preserve useful diagnostics without changing the verdict
By default -f/--fail hides error bodies for 4xx/5xx responses. To assist debugging while still failing, we demonstrate using --fail-with-body, which was added in curl 7.76.0. For example:
curl --disable -L --fail-with-body ... http://127.0.0.1:$port/unhealthy -o body_unhealthy.binWith --fail-with-body, curl still exits nonzero on 503, but saves the JSON body {"ready":false,...} to the output file (instead of discarding it). This allows us to inspect the error content in logs. However, important: preserving the body does not turn the check into a pass. We still see status 503 and ready:false, which fails our contract. The extra data is only for diagnostics. We keep stderr (curl's error message) and the saved body separate from the JSON parser input. We might log “HTTP 503: ready=false” in our test report. This rich error info can guide troubleshooting (for example, alerting that the /unhealthy service needs fixing), but the verdict remains hold.
In practice, using --fail-with-body can replace a separate --dump-header and output file, but only if the curl build supports it. We include it in our script (checking curl --version), or conditionally fall back. We do not use --fail-with-body as a way to accidentally accept errors. We also avoid any shell tricks that conflate stderr and stdout. Each artifact (headers, body, errors, exit code, JSON parse log) is recorded independently in the evidence directory.
Reconcile every expected case with the observed matrix
With all data captured, we reconcile the results. We ensure:
There are 9 case IDs (/ok, /redirect-good, etc.) × 2 modes (no-follow, follow) = 18 records.
No duplicates or missing rows.
Each record has all fields (exit_code, http_code, url, redirects, content_type, etc.) even if unknown (we can mark HTTP code blank if curl didn’t get one).
The JSON contract final pass indicator is derived solely from the predicates.
We then examine which scenarios would have been false positives under simpler rules. For example:
Exit-only gate (success if exit code = 0): would have incorrectly accepted /redirect-good, /redirect-login, /login, /wrong-json, /bad-json, /empty.
Status-only gate (exit 0 AND status 200): would have accepted /login, /wrong-json, /bad-json, but correctly rejected the redirect cases (they were 302).
Strict contract (all conditions): accepts only /ok.
Because each negative case is present in the matrix, we see that all bad cases did fail the full gate. (For example, /bad-json fails at JSON parsing; /empty fails at missing body; /wrong-json fails at the service field; etc.) In fact, a quick check of the matrix shows every “HOLD” entry except /ok.
If any case had slipped through, we would adjust the check or the test data. Here, the output aligns with expectations: our repaired gate rejects every unwanted response.
Reject a result table that silently drops bad cases
It is critical we do not leave out, say, the loop or empty cases just because we don’t care about them in the app logic. All nine are listed above so no failure is hidden. We do not drop rows for which our code marked "no body" or parse errors. Doing so would give a false sense of security. By keeping them in the table, we demonstrate that even these corner cases are handled (by rejecting them).
Repair the probe or endpoint without rewriting history
When the gate fails, we must decide who should fix what:
If the API returned a redirect or wrong data, the endpoint owner is responsible. For instance, if /redirect-login unexpectedly pointed to /login, maybe the service is misconfigured or needs authentication. The fix is on the server side (or update the contract if it’s intentional).
If the probe invocation missed a config option (e.g. we realize we should have allowed that one safe redirect), the probe maintainer updates the wrapper script and adds regression tests for that scenario.
If the failure was due to a client timeout or network issue, the infrastructure owner might increase timeouts or fix connectivity; we then rerun the probe from scratch (with a new timestamp). But crucially, a later success after a timeout does not prove the original attempt was fine; it just proves conditions improved. We keep the old evidence and label it “inconclusive” versus the new evidence after the fix.
If JSON parsing failed (like /bad-json), the service developer should correct the output format. Meanwhile, the contract might be updated if more fields are intended (but here malformed JSON is clearly a bug).
In all cases, we do not “edit history”. We version the fix: e.g. commit a change to the server code or to the probe logic, and then rerun the exact same test steps, producing a new evidence set (with different directory or timestamp). Both sets of results remain in logs. If the new run passes where the old failed, we highlight that only after the fix the contract is truly satisfied.
The outcome of a repair is a new verdict for that case, but it does not retroactively change the original run’s result. In our writing, we might say: “Initial run: FAIL; after enabling --location, second run: PASS on /redirect-good”, but acceptance of the release depends on addressing the root cause (in this example, revisiting the redirect rule).
Choose accept, repair, rerun or hold from the evidence
We now have a final decision matrix that ties every predicate to an actionable verdict. The only fully acceptable scenario is when all the strict predicates are met. The checklist is:
curl exit code = 0 (no transport error).
Final HTTP status = 200.
Number of redirects = 0 (effective URL = BASE+'/ok').
Content-Type essence = application/json.
JSON parsed successfully.
JSON fields: service=="refonte-lab", revision=="r1".
JSON field ready is the boolean true.
If any of (1)-(7) is false or unknown, the case does not satisfy the contract. In particular, if (1) is not met (exit nonzero or timeout), we cannot assume anything about (2)-(7). That result is inconclusive for the HTTP-level contract and must be investigated as a transport error.
Below is the compact version of this logic (for say logging):
Criterion | Required Value | Observed Value (example /ok) | Pass/Fail |
Exit code | 0 | 0 | PASS |
HTTP status | 200 | 200 | PASS |
Redirect count | 0 | 0 | PASS |
URL ends in /ok | yes | yes | PASS |
Content-Type has application/json | yes | yes | PASS |
JSON parse without error | yes | yes | PASS |
service == "refonte-lab" | yes | yes | PASS |
revision == "r1" | yes | yes | PASS |
ready is boolean true | yes | yes | PASS |
Only if all rows are PASS do we mark the probe as ACCEPT. Otherwise, we might HOLD or REPAIR. For example, /unhealthy had exit 22, status 503, and ready=false, resulting in multiple failures. That case would trigger either “fix the service” or “block deployment.”
The final reported decision for each endpoint is either “ACCEPT” (if strict) or “HOLD” (if not), tied to which team or process fixes it. (For negative controls, there is no “Accept”; they are only for verifying the checker itself.) In deployment, the release approver might see a summary like “No acceptance; hold release: service outputs incorrect content”.
Make negative controls part of acceptance
A proper acceptance gate also requires that every known-bad response must be rejected by the checker. In other words, our checker is only trustworthy if it does not falsely mark a wrong response as acceptable. This is why we include the loop, empty, and HTML cases: to show that the gate indeed rejects them. We can say: “The readiness checker is validated only when /ok passes and all other variant responses fail as designed.” If we found a bug where, say, /empty slipped through, we would refine the checker (e.g. require non-empty body).
Assign ownership and regression checks to the probe contract
Finally, tie each part of this contract to an owner and future-proof it. For example:
Endpoint/schema owner: perhaps the backend team owning the /ok API (e.g. the on-call engineer or service owner). They are responsible for changes to the response JSON format. Any change in fields or meanings must go through code review and update the test fixture. They are alerted if the contract fails due to schema drift (e.g. missing "revision").
Probe wrapper maintainer: the SRE or developer who wrote the curl check script. They ensure the arguments (timeouts, fail flags, no-follow etc.) stay correct. If a new requirement arises (like allowing one redirect), they must version the change and add a regression test for it. They are responsible for running this check in CI and addressing any environment differences.
Release approver: the QA or release engineer sees the final verdict. They should know that “exit=0,status=200” is not enough; they look at content results. The approver decides ACCEPT or HOLD based on the completed matrix (and may consult logs if needed).
Evidence retention: we store all request/response logs for a limited time (e.g. the CI job logs or a test report). We sanitize these logs to avoid credentials (we used none here) or proprietary data. The negative-case logs (login page HTML, wrong JSON) are retained as test artifacts, not as secrets. The team might only keep them for the duration of the release review. We do not expose sensitive info; the fixture data here is synthetic.
We also assign regression tickets: if the service intends to change the readiness output (e.g. add a “version” field), the backend owner must update the contract and corresponding tests. Conversely, if the probe logic must adapt (e.g. to follow authorized redirect), the wrapper owner records that with a test that /redirect-good now is acceptable.
In summary, each broken predicate has an owner: service fields (backend), HTTP behavior (backend), redirect policy (security/ops), and the acceptance checker itself (DevOps tester). This ensures accountability: e.g. “Why did we accept a login page last time? Oh, because we forgot to block redirects. Fix that in code.”
Develop these review habits with DevOps Engineering
A diligent DevOps or SRE engineer treats a passing health-check probe as one data point, not proof of global health. The check we built here strictly validates the exact API response, a practice aligned with good API testing/monitoring practices. To build these skills, structured learning can help. For example, Refonte Learning’s DevOps Engineering Program covers Linux scripting, CI/CD pipelines, and monitoring tools, all of which are relevant to writing robust health-check scripts like the one above. By practicing these labs and understanding how to inspect every piece of evidence, you’ll be better prepared to catch subtle probe failures in real deployments.
For structured practice in Linux scripting, CI/CD and monitoring, explore Refonte Learning’s DevOps Engineering program.
