Python developer validating gzip download sizes and SHA-256 hashes at a workstation.

A Requests Download Can Be Larger Than Its Content-Length

Wed, Oct 7, 2026

When a Python client uses requests to GET a gzip-encoded CSV over HTTP/1.1, it might see Content-Length: 52 in the headers but end up reading 1,293 bytes in memory. This is expected here, because Requests automatically decompresses gzip content by default. The important questions are: which byte sequence does our contract actually intend to use, the 1,293-byte CSV text, or the 52-byte gzip blob? And did we check the length and hash of that exact sequence?

In this article we walk through a controlled loopback server that serves a fixed CSV in various encodings and read modes. We show that status code 200 and a plausible length alone do not guarantee the right content. Instead, we explicitly choose the artifact layer (decoded CSV vs. encoded archive), then verify length and SHA-256 against an independent manifest for that layer. By treating Content-Encoding and Content-Length accurately (per HTTP semantics), we can accept or reject the downloaded artifact. We build tests in CPython 3.12.14 with requests==2.32.5 and urllib3==2.5.0 (plus certifi 2026.7.22, idna 3.20, charset-normalizer 3.5.2), run on 127.0.0.1 with a small CSV. All code runs locally; this is not a production incident analysis. We conclude with a decision table and a safe handoff policy for any artifact that fails its contract.

This byte-level verification exercise complements broader API robustness and security practices like access control and validation (which this lab does not cover). We focus narrowly on download completeness and identity. We do not assume any publisher trust from a mere HTTP 200 or matching length; instead we compute SHA-256 on the exact retained bytes and compare it to the manifest for that same layer.

Define the artifact before checking its download

First, decide what the intended artifact is. In our scenario we have two legitimate contracts:

•        Decoded CSV: The client cares about the raw CSV text (the ingestion-ready data). In this contract, the artifact is the uncompressed CSV, and the allowed received codings are declared separately. This fixture accepts an uncoded response or gzip decoded by iter_content(). A strict identity-only request remains a separate policy. The expected length and SHA-256 come from the original CSV file generated independently.

•        Encoded gzip file: The client cares about the exact gzip-compressed payload (for example, storing the raw archive). Here the artifact is the gzip bytes. The response may carry Content-Encoding: gzip read without decoding, or be an application/gzip file with no Content-Encoding. The expected length and SHA-256 come from compressing the CSV with a fixed method (e.g. gzip.compress(..., mtime=0)).

These contracts cannot share a single length or hash. For example, 1,293 ≠ 52 and their SHA-256 will differ, so we list them separately:

Contract (artifact)

Bytes included in artifact

Expected length from

Expected digest from

Permitted Content-Encoding

Decoded CSV

decoded CSV payload (1293 bytes)

Decoded manifest (local CSV)

Decoded manifest (CSV SHA-256)

No coding, or gzip decoded by the selected read mode

Encoded gzip archive

gzip-compressed bytes (52 bytes)

Encoded manifest (gzipped)

Encoded manifest (gzip SHA-256)

gzip read without decoding, or no coding for a gzip file

Each contract has an owner. The API consumer (client) owns the chosen read mode and validation code. The data/manifest owner (e.g. downstream team) provides the expected length and digest for the artifact. If the owner wants to switch from decoded to encoded or vice versa, that must be an explicit contract change. API robustness and security practices remind us that defining clear contracts is essential.

Separate content coding from application bytes

HTTP clearly separates content coding (compression) from the application data. The Content-Encoding header lists how the representation was coded. The Content-Type identifies the media type after decoding. The Content-Length is defined as the length of the transferred content in octets, i.e. the bytes on the wire before decoding, since decoding is a representation-level transformation. In our GET/200 fixture, if we see Content-Length: 52 and Content-Encoding: gzip, that means 52 bytes of compressed data were sent; once the client decodes gzip, it gets 1,293 decoded bytes.

HTTP/1.1 200 OK
Content-Type: text/csv
Content-Encoding: gzip
Content-Length: 52

[binary gzip data]

Requests will automatically decode Content-Encoding: gzip when you use response.iter_content() or response.content. Content decoding is separate from HTTP framing, which is handled by the HTTP parser before the body is exposed (we’re not using chunked encoding here). That is, Content-Length: 52 was about the compressed payload. Only after the client decompresses does it see 1,293 bytes of CSV. Conversely, if no Content-Encoding is present, then Content-Length is the number of bytes of whatever data (possibly still a .gz file if Content-Type: application/gzip) is in the payload.

Rule: Don’t compare the decimal Content-Length to a different “layer.” If you compare Content-Length=52 to the post-decode count 1,293, you are matching apples to oranges. And matching just one count doesn’t prove content is correct. (For correct content, compare a digest or exact hash. For example, Postman tests often assert status 200, but here we must verify every byte if we trust the content.)

A response header and a client iterator may describe different byte sequences

In other words, the HTTP header and the iter_content output can legitimately be different sets of bytes. Content-Length referred to gzip bytes, while iter_content (with default decoding) yielded the uncompressed CSV. Since they describe different representations, their lengths do not have to match. Nor does a matching count ensure correctness (an incomplete or malicious payload could still match a size).

For example:

•        In this fixture, if Content-Encoding: gzip, then len(body) after decompress > Content-Length in the header.

•        If Content-Encoding is absent (or identity), then len(body) will equal Content-Length (and also equals the CSV size in our fixture, 1,293).

A superficial check like len(body) == int(res.headers["Content-Length"]) is invalid in the gzip case, and might falsely flag a correct download as an error. We need to compare the correct artifact (decoded or encoded) to its expected length and hash.

Build the loopback server and independent manifest

We create a local Python fixture to serve our cases. The code below (save as requests_download_lab.py) sets up a ThreadingHTTPServer on 127.0.0.1 with fixed CSV data. It pre-computes MANIFEST entries for both the decoded and encoded artifacts. Each request path returns either the raw CSV or its gzip-compressed bytes, with appropriate headers. One path truncates the output to simulate a failure. The fixture uses gzip.compress(..., mtime=0) so the gzip timestamp is fixed within the recorded runtime.

import gzip
import hashlib
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
import importlib.metadata
import json
from pathlib import Path
import platform
import sys
import threading
import tempfile

sys.path.insert(0, str(Path(__file__).parent / "requests-deps"))
import requests
import urllib3

if not debug:
    raise RuntimeError(
        "Run this fixture without -O or PYTHONOPTIMIZE; "
        "its assertions are required."
    )
RUN_DIR = Path(
    tempfile.mkdtemp(prefix="requests-download-lab-", dir=Path(__file__).parent)
)
RESULT_PATH = RUN_DIR / "requests-preflight-results.json"
RESULT_PATH.write_text(
    json.dumps({"run_id": RUN_DIR.name, "status": "RUNNING"}) + "\n"
)
PAYLOAD = b"row_id,value\n" + b"A,10\nB,20\n"  128
ENCODED = gzip.compress(PAYLOAD, mtime=0)
REENCODED = gzip.compress(PAYLOAD, mtime=1)
ALTERED = gzip.compress(PAYLOAD.replace(b"A,10", b"A,11", 1), mtime=0)
digest = lambda b: hashlib.sha256(b).hexdigest()
MANIFEST = {
    "decoded": {"length": len(PAYLOAD), "sha256": digest(PAYLOAD)},
    "encoded": {"length": len(ENCODED), "sha256": digest(ENCODED)},
}
request_log = []


class Handler(BaseHTTPRequestHandler):
    protocol_version = "HTTP/1.1"

    def log_message(self, args):
        pass

    def do_GET(self):
        request_log.append({
            "path": self.path,
            "accept_encoding": self.headers.get("Accept-Encoding"),
        })
        variants = {
            "/identity": (PAYLOAD, None, "text/csv"),
            "/gzip": (ENCODED, "gzip", "text/csv"),
            "/gzip-file": (ENCODED, None, "application/gzip"),
            "/reencoded": (REENCODED, "gzip", "text/csv"),
            "/altered": (ALTERED, "gzip", "text/csv"),
            "/truncated": (ENCODED, "gzip", "text/csv"),
        }
        if self.path not in variants:
            self.send_error(404)
            return
        body, coding, media_type = variants[self.path]
        self.send_response(200)
        self.send_header("Content-Type", media_type)
        self.send_header("Content-Length", str(len(body)))
        self.send_header("Connection", "close")
        if coding:
            self.send_header("Content-Encoding", coding)
        self.end_headers()
        self.wfile.write(
            body[:len(body) // 2] if self.path == "/truncated" else body
        )
        self.close_connection = True


server = ThreadingHTTPServer(("127.0.0.1", 0), Handler)
thread = threading.Thread(target=server.serve_forever, daemon=True)
thread.start()
base = "http://127.0.0.1:" + str(server.server_port)
results = []
try:
    with requests.Session() as session:
        session.trust_env = False
        # Test cases: path and mode (iter for decoded, raw for encoded)
        for path, mode in [
            ("/identity", "iter"),
            ("/gzip", "iter"),
            ("/gzip", "raw"),
            ("/gzip-file", "iter"),
            ("/reencoded", "iter"),
            ("/reencoded", "raw"),
            ("/altered", "iter"),
        ]:
            with session.get(
                base + path,
                stream=True,
                timeout=(2, 2),
                allow_redirects=False,
            ) as response:
                assert response.status_code == 200
                chunks = (
                    response.iter_content(chunk_size=17)
                    if mode == "iter"
                    else response.raw.stream(amt=17, decode_content=False)
                )
                body = b"".join(chunks)
                intended = (
                    "encoded"
                    if mode == "raw" or path == "/gzip-file"
                    else "decoded"
                )
                matched = (
                    len(body) == MANIFEST[intended]["length"]
                    and digest(body) == MANIFEST[intended]["sha256"]
                )
                record = {
                    "path": path,
                    "mode": mode,
                    "content_encoding": response.headers.get("Content-Encoding"),
                    "content_length": int(response.headers["Content-Length"]),
                    "saved_length": len(body),
                    "completion": "complete",
                    "intended_manifest": intended,
                    "decision": (
                        (
                            "ACCEPT ENCODED ARTIFACT"
                            if intended == "encoded"
                            else "ACCEPT DECODED ARTIFACT"
                        )
                        if matched
                        else "HOLD INTEGRITY"
                    ),
                    "matches_decoded": (
                        digest(body) == MANIFEST["decoded"]["sha256"]
                    ),
                    "matches_encoded": (
                        digest(body) == MANIFEST["encoded"]["sha256"]
                    ),
                    "saved_sha256": digest(body),
                }
                results.append(record)
        by_case = {(r["path"], r["mode"]): r for r in results}
        # Verify the fixture acceptance cases.
        assert by_case[("/identity", "iter")]["matches_decoded"]
        assert by_case[("/gzip", "iter")]["matches_decoded"]
        assert (
            by_case[("/gzip", "iter")]["saved_length"]
            > by_case[("/gzip", "iter")]["content_length"]
        )
        assert by_case[("/gzip", "raw")]["matches_encoded"]
        assert by_case[("/gzip-file", "iter")]["matches_encoded"]
        assert by_case[("/reencoded", "iter")]["matches_decoded"]
        assert not by_case[("/reencoded", "raw")]["matches_encoded"]
        assert not by_case[("/altered", "iter")]["matches_decoded"]
        assert by_case[("/altered", "iter")]["saved_length"] == len(PAYLOAD)
        # Request with Accept-Encoding: identity (unexpected result)
        with session.get(
            base + "/gzip",
            headers={"Accept-Encoding": "identity"},
            stream=True,
            timeout=(2, 2),
            allow_redirects=False,
        ) as response:
            # Server still responds with gzip
            results.append({
                "path": "/gzip",
                "case": "identity_request_unexpected_coding",
                "mode": "not consumed",
                "completion": "not consumed: contract rejected",
                "content_encoding": response.headers.get("Content-Encoding"),
                "content_length": int(response.headers["Content-Length"]),
                "saved_length": None,
                "saved_sha256": None,
                "decision": "HOLD CONTRACT",
            })
        # Truncated case
        truncated_record = {
            "path": "/truncated",
            "mode": "iter",
            "completion": "incomplete",
            "content_encoding": None,
            "content_length": None,
            "saved_length": None,
            "saved_sha256": None,
        }
        try:
            with session.get(
                base + "/truncated",
                stream=True,
                timeout=(2, 2),
                allow_redirects=False,
            ) as response:
                assert response.status_code == 200
                truncated_record.update({
                    "content_encoding": response.headers.get("Content-Encoding"),
                    "content_length": int(response.headers["Content-Length"]),
                })
                 = b"".join(response.itercontent(chunk_size=17))
        except requests.RequestException as exc:
            results.append({
                **truncated_record,
                "exception": type(exc).__name__,
                "decision": "HOLD INCOMPLETE",
            })
        else:
            raise AssertionError("truncated framing did not raise in this baseline")
except BaseException as exc:
    RESULT_PATH.write_text(
        json.dumps({
            "run_id": RUN_DIR.name,
            "status": "FAIL",
            "all_assertions_passed": False,
            "exception": type(exc).__name__,
            "message": str(exc),
            "cases": results,
            "request_log": request_log,
        }, indent=2) + "\n"
    )
    print("Failed-run evidence: " + str(RESULT_PATH), file=sys.stderr)
    raise
finally:
    server.shutdown()
    server.server_close()
    thread.join(timeout=2)
report = {
    "python": sys.version.split()[0],
    "os": platform.platform(),
    "run_id": RUN_DIR.name,
    "status": "PASS",
    "evidence_path": str(RESULT_PATH),
    "requests": requests.__version__,
    "urllib3": urllib3.__version__,
    "dependencies": {
        p: importlib.metadata.version(p)
        for p in ["requests", "urllib3", "certifi", "idna", "charset-normalizer"]
    },
    "manifest": MANIFEST,
    "cases": results,
    "request_log": request_log,
    "all_assertions_passed": True,
}
RESULT_PATH.write_text(json.dumps(report, indent=2) + "\n")
print(json.dumps(report, indent=2))
Run the above in a disposable environment with:
python3 -m pip install --target requests-deps \
    'requests==2.32.5' 'urllib3==2.5.0'
python3 requests_download_lab.py

This produces JSON in requests-download-lab-*/requests-preflight-results.json. The report includes the environment versions and both artifact manifests. The example values are:

{
  "python": "3.12.14",
  "requests": "2.32.5",
  "urllib3": "2.5.0",
  "dependencies": {
    "certifi": "2026.7.22",
    "idna": "3.20",
    "charset-normalizer": "3.5.2"
  },
  "manifest": {
    "decoded": {
      "length": 1293,
      "sha256": "e4b0a45101b66de0dc1423b77f1148e070a26177bcb78e971b788fd18ac5f191"
    },
    "encoded": {
      "length": 52,
      "sha256": "364e1e580709494891178dc2ea81cc098e35d29707411a17c457aab5112f0128"
    }
  }
}

Use the decoded or encoded manifest for the selected contract. Record the actual interpreter, operating system, and dependency versions for each run. The fixture logs each case with content_encoding, content_length, saved_length, matches_*, and a decision. The request_log notes what Accept-Encoding the client sent (requests may include defaults). We set trust_env=False to ignore system proxies.

Reproduce the false failure with decoded iteration

Consider the simplest case /identity (no Content-Encoding). The server sends the CSV text (Content-Length:1293), and iter_content() yields 1,293 bytes. All good. Now consider /gzip with content coding: the server sends 52 bytes of gzip (Content-Length:52, Content-Encoding:gzip). If we do response.iter_content(17), Requests will decompress on the fly (because decode_content=True by default). The client body ends up 1,293 bytes and matches the decoded SHA-256. However, if we naïvely compared len(body) == int(response.headers["Content-Length"]), we would (incorrectly) see 1293 != 52 and flag an error.

with requests.Session() as s:
    s.trust_env = False
    res_id = s.get(base + "/identity", stream=True, timeout=(2,2))
    data_id = b"".join(res_id.iter_content())
    print(
        res_id.headers.get("Content-Encoding"),
        res_id.headers["Content-Length"], len(data_id)
    )
    # -> None 1293 1293
    print(hashlib.sha256(data_id).hexdigest())  # matches decoded manifest

    res_gz = s.get(base + "/gzip", stream=True, timeout=(2,2))
    data_gz = b"".join(res_gz.iter_content())
    print(
        res_gz.headers.get("Content-Encoding"),
        res_gz.headers["Content-Length"], len(data_gz)
    )
    # -> gzip 52 1293  (decoded to 1293 bytes)
    print(hashlib.sha256(data_gz).hexdigest())  # matches decoded manifest
Output (expected):
None 1293 1293
e4b0a45101b66de0dc1423b77f1148e070a26177bcb78e971b788fd18ac5f191
gzip 52 1293
e4b0a45101b66de0dc1423b77f1148e070a26177bcb78e971b788fd18ac5f191

Here, len(data_gz)==1293 while Content-Length==52. The criterion len(data)==Content-Length would fail and wrongly reject this correct download. Instead, we recognize that /gzip was expected to decode (intended manifest = "decoded") and we should compare 1,293 to the decoded manifest. The mismatch with 52 is irrelevant because 52 was only about the encoded layer.

stream=True is not a decoder switch

Note that setting stream=True only means "delay downloading until needed"; it does not disable gzip decoding. The docs emphasize using iter_content() with stream=True for large downloads, but content decoding still occurs. That is why iter_content gave us 1,293 bytes. Conversely, accessing response.raw directly lets us control decode_content. You cannot switch modes after partially reading; once you call iter_content, the gzip decoder runs. This is why our code uses a fresh new request for each mode (iterated vs raw).

Preserve the coded body when that is the contract

To accept the encoded gzip artifact (per the transport contract), we must avoid decompression. We do this by reading from response.raw.stream(..., decode_content=False). For example:

res = requests.get(base + "/gzip", stream=True, timeout=(2,2))
raw_chunks = res.raw.stream(amt=17, decode_content=False)
coded_data = b"".join(raw_chunks)
print(
    res.headers.get("Content-Encoding"),
    res.headers["Content-Length"], len(coded_data)
)
# -> gzip 52 52
print(hashlib.sha256(coded_data).hexdigest())  # matches encoded manifest
Output (expected):
gzip 52 52
364e1e580709494891178dc2ea81cc098e35d29707411a17c457aab5112f0128

Indeed, len(coded_data)==52, matching Content-Length. The SHA-256 matches the encoded manifest, not the decoded one. In this mode the saved bytes are exactly the gzip payload, so the coded manifest (manifest["encoded"]) should match. The decoded manifest is a different SHA-256 and is not relevant here. We deliberately see a mismatch between coded and decoded hashes because they are different layers. That mismatch is expected and not an error under the encoded contract.

Raw response data has already crossed the HTTP parsing boundary

It’s important to be precise: response.raw in Requests/urllib3 represents the body stream after HTTP parsing but before content-decoding. It does not include the status line, headers or TLS framing; it exposes only the raw octets of the message body. In HTTP/1.1, chunked encoding is also a framing mechanism, but our fixture has no Transfer-Encoding; if it did, response.raw.read() would automatically de-chunk it. Here, “raw” means “raw coded payload”.

We cannot mix modes on one response. Once any bytes are pulled out (e.g. via iter_content), the rest are buffered or consumed and we cannot rewind. That’s why each test case uses a separate Session.get() context. The decode_content option controls content decoding. Choose it before consuming the body and use a fresh response for the other mode.

Use re-encoding and same-length mutation as independent controls

We include two additional controls to validate our logic:

•        Re-encoded (/reencoded): We gzip the same CSV with a different mtime (1 vs 0). This yields a different gzip byte sequence (header differs), even though the decoded CSV is identical. In our preflight, both gzip bodies were 52 bytes, but with different SHA-256 (ENCODED vs REENCODED). With iter_content, the decoded CSV matches the manifest (because decoding succeeds), so under the decoded contract we accept. But raw (undecompressed) yields a different hash, so under an encoded contract we would reject (HOLD_INTEGRITY in our code) because the bytes differ from the expected encoded manifest. This proves we must check the correct layer’s hash.

•        Altered payload (/altered): We modify one number in the CSV (A,10 → A,11) but gzip with same settings. The encoded length becomes 54 bytes, while decoding still yields 1,293 bytes. However, the decoded CSV hash now differs from the original manifest. Even though len(body)==1293, the SHA-256 is not the expected one. Under the decoded contract this must HOLD_INTEGRITY. The decoded length and status code did not change, but content did. Neither compressed size nor status can replace a proper identity check. We do not assume a bad gzip header is the culprit; it’s simply the wrong CSV content.

In bullet form:

•        /reencoded: decoded SHA-256 matches original (=> ACCEPT DECODED ARTIFACT), but encoded SHA-256 does not (=> if we were using encoded contract, it would fail).

•        /altered: decoded SHA-256 does not match original (=> HOLD INTEGRITY), even though lengths match.

This shows that only a reliable hash comparison can catch a one-byte logical change. Counts or compression ratio alone are insufficient. This is not a mysterious transmission error; it’s a deliberate data change that our validation must detect.

Validate the received coding instead of trusting the request preference

We also check how the server responds to client preferences. We send Accept-Encoding: identity in the request to /gzip, expecting no compression. In our loopback, the server still returns Content-Encoding: gzip (ignoring the preference). The response headers are:

Content-Encoding: gzip
Content-Length: 52

Since we asked for identity, this violates our contract. We did not consume the body; instead we HOLD CONTRACT.

This illustrates that server choice matters. A client’s Accept-Encoding header is only a preference; the actual Content-Encoding header tells us what we received. We do not automatically decode if we insisted on identity. Instead, because the content is not in the expected form, we refuse to proceed. We do not treat this as a content corruption, but as a contract mismatch. (Do not view this as a standards compliance test; it’s simply how our client’s acceptance policy works.)

A gzip file without Content-Encoding is a separate control

Finally, we fetch /gzip-file, which returns the same 52-byte gzip content but with Content-Type: application/gzip and no Content-Encoding. In this case, Requests will not decompress, because it only decodes when Content-Encoding is set. It treats it as a generic binary payload. Thus iter_content yields the 52 gzip bytes, and the SHA-256 matches the encoded manifest. We accept under the encoded artifact contract. This shows that media type alone (e.g. .gz file) does not trigger automatic unpacking. We do not recommend manually gunzipping every payload named “.gz”; instead, the contract drives whether we decode or not.


Reconcile every case in one evidence table

The key observations from all cases can be summarized as follows:

Case (path, mode)

Received Coding

Declared Length

Method

Retained Length

Intended manifest

Digest match?

Completion/Exception

Decision

/identity, iter_content

None

1293

iter

1293

decoded

yes (decoded)

complete

ACCEPT DECODED ARTIFACT

/gzip, iter_content

gzip

52

iter

1293

decoded

yes (decoded)

complete

ACCEPT DECODED ARTIFACT

/gzip, raw.stream(decode_content=False)

gzip

52

raw

52

encoded

yes (encoded)

complete

ACCEPT ENCODED ARTIFACT

/gzip-file, iter_content

None

52

iter

52

encoded

yes (encoded)

complete

ACCEPT ENCODED ARTIFACT

/reencoded, iter_content

gzip

52

iter

1293

decoded

yes (decoded)

complete

ACCEPT DECODED ARTIFACT

/reencoded, raw.stream

gzip

52

raw

52

encoded

no (encoded)

complete

HOLD INTEGRITY

/altered, iter_content

gzip

54

iter

1293

decoded

no (decoded)

complete

HOLD INTEGRITY

/gzip, identity header (not read)

gzip

52

not consumed

–

–

–

not consumed: contract rejected

HOLD CONTRACT

/truncated, iter_content

gzip

52

iter

– (partial)

–

–

exception (ChunkedEncodingError)

HOLD INCOMPLETE

Each row names the contract in use (decoded vs encoded) to avoid confusion. For incomplete or rejected cases, we leave saved length and digest blank. This table shows which artifact bytes were kept (Retained Length) and how they matched the expected manifest. For example, /gzip with iter mode had a retained length (1293) much larger than the declared 52 bytes, but it was intentional under the decoded contract.


Keep cache freshness and artifact identity as separate questions

A successful byte-identity check says nothing about authorization or freshness. If the publisher was malicious or stale, the bytes could still match some hash. Cache validation (ETags, If-None-Match, 304 responses) is an orthogonal concern. For cache revalidation logic, see Refonte's HTTP conditional-request validation article. Here we assume the manifest (hash) we compare against is already trusted or securely obtained. If no manifest is available or if the contract’s byte layer is undefined, we must HOLD CONTRACT and not accept the download at all.

Reject a truncated response before downstream use

The /truncated path advertises a 52-byte gzip body but closes after sending 26 bytes (half of it). When iterating, requests eventually raises an exception because the end of stream arrives unexpectedly. In the example baseline, this is a ChunkedEncodingError (despite no Transfer-Encoding: chunked header). In any case, the client handles it as a RequestException. We catch it and classify the download as incomplete. The candidate artifact is not accepted or promoted.

Incomplete downloads must not be handed off to consumers. The partial bytes should be discarded. The acceptance policy is: only if all requested bytes were consumed and all hash checks passed do we accept. The underlying library exception is just a symptom; the design requirement is that an incomplete state leads to HOLD INCOMPLETE. (No hidden retries or assumptions about timeouts are made here.)

Capture an exception without promoting a partial download

In practice, the application should record this failure event and ensure the partially downloaded data is quarantined (or simply deleted). For example, one might store error details in a log or error report with context (which URL, headers, what bytes were obtained, etc.), but not copy those bytes into the "approved" directory. The operational policy is: if an error occurs, do not upgrade the old artifact with the new one. The partially downloaded blob is not allowed to become the “file of record.” Any next steps (retry logic, alerts) are higher-level design concerns outside this scope. Here we simply mark the candidate as HOLD INCOMPLETE.

Build a contract-aware length and digest gate

We can encapsulate the above rules in a simple validation function. In the function below, read_mode is "iter" or "raw", and an absent Content-Encoding is normalized to "identity" for policy comparison. Declare allowed_codings before the request and check the received header before consuming the body. The final gate repeats that check before acceptance:

import hashlib


def validate_artifact(
    received_bytes, completion, received_coding,
    expected_len, expected_hash, contract_layer,
    read_mode, allowed_codings,
):
    # Normalize absent coding only for this client's internal policy.
    coding = (received_coding or "identity").strip().lower()
    if coding not in allowed_codings:
        return "HOLD_CONTRACT"
    if completion != "complete":
        return "HOLD_INCOMPLETE"

    valid_modes = {
        ("decoded", "identity"): ("iter", "raw"),
        ("decoded", "gzip"): ("iter",),
        ("encoded", "identity"): ("iter", "raw"),
        ("encoded", "gzip"): ("raw",),
    }
    if read_mode not in valid_modes.get((contract_layer, coding), ()):
        return "HOLD_CONTRACT"

    actual_hash = hashlib.sha256(received_bytes).hexdigest()
    if len(received_bytes) != expected_len or actual_hash != expected_hash:
        return "HOLD_INTEGRITY"
    if contract_layer == "decoded":
        return "ACCEPT_DECODED_ARTIFACT"
    return "ACCEPT_ENCODED_ARTIFACT"

First verify that the received Content-Encoding is allowed by the contract and that the download completed without errors. Then confirm that the read mode yields the intended artifact layer. Decoded CSV may come from an uncoded response or decoded gzip; encoded gzip may come from a raw gzip-coded response or an application/gzip response without Content-Encoding. Use allowed_codings=("identity", "gzip") when both received forms are allowed, ("gzip",) for the raw coded-body policy, or ("identity",) for a strict identity-only policy. Compare the retained length and hash with the independent manifest, never with an expected hash derived from the received bytes. If any check fails, hold the candidate for integrity, contract, or completeness review.

Note that using Content-Length is a shortcut only for the encoded check in this fixture: we verify len(received_bytes) against the manifest’s length. We never blindly equate Content-Length to the decoded size. Instead, we treat Content-Length as metadata about the encoded body. If an oracle (manifest) is missing, we conservatively HOLD. We do not attempt to guess or transform the data beyond one layer; we only remove coding if explicitly allowed by the contract.

Recover by rebuilding the candidate under one explicit contract

If our initial choice was wrong (e.g. we read with raw but needed decoded bytes, or vice versa), we should retry with the correct mode. For example, if we did iter_content but meant to get the gzip blob, we must close the response and issue a new requests.get expecting the raw body. We cannot “undelete” the first bytes we read. Conversely, if we used raw but needed decoded, we close and reissue with default decoding.

If a previous artifact was already accepted and in use, we keep it until the new one passes. We do not overwrite it with a failing candidate. A contract change (from encoded to decoded or vice versa) effectively means the manifest must change and downstream consumers must agree on the new definition. We don’t change the publisher’s checksum to mask an unexpected result; instead we fetch again under the new assumption.

Verify the candidate before replacing the approved artifact

Before replacing an existing file, an acceptance record should be logged. For instance, you might record: contract_version, the API endpoint or URL, timestamp, the actual Content-Encoding received, the client library versions (to catch interoperability issues), the byte count and SHA-256 of the received data, and the PASS/FAIL decision. Only once all checks (as above) pass do you consider a replacement.

This lab does not implement the filesystem swap, but in a real system you could stage the new bytes in a temporary location, run a final verification (maybe even on a different machine), and then atomically rename it over the old file. Concurrency controls and error recovery would then be an application concern. Here we only guarantee that “approved artifact” is replaced only after explicit re-check under a fresh, correct contract.

Assign the decision and ownership

We arrive at the final policy decisions. The table below summarizes which conditions lead to which outcome and who is responsible:

Condition

Decision

Owner (role)

Complete, allowed coding and read mode, and expected length and hash match

ACCEPT DECODED/ENCODED ARTIFACT

API integration engineer (client implementation)

Content-Encoding not allowed by contract (unexpected)

HOLD CONTRACT

API integration / design (review contract)

Complete but length or hash mismatch at the intended layer

HOLD INTEGRITY

API integration engineer (data validation)

Download error or incomplete (e.g. truncated)

HOLD INCOMPLETE

Operations / client (retry logic)

•        ACCEPT DECODED ARTIFACT: the complete retained CSV matches the decoded length and hash under the allowed received coding and read mode. Proceed to use the CSV.

•        ACCEPT ENCODED ARTIFACT: the complete retained gzip bytes match the encoded length and hash under the allowed received coding and read mode. Use the gzip bytes as-is.

•        HOLD CONTRACT: the response coding wasn’t what our contract allowed (e.g. identity requested but got gzip). Do not consume or promote the body after detecting the mismatch. The contract itself needs review or enforcement.

•        HOLD INTEGRITY: the download completed, but the SHA-256 did not match. The data is not the expected artifact, so we refuse it (and may log the event).

•        HOLD INCOMPLETE: an exception occurred (timeout, truncated, etc.). No artifact is accepted.

Responsibilities: The API consumer (integration engineer) implements the read and validation logic. The data steward or publisher provides the correct length/hash manifest. Operations should handle exceptions, logging and safe cleanup. Failed candidates should be logged to the monitoring/observability system rather than silently ignored. The API security and observability practices suggest logging such errors and setting alerts, ensuring that an invalid or incomplete artifact doesn’t slip through.

Strengthen API testing foundations with Refonte Learning

This detailed download verification exercise ties back to fundamentals: clear contracts, rigorous error handling, and testing. Refonte Learning’s APIs Developer Fundamentals program (3 months, ~10–12 hours/week) covers these topics. It teaches REST and GraphQL design, database integration, authentication, testing, error handling and logging, versioning, and API security. Students work on projects using tools like Postman and Swagger under expert guidance. The program provides a Training Certificate (and a Certificate of Internship for those who qualify).

Always define which bytes you intend to accept and verify their length and checksum at that same layer. Don’t be fooled by a seemingly successful HTTP response; confirm explicitly whether you accepted the decoded data or the raw archive, and match against the corresponding manifest. That practice, combined with systematic testing and logging, is exactly the kind of discipline reinforced in the APIs Developer Fundamentals curriculum.