In this example, we have two small finite feeds. Each feed should contain exactly three rows with IDs A, B, C in order. We use zip(left, right, strict=True) to iterate them together, expecting three paired rows (("A",…),("A",…)), (("B",…),("B",…)), (("C",…),("C",…)). In one scenario, the consumer breaks after two pairs, so no exception is raised and no “ValueError” is seen, even though a third row exists unmatched. This raises the question: did strict mode validate both inputs were fully consumed? Strict mode raises ValueError only when one feed actually runs out of rows before the other. Here, the consumer stopped early, so zip never reached the mismatch point.
We define our complete-feed contract as follows: each feed must be finite and contain exactly A, B, C in that order, and the corresponding IDs in left and right must match positionally. In other words, zip must produce exactly three output pairs, with pair[0][0] = pair[1][0] at each position, and the ID sequence must equal ("A","B","C"). Meeting this contract requires three distinct checks: length equality, meaning both feeds exhaust together; positional alignment, meaning each emitted pair has the same ID on both sides; and domain completeness, meaning all expected IDs appear in order. A consumer that stops early examines only a prefix, so it must signal incomplete rather than accept the data. ACCEPT is reserved for a fully consumed, aligned feed with the declared domain.
1. The Paired-Feed Contract
We treat the left and right inputs as parallel data streams of keyed rows. Our independent expected domain is the sequence ("A","B","C"). We require the left-side IDs to form exactly this sequence, the right-side IDs to do the same, and each left ID to match the corresponding right ID. Readers comparing SQL and Python for data engineering can think of a one-to-one keyed pairing, but this validator does not perform a join: every key A, B, C must appear once on each side in the declared order. This separates concerns:
Length equality: Zip must exhaust both feeds simultaneously (no extra rows on either side).
Positional identity: In each output pair (a,b), a[0] == b[0] (the same ID appears on left and right).
Domain completeness: The sequence of left IDs (and right IDs) must equal ("A","B","C").
If any of these fail, the contract is not met. For example, if the consumer stops early or one feed is shorter, we cannot ACCEPT; we mark it HOLD_UNCONSUMED or REJECT_LENGTH instead. If all pairs align but the IDs are not exactly A,B,C (e.g. an unexpected ID, or missing IDs), we mark REJECT_MANIFEST. Thus a prefix consumer does not automatically satisfy the contract; it only verifies those parts it saw, not the tail.
2. Construction vs. Iteration
Simply creating a zip(left, right, strict=True) object does not consume any rows. The Python 3.12 zip documentation defines zip as lazy: elements are processed only when the iterable is advanced. The Tracked wrapper records every next() call by appending a row ID or "EOF" on exhaustion. In this fixture, constructing zip(left, right, strict=True) leaves both event logs empty. Rows begin moving only when the consumer iterates or calls next(zip_obj).
Construction still calls iter(left) and iter(right). Tracked.__iter__ returns self, as required by the Python iterator protocol, so zip receives the same one-shot cursor rather than a clone. Once next() raises StopIteration, the iterator must continue to report exhaustion. Creating a second zip over the same advanced objects therefore does not create a new pass.
Therefore, constructing zip(left, right, strict=True) alone is not a validation step. No EOF events are logged until we actually iterate. In our script we explicitly test this: right after constructed = zip(left, right, strict=True) we assert left.events == right.events == []. We mark that state as HOLD_UNCONSUMED (nothing has been consumed or verified yet). Only when we iterate (or break early) do we check the contract.
Record Next Calls, Don’t Assume
Because zip is lazy, the fixture records what its Tracked cursors actually yield. Tracked intentionally exposes no len, so the script cannot establish length in advance. The consume() function records each source advancement and each EOF request. Those observations identify the validation boundary without inventing a length the cursors do not expose. A custom iterable may perform work in iter; the no-next-call construction observation is limited to the supplied Tracked implementation.
3. A Tracked Zip-Validation Fixture
We wrote a small driver script (zip_preflight.py) to automate these cases. It defines two fixed tuples:
This fixture validates in-memory iterator pairing only. Upstream HTTP transport completeness is a separate contract: a complete response body does not prove that two local feeds have equal lengths, aligned IDs, or the expected manifest.
LEFT = (("A", 10), ("B", 20), ("C", 30))
RIGHT = (("A",100), ("B",200), ("C",300))
EXPECTED_IDS = ("A","B","C")Each row is (ID, value). The Tracked class wraps a row tuple sequence and logs each next call by appending the row’s ID or "EOF" on StopIteration.
class Tracked:
def init(self, name, rows):
self.name = name
self.rows = tuple(rows)
self.pos = 0
self.events = []
def iter(self):
return self
def next(self):
if self.pos == len(self.rows):
self.events.append("EOF")
raise StopIteration
row = self.rows[self.pos]
self.pos += 1
self.events.append(row[0])
return rowThe consume(left, right, limit=None) function runs our zip loop:
pairs, error, complete = [], None, False
try:
for pair in zip(left, right, strict=True):
pairs.append(pair)
if limit is not None and len(pairs) == limit:
break
else:
complete = True
except ValueError as exc:
error = {"type": type(exc).__name__, "message": str(exc)}We append each emitted pair to pairs.
An optional limit lets us break early (to simulate a prefix consumer).
The for-else sets complete = True only if no break occurred.
On ValueError, we record the exception info.
The decision(result) function implements our contract:
if result["error"] is not None:
return "REJECT_LENGTH"
if not result["complete"]:
return "HOLD_UNCONSUMED"
pairs = result["pairs"]
if tuple(a[0] for a,b in pairs) != EXPECTED_IDS:
return "REJECT_MANIFEST"
if any(a[0] != b[0] for a,b in pairs):
return "REJECT_ALIGNMENT"
return "ACCEPT"In this controlled fixture, REJECT_LENGTH covers a captured strict-zip mismatch because Tracked never raises ValueError itself. HOLD_UNCONSUMED means the consumer stopped early without normal completion. REJECT_MANIFEST means the left ID sequence does not match EXPECTED_IDS, REJECT_ALIGNMENT means a left/right ID pair differs, and ACCEPT means the run completed normally and all declared checks passed.
We run several test cases with fresh Tracked sources each time:
equal_complete: left=ABC, right=ABC, no limit. (Three pairs expected.)
short_right_prefix: left=ABC, right=AB, limit=2. (Break after two pairs.)
short_right_full: left=ABC, right=AB, no limit. (Run to error.)
short_left_full: left=AB, right=ABC, no limit. (Error other side.)
equal_prefix: left=ABC, right=ABC, limit=1. (One pair then break.)
reordered_right: left=ABC, right=(B,A,C), no limit. (Lengths match, but IDs misaligned.)
both_empty: left=(), right=(), no limit. (Empty feeds.)
unexpected_id: left=A,B,X; right=A,B,X, no limit. (Wrong ID.)
construction_only: test zip() constructor without iterating.
reused_exhausted: consume one mismatch, then attempt a second zip on the same iterators.
The commissioning run on October 8, 2026 used CPython 3.12.14 on Linux 6.18.44 x86_64 with glibc 2.39 and produced a JSON report. The versioned Python 3.12 documentation retrieved for the research identified 3.12.15, so the documented family and the installed runtime are recorded separately. Ten synthetic cases passed their assertions. Key results include:equal_complete: Pairs = [(("A",10),("A",100)), (("B",20),("B",200)), (("C",30),("C",300))], complete=True, no error ⇒ ACCEPT.
short_right_prefix: Pairs = first 2 rows, complete=False, no error (events saw "A","B" on each side) ⇒ HOLD_UNCONSUMED.
short_right_full: Pairs = first 2 rows, then ValueError. Events: Left saw A,B,C; Right saw A,B,EOF. ⇒ REJECT_LENGTH.
short_left_full: Pairs = first 2 rows, then ValueError. Events: Left saw A,B,EOF; Right saw A,B,C. ⇒ REJECT_LENGTH.
equal_prefix: Pairs = first 1 row, complete=False, no error ⇒ HOLD_UNCONSUMED.
reordered_right: Pairs = 3 rows, complete=True, no error (left IDs "A","B","C" match EXPECTED_IDS, but second pair is misaligned) ⇒ REJECT_ALIGNMENT.
both_empty: Pairs = (), complete=True (nothing to iterate), no error ⇒ REJECT_MANIFEST (empty is not our expected domain).
unexpected_id: Pairs = 3 rows, complete=True, no error. Left IDs are ("A","B","X") ≠ EXPECTED_IDS ⇒ REJECT_MANIFEST.
construction_only: No iteration, complete=False ⇒ HOLD_UNCONSUMED.
reused_exhausted: After a prior rejection, a second consume on the same iterators gives pairs=(), complete=True, no error. The manifest still fails ⇒ REJECT_MANIFEST.
All assertions in the script passed. The final summary was:
{"status": "PASS", "python": "3.12.14", "cases": 10, "decisions": {
"equal_complete": "ACCEPT",
"short_right_prefix": "HOLD_UNCONSUMED",
"short_right_full": "REJECT_LENGTH",
"short_left_full": "REJECT_LENGTH",
"equal_prefix": "HOLD_UNCONSUMED",
"reordered_right": "REJECT_ALIGNMENT",
"both_empty": "REJECT_MANIFEST",
"unexpected_id": "REJECT_MANIFEST",
"construction_only": "HOLD_UNCONSUMED",
"reused_exhausted": "REJECT_MANIFEST"
}}This matches our expected outcomes.
The controlled traces match the Python zip documentation and PEP 618: strict mode raises ValueError when unequal exhaustion is detected, and the check occurs at the point where ordinary iteration would otherwise stop. In short_right_full, zip fetched the extra "C" from the longer feed before the right source reported EOF. That unmatched fetch is consistent with PEP 618’s reference algorithm. A prefix break never reaches this boundary.
4. A Correct Prefix is Not Full Validation
Consider the case where zip encounters no mismatch but we break early. For example, left has A,B,C and right has A,B (so right is shorter), but we break after consuming two pairs. Both emitted IDs match and neither iterator has raised StopIteration (no "EOF" was logged). However, no exception occurs at all, because strict zip only checks lengths at the end of iteration. Our run logged:
Left events: ["A","B"]
Right events: ["A","B"]
complete=False (we broke early)
error=None
decision=HOLD_UNCONSUMED.
Similarly, if both sides were equal-length but we still break after a prefix, we get the same outcome (correct prefix but incomplete). In both cases, even though the prefix data aligns, the contract cannot be accepted; it only shows a valid prefix, not the tail. Thus a correct prefix does not imply ACCEPT. We must see full exhaustion before accepting.
The driver uses Python’s for-else semantics: complete = True is set only when the loop reaches natural exhaustion. A break skips the else suite, so complete remains False. This keeps exception, early stop, and normal completion distinct; the absence of an exception is not a completion test.
5. Unequal Lengths in Both Orders
Now let’s let zip run until one side finishes naturally, in both possible ways.
LEFT longer (RIGHT is short): Left = ("A","B","C"), Right = ("A","B"), no break. Zip yields two pairs (A,A),(B,B), then on the next iteration it fetches ("C",30) from LEFT and tries next(RIGHT). RIGHT has no more items, so Tracked logs "EOF" on Right and StopIteration is raised. We catch it, and strict zip raises ValueError. In our log: left events = ["A","B","C"], right = ["A","B","EOF"]. The decision is REJECT_LENGTH.
RIGHT longer (LEFT is short): Left = ("A","B"), Right = ("A","B","C"). Symmetrically, zip yields (A,A),(B,B), then Left runs out. We see left events = ["A","B","EOF"], right = ["A","B","C"], and we get REJECT_LENGTH.
In both cases exactly two pairs were emitted before ValueError. The mismatch check consumed the unpaired "C" from the longer side, but that row never became an output pair. The observation matches PEP 618’s reference implementation, which probes beyond the common prefix to determine that one input still has a value when the other has stopped.
These are successful negative tests, not accepted data. The common prefix is retained as evidence, but the decision remains REJECT_LENGTH. The strict-zip example in the Python documentation shows the same argument-order information in its ValueError. In this controlled fixture the message corroborates the source trace; production code should not parse exception text to prove where an arbitrary ValueError originated.
6. Check Positional Alignment Last
What about equal-length feeds where no length error happens, but the IDs misalign? Example: Left = ("A","B","C"), Right = ("B","A","C"). Strict zip finds no early end: it produces three pairs and completes normally. In logs: left events = ["A","B","C"], right = ["B","A","C"], and no ValueError. However, the second pair is ("B",20) vs ("A",200), a misalignment.
The decision function then checks IDs. The left sequence ("A","B","C") matches EXPECTED_IDS, but the first two right-side IDs are reversed, so the result is REJECT_ALIGNMENT. zip(strict=True) is not a keyed join; it pairs by position. The Python zip documentation explains that strict mode produces the same pairs as regular zip and adds unequal-length detection, not reordering or value-based matching. The related Refonte Learning article on NumPy pairwise-loss validation addresses a different known-length numerical contract; this gate relies on explicit ID equality and an independent manifest.
The key point: strict zip’s own check only covers length. Any ID mismatch is the caller’s responsibility. In our pipeline logic, detecting misalignment leads to REJECT_ALIGNMENT. We preserve the original order; we do not try to sort or realign the data to make it pass.
7. Do Not Reuse Exhausted Iterators
After an error or exhaustion, a new zip object over the same iterator objects is not a second attempt. Under the Python iterator protocol, iter(iterator) returns that same cursor, and a cursor that has reported StopIteration remains exhausted. The deliberate reuse case makes this visible:
construction_only: We created zip(left,right,strict=True) but never iterated it. Both left and right show no events, and we decide HOLD_UNCONSUMED (nothing consumed). This matches intuition: we haven’t checked anything yet.
reused_exhausted: After the short_right_full case above, both Tracked objects are already at EOF ("C","EOF" and "A","B","EOF" as their events). If we now call consume(left,right) again, zip sees immediately that the first iterator is done, so it emits no pairs and sets complete=True (normal exhaustion of a one-time iterator). The result has pairs=(), complete=True, no error. Yet the domain is still not satisfied (no IDs), so we output REJECT_MANIFEST. We explicitly test that the second attempt’s “pairs” is empty and complete=True, consistent with the iterator having nothing left.
These outcomes follow from the Python iterator protocol: iter(iterator) returns the same object, and once next() has raised StopIteration it must continue doing so. In the specific ABC/AB case, the first mismatch probe consumed C and left both cursors at the end. A longer unmatched suffix could leave later rows unconsumed, but a new zip still would not rewind either source or create an independent pass. A genuine revalidation must reopen from retained immutable inputs and use fresh cursors.
8. Distinguish Source Errors from Length Errors
In this laboratory, Tracked raises only StopIteration, so a captured ValueError can be attributed to the strict-zip unequal-length rule. An arbitrary producer may itself raise ValueError, for example while parsing a row. Preserve the original exception context and choose HOLD_PRODUCER_FAILURE unless provenance establishes a strict-zip mismatch. Do not use an exception-message parser to manufacture that attribution.
9. Use the Validated Pairs, Don’t Re-zip
Once a run passes validation, downstream code should receive the completed result, not another zip object. For these small finite inputs, materialize the emitted pairs as a tuple, require normal completion, then apply the positional-ID and manifest checks to that same tuple. The following interface sketch is proposed; it was not part of the executed commissioning driver:
# Proposed production interface; not executed by the commissioning fixture.
def validate_pairs(left, right, expected_ids=EXPECTED_IDS):
pairs = tuple(zip(left, right, strict=True)) # must finish normally
left_ids = tuple(a[0] for a, in pairs)
if leftids != expected_ids:
raise ValidationError("Manifest mismatch")
if any(a[0] != b[0] for a, b in pairs):
raise ValidationError("ID alignment mismatch")
return pairsThe returned tuple is the result associated with ACCEPT. Downstream code must receive that exact tuple, not a fresh zip over spent or changed sources. At a higher boundary, classify ValueError only when producer provenance is controlled; otherwise preserve the failure for HOLD_PRODUCER_FAILURE. The immutability claim here is limited to a tuple containing immutable strings and integers, not arbitrary nested mutable record graphs.
Treat the accepted tuple and its evidence as one result. Consume the inputs once, retain the pairs, and do not ask downstream code to revalidate the same one-shot cursors.
10. Edge Cases: Empty and Missing IDs
Finally, consider the “trivial” cases:
Both empty: If left=(), right=(), zip(strict=True) immediately terminates (no pairs, complete=True). No exception is raised (empty iterators end together). But our domain is ("A","B","C"), not (). So we flag REJECT_MANIFEST. An empty feed would only be acceptable if our contract explicitly allowed an empty result, which it does not here. This highlights the same-omission risk: absence of data could hide missing expected IDs, so we rely on the independent manifest check.
Unexpected ID: If both sides contain ("A","B","X"), zip happily emits three pairs and completes, no length error. But the left ID sequence is ("A","B","X") which ≠ ("A","B","C"). We must reject this too: REJECT_MANIFEST. (Note alignment is fine here, but the values don’t match our expected set.)
In both cases zip did its job (no length mismatch), but the final data violates our declared domain. That’s a logical failure, not a technical one, so we use REJECT_MANIFEST. If truly empty feeds were valid in some context, we would need a different policy or contract; our fixed contract doesn’t allow it.
Conclusion
zip(strict=True) provides a length check during iteration, but only a consumer that reaches the end-of-input boundary exercises the whole-feed contract. An early break remains HOLD_UNCONSUMED even when every emitted pair is correct. In the declared two-input, one-extra-row fixture, mismatch detection consumed the unmatched C. PEP 618 does not imply draining an arbitrary remaining suffix, and neither strict zip nor a later zip object rewinds a source.
For downstream code, consume the finite inputs once, retain the completed tuple of pairs, and pass that same accepted object onward. Do not re-zip or revalidate the same cursors: Python iterators do not reset. The acceptance evidence establishes simultaneous completion, positional ID agreement, and the declared A, B, C manifest; it does not establish arbitrary business-value accuracy or external provenance.
A different application may define a legitimate prefix-only contract. Under this article’s complete-feed policy, however, an interrupted run remains HOLD_UNCONSUMED and its partial pairs stay withheld from downstream use. A later comparison must start from fresh, independent cursors over retained inputs.
For wider context, Refonte Learning’s articles on reliable data pipelines and cloud-native data engineering place local validation gates within broader pipeline ownership. The Data Engineering program runs for three months at 12–14 hours per week and covers data warehousing and ETL, pipeline design, Hadoop, Spark, streaming and batch ingestion, and governance. The program page describes projects, educational mentorship, and internship opportunities, with a Training Certificate and a Certificate of Internship after successful completion. Review the Data Engineering program for current admission details and curriculum information.
