The same two overlapping polygons with different class codes (10 and 20) can produce different raster labels when burned in opposite order. In this fixed 4×5 grid example, sixteen cells are covered by the union of the two rectangles, yet two cells can end up with unexpected labels after rasterization. By default, Rasterio’s rasterize follows a painter’s algorithm: later geometries in the input list overwrite earlier ones. Thus swapping the input sequence changes which feature gets priority in the overlap, even though the shapes and codes themselves never changed. This isn’t a bug in mathematics but a lack of an explicit overlap rule: our exclusive-label problem requires a policy decision beyond the raw algorithm.
In a real pipeline, the domain expert (e.g. a land cover specialist) must decide which class “wins” in overlaps. Here we adopt a sample policy, higher priority wins, that makes class 10 outrank class 20. This policy is not derived from the code values themselves but from an agreed hierarchy.
The key question is: Do the overlapping cells follow this policy? We will freeze all other factors (grid, coordinates, CRS label, affine transform, class definitions) and test both input orders. Using an independent per-cell oracle, we compare raw raster outputs against the approved outcome. We validate whether each pixel’s final label matches domain policy, or if the output must be rejected or rebuilt. The grids below are independently defined expectations; an observed Rasterio result requires running the library engine and inspecting its report.
1. Define the label decision before testing the rasterizer
We assume that exactly one class label must occupy each overlapping cell (exclusive labels). Our two axis-aligned rectangles both have positive-area overlap, so certain grid cell centers lie inside both geometries. That means we need a rule to pick one label per overlapped cell. The raw class codes (10 and 20) are identifiers, not numeric ranks. In our example, we explicitly set class 10 > class 20 by policy (the domain owner approves that “higher priority wins”). This is an arbitrary convention for illustration, not an inherent GIS rule. A different project might choose the opposite or use a tie-breaker by ID, but we do not assume any hidden rule. By default, Rasterio has no built-in concept of priority between classes; it simply uses input sequence. Thus our job is to fix a clear policy before evaluating the raster output.
The fixed inputs are: two polygons (A and B) with codes 10 and 20 and assigned priorities 2 and 1, respectively; a constant grid extent and transform; and no coordinate or CRS conversion issues (the coordinates are already in EPSG:3857 and the affine is given). We will record a precise source manifest listing each feature’s ID, geometry and code. Then we plan two raw tests (order A→B and B→A) to show how polygon order alone changes the result. Finally, we enforce our policy by sorting features by priority before rasterization. Only after running all cases do we decide: if the output labels violate the approved policy (two cells are wrong in one raw order), we reject or rebuild; if a raw result just happens to match policy (as in one order here), we note it as a diagnostic check, not as proof that any order is fine.
This approach ensures that we are testing the question: “Does the final label match the authorized policy?” rather than “Did we just happen to feed the shapes in the right order?”. We rely on a literal cell-by-cell oracle of membership, not on any assumed default hierarchy. With this exclusive-label contract in place, we can proceed to inspect how Rasterio’s overwrite mechanism works.
2. Read the overwrite contract at the correct scope
Rasterio’s documentation describes its rasterization as a painter’s algorithm: geometries are burned in sequence and each new feature can overwrite previous pixel values. In other words, by default merge_alg=MergeAlg.replace means “last value wins.” (There is also an ADD mode, but that sums values instead of overwriting, which serves a different semantic purpose.) We explicitly use all_touched=False so that only pixels whose center lies inside a polygon are burned (the default behavior). Note that GDAL’s CLI warns that enabling all-touched can make results order-dependent at polygon edges; in our test we avoid that edge ambiguity entirely by using integer bounds and half-integer centers.
In code terms, the rasterio.features.rasterize call accepts an iterable of (geometry, value) pairs, an output shape, fill value, data type, transform, all_touched, merge_alg, and skip_invalid. The signature (for Rasterio 1.4.4) shows all_touched=False and merge_alg=MergeAlg.replace by default. We also set skip_invalid=False so that any invalid geometry would raise an error (though our simple rectangles are valid). The critical point is: Rasterio itself makes no decision about class priority. It only follows the given list order. Therefore, we must not rely on input order as policy. Later, we will sort the features by the approved priority and pass that to rasterize, ensuring that the final output does obey the desired rule.
Geometry coverage and class precedence are separate decisions
To clarify, the fact that a grid cell lies under both polygons (geometric coverage) is one issue, and selecting its single output label (class precedence) is another. For our two rectangles, the cells at (row 1, column 2) and (row 2, column 2) are in the positive-area intersection. Under an exclusive-label contract, we still must assign one class to each. By default MergeAlg.replace will simply paint whatever value is encountered last. In a different setting one could use MergeAlg.add to accumulate numeric values, but that yields 30 (10+20) in the overlap, a “heatmap” result rather than an exclusive label. We mention it later for contrast (Section 8), but here it’s clear we want a single class, not a sum. Thus, we separate the questions: we first note which cells are covered by A or B (or both), and only then apply the priority rule to decide which class actually gets burned.
3. Freeze the grid and publish the independent source manifest
We fix a 4×5 uint8 grid (fill 0) with an Affine transform Affine(1, 0, 0, 0, -1, 4). This means the pixel centers have integer+0.5 coordinates from (0.5,3.5) in the upper-left to (4.5,0.5) in the lower-right. Our two polygons are given by the canonical manifest below (in EPSG:3857 for convention). Rasterio’s rasterize uses the transform, not the CRS, so we do not perform any coordinate conversion here. Importantly, no cell center falls exactly on a boundary, so all coverages are unambiguous.
The canonical feature definitions (ID, code, priority, bounds, geometry) are:
CANONICAL = (
{"id": "A", "code": 10, "priority": 2,
"bounds": (0, 1, 3, 4),
"geometry": {"type": "Polygon",
"coordinates": [[(0, 1), (3, 1), (3, 4), (0, 4), (0, 1)]]}},
{"id": "B", "code": 20, "priority": 1,
"bounds": (2, 0, 5, 3),
"geometry": {"type": "Polygon",
"coordinates": [[(2, 0), (5, 0), (5, 3), (2, 3), (2, 0)]]}}
)From these, we compute membership by geometry bounds (not by calling any library):A_CELLS = ((0,0), (0,1), (0,2),
(1,0), (1,1), (1,2),
(2,0), (2,1), (2,2))
B_CELLS = ((1,2), (1,3), (1,4),
(2,2), (2,3), (2,4),
(3,2), (3,3), (3,4))
OVERLAP = ((1,2), (2,2))These lists show each (row, col) of the covered cells. A covers 9 cells, B covers 9 cells, their intersection is exactly the two cells in OVERLAP, and the union is 16 cells total. We also publish the approved label grid (per domain policy), and the two candidate outcomes:APPROVED = ((10,10,10, 0, 0),
(10,10,10,20,20),
(10,10,10,20,20),
( 0, 0,20,20,20))
RAW_AB = ((10,10,10, 0, 0),
(10,10,20,20,20),
(10,10,20,20,20),
( 0, 0,20,20,20))
ADDITIVE = ((10,10,10, 0, 0),
(10,10,30,20,20),
(10,10,30,20,20),
( 0, 0,20,20,20))APPROVED is our target result: code 10 wherever A covers but B does not, code 20 wherever B covers but A does not, and code 10 in the overlap (since A has higher priority).
RAW_AB is what we expect if we rasterize [A,B] in that order (A first, then B overwrites). Note it has 20 in the overlap cells at (1, 2) and (2, 2).
ADDITIVE shows what would happen if we used MergeAlg.add: the overlap cells become 30 (10+20).
We keep the source manifest (IDs, geometries, codes, priorities) separate from these computed grids, ensuring the expectations are not derived by running Rasterio but by our independent calculations. This completes our literal oracle of membership and the approved output.
Author the expected cell membership before generating candidates
Before running any rasterization, we author the literal cell grid and membership lists. These become our truth tables. For example, a Python pre-flight could compute membership(bounds) and verify it equals A_CELLS and B_CELLS. It can then verify the overlap set and union count. These checks ensure our fixture truly represents the stated geometry contract. We record this manifest (CANONICAL) alongside the approved grid (APPROVED). Only after these are fixed do we let the driver code generate candidate rasters to compare.
4. Run one driver with explicitly different evidence modes
Below is the complete test driver, rasterize_priority_fixture.py, which executes both the arithmetic checks and (optionally) the actual Rasterio calls. The driver is written to never call Rasterio when using the “arithmetic” engine. It prints a unique JSON report path and sets status flags depending on whether each case meets the literal expectations.
"""Two explicit engines: arithmetic evidence is never Rasterio execution."""
import argparse
import copy
import hashlib
import json
from pathlib import Path
import platform
import tempfile
def require(condition, message):
if not condition:
raise RuntimeError(message)
A_CELLS = ((0, 0), (0, 1), (0, 2), (1, 0), (1, 1), (1, 2),
(2, 0), (2, 1), (2, 2))
B_CELLS = ((1, 2), (1, 3), (1, 4), (2, 2), (2, 3), (2, 4),
(3, 2), (3, 3), (3, 4))
OVERLAP = ((1, 2), (2, 2))
APPROVED = ((10, 10, 10, 0, 0),
(10, 10, 10, 20, 20),
(10, 10, 10, 20, 20),
(0, 0, 20, 20, 20))
RAW_AB = ((10, 10, 10, 0, 0),
(10, 10, 20, 20, 20),
(10, 10, 20, 20, 20),
(0, 0, 20, 20, 20))
ADDITIVE = ((10, 10, 10, 0, 0),
(10, 10, 30, 20, 20),
(10, 10, 30, 20, 20),
(0, 0, 20, 20, 20))
CANONICAL = (
{"id": "A", "code": 10, "priority": 2, "bounds": (0, 1, 3, 4),
"geometry": {"type": "Polygon", "coordinates": [[
(0, 1), (3, 1), (3, 4), (0, 4), (0, 1)]]}},
{"id": "B", "code": 20, "priority": 1, "bounds": (2, 0, 5, 3),
"geometry": {"type": "Polygon", "coordinates": [[
(2, 0), (5, 0), (5, 3), (2, 3), (2, 0)]]}},
)
CASE_NAMES = {"raw_AB", "raw_BA", "add", "sorted_AB", "sorted_BA",
"missing_priority", "tied_priority", "duplicate_id", "missing_feature"}
def membership(bounds):
xmin, ymin, xmax, ymax = bounds
return tuple((r, c) for r in range(4) for c in range(5)
if xmin < c + 0.5 < xmax and ymin < 4 - (r + 0.5) < ymax)
def arithmetic_burn(features, additive):
grid = [[0] 5 for in range(4)]
for feature in features:
for (r, c) in membership(feature["bounds"]):
grid[r][c] = grid[r][c] + feature["code"] if additive else feature["code"]
return tuple(map(tuple, grid))
def prioritygate(features):
ids = [f.get("id") for f in features]
if len(ids) != 2 or set(ids) != {"A", "B"}:
return "HOLD_FEATURE_INVENTORY", ()
baseline = {f["id"]: f for f in CANONICAL}
for f in features:
if any(f.get(k) != baseline[f["id"]][k] for k in ("geometry", "bounds", "code")):
return "HOLD_FEATURE_INVENTORY", ()
ranks = [f.get("priority") for f in features]
if (any(type(p) is not int for p in ranks) or len(set(ranks)) != 2
or any(f["priority"] != baseline[f["id"]]["priority"] for f in features)):
return "HOLD_PRIORITY", ()
return "READY", tuple(sorted(features, key=lambda f: f["priority"]))
def differences(left, right):
return tuple((r, c) for r in range(4) for c in range(5)
if left[r][c] != right[r][c])
def main():
parser = argparse.ArgumentParser()
parser.add_argument("--engine", choices=("arithmetic", "rasterio"), required=True)
engine = parser.parse_args().engine
folder = Path(tempfile.mkdtemp(prefix="rasterize-priority-"))
report = {"status": "FAIL", "engine": engine, "stage": "startup",
"python": platform.python_version(), "platform": platform.platform(),
"source_sha256": hashlib.sha256(Path(__file__).read_bytes()).hexdigest(),
"rasterio_calls_started": 0, "rasterio_calls_completed": 0,
"rasterio_execution": "NOT_EXECUTED", "cases": {},
"canonical_manifest": copy.deepcopy(CANONICAL),
"parameters": {"shape": [4, 5], "dtype": "uint8", "fill": 0,
"affine": [1, 0, 0, 0, -1, 4], "all_touched": False,
"skip_invalid": False, "declared_crs": "EPSG:3857"}}
try:
if engine == "rasterio":
report["stage"] = "import_and_version_gate"
import numpy
import rasterio
from rasterio.features import rasterize
from rasterio.enums import MergeAlg
from affine import Affine
report["versions"] = {"numpy": numpy.__version__, "rasterio": rasterio.__version__,
"gdal": rasterio.__gdal_version__}
require(rasterio.__version__ == "1.4.4", "This fixture targets Rasterio 1.4.4")
report["stage"] = "independent_membership"
require(membership(CANONICAL[0]["bounds"]) == A_CELLS, "A cell manifest mismatch")
require(membership(CANONICAL[1]["bounds"]) == B_CELLS, "B cell manifest mismatch")
require(tuple(sorted(set(A_CELLS) & set(B_CELLS))) == OVERLAP, "overlap mismatch")
require(len(set(A_CELLS) | set(B_CELLS)) == 16, "coverage mismatch")
report["literal_membership"] = {"A": A_CELLS, "B": B_CELLS, "overlap": OVERLAP}
report["literal_approved_grid"] = APPROVED
def draw(features, additive):
if engine == "arithmetic":
return arithmetic_burn(features, additive)
report["rasterio_execution"] = "ATTEMPTED"
report["rasterio_calls_started"] += 1
array = rasterize(
[(f["geometry"], f["code"]) for f in features],
out_shape=(4, 5), fill=0, dtype="uint8", transform=Affine(1, 0, 0, 0, -1, 4),
all_touched=False, skip_invalid=False,
merge_alg=MergeAlg.add if additive else MergeAlg.replace)
report["rasterio_calls_completed"] += 1
require(array.shape == (4, 5) and str(array.dtype) == "uint8", "array contract")
return tuple(tuple(int(v) for v in row) for row in array)
def run(name, features, expected, verdict, sort=False, additive=False):
report["stage"] = name
entry = {"input_manifest": copy.deepcopy(features), "status": "INCOMPLETE",
"merge_alg": "ADD" if additive else "REPLACE", "policy_verdict": verdict}
report["cases"][name] = entry
if sort:
state, ordered = priority_gate(features)
entry["gate"] = state
require(state == "READY", "approved source was blocked")
require([f["id"] for f in ordered] == ["B", "A"], "wrong burn order")
else:
ordered = features
entry["gate"] = "DIAGNOSTIC_BYPASS_ON_KNOWN_MANIFEST"
entry["burn_order"] = [f["id"] for f in ordered]
result = draw(ordered, additive)
entry["grid"] = result
entry["difference_from_approved"] = differences(result, APPROVED)
require(result == expected, name + " differed from literal expectation")
entry["status"] = "MATCHES_LITERAL_EXPECTATION"
return result
ab = copy.deepcopy(CANONICAL)
ba = tuple(reversed(copy.deepcopy(CANONICAL)))
first = run("raw_AB", ab, RAW_AB, "REJECT_PRIORITY")
second = run("raw_BA", ba, APPROVED, "MATCHES_BY_ORDER")
require(differences(first, second) == OVERLAP, "nonoverlap cell changed")
run("add", ab, ADDITIVE, "REJECT_CLASS_POLICY", additive=True)
run("sorted_AB", ab, APPROVED, "ACCEPT_FIXTURE_POLICY", sort=True)
run("sorted_BA", ba, APPROVED, "ACCEPT_FIXTURE_POLICY", sort=True)
rejected = {
"missing_priority": ({*ab[0], "priority": None}, ab[1]),
"tied_priority": ({**ab[0], "priority": 1}, ab[1]),
"duplicate_id": (ab[0], ab[0]), "missing_feature": (ab[0],),
}
for name, features in rejected.items():
report["stage"] = name
entry = {"input_manifest": copy.deepcopy(features), "status": "INCOMPLETE",
"rasterization_called": False}
report["cases"][name] = entry
before = report["rasterio_calls_started"]
state, ordered = priority_gate(features)
entry["policy_verdict"] = state
expected = "HOLD_PRIORITY" if "priority" in name else "HOLD_FEATURE_INVENTORY"
require(state == expected and not ordered, "invalid source accepted")
require(report["rasterio_calls_started"] == before, "HOLD called rasterize")
entry["status"] = "EXPECTED_HOLD"
require(set(report["cases"]) == CASE_NAMES, "case set incomplete")
if engine == "rasterio":
require(report["rasterio_calls_started"] == report["rasterio_calls_completed"] == 5,
"library call count incomplete")
report["rasterio_execution"] = "COMPLETED_FIVE_LIBRARY_CALLS"
report["status"] = "RASTERIO_FIXTURE_PASS"
else:
require(report["rasterio_calls_started"] == 0, "arithmetic mode called library")
report["status"] = "ARITHMETIC_PASS_RASTERIO_NOT_EXECUTED"
report["stage"] = "complete"
except Exception as exc:
report["error"] = {"type": type(exc).__name__, "message": str(exc)}
raise
finally:
output = folder / "results.json"
output.write_text(json.dumps(report, indent=2) + "\n", encoding="utf-8")
print(json.dumps({"status": report["status"], "engine": engine,
"report": str(output)}, indent=2))
if name == "__main__":
main()Arithmetic preflight is not a Rasterio result
Run this script first in “arithmetic” mode (for example python rasterize_priority_fixture.py --engine arithmetic). It will perform all the internal checks above without importing Rasterio, verifying membership, overlapping cells, and grid values purely by Python arithmetic. In this mode the report’s "status" should be "ARITHMETIC_PASS_RASTERIO_NOT_EXECUTED", and "rasterio_calls_started" will remain 0. It confirms our expected matrices and priority logic are mathematically consistent.
In a separate environment with Rasterio 1.4.4 installed, you would run --engine rasterio. The code then records actual library versions (rasterio.__version__ == 1.4.4) and calls rasterio.features.rasterize five times. It logs each output grid, differences vs APPROVED, and ensures exactly five calls complete. If all cases match expectation, it sets "status": "RASTERIO_FIXTURE_PASS" and "rasterio_execution": "COMPLETED_FIVE_LIBRARY_CALLS". Any deviation (wrong version, missing case, unexpected array shape) would raise an error, halting the test. The script writes engine, platform, parameters, burn order, and outputs to the detailed JSON report, then prints a summary containing only status, engine, and the unique report path. This comprehensive evidence distinguishes arithmetic logic from actual library behavior.
5. Reverse the raw order and inspect the two changed cells
Comparing the expected raw outputs of A→B vs B→A makes clear which cells flipped. Burning A first then B produces the RAW_AB grid with values of 20 at (1,2) and (2,2). Reversing the sequence (burn B then A, RAW_BA) yields the APPROVED grid (values of 10 at those cells). In other words, the two overlap cells are exactly where Rasterio’s overwriting changed the label from 20 to 10. No other cells differ. The test driver verifies that differences(first, second) == OVERLAP, meaning the only mismatches are the two overlapped positions. In our reporting schema, raw_AB receives "REJECT_PRIORITY" (because it violates the approved priority rule), while raw_BA is marked "MATCHES_BY_ORDER". However, this luck of order alone should not be mistaken for correct practice: raw_BA only matched our policy because its input order happened to align with the intended priority. The driver deliberately bypasses the priority gate for these raw cases (since the manifest is known) to isolate the effect of ordering. The takeaway is simply that exactly two cells differ (the overlap) when we swap the polygon order, confirming the central issue.
6. Identify the checks that can pass while priority is wrong
Simply checking that the output image has the right size, grid coordinates, or count of covered cells is not enough to catch the failure. Both RAW_AB and RAW_BA images cover 16 nonzero cells and share the same spatial extent. Their sums and means do differ in this fixture: RAW_AB sums to 250 with a mean of 12.5 across all 20 cells, while RAW_BA sums to 230 with a mean of 11.5. Those aggregates still do not identify which cells violate policy. Yet in RAW_AB the overlapping cells have class 20 instead of the policy’s class 10. We need a cell-by-cell check.
Consider how misleading a simple coverage check can be. Both raw outputs cover all the intended cells (16 of them) so an equal-coverage test would pass. Both have the same class codes overall (10 and 20) just arranged differently. Only by comparing each cell’s label against the approved value do we detect the two wrong assignments. This is why the driver records the entire grid and explicitly computes difference_from_approved. Even a duplicated feature or reshuffled attribute might leave the output array looking valid (same shape, same set of values) while still being wrong in detail. Thus the independent cell ledger, our oracle, is decisive. Any change or duplication in the source that doesn’t alter coverage or totals could slip through naïve checks, which is why we also gate by manifest inventory.
Equal coverage does not establish correct class assignment
Imagine two labels swapped or a duplicate feature inserted: you could still end up with 16 covered cells and the expected set of values, yet violate policy. The driver’s duplicate_id case specifically supplies (A, A), omitting B, and is held before rasterization. A different case, inserting an extra copy in (A, A, B), could leave the raw output unchanged while still violating the source manifest. We therefore do not accept an output just because dimensions, CRS, and covered-cell count are correct. Instead, we insist on verifying each overlap cell against the approved class (and ensuring no extra or missing cells).
7. Apply the approved priority before burning any cells
With the policy “higher priority wins”, our code then sorts the two features by their priority (lowest first, meaning B then A since A’s priority 2 outranks B’s 1). We run the driver with sort=True. In both sorted_AB (original A,B order) and sorted_BA (original B,A) cases, the actual burn order becomes B→A. Both outputs match the APPROVED grid exactly, and are flagged "ACCEPT_FIXTURE_POLICY". This demonstrates that applying the domain rule before rasterization yields a deterministic, correct result independent of the initial list order.
Crucially, we maintain the pairing of each feature’s ID, code and approved priority. We do not allow ourselves to arbitrarily change priorities or swap codes as a quick fix. Even if one were tempted to say “just switch the codes of the polygons”, that would violate the source contract. The priority gate strictly enforces the canonical priority values. If someone had put priority 1 on A and priority 2 on B (or missing it entirely), we would hold the case. In our setup, numeric labels (10, 20) do not equal numeric priority (indeed we set priority 2 for code 10), so sorting by code alone would actually pick the wrong “winner” in general. The correct approach is to rely only on the explicit priority field, as we have done.
Approve ties and missing ranks as policy decisions, not sorting accidents
In this fixture, missing or equal priorities cause a hold. The driver’s gate logic is clear: if priorities are None, equal, or non-integers, we return a "HOLD_PRIORITY" status before rasterizing. Why not just break ties by ID or arbitrary order? Because that would secretly embed an unapproved rule. A genuine policy must be defined by the data owner. Some users might accept a stable sort or a label rule, but here we treat ties as incomplete policy. This is intentional: if the domain decides to allow equal priority or a specific tie-break (for example, always let ID “A” win in a tie), that should be explicitly recorded as part of the policy, not silently assumed. As it stands, any tie or missing priority in our test manifest is a blocker (HOLD_PRIORITY), requiring the domain to clarify the rule.
8. Use ADD only when the product contract calls for accumulation
The merge_alg=MergeAlg.add mode is available (also via GDAL’s -add flag), but it represents a different use-case: numeric accumulation. In our example, using ADD gives the ADDITIVE grid with values of 30 in the overlap cells, which clearly violates the exclusive-label vocabulary. We reject this mode here ("REJECT_CLASS_POLICY" in the driver) because our class definitions are categorical, not counts. However, ADD mode is not inherently unsafe; it would be appropriate if the desired outcome was a “heatmap” or an accumulated count, as noted in Rasterio’s docs and GDAL’s guide. The key is contractual: unless accumulation was intended and approved (with its own expected result matrix), we do not use ADD. For exclusive classification, REPLACE is the correct semantic choice.
9. Hold incomplete or ambiguous source inventories before rasterization
The driver also tests that our input manifest exactly matches the canonical contract. Four negative cases are explored:
missing_priority or tied_priority: Here one feature’s priority is None or equal, triggering "HOLD_PRIORITY". We do not even attempt a burn.
duplicate_id or missing_feature: Here the list of features is wrong (e.g. one geometry appears twice, or B is absent). This yields "HOLD_FEATURE_INVENTORY". Again, no rasterization call is made.
Additionally, the code cross-checks that each feature’s geometry, bounding box and code match the canonical values. Any deviation would also result in "HOLD_FEATURE_INVENTORY". These checks are specific to our fixed two-rectangle fixture; in general they represent a narrow schema validation rather than a full geojson sanity check. We simply ensure that exactly two known features (IDs A and B) are present with the expected properties. If not, the output is held for the data steward to fix the inventory. The lesson is that source completeness and clarity are as important as the raster output itself.
10. Capture evidence that identifies the operation actually performed
Every run of the fixture must produce a detailed report. This JSON evidence includes: the engine used ("arithmetic" or "rasterio"), a SHA-256 of the source code, Python and platform versions, and (for library mode) the actual Rasterio, GDAL, and NumPy versions. It also records all rasterize parameters (shape, dtype, fill, transform, all_touched, skip_invalid, etc) and the exact input manifest. For each case, the report logs the burn order (ids of features in order), the resulting grid array, and the list of differences from the approved grid. The overall status remains "FAIL" until all expected steps complete. Any exception or missing data writes an "error" with type and message.
For a successful rasterio run, we explicitly require "rasterio_calls_started": 5 and "rasterio_calls_completed": 5; if these do not match, the report will indicate failure. The "status" transitions to "RASTERIO_FIXTURE_PASS" only when all checks (including the four holds) have passed. The folder path is randomized, ensuring each execution’s results are fresh. Importantly, passing a “negative control” hold case (e.g. missing ID) does not produce a releasable raster; it simply shows the system correctly held it. Only when the final status is RASTERIO_FIXTURE_PASS (with all five rasterization cases matching their literal expectations and all four holds recorded) can we accept the two sorted outputs under this fixture’s policy. The raw diagnostic and ADD outputs retain their individual rejection or diagnostic verdicts. Stale or old JSON files cannot be used as proxies; we inspect the new, printed report path to confirm the current run’s outcome.
11. Rebuild the derived labels without editing the source policy
Once the evidence clearly identifies the issue (input order not matching priority), we can fix the output by applying the approved rule. We do not alter the source policy (priority values); we only reorder the burn. Practically, this means rerunning the rasterize step with the sorted manifest (B then A). Doing so regenerates the class grid exactly as in APPROVED. We then re-run the full real-mode suite (engine=rasterio) to verify that now all cell values agree with expectations. It’s not enough to fix just one example; we validate the entire fixture again.
In general, any downstream products built from the rejected output (e.g. a classification map used by a model) would need the same treatment: identify that two pixels were wrong, and regenerate them by rerunning rasterization. We do not “patch” the array by hand, because that breaks reproducibility. Instead, we derive the corrected labels using the same geometry/grid contract. For audit, we report that the rebuilt grid differs from the bad raw output exactly at the overlap cells (the same OVERLAP list).
Notably, this approach parallels other EO validation practices. For example, verifying a delivered GeoTIFF’s validity mask requires comparing every valid-cell coordinate against an oracle, rather than just matching aggregate stats. Here we similarly revalidate each pixel. If the raster were part of a delivered dataset, file checksums and mask validation would also matter; see our GeoTIFF delivery article for mask examples. Those checks are outside this in-memory test. The key is: rebuild the labels algorithmically from the approved source, then treat the corrected raster as the new reference.
Regenerate from the approved manifest and recheck every cell
After enforcing the approved manifest order and completing the checks, the rebuilt raster can replace the rejected output. We load it into a fresh array and compare it to APPROVED. We confirm that the only differences from the previously rejected raw_AB image are exactly the two cells in OVERLAP. Everything else remains the same. This shows that the root cause was indeed ordering, and that no other policy change was needed. In practice, having this evidence allows one to “relabel” the two wrong cells in the official dataset: by rerunning the authorized rasterization, those cells now carry the correct class. Any derived artifacts (e.g. a valid-mask or report) should then be regenerated from this fixed raster.
At no point do we modify the original geometry or class definitions; we only enforce the approved source order. The result is an output that fully satisfies the policy contract. We can then close the loop by recording that the faulty raster was replaced (if it had been published or passed downstream) with the corrected version. This is analogous to how one would recalculate a validity mask from a reference mask rather than guessing values: use the approved source to regenerate any missing or wrong elements.
12. Turn the evidence into a bounded release decision
Based on the combined evidence, we make a clear go/no-go decision for each case. The following summary (Decision matrix) guides the release:
Condition | Outcome |
Incorrect labels in overlap cells | Reject/Rebuild: rerun using approved priority. |
Favorable match from unsorted input | Diagnostic only; do not release as-is. |
Sorted order case and full suite passed | Accept: fixture contract satisfied. |
Missing or tied priorities | Hold: domain must define policy. |
Duplicate IDs or missing features | Hold: fix source inventory. |
MergeAlg=ADD with overlap values | Reject: incompatible with exclusive labels. |
Arithmetic-only checks passed (no library) | Pending library run; cannot accept. |
Crucially, passing a negative control does not authorize its output. For example, the driver correctly holds a duplicate-ID input. There is no duplicate raster to release because the gate stops the burn; we record that the input was properly rejected. Only when an output follows the approved priority path, matches the approved grid, and is supported by a completed Rasterio suite do we mark it as acceptable. Outside the scope of this fixture, other factors (like file-level checksums or masks) would come into play, but within this test we keep the decision strictly about source, order and pixel labels.
13. Assign ownership and carry the contract into a real pipeline
In an operational workflow, assign responsibility for each aspect of this process. The domain owner or subject-matter expert (e.g. a land-use manager) approves the policy (which class has priority in overlaps). The data steward or GIS administrator maintains the source manifest (ensuring correct geometries, IDs, etc.). The GIS engineer implements the rasterization and runs these checks, generating the cell ledger and verifying outputs. Finally, a reviewer or manager makes the release decision based on the collected evidence. This separation of duties ensures no one person silently changes priorities or ignores missing data. (For context on these roles, see the career-path comparison; for example, GIS analysts focus on implementation while remote-sensing scientists may set domain rules.)
Remember that this fixture is bounded: it uses only two simple rectangles on a fixed grid, no reprojection, no geometry fixing, no boundary-point ambiguity, and no model training analysis. Real-world data might involve many polygons, nested features, or different spatial queries; each would require its own validation plan. If a query (e.g. SQL ordering) or data structure changed, the feature burn sequence could change, so our validation would have to adapt accordingly. But in every case, the principle holds: have an approved policy, record the manifest, and compare the rasterization result exactly to the approved outcome. Without an independent oracle, field measurements or known references would be needed, but that is beyond this example. Here we simply enforce that the declared source (IDs, polygons, codes, priority) yields the expected labels.
14. Connect the exercise to reproducible Earth-observation practice
This lab exemplifies the kind of rigorous preprocessing and validation that modern remote-sensing workflows demand. It illustrates a reproducible pipeline step: starting from an explicit source manifest, applying a documented rule (overlap priority), and verifying the output at the pixel level. These skills connect with the broader preprocessing and project work described in Refonte Learning’s Remote Sensing Scientist/Engineer Program. By practicing this exercise, engineers build confidence that they can trace each output label back to its input definitions and policy.
Refonte’s program (three months, approximately 10–12 hours per week) covers GDAL/Rasterio, image preprocessing, and validation techniques, culminating in a capstone project and code repository. Its audience includes learners with a STEM background who are comfortable with Python and numerical thinking; no prior EO experience is strictly required, but a working knowledge of matrices/signals is recommended. Admission requires working toward a bachelor’s or higher-level degree. Upon successful completion of the program, Refonte Learning offers a Training Certificate and a Certificate of Internship. This exercise of connecting approved source policy to each cell’s label provides practice relevant to that emphasis on reproducible projects.
This exercise shows how fixing the grid, bounding the input manifest, and sorting features by an approved priority rule produces a deterministic, validated raster. By maintaining evidence at each step, a team can confidently release only those rasters that certifiably obey domain policy, keeping data delivery both correct and reproducible.
