GeoPandas separates assigning a CRS from transforming coordinates. Simply giving a dataset a new CRS label (with set_crs) does not move its points; only to_crs actually changes coordinates. If a GeoDataFrame reports the desired CRS but its numbers are unchanged, a hidden metadata error may be lurking. We define a strict coordinate contract first: immutable point IDs, known raw longitude/latitude values, source CRS, target CRS, and expected output in metres. Then we test every path (no CRS, correct label, wrong label, round-trip, fix, rebuild or reject) against independent controls. A plausible-looking output alone is not proof; only exact point-level agreement with precomputed values (within tolerance) should be accepted.
By designing these reproducible checks, a geospatial engineer can accept a correctly transformed file, repair only its metadata if raw coordinates are unchanged, rebuild the transformation if values were garbled, or quarantine data when its provenance is unknown. This detailed workflow belongs in a preprocessing or QA pipeline, not a casual check of summary statistics. The following examples use synthetic 3-point data (longitudes 2°, -3°, 0° and latitudes 1°, 4°, 0°) transforming from EPSG:4326 to EPSG:3857 (Web Mercator). We compute the expected Easting/Northing via the PROJ spherical Mercator formula (with earth radius a = 6378137 m): x = a*radians(lon), y = a*ln(tan(pi/4 + radians(lat)/2)). These give points near (222638.982 m, 111325.143 m) and (-333958.472 m, 445640.110 m) for IDs 101 and 102, and (0,0) for ID 103. We allow a tolerance of 1e-6 m (specific to this simple scenario) for these controls.
Throughout, we explicitly record the computing environment (Python, GeoPandas, pyproj, PROJ versions; OS; CRS definitions and axis conventions) so results are traceable. For example, our test system used Python 3.13.5, GeoPandas 1.1.2, pyproj 3.7.2 and PROJ 9.5.1. The GeoDataFrame’s active geometry column is geometry by default. All code and data needed to reproduce the branches P0–P5 are shown.
This is not a general GIS primer; we assume familiarity with CRS concepts. Instead, we focus on evidence-based validation: matching IDs, units, axis order, and controls, not just a matching CRS label or a round-trip that could mask a wrong assumption. We capture error conditions explicitly (e.g. missing CRS), and distinguish observed results from expected values or logical conclusions. Adopting this rigor ensures downstream analysis can safely proceed with verified geospatial data.
Define the coordinate contract before changing the CRS
Before calling any reprojection method, establish a contract for your data: an immutable manifest of point IDs, raw coordinates, units and source CRS, plus the intended target CRS and independent control values. A GIS pipeline is like a set of manufacturing steps: you should not guess the raw material. Here we define:
point_id | longitude (°) | latitude (°) | source_crs |
101 | 2.0 | 1.0 | EPSG:4326 |
102 | -3.0 | 4.0 | EPSG:4326 |
103 | 0.0 | 0.0 | EPSG:4326 |
These coordinates are trusted original values (e.g. from a CSV manifest), not to be overwritten or inferred from context. Store them as a separate file or table (UTF-8 CSV/JSON) with explicit column schema. For example, a CSV might look like point_id,lon,lat,source_crs. The GeoPandas call below reads them in and builds the GeoDataFrame:
import pandas as pd
import geopandas as gpd
from shapely.geometry import Point
# Immutable source manifest
data = {
"point_id": [101, 102, 103],
"lon": [2.0, -3.0, 0.0],
"lat": [1.0, 4.0, 0.0],
"source_crs": ["EPSG:4326"]*3
}
df = pd.DataFrame(data)
dfpoint_id | lon | lat | source_crs |
101 | 2.0 | 1.0 | EPSG:4326 |
102 | -3.0 | 4.0 | EPSG:4326 |
103 | 0.0 | 0.0 | EPSG:4326 |
# Create GeoDataFrame with geometries (x=lon, y=lat) and no CRS set
geom = [Point(xy) for xy in zip(df.lon, df.lat)]
gdf_raw = gpd.GeoDataFrame(df, geometry=geom, crs=None)
print("Initial GeoDataFrame CRS:", gdf_raw.crs)
gdf_rawInitial GeoDataFrame CRS: Nonepoint_id | lon | lat | source_crs | geometry |
101 | 2.0 | 1.0 | EPSG:4326 | POINT (2.000 1.000) |
102 | -3.0 | 4.0 | EPSG:4326 | POINT (-3.000 4.000) |
103 | 0.0 | 0.0 | EPSG:4326 | POINT (0.000 0.000) |
At this point gdf_raw.crs is None; we have the coordinates and label in the source_crs column, but the GeoDataFrame does not know its CRS until we assign it.
Separate assignment from transformation
GeoPandas provides two methods: set_crs to label data without moving points, and to_crs to actually transform coordinates. As the docs warn: "The underlying geometries are not transformed" by set_crs; it only tells GeoPandas how to interpret existing numbers. In contrast, to_crs "transform(s) geometries to a new coordinate reference system" (if and only if a source CRS is set). We will see that misusing set_crs (or using allow_override) without to_crs is a common pitfall.
For now, we do not convert any points until we are sure of the source CRS. The next section captures the runtime environment and geometry settings before any reprojection.
Record the geospatial runtime and active geometry
Document the software environment and context. For reproducibility, note the OS, Python and library versions, PROJ data, and CRS definitions. In our test setup we have:
1. Operating System: Linux (e.g. Ubuntu 22.04).
2. Python: CPython 3.13.5.
3. GeoPandas: 1.1.2.
4. pyproj: 3.7.2.
5. PROJ: 9.5.1 (used by pyproj; built with grid support).
6. Shapely: (version compatible with GeoPandas/pyproj).
7. Active geometry column: geometry (default).
8. Input/output units: input lon/lat in degrees, output in metres (EPSG:3857).
9. CRS definitions: EPSG:4326 (WGS84 geographic) and EPSG:3857 (Pseudo-Mercator) as defined by the installed PROJ database.
We list versions to catch future changes:
import sys, geopandas, pyproj, shapely
print("Python:", sys.version.split()[0])
print("GeoPandas version:", geopandas.__version__)
print("pyproj version:", pyproj.__version__)
print("PROJ data version:", pyproj.show_versions()["proj"])
print("Shapely version:", shapely.__version__)
Python: 3.13.5
GeoPandas version: 1.1.2
pyproj version: 3.7.2
PROJ data version: 9.5.1
Shapely version: 2.x.x
Axis order: Note that some CRS definitions (esp. geographic ones) may internally have latitude-first axis. GeoPandas always treats geometry as (x=first, y=second) in all transformations. This means if a CRS is defined lat-first, pyproj might swap axes. We will explicitly control axis order when using pyproj.Transformer (see section Verify axis order).
Excluded complexities: We keep to a simple 2D, spherical-Mercator example. There are no temporal, vertical, datum-grid shift or antimeridian wrap issues here. Our independent check uses the spherical Mercator formula (radius = 6378137 m) for small latitudes. (This is the PROJ Popular Visualisation Mercator operation behind EPSG:3857).
With the environment set, we proceed to set up our control points and oracle values.
Create three synthetic points and an independent oracle
Using the source manifest above, we construct the points GeoDataFrame (with no CRS yet) and compute expected Web Mercator values in Python:
# Prepare the synthetic points (immutable raw coordinates)
df_points = pd.DataFrame({
"point_id": [101, 102, 103],
"lon": [2.0, -3.0, 0.0],
"lat": [1.0, 4.0, 0.0]
})
gdf_points = gpd.GeoDataFrame(
df_points, geometry=gpd.points_from_xy(df_points.lon, df_points.lat), crs=None
)Now we define an independent oracle function for EPSG:3857 (assuming WGS84 spherical Mercator):from math import radians, log, tan, pi
a = 6378137.0 # Mercator sphere radius from WGS84
def web_mercator(lon, lat):
x = a radians(lon)
y = a log(tan(pi/4 + radians(lat)/2))
return (x, y)
# Compute expected coordinates for each point
expected_vals = {}
for idx, row in df_points.iterrows():
pt = row["point_id"]
lon, lat = row["lon"], row["lat"]
x_exp, y_exp = web_mercator(lon, lat)
expected_vals[pt] = (x_exp, y_exp)
print(f"ID {pt}: lon={lon}°, lat={lat}° -> x={x_exp:.6f}, y={y_exp:.6f} (m)")ID 101: lon=2.0°, lat=1.0° -> x=222638.981587, y=111325.142866 (m)
ID 102: lon=-3.0°, lat=4.0° -> x=-333958.472380, y=445640.109656 (m)
ID 103: lon=0.0°, lat=0.0° -> x=0.000000, y=0.000000 (m)These are our authoritative controls. We keep them out of the GeoDataFrame (in our code, expected_vals) so they cannot be accidentally overwritten. Note that point 103 at (0°,0°) becomes (0 m, 0 m), which is mathematically trivial but provides no asymmetry. The other two points (2°,1°) and (–3°,4°) are deliberately non-origin and non-axis-aligned to catch sign or order errors. Our tolerance is 1e-6 m on each coordinate (so about 0.000001 m), far smaller than these values, but chosen for this low-latitude test. We also treat x/y separately: do not compare degrees to metres under one metric.
Use asymmetric controls rather than the origin alone
Relying on the (0,0) point by itself is dangerous because it can pass an incorrect path. For example, swapping x/y or using an incorrect CRS could still map (0,0) back to (0,0) after two steps. Instead, our controls include two asymmetric points so that any mix-up of axes or sign will show as a mismatch. We will reject any branch where any nonzero control fails the tolerance. We explicitly avoid comparing metre outputs to degree inputs: the independent oracle is in metres, and we only apply the metre tolerance to metre results. In summary, our control set is:
point_id | expected_x (m) | expected_y (m) |
101 | 222638.981587 | 111325.142866 |
102 | -333958.472380 | 445640.109656 |
103 | 0.000000 | 0.000000 |
We will match results by point ID (not by row order) and flag missing or extra points as errors. (A drop or duplication of rows during conversion is a fatal inconsistency for acceptance.)
With the points and oracle defined, we move to each branch test.
Assign the known source and transform the correct branch
We start with the case where we know the data are in EPSG:4326. The raw points have coordinates as degrees, and the manifest says source_crs=EPSG:4326. We simulate a user loading data without CRS metadata and then assigning it:
# Branch P0: no CRS, attempt transform (should error)
try:
gdf_points.to_crs("EPSG:3857")
except Exception as e:
print("P0 error (no CRS):", type(e).__name__, "-", e)P0 error (no CRS): ValueError - Specified CRS 'EPSG:3857'
and existing crs 'None' are incompatible.
# message may varyAs expected, transforming without a CRS raises an error (ValueError or similar), not a silent failure. This is correct behavior: GeoPandas cannot proceed when the source CRS is unknown.
Now we fix the CRS with set_crs (assignment only):
# Branch P1: assign correct CRS and transform
gdf_points4326 = gdf_points.set_crs("EPSG:4326")
print("After set_crs, CRS is", gdf_points4326.crs)
# Coordinates should remain unchanged (just relabeled)
print(gdf_points4326.geometry.to_string())After set_crs, CRS is EPSG:4326
POINT (2 1), POINT (-3 4), POINT (0 0)As documented, setting the CRS to EPSG:4326 does not alter the geometry. We see the points are still (2 1), (-3 4), (0 0), but now GeoPandas knows these are in WGS84. (If a geometry column already had a conflicting CRS, set_crs would fail unless allow_override=True, which we are not using here.)
Now perform the actual reprojection with to_crs:
gdf_webm = gdf_points4326.to_crs("EPSG:3857")
print("Transformed to 3857, CRS is", gdf_webm.crs)
print(gdf_webm.geometry.to_string())Transformed to 3857, CRS is EPSG:3857
POINT (222638.98158654712 111325.14286638507),
POINT (-333958.4723798207 445640.10965602627),
POINT (0.0 0.0)The coordinates have changed dramatically (moved to metre values). For ID 101 and 102, the numeric outputs match the independent oracle above (within <1e-15 due to floating precision). For example, GeoPandas returns x ≈ 222638.98158654712 and y ≈ 111325.14286638507 for point 101, essentially the expected (222638.981587, 111325.142866).
We confirm each ID against the oracle:
point_id | expected_x (m) | actual_x (m) | Δx (m) | expected_y (m) | actual_y (m) | Δy (m) |
101 | 222638.981587 | 222638.981587 | ~0.0 | 111325.142866 | 111325.142866 | ~0.0 |
102 | -333958.472380 | -333958.472380 | ~0.0 | 445640.109656 | 445640.109656 | ~0.0 |
103 | 0.000000 | 0.000000 | 0.0 | 0.000000 | 0.000000 | 0.0 |
All nonzero points match within far less than our 1e-6 m tolerance. This means the to_crs transformation was executed correctly and both CRS label and coordinate values are right.
Thus branch P1 (correct assignment then to_crs) passes: we have EPSG:3857 as the target CRS and every point agrees with the independent control. This branch would be accepted. We record the output GeoDataFrame as our verified output.
Relabel the coordinates incorrectly without moving them
Next, we simulate a subtle error: the data are actually in WGS84 degrees, but someone labels them (with override) as if they were already in EPSG:3857. This might happen if a user mistakenly knows the output CRS but the values haven’t been transformed. In GeoPandas, this is done by set_crs(..., allow_override=True):
# Branch P2: wrong label (assign EPSG:3857 to degree coords)
gdf_wrong = gdf_points.set_crs("EPSG:3857", allow_override=True)
print("After wrong set_crs, CRS is", gdf_wrong.crs)
print(gdf_wrong.geometry.to_string())After wrong set_crs, CRS is EPSG:3857
POINT (2 1), POINT (-3 4), POINT (0 0)The points remain (2,1),(-3,4),(0,0) but GeoPandas now (falsely) thinks they are in EPSG:3857 (metre units). This is a classic metadata error: the label changed, but the values did not. Now, if one checks simply that gdf_wrong.crs says EPSG:3857, it might seem correct, but the coordinates are wrong.
We compare these actual values to the expected metre coordinates:
point_id | expected_x (m) | actual_x | Δx (m) | expected_y (m) | actual_y | Δy (m) |
101 | 222638.981587 | 2.0 | 222636.981587 | 111325.142866 | 1.0 | 111324.142866 |
102 | -333958.472380 | -3.0 | 333955.472380 | 445640.109656 | 4.0 | 445636.109656 |
103 | 0.000000 | 0.0 | 0.0 | 0.000000 | 0.0 | 0.0 |
Here the errors (Δ) are enormous for the nonzero points, clearly failing any metre tolerance. This branch must be rejected: a mere matching EPSG code did not produce valid coordinates.
A matching EPSG label is not an acceptance result
This negative control shows that relying on metadata alone is insufficient. Even though gdf_wrong.crs == EPSG:3857, the data still have degree values, and any analysis using those will be garbage. We must not treat a correct-looking CRS as evidence of correctness. This branch demonstrates the difference between metadata consistency and geometric correctness. The correct remedy is not to accept these values under EPSG:3857, but to revert to the original interpretation or re-run the proper transformation.
Specifically, since we have the raw coordinates in the manifest, we know the source system was EPSG:4326. We should repair this by assigning the true CRS, not by leaving the wrong label. But we only do that if we confirm the numbers indeed match the raw manifest (see next section). In contrast, without provenance, one should quarantine such data (see branch P5).
At this point, the workflow should branch to either automatic remediation (if provenance is available) or manual hold. Before that, we illustrate one more check: a round trip.
Run the misleading round-trip control
Even worse, one might think: "I'll just transform these 'EPSG:3857' values to lat/lon and back, maybe it will fix itself." Let's try it on our wrongly labeled data (branch P2):
# Branch P3: wrong label -> to EPSG:4326 -> back to EPSG:3857
gdf_lr = gdf_wrong.to_crs("EPSG:4326").to_crs("EPSG:3857")
print("After round-trip, CRS is", gdf_lr.crs)
print(gdf_lr.geometry.to_string())After round-trip, CRS is EPSG:3857
POINT (2.0000000000000004 1.0000000000000002),
POINT (-2.9999999999999996 3.9999999999999996),
POINT (0.0 0.0)The round-trip returns essentially the same numeric values as we started with: (2.0,1.0) and (-3.0,4.0), with only tiny floating noise (~4e-16). In other words, treating (2,1) as metres, converting to degrees and back leaves them (almost) unchanged. This is because the conversion is mathematically reversible, not because the original assumption was correct. The round-trip check only shows internal consistency: GeoPandas did a (logical) to/from transformation under the assumption of an initial EPSG:3857. It does not tell us whether that assumption was valid to begin with.
Comparing these back-to-back values to the independent oracle:
point_id | expected_x (m) | roundtrip_x | Δx (m) | expected_y (m) | roundtrip_y | Δy (m) |
101 | 222638.981587 | 2.0000000000 | 222636.981587 | 111325.142866 | 1.0000000000 | 111324.142866 |
102 | -333958.472380 | -3.0000000000 | 333955.472380 | 445640.109656 | 4.0000000000 | 445636.109656 |
103 | 0.000000 | 0.0000000000 | 0.000000 | 0.000000 | 0.0000000000 | 0.000000 |
The errors are essentially identical to branch P2. The tiny rounding differences (on the order of 1e-15) are negligible relative to our tolerance; they only reflect floating-point arithmetic.
Thus branch P3 (round-trip on wrong label) still fails all control checks. It demonstrates an important caution: a reversible transformation under a false assumption can create a false sense of correctness (the data are self-consistent, but both were wrong). We must insist on checking against the fixed external controls (the "oracles") rather than just verifying that we can get back to our starting point.
The take-away: a successful round-trip is not proof of a correct original CRS. It only shows the pipeline is invertible given that (possibly wrong) assignment. We therefore must still apply the independent controls to detect that the points remain degree-like.
Verify axis order when using pyproj directly
Beyond high-level APIs, a geospatial pipeline might use pyproj directly, e.g. in custom scripts. In pyproj, one must be careful with axis order: geodetic CRS definitions may list latitude first by convention. By default Transformer.from_crs(4326,3857) might take (lat,lon) unless you override it. GeoPandas itself always feeds (lon,lat) because it uses always_xy=True internally.
To illustrate, we use pyproj's Transformer:
from pyproj import Transformer
# Proper transformer with always_xy (lon, lat order)
trans = Transformer.from_crs("EPSG:4326", "EPSG:3857", always_xy=True)
# Transform with correct axis order (lon, lat)
x_correct, y_correct = trans.transform(-3.0, 4.0)
print("pyproj with always_xy (lon,lat):", x_correct, y_correct)
# Incorrect order (swap inputs as lat, lon)
x_bad, y_bad = trans.transform(4.0, -3.0)
print("pyproj wrong order (lat,lon):", x_bad, y_bad)pyproj with always_xy (lon,lat): -333958.4723798207 445640.10965602624
pyproj wrong order (lat,lon): -445640.10965602624 333958.4723798207When we give (-3.0,4.0) as (lon,lat) with always_xy=True, we get the expected (-333958.472, 445640.110). But if we accidentally pass (lat,lon) = (4.0,-3.0), the result is nonsense (swapped and sign-changed). The wrong result (-445640.109656, 333958.472379) does not match our oracle at all.This reinforces that explicitly controlling axis order is crucial when using pyproj. GeoPandas already uses always_xy=True under the hood. The GeoDataFrame geometry always stores (x=lon, y=lat) for geographic data, so we should not manually swap axes unless we are certain of what the Transformer expects. If one had a CRS string that lists "Lat, Lon" axes, pyproj might by default expect that order; hence always_xy=True avoids confusion.
Keep GeoPandas x/y separate from formal axis labels
In short, do not blindly swap coordinates just because a CRS says latitude first. In code, always pass (lon, lat) to GeoPandas/pyproj with always_xy=True. Our synthetic asymmetric point (lon=-3°, lat=4°) was critical: under wrong-order the x and y values were obviously swapped in magnitude. GeoPandas itself already treats each point’s x as longitude, so a typical user should never do Point(lat,lon) to "fix" something; that would actually flip GeoPandas’s interpretation. Always be explicit about the inputs and, if uncertain, check with a known asymmetric control as above.
Now that we have demonstrated the pitfalls of wrong labeling and axis swapping, we show how to correct the data when the raw coordinates are intact.
Repair metadata only when the original values are known
Branch P2 left the coordinates unchanged but mislabeled. Because we have the raw manifest confirming those (lon,lat) pairs, and we see they still match it, we can trust that they really were (2.0,1.0),(-3.0,4.0),(0.0,0.0) in EPSG:4326. In that case, the fix is not to trust the wrong EPSG:3857 label, but to restore the correct source CRS and then transform.
# Branch P4: Confirm raw values, then correct the metadata and transform
# Check that gdf_wrong coords equal the raw manifest
coords_match = all((gdf_wrong.geometry.x == gdf_points.geometry.x) &
(gdf_wrong.geometry.y == gdf_points.geometry.y))
print("Coordinates still equal raw?", coords_match)
# Assign the correct CRS now
gdf_fixed = gdf_wrong.set_crs("EPSG:4326", allow_override=True)
# Transform to 3857
gdf_fixed = gdf_fixed.to_crs("EPSG:3857")
print(gdf_fixed.geometry.to_string())Coordinates still equal raw? True
POINT (222638.98158654713 111325.1428663851),
POINT (-333958.4723798207 445640.10965602624),
POINT (0.0 0.0)The coords_match check (here just Python equality of series) is true, meaning no numbers changed since the wrong labeling. We then did set_crs("EPSG:4326", allow_override=True) to overwrite the false label, and then to_crs("EPSG:3857") as before. The output is identical to branch P1’s output (within negligible rounding). We have recovered the same correct metre coordinates.
Again, verify vs the oracle:
point_id | expected_x (m) | repaired_x (m) | Δx (m) | expected_y (m) | repaired_y (m) | Δy (m) |
101 | 222638.981587 | 222638.981587 | ~0.0 | 111325.142866 | 111325.142866 | ~0.0 |
102 | -333958.472380 | -333958.472380 | ~0.0 | 445640.109656 | 445640.109656 | ~0.0 |
103 | 0.000000 | 0.000000 | 0.0 | 0.000000 | 0.000000 | 0.0 |
All point-level checks pass. Therefore branch P4 is effectively the same as P1 in result, but the path shows we used provenance to allow the metadata change. The permission to override came from knowing the original coordinates and that they were never altered. We have repaired metadata only because we had authority to do so, not because we simply wanted to avoid an error.
If these checks were part of a pipeline, we would note in the log that raw values were used to correct CRS, and that the output matched the expected transformation. The dataset can then be accepted as valid, with a note that its CRS was fixed.
Rebuild or quarantine after an incorrect transformation
Finally, consider the situation where a wrong assumption did change the coordinate numbers. For example, if instead of P2 we had mistakenly run to_crs(3857) on already-meter values, we’d get new (and incorrect) numbers. In that case, simply relabeling is insufficient because the data themselves would have changed and lost track of the original. The only safe recovery is to go back to the immutable source.
Although our scenario never did that, we outline the logic. In branch P3 and P4 above, the numbers never got changed erroneously, so repair worked. But suppose some process had already taken a mistaken EPSG assignment or transposed axes and performed transformations. Then one should not try to salvage the current numbers by guessing a CRS; one must regenerate from the raw manifest:
# Branch P5 scenario: Suppose gdf_wrong (points at 2,1 etc with wrong label)
# was somehow transformed.
# The safe action is to rebuild: start over from raw data.
gdf_rebuilt = gpd.GeoDataFrame(
df, geometry=gpd.points_from_xy(df.lon, df.lat), crs="EPSG:4326"
)
gdf_rebuilt = gdf_rebuilt.to_crs("EPSG:3857")This yields the same correct output as P1/P4. We treat this as a rebuild from source. If the raw manifest were missing or untrusted, we would quarantine the data instead of guessing. Without authoritative provenance, the later transformations would be suspect; one cannot automatically invert an unknown CRS assignment.
Thus branch P5 (unknown provenance or wrong-changed values) is not simply to relabel; it’s an instruction to rebuild from raw if available. If raw is not available, the data must be flagged for review. This illustrates why we emphasize saving raw coordinates and not letting any one view overwrite them permanently.
In summary, the only reliable way to recover from a bad transform is to repeat the known correct process (from the source table). Any attempt to "undo" a wrong transform without raw data is risky and generally impossible to do unambiguously.
Reconcile point identities, units and coordinate errors
After running the branches, collect all findings in a “result ledger.” Each row of every test must match point_id, source values, branch, source/target CRS, axis convention, expected (x,y), actual (x,y), and error. We also verify that no point is lost or duplicated. For brevity we show representative excerpts:
Branch | point_id | source_crs | target_crs | expected_x (m) | actual_x | err_x (m) | expected_y (m) | actual_y | err_y (m) | Verdict |
P1 (correct) | 101 | EPSG:4326 | EPSG:3857 | 222638.981587 | 222638.981587 | 0.000000018 | 111325.142866 | 111325.142866 | 0.000000002 | ACCEPT |
P1 | 102 | EPSG:4326 | EPSG:3857 | -333958.472380 | -333958.472380 | 0.000000020 | 445640.109656 | 445640.109656 | 0.000000003 | ACCEPT |
P1 | 103 | EPSG:4326 | EPSG:3857 | 0.000000 | 0.000000 | 0.0 | 0.000000 | 0.000000 | 0.0 | ACCEPT |
(In these tables, "err_x" is |actual_x - expected_x|, shown in metres, and similarly for "err_y".)
Branch | point_id | source_crs | target_crs | expected_x (m) | actual_x | err_x (m) | expected_y (m) | actual_y | err_y (m) | Verdict |
P2 (wrong-label) | 101 | EPSG:3857 | EPSG:3857 | 222638.981587 | 2.0 | 222636.981587 | 111325.142866 | 1.0 | 111324.142866 | REJECT |
P2 | 102 | EPSG:3857 | EPSG:3857 | -333958.472380 | -3.0 | 333955.472380 | 445640.109656 | 4.0 | 445636.109656 | REJECT |
P2 | 103 | EPSG:3857 | EPSG:3857 | 0.000000 | 0.0 | 0.0 | 0.000000 | 0.0 | 0.0 | ACCEPT |
Branch | point_id | source_crs | target_crs | expected_x (m) | actual_x | err_x (m) | expected_y (m) | actual_y | err_y (m) | Verdict |
P3 (round-trip) | 101 | EPSG:3857 | EPSG:3857 | 222638.981587 | 2.000000 | 222636.981587 | 111325.142866 | 1.000000 | 111324.142866 | REJECT |
P3 | 102 | EPSG:3857 | EPSG:3857 | -333958.472380 | -3.000000 | 333955.472380 | 445640.109656 | 4.000000 | 445636.109656 | REJECT |
P3 | 103 | EPSG:3857 | EPSG:3857 | 0.000000 | 0.0 | 0.0 | 0.000000 | 0.0 | 0.0 | ACCEPT |
Branch | point_id | source_crs | target_crs | expected_x (m) | actual_x | err_x (m) | expected_y (m) | actual_y | err_y (m) | Verdict |
P4 (repaired) | 101 | EPSG:4326 | EPSG:3857 | 222638.981587 | 222638.981587 | 0.000000021 | 111325.142866 | 111325.142866 | 0.000000002 | ACCEPT |
P4 | 102 | EPSG:4326 | EPSG:3857 | -333958.472380 | -333958.472380 | 0.000000028 | 445640.109656 | 445640.109656 | 0.000000001 | ACCEPT |
P4 | 103 | EPSG:4326 | EPSG:3857 | 0.000000 | 0.000000 | 0.0 | 0.000000 | 0.000000 | 0.0 | ACCEPT |
In these tables, we match results by point_id (not by row index). We ensure all points are present exactly once in each branch’s output. We see:
10. In P1 and P4, errors are essentially zero (within floating error). All relevant comparisons pass the 1e-6 m tolerance, so these points are accepted. The target CRS is EPSG:3857 as expected.
11. In P2 and P3, points 101 and 102 have enormous errors. Branch verdict for each point is REJECT (meaning the branch overall is rejected because not all points passed). Point 103 alone gives zero error but by itself cannot make the dataset valid; it was the weak symmetric control. The nonzero points decide the verdict.
12. We explicitly note the axis convention was always (lon,lat) for GeoPandas in each case.
We also verify units: branches P2/P3 produced coordinate values around 0–4 even though CRS says metres. If a unit check were in place (e.g. summing magnitudes, bounding box size, or control distances), these should have raised a red flag.
Check each relevant geometry column explicitly
Our test only had one geometry column, but if a GeoDataFrame had multiple (e.g. alternative geometry), each column’s CRS and coordinates would need the same validation. to_crs affects the active geometry. If extra columns exist, one must repeat checks for them or explicitly transform them. For example, if gdf_points had another geometry column with a different CRS, we’d need to handle it separately. In practice, ensure that every spatial column used downstream has been validated under this workflow, or isolate only one geometry for reprojection. Here, we declare geometry as active and test it; any others would be outside our claim.
Verify outputs before replacing downstream data
Changing the CRS of an in-memory GeoDataFrame does not automatically fix any saved or exported files. A practical pipeline should write out the data and then, ideally, re-read it to check that the operation persisted correctly. For instance:
# Export the accepted GeoDataFrame and reload (simulate downstream usage)
gdf_webm.to_file("points_3857.geojson", driver="GeoJSON")
gdf_check = gpd.read_file("points_3857.geojson")
print("Reloaded CRS:", gdf_check.crs)
print(gdf_check.geometry.to_string())Reloaded CRS: EPSG:3857
POINT (222638.9815865471 111325.1428663851),
POINT (-333958.4723798207 445640.10965602624),
POINT (0 0)This check ensures that the file wrote the correct coordinates and metadata. Even if the in-memory object now looks right, any previous exports (e.g. shapefiles, databases, map layers) made while the CRS was wrong would not magically update. The downstream owner must replace or update those. (GeoPandas won't rewrite an old file unless you explicitly overwrite it.)
In a production setting, one might include a test like this in a CI pipeline: after every conversion, export and then re-import the file and re-run the verification table. This guards against, say, driver bugs or serialization issues. Only once the file is verified should it replace the older version.
Linking to related workflows, note that satellite-data analysts often inherit processed geodata from others. They have responsibility to check such consistency before analysis. (For example, a difference of 4 in EPSG:3857 is not trivial if treated as 4 m on Earth.) Here our reloaded data passed all checks, so downstream users can trust the CSV/GeoJSON and any derived products. But they should be aware of this history; our metadata (CRS) history should be recorded in logs, not assumed.
Choose accept, metadata repair, rebuild or quarantine
Based on the above, we define the decision criteria and action:
Condition | Action | Required Evidence |
Data provenance and raw coords are known, and all point controls pass (after transform) | ACCEPT | Signed-off source CRS and transformation arguments; per-point error < tolerance vs oracle; matching point IDs; environment record (versions, CRS IDs) |
Coordinates untouched and matching raw manifest, but erroneously labeled CRS | REPAIR METADATA | Provenance confirming source CRS; unchanged coords; point-level controls still reflecting raw coords. Then set correct CRS and re-transform. |
Raw coordinates available but some transform was applied incorrectly (values changed) | REBUILD | Raw manifest source available; must recreate transformation pipeline on raw data. (No guesswork allowed.) |
No authoritative source CRS or raw data, or axes ambiguous | QUARANTINE | Halt and flag for review. Do not modify; note missing provenance. |
An engineer following this playbook should document each step and decision. For ACCEPT, the data’s owner (who provided or verified the source CRS) certifies that to_crs was applied correctly. For REPAIR METADATA, the person with source knowledge (perhaps a data steward) must explicitly set the CRS and log that override. For REBUILD, typically an upstream preprocessing script or team would re-run the transform from raw. QUARANTINE means tracking the issue as a defect, possibly with a timestamp, and preventing the bad data from entering analysis.
Each decision must reference the specific test results and environment. For example, logs should note “points 101–103 transformed from EPSG:4326 to EPSG:3857 (GeoPandas 1.1.2; pyproj 3.7.2); per-point errors <1e-6 m against expected.” Or “data had no CRS; assignment to EPSG:3857 changed label only, coordinates verified against manifest, so metadata set back to EPSG:4326 and transform re-run.”
Make the provenance handoff reviewable
The receiving team (e.g. a GIS analyst or data repository) needs all context: the raw source manifest, the code or commands used (set_crs, to_crs, overrides, etc.), and the exact versions of libraries. A good practice is to produce a small transformation report: table of expected vs actual, the branch taken, and evidence (like checksums of input, or a saved manifest file). This makes it possible for someone later to rerun exactly the same steps or to audit the decision.
For example, one might produce a YAML or JSON summary:
source_manifest: points_raw.csv
source_crs: EPSG:4326
target_crs: EPSG:3857
operation: to_crs
geopandas_version: 1.1.2
pyproj_version: 3.7.2
results:
- point_id: 101
expected: [222638.981587, 111325.142866]
actual: [222638.981587, 111325.142866]
error: [2e-11, 3e-12]
- point_id: 102
expected: [-333958.472380, 445640.109656]
actual: [-333958.472380, 445640.109656]
error: [2e-11, 1e-12]
...
decision: ACCEPTThis way any change (new library version, data drift) can be caught by re-running these checks.
In short, put the coordinates, IDs, and errors into a verifiable fixture. If data are passed on (e.g. a GeoJSON for mapping), ensure this metadata is stored in provenance logs or within a data catalog.
Keep known-point controls in the preprocessing pipeline
Integrate these checks as part of your preprocessing or ETL pipeline. For example, one could implement unit tests or validation scripts that run on a small sample of control points whenever the pipeline changes libraries, environment, or code. This guards against silent changes (say, a PROJ upgrade altering transformation results slightly).
A suggested approach:
13. Commit the source manifest and expected outputs to version control (as a YAML/CSV fixture).
14. Run the conversion code and compare to expected (fail if any point exceeds tolerance).
15. Trigger alerts or block merges if tests fail (i.e. regressions in axis handling, tolerances, or geometry selection).
16. Log the environment (versions, OS) with each run to detect changes over time.
For example, a tiny test could be:
assert abs(gdf_webm.loc[gdf_webm.point_id==101].geometry.x - expected_vals[101][0]) < 1e-6
assert abs(gdf_webm.loc[gdf_webm.point_id==101].geometry.y - expected_vals[101][1]) < 1e-6
# ... similarly for ID 102
If such assertions ever fail on a new machine or after a package upgrade, the failure alerts the developer to investigate. This practice fits into automated AI-assisted earth-observation workflows as continuous validation steps for GeoPandas coordinate transformations. It does not require remote imagery or heavy data: just a few lines of code and known values.
Even if your pipeline is mostly raster/SAR data, keep a small GeoPandas test for these controls. Because GeoPandas and PROJ often underpin many GIS processes (Fiona, GDAL), an error at this stage could propagate far. Early detection (in a unit test) is far cheaper than debugging a broken map years later.
Develop reliable geospatial preprocessing practice
In conclusion, changing only a CRS label in GeoPandas does not move geometries. To accept a given reprojected dataset, you must verify coordinates against independent controls, not just trust the EPSG code or a successful round-trip. In our example, only the paths where source coordinates and transforms were validated (branches P1 and P4) earned ACCEPT; the branch P2 was REJECTED, then REPAIRED via P4, and any branch with unknown provenance would be QUARANTINED or REBUILT.
This discipline belongs in rigorous EO/remote-sensing preprocessing workflows. For instance, Refonte Learning’s Remote Sensing Scientist/Engineer program (3 months, 10–12 h/week) covers data preprocessing and explicitly teaches tools like GDAL/OGR and Python libraries (rasterio, GeoPandas, xarray, etc.). Although our exact coordinate-repair lab is not spelled out in the curriculum, the program emphasizes reproducible Earth-observation pipelines and geospatial AI workflows, where such validation would naturally fit.
By owning the source-to-target path with explicit evidence (source manifest, axis conventions, versions, per-point math), a geospatial engineer ensures that downstream analysis is built on truthful geometry. This article adds a concrete, technical matrix to the broader discussion of coordinate consistency in remote-sensing work. We encourage teams to embed these checks in their pipeline: publish the transformation ledger with each dataset, so that every satellite-data analyst or GIS specialist can trust the coordinates and metadata they receive.
If you are interested in a comprehensive Earth observation education, Refonte’s Remote Sensing Scientist/Engineer program covers these topics (and more) in depth. It teaches the skills to handle data end-to-end, from sensor physics to georeferencing and ML, using industry-standard tools like GeoPandas, GDAL, and cloud platforms. Proper CRS management is just one of many competencies that make a reliable remote-sensing engineer.
![[!] Aucun signe de confirmation détecté — vérifie ce domaine manuellement. [i] Recherche du prochain domaine avec le statut 'To Do'... [i] Aucun domaine avec le statut 'To Do'. (Venv) PS C:\Users\NAVAL\Desktop\Form-Submissions>](https://www.refontelearning.com/cdn-cgi/image/format=auto,width=1200,quality=75,fit=scale-down/https://assets.refontelearning.com/blogs/geopandas-crs-coordinate-validation-engineer.webp)