Data scientist validating scikit-learn feature names and column order before model prediction

Same Shape, Wrong Features: Protect scikit-learn Column Order

Last updated: Mon, Sep 28, 2026

A prediction can be numerically valid, shape-valid, and still be wrong because the right values reached the wrong feature coordinates. That is the acceptance question here: for every request, did the value intended for units reach the coordinate learned for units, and did adjustment reach the coordinate learned for adjustment? A (1, 2) array cannot answer that question by shape alone.

This playbook freezes the fitted mathematics and changes only the inference boundary. The synthetic fixture has two numeric features, six rows, and the exact relation y = 10*units + adjustment. For units=2 and adjustment=7, the independent known answer is 27. A reversed plain array contains the same two values and the same shape but means (7, 2) positionally, producing 72 with the same coefficients. The defect is therefore feature identity, not model quality.

The acceptance strategy has three parts: reproduce the unsafe control, observe what the pinned scikit-learn installation actually does with names and arrays, and put a fail-closed application contract in front of a name-based ColumnTransformer pipeline. The result is deliberately bounded. Passing these tests establishes that this packaged path preserves the declared two-feature mapping under the tested cases. It does not prove that upstream systems computed the right business meaning, used the right physical units, or remained semantically correct outside this fixture.

Define feature identity before accepting a prediction

Feature identity is a mapping, not a width. The training contract for this fixture is ("units", "adjustment"). Coordinate zero means units; coordinate one means adjustment. If serving sends [7, 2] while the fitted coefficient vector is approximately [10, 1], the estimator computes 10*7 + 1*2 = 72. Nothing about the array shape is malformed. The values are not corrupted. Their identities were lost before prediction.

That is a different problem from ordinary model-performance evaluation. Refonte’s broader guide to model evaluation and baseline discipline addresses splits, baselines, metrics, calibration, and generalization. This playbook assumes the fitted relationship is already frozen and asks whether serving preserves the training coordinates.

Treat three artifacts differently. The normative specification is the local feature contract: exact names, order, meanings, dtypes, schema revision, and preprocessing/model identity. The API references for scikit-learn and pandas document library behavior, but they do not define your business semantics. The client implementation is the adapter and pipeline that enforce the local contract. An explanatory guide can suggest practices; it does not replace either the normative contract or executable acceptance evidence.

scikit-learn’s LinearRegression API records both n_features_in_ and, when fitted input exposes all-string feature names, feature_names_in_. Those attributes capture different information: count versus names. They are evidence about fitting, not a universal promise that arbitrary input will be realigned automatically.

Acceptance therefore begins with one sentence that a reviewer can falsify: for schema feature-contract-v1, every prediction must bind units to trained coordinate zero and adjustment to trained coordinate one before the estimator executes.

That sentence also fixes the unit of acceptance. A request is not approved because its values are “reasonable,” because the estimator returns a scalar, or because the feature count matches. Approval requires a traceable correspondence from declared feature identity to fitted coordinate. This makes review concrete: a feature owner can challenge the semantic declaration, a serving engineer can challenge the adapter behavior, and a model owner can challenge the fitted artifact identity without collapsing all three questions into a vague “model works” judgment.

Keep the scope deliberately narrow. The fixture does not estimate production accuracy and does not compare algorithms. It is a numerical interface test. That narrowness is useful because it converts a class of hard-to-see inference bugs into a deterministic acceptance condition with an independently calculable answer.

Freeze the model and runtime contract

Do not investigate a serving-order defect while also changing coefficients, preprocessing, dependencies, warnings handling, or fixture data. Freeze enough state that a difference in output can be attributed to the input boundary.

The local execution used for the observed trace in this article was one Python process with Python 3.13.5, scikit-learn 1.8.0, pandas 2.2.3, NumPy 2.3.5, SciPy 1.17.0, joblib 1.5.3, threadpoolctl 3.6.0, on Linux-6.18.44-x86_64-with-glibc2.41. The related Python data-science tools are broader than this contract test; here, version identity matters because validation and warning behavior can change across releases.

The scikit-learn stable documentation resolved to version 1.9.1 when checked on September 28, 2026, while this lab remains pinned to 1.8.0. That is exactly why a stable documentation selector is not an environment lock and an access date is not a release date. Current stable documentation still describes the same concepts for feature names, counts, and column selection, but the observed warning/exception trace below is scoped to the pinned 1.8.0 process.

Use an explicit dependency lock:

python==3.13.5
scikit-learn==1.8.0
pandas==2.2.3
numpy==2.3.5
scipy==1.17.0
joblib==1.5.3
threadpoolctl==3.6.0

Record these revisions with the bundle: fixture feature-order-fixture-v1, schema feature-contract-v1, preprocessor named-column-selector-v1, and model linear-regression-fixture-v1. Define units as a synthetic quantity and adjustment as a synthetic additive adjustment; both are unitless in this fixture. The release warnings policy is also versioned: capture warnings during evidence collection, but promote warnings to errors at the release gate. That is an application policy, not scikit-learn’s default.

Record the lock in two forms. Keep the human-readable manifest above, and also persist the package-manager output used to recreate the environment. A suitable invocation is python -m pip freeze, stored beside the artifact rather than pasted selectively into an incident ticket. The purpose is not to prove every transitive dependency caused the behavior; it is to make the environment reconstructable when library validation changes.

The warnings policy deserves the same treatment. During evidence collection, warnings.catch_warnings(record=True) plus simplefilter("always") exposes what the library emitted without hiding it. During release acceptance, simplefilter("error") makes an unexpected warning fail the gate. Those are two intentionally different modes. Do not globally mutate warnings behavior in a notebook and then forget which policy produced the evidence.

Finally, preserve the fitted numerical state. For this fixture, each path records coefficients, intercept, n_features_in_, and whether feature_names_in_ exists. The point is not that coefficients are secret internals; it is that a reviewer must be able to see that the unsafe and corrected comparisons used the same mathematical relationship rather than different fitted models disguised as an input experiment.

Build a two-feature known-answer model

The fixture is intentionally small enough to inspect without another oracle. Six rows define the exact plane y = 10*units + adjustment:

units

adjustment

target

0

0

0

1

0

10

0

1

1

1

1

11

2

3

23

3

2

32

Fit one distinct model per explicitly labeled path. Path A is positional and trained on an ndarray. Path B is trained directly on a named DataFrame. Path C is the corrected name-based ColumnTransformer/Pipeline path developed later. All three use LinearRegression(fit_intercept=False) so the zero intercept is part of the declared fixture rather than an approximately estimated nuisance term. scikit-learn documents LinearRegression as ordinary least squares and exposes fitted coefficients, intercept, feature count, and feature names where applicable. The official LinearRegression reference is the primary API source.

import numpy as np
import pandas as pd
from sklearn.linear_model import LinearRegression

TRAINING = pd.DataFrame(
    {
        "units": [0, 1, 0, 1, 2, 3],
        "adjustment": [0, 0, 1, 1, 3, 2],
    },
    dtype="float64",
)
TARGET = np.array([0, 10, 1, 11, 23, 32], dtype="float64")

path_a = LinearRegression(fit_intercept=False).fit(
    TRAINING.to_numpy(), TARGET
)
path_b = LinearRegression(fit_intercept=False).fit(
    TRAINING, TARGET
)

np.testing.assert_allclose(
    path_a.coef_, [10.0, 1.0], rtol=0, atol=1e-12
)
np.testing.assert_allclose(
    path_b.coef_, [10.0, 1.0], rtol=0, atol=1e-12
)
np.testing.assert_allclose(
    path_a.intercept_, 0.0, rtol=0, atol=1e-12
)
np.testing.assert_allclose(
    path_b.intercept_, 0.0, rtol=0, atol=1e-12
)

In the observed pinned process, Path A fitted coefficients were [10.0, 1.000000000000001] and intercept 0.0; Path B exposed feature_names_in_ == ["units", "adjustment"]. The tolerance acknowledges ordinary floating-point representation without weakening the exact feature-order contract.

Derive 27 before running the estimator

Do not let the estimator under test define its own expected answer. The query is units=2, adjustment=7. The independent fixture equation gives:

10 * 2 + 7 = 27

That 27 is a mathematically expected result, established before prediction. It is not “what the model happened to output.” The acceptance test may compare predictions to 27 within a tight floating-point tolerance because the synthetic contract itself defines the expected mapping.

Create both labeled orders:

CANONICAL = pd.DataFrame(
    {"units": [2.0], "adjustment": [7.0]}
)
PERMUTED = CANONICAL.loc[:, ["adjustment", "units"]]
EXPECTED = 27.0

Both frames describe the same labeled record. Their presentation order differs; their feature identities do not.

Reverse the array without changing its shape

The failing control is one line at the DataFrame-to-array boundary:

canonical_array = CANONICAL.to_numpy()
reversed_array = PERMUTED.to_numpy()

assert canonical_array.shape == (1, 2)
assert reversed_array.shape == (1, 2)

print(canonical_array.tolist())  # [[2.0, 7.0]]
print(reversed_array.tolist())   # [[7.0, 2.0]]

np.testing.assert_allclose(
    path_a.predict(canonical_array),
    [27.0],
    atol=1e-12,
)
np.testing.assert_allclose(
    path_a.predict(reversed_array),
    [72.0],
    atol=1e-12,
)

The observed pinned result was 27 for [[2.0, 7.0]] and 72 for [[7.0, 2.0]], with both arrays shaped (1, 2). That is the intentionally failing control: the model is unchanged; the coefficient vector is unchanged; the two scalar values are unchanged as a set; only their positions changed.

pandas documents DataFrame.to_numpy() as converting a DataFrame to a NumPy array and describes dtype/copy behavior. The official DataFrame.to_numpy reference returns an ndarray; the pinned pandas 2.2.3 reference states the same conversion contract. A plain ndarray does not retain pandas column labels as a feature-identity contract.

This is the core NumPy feature order prediction error: numeric dimensions remain acceptable while semantic coordinates are silently different. The issue is not that NumPy is defective. The issue is that converting a labeled structure into an unlabeled positional structure discards information your serving path may still need.

Why a feature count cannot identify a feature

n_features_in_ == 2 answers “how many fitted coordinates?” It cannot answer “which semantic feature belongs at coordinate zero?” Two permutations share the same width.

feature_names_in_ can record the names seen at fit when the estimator receives suitable labeled input. scikit-learn documents that distinction explicitly for LinearRegression. But names attached to the fitted estimator do not magically attach labels to a later ndarray. Once the serving caller supplies [[7, 2]], the values are positional unless some separate, authenticated order contract establishes otherwise.

Therefore a shape-only release check is insufficient. The minimum acceptance evidence includes names before label removal, canonical names after any selection/reordering step, and the actual coordinate order entering the estimator.

Test the named-estimator boundary honestly

Path B answers a different question: what happens when LinearRegression was fitted directly on the named DataFrame and then receives several input forms? Do not summarize this as “scikit-learn handles feature names.” Record the exact installed-version behavior.

The evidence code captures warnings rather than suppressing them:

import warnings

def observe(fn):
    with warnings.catch_warnings(record=True) as caught:
        warnings.simplefilter("always")
        try:
            value = fn()
            return {
                "status": "returned",
                "value": np.asarray(value).tolist(),
                "warnings": [
                    (
                        type(w.message).__name__,
                        str(w.message),
                    )
                    for w in caught
                ],
            }
        except Exception as exc:
            return {
                "status": "raised",
                "exception": type(exc).__name__,
                "message": str(exc),
                "warnings": [
                    (
                        type(w.message).__name__,
                        str(w.message),
                    )
                    for w in caught
                ],
            }

b1 = observe(lambda: path_b.predict(CANONICAL))
b2 = observe(lambda: path_b.predict(PERMUTED))
b3 = observe(
    lambda: path_b.predict(PERMUTED.to_numpy())
)

Observed with scikit-learn 1.8.0:

Path B input

Observed result

Warning/exception

Acceptance meaning

canonical DataFrame units, adjustment

about 27

none

acceptable

permuted DataFrame adjustment, units

no prediction

ValueError: names must be in fit order

library rejects this permutation

permuted ndarray [[7,2]]

72

UserWarning about missing valid feature names

unsafe if warnings are ignored

This is exactly why the article does not claim universal permutation invariance for a direct named estimator. In the pinned 1.8.0 client implementation, a permuted DataFrame is rejected rather than automatically aligned. The array still reaches prediction after a warning and produces the wrong value. Current scikit-learn documentation describes validation infrastructure that can set or check names and counts, but behavior must be treated as versioned and estimator-specific. The official validate_data reference says it validates input and sets or checks feature names and counts.

The distinction between rejection and alignment matters operationally. A direct estimator rejecting a permuted DataFrame is useful because it prevents a silent wrong prediction, but it does not provide permutation invariance. If a serving API promises that labeled columns may arrive in any order, then “raises on permutation” still fails that API contract. The application must either canonicalize labels before calling the estimator or use a preprocessing path that selects by name.

Conversely, a warning on an ndarray is weaker than rejection. In a batch job or service where warnings are not surfaced as failures, the caller may receive 72 and continue normally. That is why acceptance should store both the numeric return and the warning/exception channel. Looking only at exceptions would miss the dangerous case; looking only at predictions would miss the library’s warning signal.

This also explains why test expectations should name the exact path. “Permuted input should work” is ambiguous. For Path B in the pinned environment, the correct expected behavior is rejection of a permuted DataFrame. For guarded Path C, the expected behavior is successful canonicalization and a prediction of 27. Different paths can have different acceptable behavior as long as each is explicitly specified and the production path is the guarded one.

Separate library checks from application policy

A library warning is not a serving contract. The default warning machinery did not prevent the observed 72. For the release gate, this playbook intentionally turns warnings into errors:

with warnings.catch_warnings():
    warnings.simplefilter("error")
    path_b.predict(PERMUTED.to_numpy())

In the pinned process, the UserWarning about invalid feature names becomes an exception under that explicit policy.

That stricter treatment is policy owned by the application. It must not be described as scikit-learn’s default.

Likewise, feature_names_in_ is useful evidence that fitting occurred with named features, while n_features_in_ records only the count. Neither attribute verifies that an upstream producer used the right measurement units, computed the right concept, or attached honest labels. Those are separate contract dimensions.

Build the name-based comparison pipeline

The corrected comparison makes selection by name explicit before the estimator. scikit-learn documents that ColumnTransformer interprets string selectors as DataFrame column names, integers as positions, and orders transformed output according to the transformer specification. The official ColumnTransformer reference also documents remainder="drop" and the treatment of unspecified or extra columns.

Fit a third, distinct model for Path C:

from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.linear_model import LinearRegression

selector = ColumnTransformer(
    transformers=[
        (
            "contract",
            "passthrough",
            ["units", "adjustment"],
        ),
    ],
    remainder="drop",
    verbose_feature_names_out=False,
)
selector.set_output(transform="pandas")

path_c = Pipeline(
    [
        ("select", selector),
        (
            "model",
            LinearRegression(fit_intercept=False),
        ),
    ]
).fit(TRAINING, TARGET)

np.testing.assert_allclose(
    path_c.named_steps["model"].coef_,
    [10.0, 1.0],
    rtol=0,
    atol=1e-12,
)

assert (
    path_c.named_steps["select"]
    .get_feature_names_out()
    .tolist()
    == ["units", "adjustment"]
)

Now verify transformed identity, not just predictions:

t1 = path_c.named_steps["select"].transform(CANONICAL)
t2 = path_c.named_steps["select"].transform(PERMUTED)

assert list(t1.columns) == ["units", "adjustment"]
assert list(t2.columns) == ["units", "adjustment"]

np.testing.assert_allclose(
    t1.to_numpy(),
    [[2.0, 7.0]],
)
np.testing.assert_allclose(
    t2.to_numpy(),
    [[2.0, 7.0]],
)

np.testing.assert_allclose(
    path_c.predict(CANONICAL),
    [27.0],
    atol=1e-12,
)
np.testing.assert_allclose(
    path_c.predict(PERMUTED),
    [27.0],
    atol=1e-12,
)

Both DataFrame orders returned about 27 in the observed 1.8.0 process. That is the required pipeline permutation invariance test for labeled inputs.

One subtlety is non-negotiable: remainder="drop" is not an “extra columns are forbidden” rule. The documentation states that unspecified columns are dropped, and extra DataFrame columns not seen during fit are excluded from transformation. In the pinned process, an unguarded frame containing debug=999 still produced 27 because debug was dropped. That behavior is useful inside a controlled transformer but unacceptable as the sole application contract. Reject extras before the pipeline.

The comparison path uses passthrough deliberately. There is no scaling, imputation, encoding, or learned transformation that could obscure the identity question. ColumnTransformer is serving here as an explicit name-to-position compiler: select units, then adjustment, preserve those values, and emit them in that declared order. Setting output to pandas keeps transformed labels inspectable so the acceptance test can assert names as well as values.

That design also produces a useful failure for unlabeled input. Because the selector uses string column names, passing a plain ndarray directly to Path C raises a ValueError in the pinned installation stating that string column selection is supported only for DataFrames. That library failure is helpful defense in depth, but the application adapter should reject the array earlier with a contract-specific message. The adapter is where the system expresses policy; the transformer is where the fitted preprocessing expresses name-based selection.

Add a fail-closed serving adapter

The serving adapter owns the contract that scikit-learn should not be expected to infer. It requires a pandas DataFrame, all-string unique names, the exact allowed feature set, the declared schema revision, numeric non-boolean dtypes, finite values, and finally the canonical order. Only after all checks pass may prediction execute.

from dataclasses import dataclass
from pandas.api.types import (
    is_bool_dtype,
    is_numeric_dtype,
)

SCHEMA_REVISION = "feature-contract-v1"

@dataclass(frozen=True)
class FeatureSpec:
    name: str
    meaning: str

@dataclass(frozen=True)
class FeatureContract:
    schema_revision: str
    features: tuple[FeatureSpec, ...]
    dtype: str = "float64"

    @property
    def order(self) -> tuple[str, ...]:
        return tuple(
            feature.name
            for feature in self.features
        )

CONTRACT = FeatureContract(
    schema_revision=SCHEMA_REVISION,
    features=(
        FeatureSpec(
            "units",
            "Synthetic quantity; unitless in this fixture.",
        ),
        FeatureSpec(
            "adjustment",
            "Synthetic additive adjustment; unitless in this fixture.",
        ),
    ),
)

class ContractError(ValueError):
    pass

def adapt_dataframe(
    X,
    *,
    schema_revision: str,
) -> pd.DataFrame:
    if not isinstance(X, pd.DataFrame):
        raise ContractError(
            "Expected a pandas DataFrame; "
            "unlabeled arrays are not accepted."
        )

    columns = list(X.columns)

    if not all(
        isinstance(column, str)
        for column in columns
    ):
        raise ContractError(
            "All feature names must be strings."
        )

    if len(set(columns)) != len(columns):
        raise ContractError(
            "Duplicate feature names are not allowed."
        )

    if schema_revision != CONTRACT.schema_revision:
        raise ContractError(
            f"Schema revision mismatch: "
            f"got {schema_revision!r}, "
            f"expected "
            f"{CONTRACT.schema_revision!r}."
        )

    expected = set(CONTRACT.order)
    received = set(columns)

    missing = sorted(expected - received)
    extra = sorted(received - expected)

    if missing or extra:
        raise ContractError(
            f"Feature set mismatch: "
            f"missing={missing}, extra={extra}."
        )

    for name in CONTRACT.order:
        series = X[name]
        if (
            is_bool_dtype(series.dtype)
            or not is_numeric_dtype(series.dtype)
        ):
            raise ContractError(
                f"Feature {name!r} must have "
                f"a real numeric dtype; "
                f"got {series.dtype}."
            )

    canonical = X.loc[
        :, list(CONTRACT.order)
    ].astype(
        CONTRACT.dtype,
        copy=False,
    )

    if not np.isfinite(
        canonical.to_numpy(copy=False)
    ).all():
        raise ContractError(
            "All feature values must be finite; "
            "NaN and +/-inf are rejected."
        )

    return canonical

An ndarray is rejected by default because its feature identity is ambiguous at this boundary. A system that must accept arrays needs a separate, authenticated and versioned order contract proving what each position means. Never guess that contract, never alphabetically sort values, and never infer it from array width.

Validate before discarding labels

The order of operations matters. First validate labels; then reorder a trusted labeled frame; only then allow a downstream component to materialize an array if required.

Wrong implementations often reverse those steps: df.to_numpy() first, followed by a positional “fix.” At that point, duplicate names, an unexpected debug column, or a missing feature may already have been flattened into a shape that looks plausible. The safe adapter preserves labels until it has proven the exact set and revision.

The adapter also does not refit. Prediction is a pure serving operation against the packaged fitted path. If an input fails the contract, the result is rejection, not imputation, renaming, filling, model adaptation, or a best-effort reorder.

Challenge names, dtypes and extra columns

A release gate needs negative cases that try to cross the boundary, plus evidence that the estimator was never invoked. Wrap the already-fitted Path C estimator with a call counter; this is a spy around prediction, not another fitted model.

class PredictCounter:
    def init(self, estimator):
        self.estimator = estimator
        self.calls = 0

    def predict(self, X):
        self.calls += 1
        return self.estimator.predict(X)

guarded_estimator = PredictCounter(path_c)

def guarded_predict(
    X,
    *,
    schema_revision: str,
):
    canonical = adapt_dataframe(
        X,
        schema_revision=schema_revision,
    )
    return guarded_estimator.predict(canonical)

np.testing.assert_allclose(
    guarded_predict(
        PERMUTED,
        schema_revision=SCHEMA_REVISION,
    ),
    [27.0],
    atol=1e-12,
)

Exercise all required failures:

negative_cases = {
    "unlabeled_array":
        np.array([[2.0, 7.0]]),

    "missing_adjustment":
        pd.DataFrame({"units": [2.0]}),

    "extra_debug":
        pd.DataFrame(
            {
                "units": [2.0],
                "adjustment": [7.0],
                "debug": [1.0],
            }
        ),

    "duplicate_names":
        pd.DataFrame(
            [[2.0, 7.0]],
            columns=["units", "units"],
        ),

    "numeric_names":
        pd.DataFrame(
            [[2.0, 7.0]],
            columns=[0, 1],
        ),

    "string_value":
        pd.DataFrame(
            {
                "units": [2.0],
                "adjustment": ["7"],
            }
        ),

    "nan_value":
        pd.DataFrame(
            {
                "units": [2.0],
                "adjustment": [np.nan],
            }
        ),

    "inf_value":
        pd.DataFrame(
            {
                "units": [2.0],
                "adjustment": [np.inf],
            }
        ),
}

for label, bad_input in negative_cases.items():
    before = guarded_estimator.calls

    try:
        guarded_predict(
            bad_input,
            schema_revision=SCHEMA_REVISION,
        )
    except ContractError:
        assert guarded_estimator.calls == before
    else:
        raise AssertionError(
            f"{label} unexpectedly reached predict"
        )

The pinned execution rejected every case with the call counter remaining 1->1: the one prior call was the valid permuted DataFrame, and no negative case incremented it. The observed reasons were, respectively, unlabeled input, missing feature, extra feature, duplicate names, non-string names, object/string dtype, and nonfinite values. A wrong schema revision was also rejected before prediction.

This pre-predict proof is stronger than merely asserting that a bad request eventually raised. An estimator could raise only after partially processing input, or a permissive preprocessor could silently discard the offending column. The spy establishes the application boundary itself stopped the request.

Capture the complete feature-to-coordinate ledger

When feature identity is the failure mode, “shape=(1,2)” is inadequate evidence. Record the mapping at each boundary in a ledger that can be reconciled with the deployed artifact.

Ledger field

Example in this fixture

Why it matters

incoming names

["adjustment","units"]

proves caller presentation order

incoming values

[[7.0,2.0]] with labels

preserves name/value association

canonical names

["units","adjustment"]

proves adapter order

canonical values

[[2.0,7.0]]

proves value-to-name mapping

transformed names

["units","adjustment"]

proves preprocessor output identity

shape / dtype

(1,2) / float64

secondary structural evidence

schema revision

feature-contract-v1

binds input to contract

preprocessor revision

named-column-selector-v1

binds ordering logic

model revision

linear-regression-fixture-v1

binds fitted coordinates

fixture revision

feature-order-fixture-v1

binds known-answer evidence

dependency lock

exact versions

scopes observed behavior

evidence ID

manifest SHA-256

identifies the acceptance record

This ledger is the operational version of avoiding common ML project pitfalls: rather than another generic pitfalls list, it ties one concrete defect to evidence at the serving boundary.

Do not turn the ledger into a confidential feature dump. This article uses only synthetic values. In production, log schemas, revisions, hashes, and bounded diagnostic samples according to privacy and security policy. The minimum useful evidence is enough to prove identity and ordering without retaining sensitive payloads unnecessarily.

Reconcile the ledger rather than merely collecting it. For the known-answer request, incoming names may be ["adjustment","units"], but the incoming labeled association must still be adjustment=7 and units=2. The canonical stage must then report ["units","adjustment"] with values [[2,7]]. The transformed stage must preserve that same order and values. If any stage reports the right shape but a different association, stop there; a later prediction equal to 27 could be accidental for another fixture.

Evidence identity should cover the configuration that gives the evidence meaning. A manifest hash is useful only if the unhashed manifest is retained too. Keep the hash, manifest contents, packaged artifact identifiers, script revision, and invocation together. A bare hash without the material it identifies is not reproducibility; it is only a checksum string.

The ledger is also the starting point for incident scoping. If historical records preserve caller revision and schema revision, the team can often determine which requests were exposed to positional ambiguity without replaying confidential raw features. When those records are absent, the uncertainty itself is evidence and should push the decision toward hold rather than retrospective certification.

A permutation test is not a semantic-units test

A labeled permutation test answers: “Given trustworthy names and values, does presentation order change the model input?” It does not answer: “Does the label truthfully describe the value?”

Suppose an upstream producer sends a number measured in a different physical unit but still labels it units. The adapter sees the expected string, numeric dtype, finite value, and schema revision. It cannot discover the semantic mismatch from the label alone. The same is true if the producer changes a business definition while retaining the old feature name.

Therefore store feature meaning and units alongside order, and version the schema when semantics change. Feature-name validation catches identity loss caused by missing, extra, duplicated, unlabeled, or reordered columns; it cannot certify upstream truth.

Keep training and serving contracts paired

The deployable unit is not just model.pkl. It is the fitted estimator plus the preprocessing path, adapter, feature manifest, dependency lock, fixture revision, warnings policy, and test invocation that jointly define how values reach coordinates.

A minimal manifest for this lab is:

{
  "fixture_revision": "feature-order-fixture-v1",
  "schema_revision": "feature-contract-v1",
  "preprocessor_revision": "named-column-selector-v1",
  "model_revision": "linear-regression-fixture-v1",
  "feature_order": ["units", "adjustment"],
  "warnings_policy":
    "capture-default-for-evidence; promote-to-error-at-release-gate",
  "test_invocation":
    "python feature_order_acceptance.py"
}

Add exact dependency versions, platform, feature meanings, and an evidence identifier derived from the manifest. In the observed run, the manifest evidence ID was sha256:b8f5732b5f428731a8c10687683ece566d3d574fe41ada65a102afdc15e0fdf9.

This pairing explains how an unchanged fitted artifact can become unsafe. A caller that previously produced ["units","adjustment"] might be refactored to build ["adjustment","units"] before calling to_numpy(). The model file hash can remain identical while the effective numerical interface changes. Artifact immutability is therefore necessary but not sufficient.

The contract order should be sourced from a versioned manifest, not rediscovered from whichever DataFrame happens to arrive. Conversely, do not treat the manifest as permission to reorder arbitrary unlabeled arrays. The manifest can safely reorder a labeled frame because name/value associations are still visible; for arrays, the original associations are already absent unless a separate trusted order contract accompanies them.

Persist the manifest mechanically:

import hashlib
import json
import platform

manifest = {
    "fixture_revision":
        "feature-order-fixture-v1",

    "schema_revision":
        "feature-contract-v1",

    "preprocessor_revision":
        "named-column-selector-v1",

    "model_revision":
        "linear-regression-fixture-v1",

    "feature_order":
        ["units", "adjustment"],

    "feature_meanings": {
        "units":
            "Synthetic quantity; "
            "unitless in this fixture.",

        "adjustment":
            "Synthetic additive adjustment; "
            "unitless in this fixture.",
    },

    "dependency_lock": {
        "python": "3.13.5",
        "scikit-learn": "1.8.0",
        "pandas": "2.2.3",
        "numpy": "2.3.5",
        "scipy": "1.17.0",
        "joblib": "1.5.3",
        "threadpoolctl": "3.6.0",
    },

    "platform":
        platform.platform(),

    "test_invocation":
        "python feature_order_acceptance.py",
}

payload = json.dumps(
    manifest,
    sort_keys=True,
).encode()

manifest["evidence_id"] = (
    "sha256:"
    + hashlib.sha256(payload).hexdigest()
)

The locally observed manifest hash for the recorded environment was sha256:b8f5732b5f428731a8c10687683ece566d3d574fe41ada65a102afdc15e0fdf9. That identifier is evidence for this synthetic run only; copying it into a different package would not make the new package equivalent. Generate a new evidence identity whenever the manifest changes.

Repair an adapter without silently retraining

When evidence shows a positional caller is wrong, repair the boundary first. Do not change model coefficients simply to make the wrong order produce familiar outputs.

The bounded repair sequence is:

  1. Hold the affected serving path.

  2. Preserve the failing request representation, caller revision, model revision, schema revision, warning trace, and known-answer evidence.

  3. Restore labeled input at the earliest trustworthy boundary.

  4. Validate exact names, dtypes, finiteness, and schema revision.

  5. Reorder only after validation, using the contract order.

  6. Re-run the known-answer case in canonical and permuted DataFrame orders.

  7. Re-run every negative case and prove zero additional estimator calls.

  8. Assess which historical predictions could have used the wrong positional order.

Historical assessment should be evidence-driven. If request logs or caller revisions prove an interval used [adjustment, units] where [units, adjustment] was expected, mark that interval affected. If the historical order cannot be established, do not retroactively declare predictions correct from shape alone.

“Repair” here means repairing the adapter or caller while preserving the frozen fitted artifact. “Rebuild” is reserved for cases where the authoritative original feature mapping or deployable package cannot be recovered with confidence. That decision may require reconstructing a new governed bundle from authoritative sources; this article intentionally does not prescribe a retraining strategy.

The stop condition is simple: if feature identity cannot be proven, inference remains on hold. Availability pressure is not evidence that coordinate zero suddenly means the right thing.

Release the guarded path with rollback evidence

Notebook success is insufficient. Acceptance must run against the packaged artifact and the same entry point deployment will call. This is where model validation practices remain useful context, but the release criterion here is narrower: preservation of feature identity at inference, not another holdout or cross-validation exercise.

The release package must pass four groups of checks. First, the independent equation gives 27. Second, canonical and permuted labeled requests both return 27 through the guarded Path C. Third, the transformed names and values are exactly ["units","adjustment"] and [[2,7]] for both presentation orders. Fourth, every negative case is rejected before estimator invocation, including the wrong schema revision.

Run the script under the locked dependencies and persist the manifest, exit code, captured warnings, exception text, and artifact identities. A release reviewer should be able to distinguish documented behavior from official API references, mathematically expected results from the fixture equation, and actually observed output from the pinned execution. Do not relabel an expected trace as observed if the packaged test was not executed.

Stage the change behind the narrowest integration boundary that owns input adaptation. That limits rollback scope and avoids teaching unrelated systems about model internals.

Package-level acceptance should also test the boundary that deployment actually exposes. If a service wrapper deserializes JSON into a DataFrame, test that wrapper; if a batch job reads a file and constructs columns, test that constructor. The notebook objects are only reference components. A common failure pattern is that the notebook preserves DataFrame labels while the packaging layer converts to an ndarray for convenience. The acceptance script must sit after that convenience code, not before it.

Persist the negative-case outcomes as named assertions rather than a single Boolean. Reviewers should be able to see that missing_adjustment, extra_debug, duplicate_names, numeric_names, string_value, nan_value, inf_value, unlabeled arrays, and wrong schema revisions each failed before predict. A future refactor may accidentally weaken one case while leaving the aggregate test green if the test suite is too coarse.

The release artifact should therefore carry both positive and negative known answers. Positive evidence proves that valid labeled permutations converge to the same coordinate mapping. Negative evidence proves that ambiguous or malformed requests do not get a chance to produce plausible-looking numbers.

Rollback must keep feature identity enforceable

A rollback target must be a previously verified model/adapter/manifest pair. Rolling back only the model while leaving a positional caller in place recreates the same class of defect.

If no verified pair is available, hold inference. Do not restore availability by bypassing the adapter, downgrading contract errors to warnings, or accepting arbitrary ndarrays. The rollback objective is not merely “old code”; it is a previously accepted end-to-end mapping from feature identity to fitted coordinate.

Store rollback evidence before release: artifact IDs, schema revision, dependency lock, known-answer output, negative-test results, and the exact invocation. That makes rollback an executable state transition rather than an emergency guess.

Apply release, repair, hold and rebuild decisions

Use the evidence to make one of four explicit decisions. “Release” means the packaged guarded path meets the contract. “Repair” means the fitted artifact is still usable but an adapter/caller defect has a bounded fix. “Hold” means identity is unresolved or a required gate failed. “Rebuild” means the original authoritative mapping or bundle cannot be trusted or recovered; it does not mean silently retraining inside a serving incident.

Condition

Required evidence

Decision

Accountable owner

Required action

Named input, both DataFrame orders return 27; transformed names/values match; all negatives stop pre-predict

packaged known-answer and spy assertions pass

Release

serving release owner

deploy guarded pair after model-owner and feature-owner sign-off

Wrong-position ndarray path is identified; authoritative names/order are recoverable; fitted coefficients remain valid

failing 72 control plus recoverable caller/contract mapping

Repair

serving input-contract owner

restore labels, validate, reorder, rerun full gate

Warning, schema mismatch, missing/extra/duplicate names, nonnumeric/nonfinite input, or unexplained prediction mismatch

any acceptance assertion fails

Hold

release owner

stop inference on affected path; preserve evidence

Historical or current feature order cannot be established from trustworthy artifacts

no authoritative mapping to fitted coordinates

Rebuild

model artifact owner

reconstruct governed bundle from authoritative source; keep service held until accepted

Previously verified pair exists and current rollout fails

rollback manifest and known-answer evidence are intact

Roll back

deployment owner

restore the complete verified pair, not model alone

Historical predictions may have used wrong order but interval is uncertain

incomplete request/caller evidence

Hold assessment

feature owner

bound uncertainty; do not certify unknown records

The release gate can be represented as a small decision flow:

post-release

The operational discipline behind this matrix is consistent with strong Python engineering foundations: explicit interfaces, versioned dependencies, assertions, and reproducible execution matter as much as the estimator call itself.

Two sign-offs are mandatory for release. The model owner confirms the packaged fitted identity, coefficients, fixture, and dependency lock. The feature owner confirms the schema names, order, meanings, and revision. The serving release owner remains accountable for the deployment decision and rollback readiness.

Make the difference between repair and rebuild explicit in the ticket. A repair is justified when trustworthy evidence says the fitted artifact’s intended mapping is known and only the serving boundary violated it. The repaired package must reproduce 27 for both labeled orders without changing the coefficients to “fit” the bug. A rebuild is required when the authoritative mapping itself is missing, contradictory, or inseparable from an untrusted package. In that state, there is no safe basis for asserting what coordinate zero was supposed to mean.

Historical impact gets a similarly conservative decision. When caller code, request schemas, or serialized payloads establish order, affected predictions can be bounded. When only shape is known, the correct state is unknown, not “probably fine.” That distinction is uncomfortable but operationally important: uncertainty should survive into the incident record instead of being erased by a retrospective assumption.

The matrix should be executable as policy. Every row maps an evidence state to an action and owner. That prevents a release meeting from inventing new acceptance standards after seeing the result. It also keeps business availability pressure from silently changing a technical contract: if identity cannot be proven, the route is hold or rollback, not best-effort prediction.

The strongest stop condition is unresolved identity. A numeric prediction is not acceptable merely because no exception occurred. In this fixture, 72 is a perfectly ordinary floating-point number produced by a perfectly valid matrix multiplication. Acceptance comes from proving that each scalar reached the coordinate it was trained to occupy.

Turn input-contract testing into durable ML practice

Make this test part of the artifact lifecycle, not an incident-only notebook. Every change to the adapter, schema, preprocessor, dependency lock, serialization package, or request-construction code should rerun the known-answer fixture, both DataFrame permutations, transformed-name assertions, and every pre-predict rejection case. Preserve the evidence record with the deployable pair.

Keep the claim narrow. This lab demonstrates one failure mode: same shape, different feature identity. It does not validate training data quality, leakage, cross-validation design, drift, categorical encoding, feature stores, GPU serving, or model-format migration. It also cannot detect a semantically wrong value that carries the correct name. Those controls belong elsewhere.

For readers building the broader foundations around this kind of acceptance work, Refonte Learning’s Data Science & AI page currently lists Python, Pandas, NumPy, and scikit-learn among its tools, and it includes machine learning/predictive modeling and industry-project work. Those confirmed foundations are relevant to understanding this playbook; the page does not establish that this exact feature-order serving lab is a taught module, so no such claim should be inferred.

The durable rule is simpler than any framework choice: preserve feature identity until you have validated it, bind the contract to the fitted artifact, and refuse to predict when the mapping cannot be proven.