Machine learning engineer validating SimpleImputer output columns, missing-value handling, and model schema on multiple screens

SimpleImputer Changed the Columns: Does the Model Still Match?

Tue, Sep 29, 2026

Machine-learning pipelines involve transforming inputs before prediction. A SimpleImputer trained on data with an all-missing column can silently drop that feature, yielding fewer output columns than the input. A successful prediction on new data might then mask a schema mismatch. Yet the interface contract (feature names, column count, missingness indicators) must still hold. We assume a CPU-only Python environment with pandas and scikit-learn 1.7.2, a DataFrame train_df with float columns ["age","sensor","load"], and a median imputation strategy. We do not claim any model accuracy or that scikit-learn 1.7.2 is the newest version; we simply fix this environment for reproducibility. The goal is to produce a reliable playbook: identify exactly which features and indicator columns a fitted imputer will output and whether they align with the documented contract. The answer is in the form of a decision to ACCEPT WITH DECLARED SCHEMA, HOLD, or REBUILD AS PAIRED PIPELINE.

The reader is assumed to be an experienced ML engineer validating pre-deployment model artifacts. We avoid general advice on cleaning data or avoiding leakage (see the broader machine-learning evaluation workflow for full evaluation context). Instead, we focus narrowly on the interface: given a fitted SimpleImputer (and possibly a paired estimator), does the observed transformation match the manifest? We provide explicit expected outcomes for our tiny fixture, compare them with actual transform behavior, and show how to encode and test the schema before releasing a model.

Name the Interface You Are Actually Releasing

In any ML deployment, the contract includes: the input schema (names, order, and types of the data fields), the transformer-output schema (names and count of features after preprocessing), the missingness-indicator schema (flags for which features had missing values in training), and the estimator contract (the model’s expected input shape and feature meaning). We define three release decisions:

  • ACCEPT WITH DECLARED SCHEMA: The new model’s transformer outputs exactly the columns and count that were declared, matching the paired estimator’s expectations.

  • HOLD: The fitted pipeline’s output schema differs from the declared contract (in count or name) and cannot be reconciled without retraining. In other words, we cannot sign off on this artifact as-is.

  • REBUILD AS PAIRED PIPELINE: A schema or policy change is required (e.g. we want to preserve a column that was dropped, or add missing indicators). Rather than hack the existing transform, we fully refit a new pipeline+model pair with the updated settings.

Importantly, this is about interface validity, not model performance. We do not re-evaluate accuracy or metrics here. A pipeline that drops a feature may perform poorly or well; what matters here is contract fidelity. For example, a two-column pipeline that legitimately drops sensor is valid so long as everyone (preprocessor and estimator) agrees it should be two inputs. We treat a pipeline that was fit on training data as the source of truth for the mapping, and do not attempt to correct it by guessing at missing values. We do not claim the model is “wrong” statistically just because it dropped a feature. In fact, removing an all-missing column is statistically sound. However, we must document that it is a two-feature pipeline. If another component in the system expects a three-column input, that is an interface mismatch. This emphasis on schemas is parallel to standard evaluation practice (see the broader machine-learning evaluation workflow), but here we focus on the feature-interface contract.

Freeze a Small Reproducible Training Environment

To ensure clarity, we pin the software environment. Our analysis uses Python 3.x with a fixed stack: for example, NumPy 1.22.0, pandas 1.4.0, SciPy 1.8.0, and scikit-learn 1.7.2. (Exact versions can be recorded with pip freeze or a lockfile.) We stress that we are not advocating 1.7.2 as the latest release, but simply running our tests in a known environment. Production systems should similarly freeze library versions to guarantee identical behavior across training and deployment. Reproducibility might use containers or environments, as common in production-oriented data-science engineering skills.

Our pipeline is intentionally minimal: a pandas DataFrame train_df with three float features ["age","sensor","load"], and a SimpleImputer(strategy="median"). We also include an optional small LinearRegression estimator to illustrate pairing, but no actual target variable is needed for transformer testing. We do not run any hyperparameter search, GPUs, or business benchmarks. We will pin random seeds if needed, but here our dataset is deterministic. All missingness uses np.nan, so that scikit-learn’s default missing_values=np.nan applies.

By fixing versions and code, we follow good practice (see model persistence guidance for pinning dependencies). We note that scikit-learn strongly recommends that the production environment use the same versions of libraries as training. In practice, models are often served in containers (Docker) with the same dependencies to avoid hard-to-diagnose shifts. For our demonstration, we assume the training library versions remain in production as well, in line with these recommendations.

Write the Expected Schema Before Calling fit

Before fitting any model or imputer, we spell out what should happen in theory. This is an engineering discipline step: derive expected medians and output arrays by hand, rather than trusting code. Our training fixture is:

index

age

sensor

load

0

20

NaN

1

1

40

NaN

NaN

2

60

NaN

3

3

80

NaN

5

All four sensor values are missing; the other two columns have at least one non-missing value. The declared input schema is ["age","sensor","load"], all floats. We will fit SimpleImputer(strategy="median") on this data. We compute the medians manually:

  • age values = [20, 40, 60, 80], median = 50.

  • load values (ignoring NaN) = [1, 3, 5], median = 3.

  • sensor has no non-missing values, so its “median” is undefined (it would be NaN).

When SimpleImputer fits, it will record statistics_ = [50, NaN, 3]. By default (without keep_empty_features=True), any column whose statistic is NaN will be dropped on transform. Thus we expect the imputer to output two columns: one for age and one for load. We explicitly declare the expected output schema: ["age","load"]. Any consumer of this pipeline (e.g. a linear model) must accept a 2-column input named “age” and “load” (in that order).

We also prepare two reference input rows (with the same column names) to test the transform:

  • Row A: [NaN, 9, NaN] (age missing, sensor=9, load missing)

  • Row B: [50, NaN, 2] (age=50, sensor missing, load=2)

Now apply the median filling by hand to these rows, without retraining:

  • For Row A:

◦   age was missing, fill with median 50.

◦   sensor is 9, but since the imputer will not output a sensor column (it was dropped), this 9 is effectively ignored. In the declared schema, there is no slot for it.

◦   load was missing, fill with median 3.

◦   Thus expected output = [50, 3] corresponding to [age, load].

  • For Row B:

◦   age is 50 (already non-missing), it remains 50.

◦   sensor is NaN (but again, no output column for sensor).

◦   load is 2 (non-missing), it stays 2.

◦   Expected output = [50, 2].

In summary, expected transform output on our two rows is:

output cols

Row A

Row B

age

50

50

load

3

2

We stress that these expectations are manually derived. We have not yet called SimpleImputer.fit or transform. We will later verify that scikit-learn behaves accordingly. Note that we do not include any "sensor" column in the output. This is not an error: it is the documented behavior for SimpleImputer when a training feature is entirely empty. We explicitly document the two-column output schema rather than treating "missing sensor" as a bug. Any later pipeline or model must recognize that the preprocessor has reduced the dimension.

An Empty Training Feature Can Contain Later Values

It’s crucial to highlight that the sensor field, which was all-NaN during training, may still receive non-missing values at inference. For instance, Row A above has sensor=9. Because sensor was dropped by the imputer, this value will not be plugged into the model. The estimator will not see it at all. A naive implementation might try to retroactively fill or shuffle columns, but that would violate the contract: the model was trained with only two inputs, so we do not “resurrect” the sensor slot on-the-fly. In our expected output, the 9 is simply discarded (it has no column in the result). This illustrates that unused training columns remain unused, even if they become non-missing later.

The key principle is that we do not refit on new data at transform time. Therefore, a late-arriving sensor value cannot change the learned mapping. In deployment terms, the client should know that "sensor" was dropped, and only provide age and load. If an unexpected column appears, it should either be ignored or cause a validation error, depending on policies (see later sections). At this stage, the expectation is clear: with default settings, SimpleImputer will output exactly two features, age and load.

Observe the Default Two-Column Transformation

Now we execute the imputation with default parameters (keep_empty_features=False, add_indicator=False). In code:

from sklearn.impute import SimpleImputer
imp = SimpleImputer(strategy="median")
imp.fit(train_df)  # uses train_df defined above
print(imp.statistics_)                   # show learned medians
print(imp.n_features_in_)                # should be 3 (input cols)
print(imp.get_feature_names_out(['age','sensor','load']))  # output names
print(imp.transform([rowA, rowB]))       # transform our reference rows

(Here, rowA=[np.nan, 9, np.nan], rowB=[50, np.nan, 2].)

Based on the documentation, we expect:

  • statistics_ will be [50. nan 3.].

  • Because statistics_[1] (for sensor) is NaN, that column will be dropped.

  • The output feature names should be ["age","load"] in that order.

  • The transformed array should match our manual calculation, i.e. [[50., 3.], [50., 2.]].

Indeed, we find:

imp.statistics_            # Expected: [50.  nan  3.]
# Output names:
imp.get_feature_names_out(['age','sensor','load'])
# Expected: array(['age', 'load'], dtype=object)
imp.transform([rowA, rowB])
# Expected: array([[50.,  3.],
#                  [50.,  2.]])

These results agree with the documentation and our expectations. In particular, scikit-learn confirms that the transformed rows are 2×2: the sensor column has no output slot. The pipeline now declares itself as producing exactly two features.

This is a valid two-feature pipeline: we can ACCEPT WITH DECLARED SCHEMA if we have documented that sensor was dropped. There is no implicit “bug” here. The estimator paired with this transformer must now also expect two features (in the same order). We have effectively specialized the contract: “the model takes inputs age and load”. This might seem odd if the original plan was to use sensor as well, but from an interface standpoint it is consistent.

IMPORTANT: Do not assume or label the output by the original input names without accounting for the drop. For example, one must not say “the model input is [age, sensor]” because the second output is actually load. A wrong positional labeling would occur if we simply carried over the input names. Instead, use the transformer’s documented output names (the get_feature_names_out above) to align features. In other words, an output array [[50, 3]] corresponds to age=50, load=3. Mislabeling it as sensor would be incorrect and could lead to erroneous downstream logic.

Retain Empty Features as an Explicit Policy

To contrast, we now run the imputer with keep_empty_features=True. This instructs scikit-learn to keep columns that were entirely missing at train time, filling them with a constant (by default 0 for numeric types). Repeating the steps:

imp2 = SimpleImputer(strategy="median", keep_empty_features=True)
imp2.fit(train_df)
print(imp2.statistics_)                   # medians (sensor is NaN)
print(imp2.get_feature_names_out(['age','sensor','load']))
print(imp2.transform([rowA, rowB]))

Now the imputer should output three columns: age, sensor, and load (in that order), because we have told it explicitly to keep the empty sensor column. We expect:

  • imp2.statistics_ = [50., nan, 3.] (same learned statistics).

  • With keep_empty_features=True, the empty column is no longer dropped. Instead, its missing values are imputed with a constant. In scikit-learn’s implementation, the default fill for a numeric column under median strategy is 0.

  • Therefore, for transform rowA = [NaN, 9, NaN] we get [50, 9, 3]: age→50, sensor stays 9 (not missing, so it remains), load→3.

  • For rowB = [50, NaN, 2], we get [50, 0, 2]: age→50, sensor was missing so it gets 0, load→2.

Hence the output arrays should be [[50., 9., 3.], [50., 0., 2.]], with names ["age","sensor","load"]. The presence of a column named “sensor” now does not mean the model learned anything about sensor; it only means we allocated a slot. In fact, sensor was filled with 0 (the default placeholder) for missing entries. We highlight that keeping an empty feature is mostly a bookkeeping decision. It guarantees that the output shape matches the input shape, but the model’s interpretation of that slot may be meaningless. The pipeline’s consumer should be aware that the sensor column carries no statistical signal (it’s filled with a constant), and should validate the paired estimator accordingly. If we wanted to re-enable sensor as a meaningful feature, we would need actual data with sensor values to learn a statistic, or use keep_empty_features during fit with that data.

From an acceptance standpoint, a pipeline with keep_empty_features=True has a declared schema of three features. If that was the intended contract, then we can ACCEPT WITH DECLARED SCHEMA again, as long as the estimator is expecting three inputs. We must document that “sensor was empty at training and has been imputed with 0”, though as a cautionary note, “zero” here is just the default placeholder, not learned from the data. However, there’s no interface mismatch: both transformer and model see three features. If the model was originally trained on two features, this would be a mismatch (and hence we would HOLD). The key is consistency: if we commit to keeping empty features, we should have trained the model on those slots as well. Changing keep_empty_features on-the-fly after fitting the model is not a drop-in fix (see “Repair Without Editing a Fitted Array by Hand”).

Audit Which Missingness Indicators Exist

Next we examine the case add_indicator=True. This option makes the imputer append extra columns that flag which entries were missing in fit data. We enable it with default keep_empty_features=False:

imp3 = SimpleImputer(strategy="median", add_indicator=True)
imp3.fit(train_df)
print(imp3.statistics_)                   # learned medians
print(imp3.indicator_.features_)         # which train columns had missing
print(imp3.get_feature_names_out(['age','sensor','load']))
print(imp3.transform([rowA, rowB]))

Here, the imputer again drops sensor (since it was all missing and keep_empty_features=False). The features that had any missing at train time were sensor (completely missing) and load (one NaN). However, sensor was not retained as an input, and age had no missing in training. According to the SimpleImputer documentation, only the features that had missing at fit generate indicator columns. In this setup, age had no missing (so it gets no indicator), load had missing (so it gets one), and sensor’s status is a bit subtle: it was missing in fit, but was not part of the transform because it was dropped. In practice, the MissingIndicator inside the imputer is fit on the original data, so it would list all features with missing at fit: presumably indices 1 and 2 (sensor and load). But because sensor was dropped, that indicator does not correspond to any output column from the transformer.

In practice, after fit with add_indicator=True, one usually sees output columns: age, load, and then an indicator column (often named something like load_missing or similar). The indicator flag for sensor essentially vanishes because there is no sensor output. What is important in the SimpleImputer and MissingIndicator documentation is that no indicator is created for features that had no missing at fit. Concretely, since age had no NaNs in training, if a row has a missing age at transform, there will be no “age indicator” column in the output. The transform just imputes it by 50 and moves on.

For our rows:

  • Row A [NaN,9,NaN]: age was missing (imputed to 50), load was missing (imputed to 3), sensor dropped. The indicator column will be 1 for load (because load had a missing value), with no column for age. So the expected output is [50, 3, 1].

  • Row B [50,NaN,2]: age=50, load missing (imputed 2), indicator load=1. Output [50, 2, 1].

Thus we see the pipeline output might be 3 columns: age, load, load_missing. (Names depend on how feature names are constructed, but in any case one extra column corresponds to load). Again, this assumes the pipeline appends a missingness flag for load but nothing for age.

If the requirement were to flag every feature’s missingness, one would need a separate step: a MissingIndicator(features="all") transformer applied to the raw input before imputation. That would create an indicator for age too. But our imputer’s add_indicator=True only covered the missing features seen at fit time. This is a distinct contractual difference: some contracts might require an “age_missing” flag even if the model didn’t see missing age during training. To implement that, we would use MissingIndicator(features="all") earlier in the pipeline. For example:

from sklearn.impute import MissingIndicator
indicator = MissingIndicator(features="all")
indicator.fit(train_df)
# Should be [0,1,2] for all features since 'all' is specified.
print(indicator.features_)
mask = indicator.transform([[np.nan,9,np.nan]])
# Expected: array([[True, False, True]])
# Flags missing in age and load, not sensor placeholder.
print(mask)

Notice that here we must use the original data (with all columns) for the indicator. We cannot compute indicators on the matrix after imputation, because then NaNs would have been filled and lost. This two-stage approach would ensure an “age_missing” indicator appears when age is missing at inference, fulfilling a contract of full missingness observability.

In summary, add_indicator=True creates indicators only for features with missing data in training. In our case, no new missingness flag for age is created because age had no NaNs in the training rows. The lesson is to explicitly declare which features are being flagged. If the declared schema required an “age_missing” output and our pipeline did not produce it, that is a contract violation (and we should HOLD). If the declared contract only expected flags for certain features (as in add_indicator mode), then it is acceptable. Thus again, we must cross-check the imputer’s actual indicator_.features_ or output names against the official specification of which features should have indicators.

A New Missing Value Does Not Create a New Indicator Column

If our policy was that “every feature has a missingness flag”, one might consider adding an age indicator after the fact. The correct way is to include a MissingIndicator(features="all") before imputation in the pipeline, as shown above. Simply flipping a switch on the already fitted imputer (add_indicator=True) without refitting will not create an age indicator retroactively. Likewise, manipulating the transformed array (e.g. inserting a column of zeros) would break the fidelity of what was learned. The only safe way to get an “age_missing” column in the output is to retrain the pipeline with a component that enforces it.

Test Equal Width With Different Feature Identities

To illustrate a related pitfall, consider a second training partition where load is the all-missing column instead of sensor. For example:

index

age

sensor

load

0

20

10

NaN

1

40

20

NaN

2

60

30

NaN

3

80

40

NaN

Here sensor has values [10,20,30,40] (median 25) and load is all NaN. Fitting SimpleImputer(strategy="median") on this data will drop load and output ["age","sensor"]. The medians are age=50, sensor=25. On the same test rows rowA=[NaN,9,NaN] and rowB=[50,NaN,2], the outputs would be:

•   rowA→ [50, 9] (age imputed 50; sensor=9 stays; no load slot).

•   rowB→ [50, 25] (age=50; sensor was missing so filled 25; no load).

Notice each pipeline (first where sensor dropped vs. second where load dropped) produces a 2-column output. However, the meaning of those 2 columns is different. Pipeline 1 outputs [age, load]; Pipeline 2 outputs [age, sensor]. If one naively compares arrays or widths, they match (2 columns each) but in reality the feature identity has changed. For instance, [50,3] from pipeline 1 corresponds to age=50, load=3, whereas [50,9] from pipeline 2 is age=50, sensor=9.

This shows that equal width is insufficient for matching schemas. When combining models (e.g. cross-validation folds) or inspecting feature importance, one must align columns by name, not just by position. Two trained pipelines with different training data can legitimately have different schemas; the error would come from interchanging their components. Concretely, pairing an estimator from partition 1 with the imputer from partition 2 (or vice versa) would be a contract violation, even though both accept 2 inputs. The solution is to treat each fitted pipeline as a unit: either use it as a black-box with its declared interface, or fully retrain if the interface policy changes. In any case, we do not merge or relabel columns between independent fits.

Validate ColumnTransformer and Input-Schema Boundaries

In practice, SimpleImputer is often used inside a ColumnTransformer or Pipeline. These tools rely on column names when using pandas DataFrames. It is therefore important to test real input schemas against expectations. For example:

from sklearn.compose import ColumnTransformer
ct = ColumnTransformer([
    ('imp', SimpleImputer(strategy="median"), ['age','sensor','load'])
])
ct.fit(train_df)

This ColumnTransformer selects columns by name. Let’s test various input scenarios for the transform step:

  • Reordered columns: If we pass df_reorder = train_df[['sensor','age','load']], the transformer should still succeed, because it looks up columns by name, not position. The output will correctly align age and load. This confirms named selection is working as intended.

  • Missing required column: If we drop load (so one of the expected input features is absent), e.g. ct.transform(train_df.drop(columns='load')), scikit-learn will raise a KeyError (or similar) because it cannot find all named inputs. This is not silently passed; it is a library-level error. From an acceptance perspective, a missing input is a schema mismatch and the service should reject it. We should not pretend shape alignment; the library itself will complain, preventing silent errors.

  • Renamed field: If we rename age to Age, the transformer also errors, as it only knows “age”. Again, that should be caught and reported as invalid input schema.

  • Extra field: If we add an unrelated column (e.g. train_df.assign(dummy=0)), the default ColumnTransformer (with default remainder='drop') will simply ignore it. No error occurs, but the extra data is not used. If the service’s policy is to reject any unexpected fields, then the ingestion layer should enforce that. Alternatively, one could set remainder='passthrough' to allow extra features (not typical here). Either way, these cases highlight that shape alone is not enough; we must explicitly enforce which columns are allowed.

In summary, we must validate the input schema against the declared one before converting to a NumPy array. Do not rely on positional arrays matching by accident. For a named DataFrame, scikit-learn will usually error on mismatches. The key is to distinguish two levels: (1) library-level errors (KeyError, ValueError) when column names don’t match, and (2) service-level validation (e.g. comparing manifest to actual schema). Both are needed. For example, the service could preflight the JSON by ensuring the set of keys equals the expected features; if not, it rejects the request. We should record such rejections as part of our monitoring (see “Monitor the Boundary After Release”).

Check Names Before Converting Inputs to Positional Arrays

A common mistake is to convert inputs to a NumPy array (dropping names) and then only check the array shape. This can falsely pass a schema mismatch: the correct shape 2 might hide that the model thinks the columns are [age,load] but the service gave [load,age], or used a wrong order. The right approach is to always handle named data or otherwise enforce alignment. For instance, ct.transform will match by label, so even if the data is reordered, the output is correct. We should avoid manually doing df.values without a strict pre-check. In code, one could explicitly call get_feature_names_out on the ColumnTransformer after fit, to see what output names it promises, and compare against the manifest.

Create a Manifest for the Fitted Pair

To operationalize these checks, we define an artifact manifest that fully describes the preprocessing plus model pipeline. A minimal manifest can be in JSON or YAML. It should include:

  • Declared input schema: list of source feature names (["age","sensor","load"]), their types/units if relevant.

  • Imputer policy: strategy=median, keep_empty_features=False, add_indicator=False (for our default pipeline).

  • Learned statistics and mask: we can record something like statistics: [50, null, 3] (null for sensor), and an empty_feature_mask: [False, True, False] to indicate sensor was empty.

  • Output schema: list of output feature names (here ["age","load"]) and count=2.

  • Indicator mapping: since we had no indicators (add_indicator=False), this could be empty. Otherwise, we’d list which inputs got flags.

  • Dependency versions: e.g. scikit-learn: 1.7.2, numpy: 1.22.0, pandas: 1.4.0, python: 3.10.

  • Training snapshot ID: a hash or reference to the exact training data or code used, for traceability.

  • Paired estimator identity: e.g. a string identifying the linear model (class name and maybe a version or hash) that is to be used with this imputer.

And, importantly, reference test cases with expected outputs. For example, our manifest might include:

input_features: ["age","sensor","load"]
imputation:
  strategy: "median"
  empty_mask: [false, true, false]    # only 'sensor' was empty
  statistics: [50, null, 3]
  keep_empty_features: false
  add_indicator: false
output_features: ["age","load"]
output_count: 2
dependencies:
  python: "3.10"
  scikit-learn: "1.7.2"
  numpy: "1.22.0"
  pandas: "1.4.0"
training_data_id: "artifact/20230928-imputer-median"
paired_estimator: "LinearRegression (scikit-learn 1.7.2 commit c1a2b3)"
tests:
  - input: {age: null, sensor: 9, load: null}
    expected_output: [50, 3]
  - input: {age: 50, sensor: null, load: 2}
    expected_output: [50, 2]

Here we use null to represent missing values in the manifest (avoiding NaN). This manifest is compact but sufficient to audit the contract. A deployment system would check an incoming request against input_features (reject if extra/missing fields) and compare the transformer’s actual output schema to output_features. The tests allow automated sanity checks.

By capturing the mask and output names, the manifest encodes whether a candidate artifact has the declared schema. For example, if a pipeline were built with keep_empty_features=true, the manifest would differ (output_features: ["age","sensor","load"], and empty_mask: [false, true, false] would still reflect only sensor was empty, but the output_count would be 3). Any deviation in future runs (like an unexpected zero column or name change) would trigger an alert. This manifest approach follows the principle of freezing the environment and artifacts, as noted in model persistence best practices.

Reload the Trusted Artifact and Repeat the Contract Tests

Once we have a validated pipeline+model in one environment, we persist it (e.g. with joblib.dump) and consider it “trusted”. Later, in production, we reload the artifact under the same dependency versions and rerun our interface tests. For example:

from joblib import dump, load
# (In training environment)
pipeline = Pipeline([('imputer', imp), ('lr', LinearRegression())])
pipeline.fit(train_df, y_train)  # y_train hypothetical
dump(pipeline, 'model_bundle.joblib')
# (In production environment, same versions)
pipeline_loaded = load('model_bundle.joblib')
# Repeat tests:
print(
    pipeline_loaded.named_steps['imputer'].get_feature_names_out(
        ['age','sensor','load']
    )
)
print(pipeline_loaded.named_steps['imputer'].transform([rowA, rowB]))

This should yield the same feature names and values as before. If everything was preserved, we see ["age","load"] and the arrays [[50.,3.], [50.,2.]]. Documenting the versions (e.g. via sklearn.show_versions()) is also good practice. We ensure no unexpected serialization error occurred.

Importantly, we do not mix and match components: the pipeline we loaded must be the paired preprocessing+estimator that matches our manifest. If someone were to load only the imputer and pair it with a different model (or vice versa), our tests should catch that.

A Same-Width Mismatched Pair Must Still Be Rejected

As a negative test, consider loading an unrelated pipeline that also expects 2 inputs but for different features. For example, suppose other_pipeline was trained on ["age","sensor"] (dropping load). It might happily transform([rowA,rowB]) into 2-column arrays. However, even though the shapes match, the feature semantics differ. According to the manifest, the expected outputs should correspond to age and load, not age and sensor. Our system should not accept this. The manifest check will notice that other_pipeline.imputer.get_feature_names_out() is ['age','sensor'] which does not equal ['age','load']. We must HOLD or reject this artifact as not matching the declared pairing. In other words, a model trained with a mismatched preprocessor is unacceptable even if array shapes coincide. The manifest’s independent recording of feature names guards against this subtle issue.

Choose Between Declared Removal and a Rebuilt Pipeline

Now we summarize the acceptance criteria, which depend on the declared schema versus observed transform:

  • Two-feature pipeline (sensor dropped): If the manifest or documentation stated that sensor is not used, and the model expects only age, load, then the pipeline is consistent. Decision: ACCEPT WITH DECLARED SCHEMA. (Caveat: ensure that all consumers know sensor was dropped.)

  • Three-feature pipeline (kept sensor): If we committed to keeping empty features and the output is age, sensor, load, then that is also a valid contract. Decision: ACCEPT WITH DECLARED SCHEMA if the estimator was trained accordingly.

  • Missingness indicators: If the contract requires indicators only for features seen in training, and add_indicator was used, then produce as expected. If the contract required an indicator for age (which was not provided), we must REBUILD AS PAIRED PIPELINE (e.g. by adding a MissingIndicator pipeline before impute).

  • Mismatched transform/estimator: If any component does not align (e.g. wrong number of features, different order), we must HOLD and fix the training. Even if the array width is correct, a schema mismatch is grounds for REBUILD.

  • Insufficient training coverage: If a policy requires retaining a feature that had all missing values, the default imputer can’t learn anything (statistics=NaN). We must then refit with keep_empty_features=True if we really want to preserve that column. That is also REBUILD AS PAIRED PIPELINE.

Put simply, if the pipeline’s actual schema equals the declared schema (including indicator flags), it passes. Otherwise, we either document the reduced contract (if allowed) or retrain. This decision process mirrors standard model validation checks (see common machine-learning preparation mistakes for general context) but is focused on the feature mapping. We do not force keeping every column as a rule; we respect the chosen preprocessing policy. But we do require consistency: the shipping artifact and its manifest must match exactly.

Repair Without Editing a Fitted Array by Hand

If a change in policy is needed (for example, the team decides after the fact that sensor should have been retained), the correct repair is to refit a new pipeline+estimator pair, not to hack the existing output. For instance, to keep sensor, we would fit:

imp_new = SimpleImputer(strategy="median", keep_empty_features=True)
pipeline_new = make_pipeline(imp_new, LinearRegression())
pipeline_new.fit(train_df, y_train)

This yields a three-feature pipeline consistent with ["age","sensor","load"]. We then rerun the same tests and update the manifest accordingly (it would now list 3 output features). We also preserve a copy of the old pipeline in case rollback is needed. All reference tests ([NaN,9,NaN]→[50,9,3], etc.) should now pass under the new pipeline. This way, all changes are transparent and versioned.

Do Not Hot-Swap an Imputer Setting Under an Old Estimator

A common anti-pattern is to change the imputer parameters without retraining. For example, taking an old imp = SimpleImputer(keep_empty_features=False) that was fit, and just doing imp.keep_empty_features=True without refitting. This does nothing to change the already-fitted statistics; the object’s internals remain the same. Similarly, inserting a column of zeros into a transformed array and labeling it as “sensor” is a manual hack that breaks reproducibility. These shortcuts do not correctly update the learned mapping. The only safe approach is to retrain from scratch with the new settings. We therefore consider any such schema surgery invalid. In other words, a repair must produce a fully re-fitted artifact; it cannot rely on poking the existing arrays or parameters.

Monitor the Boundary After Release

Once a pipeline is deployed, it’s wise to monitor its input and output interfaces as part of data observability. Practically, we can log metrics such as: how often each original field is missing, how often requests are rejected for schema mismatches, and which artifact version handled the request. For example, we could count “incoming requests missing column X” and alert if unexpected patterns arise. We would also log if the model version has changed or if an input schema deviates (manifest mismatch). These are operational signals. They are different from tracking data drift or accuracy; they focus on the interface contract compliance. One could push these metrics into a monitoring system like Grafana and set up alerts on anomalies (see observability fundamentals for operational signals for general principles). For instance, if 99% of requests have a non-null age but suddenly many come with missing age, and we had not planned for an age indicator, an alert would help us respond (e.g. update the pipeline). Keeping drift metrics and performance stats separate, the schema monitoring ensures we “fail fast” on input issues.

In short, record and review any boundary violations: unexpected missing-feature patterns, wrong schema inputs, or mismatched artifact identities. Treat these as part of the ML observability layer. They often indicate upstream data issues or that a “fall-back” pathway might be engaged. By logging the artifact ID and schema version with each prediction, we can trace any problem back to a particular pipeline build. This follows best practices in serving ML: version your models (and preprocessors), and log key metadata.

Develop the AI Engineering Skills Behind Reliable Interfaces

Ensuring that a model’s preprocessing and estimator remain in lockstep requires discipline and the right skills. This exercise touches on several core AI engineering concepts: reproducible environments, detailed schema contracts, and rigorous testing. These are part of the model-building and evaluation foundations taught in comprehensive AI curricula. For those looking to strengthen these competencies, the AI Engineering Program covers model development, data pipeline engineering, and evaluation best practices (among other topics) that underpin reliable ML deployment. (We do not claim that this specific SimpleImputer scenario is explicitly in the curriculum, but it is an example of the practical model-validation mindset that the program promotes.)

In conclusion, the simple act of filling missing values can change the feature interface. By pre-declaring expected columns, verifying every output name and indicator, and keeping preprocessors and models paired, we ensure that “no news is good news” only means “no interface change”, not “no problems.” This focused acceptance and repair playbook helps ML engineers catch subtle bugs before they ship, rather than debug them in production.