Two data analysts reviewing pandas report charts and missing-value provenance on a laptop in a modern office.

Which Zeros Did pandas reindex Add to Your Report?

Wed, Oct 7, 2026

When you reindex a pandas Series to match a new set of target labels, some positions get NA by default. If you supply a fill_value=0, those introduced slots become zeros. But existing missing values in the original data remain unchanged.

Consider a simple five-row report where the source has records for labels A–C (A=5, B=<NA>, C=0) and the required target labels are A–E. The operational question is: Which zeros came from the source versus which were inserted by pandas, and is each zero authorized by our policies? The evidence is a label-by-label provenance ledger (source membership, values, and a declared structural-zero policy). The reviewer must decide only whether to accept the transformation as approved (with original unknowns intact), repair the unauthorized zero fill, or hold until provenance is resolved. This focuses on the local transformation contract (source snapshot, target list, and approved zero slots) rather than any broader data completeness.

This task is narrower than general cleaning advice. Here we work with a frozen example: a short pandas program on a fixed environment (Python 3.12.14, pandas 2.2.3 and NumPy 2.3.5, with a supplied local synthetic preflight dated October 7, 2026). The input Series uses pandas’ nullable Int64 dtype so that original missing values are pd.NA rather than zeros. We do not infer or fill any data by statistics or machine learning; we simply apply an explicitly documented rule: label D is an approved structural zero, with a reference reason string, and no other new labels are. Every row from the source (A, B and C) must survive as-is.

In this view, the approved output should be [5, NA, 0, 0, NA] for labels A–E. That is, B’s unknown stays unknown, C’s 0 stays 0, D becomes 0 (by policy), and E remains missing (no policy). A tempting but naive use of fillna(0) would wrongly convert B’s NA to 0. Instead we reindex first (which aligns and marks new rows with NA) and then assign zeros only to policy-approved labels. Along the way we keep a membership ledger to prove each value’s origin and authority. If any check fails (unauthorized zeros, changed existing values, wrong target, etc.), we flag it. The end result is a clear accept/repair/hold decision based on the ledger, not just checking totals or blindly trusting pandas.

Define the report transformation you are approving

Source and target: The raw report has three records:

  • A = 5 (observed)

  • B = Unknown (None → pd.NA)

  • C = 0 (observed)

The approved target labels, in fixed order, are A, B, C, D and E. This means the finalized report should have five rows, one per label in that order. We use pandas’ nullable integer type (Int64) so that missing data is represented by pd.NA. We never coerce or drop original records; every source row remains. If a source value is missing, it stays missing under this contract. We do allow adding new rows, D and E, as placeholders to fit the target index, but only with values sanctioned by policy.

Change policy: Our only approved structural zero is for D. Label D is absent from the source, and the stipulated fixture policy authorizes reporting D’s count as zero, with a documented reason. Label E is also absent from the source, but we have no approval for E; its absence remains unresolved. In other words:

  • Structural zero for D (approved): we will fill D with 0, citing "fixture-policy-v1: structural zero approved for D".

  • No policy for E: We will not assume E=0. Its count remains unknown (NA), indicating an unresolved measurement for an introduced label.

We write the question to the data owner as: “Accept only the transform that preserves all original values (including B’s NA and C’s 0) and adds a zero for D per the approved policy. Any zeros beyond that (or changes to A/B/C) must be flagged.” This is a tight contract on the reindex step, not a judgment of the numbers themselves.

The context is a standard analytics workflow where data cleaning adheres to explicit rules. As in our Python data-cleaning workflow discussion, we link each change to business rules rather than whims of code. Here, pandas’ reindex fill-value behavior adds zeros in any new slot when fill_value=0, but we authorize only D, not E. Any change outside the policy should be reviewed or rebuilt from the preserved source.

Require independent authority for each structural zero

Business truth does not come from code. The fact that reindex(..., fill_value=0) puts a zero in E’s row does not make E=0 correct. Only an independent policy can authorize that assignment; this fixture does not verify the policy’s business truth. That is why we keep a ZERO_AUTHORITY = {"D": "fixture-policy-v1: structural zero approved for D"} mapping. It explicitly ties the structural-zero decision to labels not in the source. If the code creates any other zero, such as at E, we mark it as “introduced without approval.”

We never accept structural-absence approval for a label already in the source. B is present but unknown, so it cannot qualify as an absent label. The fixture’s validate_contract function requires ZERO_AUTHORITY keys to be absent from the original labels. If someone tries to approve B through this policy, the validator rejects it. Likewise, the policy must name only introduced target labels; a blank or missing reason is invalid. Every structural-zero assignment must be independently documented. A library function or default setting never substitutes for a business rationale. All relevant policy context, including fixture-policy-v1, must travel with the transformation.

Freeze the source, target index and software versions

First, we lock in our inputs and environment. The preserved source records, target order and policy are exactly:

ROWS = [("A", 5), ("B", None), ("C", 0)]
TARGET = ["A", "B", "C", "D", "E"]
ZERO_AUTHORITY = {"D": "fixture-policy-v1: structural zero approved for D"}

Here None stands for an unknown measurement, which pandas will read as pd.NA in an Int64 Series. We do not use any other placeholder (no strings like "unknown" or numeric codes) to ensure type safety. The target index is exactly ["A","B","C","D","E"] in that order. We record the zero-policy as version "fixture-policy-v1" for reference.

We also record the exact software versions and the date for reproducibility. For a local environment with Python 3.12 already available:

python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install pandas==2.2.3 numpy==2.3.5
python reindex_provenance.py

The supplied local synthetic preflight completed successfully on October 7, 2026, using Python 3.12.14, pandas 2.2.3 and NumPy 2.3.5. The fixture captures a snapshot of the source Series before any transformations, using raw.copy(deep=True), and checks that the source remains unchanged. No external dataset is needed. The literal source values and membership are preserved separately from what pandas does. A different environment requires its own fresh run and recorded versions.

Keep nullable integers explicit

We make the data type explicit: raw = pd.Series([5, None, 0], index=["A", "B", "C"], dtype="Int64", name="count"). The nullable integer dtype Int64 uses pd.NA, displayed as <NA>, for missing values rather than floating-point NaN. This preserves the integer nature of observed counts. The pandas missing-data guide explains the missing-value sentinels used by different dtypes; this article uses its pandas 2.2.3 behavior.

B’s value therefore stays <NA> rather than converting to a float. We use Python None when defining the ROWS list; pandas represents it as pd.NA in this Int64 Series. We do not normalize missing values to zero or drop them. In the approved branch, the only new zero is the one authorized by ZERO_AUTHORITY for D.

Build the membership ledger before any fill

Before letting pandas align indices or fill values, we record an independent manifest of what the source contains. This is a simple table listing each target label, whether it was in the original source, what the raw value was, and what value we expect in the final report, plus a note on provenance. In our case it looks like this:

Label

In original source

Source value

Expected final value

Provenance

A

Yes

5

5

Observed positive measurement

B

Yes

Unknown

Unknown

Existing unknown measurement

C

Yes

0

0

Observed zero

D

No

No source record

0

Introduced, approved zero

E

No

No source record

Unknown

Introduced, unresolved absence

Each row is defined before pandas touches the data. For example, B’s source value is unknown: it is represented as Python None, exported as JSON null, and displayed as <NA> by pandas. It is not a numeric value. C’s source value is zero, so its expected final value is also zero. D has no source record, shown by “No” under “In original source,” and gets zero by policy. E has no source record and no policy, so its final value stays unknown.

This ledger separates membership from value. A missing value alone (like B vs E) is not enough to know if it was originally there. That’s why we explicitly flag “Yes” or “No” in the source. Note how C and D both end up with numeric 0 in the final plan, but C was “observed zero” in source while D is an “introduced, approved zero.” Conversely, B and E would both appear as NA after reindexing, but B is flagged “present in source” while E is not. We will never guess membership from pandas output alone (e.g. using isna()). For clarity, the manifest uses the Python literal None (to be printed as JSON null) to denote an unknown. In our code we keep this independent of the Series object’s state.

Separate a present unknown from an absent label

It’s crucial not to conflate B and E. Both have no numeric value in output, but B’s missingness is original (source member with unknown measurement) and E’s missingness is due to absence (no source record). A single NA cell cannot encode both stories. By checking the source_lookup dictionary (built from ROWS), the code sets state = "source_unknown" for B and state = "introduced_unknown" for E. In practice, after reindex, B’s value will be <NA>, but we know from the ledger it came from source. Conversely, E’s cell will also be <NA> if we never fill it, but we mark it as unresolved absence.

Similarly, C and D both have 0 in the final plan, but we must treat them differently. C’s zero was measured; D’s zero is synthetic. We never let pandas’ own output or isna() logic blur this. Instead, the ledger explicitly captures ("C", True, 0, "source_observed") vs ("D", False, None, "introduced_approved_zero"). (Here None under source_value means “no source record.”)

Use plain reindex to establish the baseline

Next we do a straight reindex without any fill. In code:

plain = raw.reindex(TARGET)

This aligns rows to A–E. The pandas reindex documentation explains that positions without a matching original label receive a missing value when no fill value is supplied. The result in the supplied fixture is:

  • A → 5 (carried over)

  • B → <NA> (existing unknown remains unknown)

  • C → 0 (carried over zero)

  • D → <NA> (new index, no source)

  • E → <NA> (new index, no source)

The values of plain are therefore [5, NA, 0, NA, NA]. We save this Series as a baseline. Nothing about the structural-zero policy has happened yet: D and E are just missing. This step is not an error; it shows pandas’ alignment behavior. The important fact is that A, B and C stayed exactly as they were, and no fill-value logic has run. We use plain later to compare with filled versions. The fixture also checks raw.equals(raw_snapshot) to ensure the input was not modified.

Inspect the zeros introduced by fill_value

Now we apply the fill value for reindex:

fill_during_reindex = raw.reindex(TARGET, fill_value=0)

With fill_value=0, pandas fills each newly introduced position with zero, as illustrated by the documented reindex fill-value examples. The resulting values are:

  • A → 5

  • B → <NA> (unchanged)

  • C → 0

  • D → 0 (new zero added by reindex)

  • E → 0 (new zero added by reindex)

The values of fill_during_reindex are [5, NA, 0, 0, 0]. The zeros at D and E are new: pandas filled both because we asked for zero. This shows precisely why we need a policy. D’s zero is authorized, but E’s is not. The fill_value parameter treats the new slots alike in this example; the application’s policy is a stricter filter.

Importantly, B stays <NA>. The documented distinction between index-based filling and existing missing data explains why an existing unknown is not replaced by the reindex fill value. Some users might expect fill_value to replace every missing measurement, but that is not what this operation does. Preserving B’s unknown is required by our contract.

Preserve the unknown already stored at B

By construction, B was an original label with an unknown count. Even after reindex(..., fill_value=0), B is <NA>. This agrees with the reindex alignment rules: B already has a matching source label, so the fill value for introduced positions does not replace its stored unknown. In the supplied fixture, pandas writes zeros for D and E, not for B. That matches the independent expectation that B’s unknown must not be converted to zero automatically.

If an analyst thought fill_value=0 would “deal with all missing” (like a naive fix), they’d be wrong. Filling B to 0 would violate the idea that B’s measurement is truly unknown. Our separate mapping (approved = plain.copy(); approved.loc["D"] = 0) will explicitly set only D.

The contract check on fill_during_reindex identifies the issue: D’s zero is authorized, but E’s is not. The fixture reports defects(fill_during_reindex) == ["E"], meaning that E violates the final transformation contract. The plain branch remains a valid intermediate baseline, but the final-contract checker identifies D because its approved zero has not yet been applied. E’s unresolved absence is allowed to remain unknown. This separates library behavior from the application’s narrower policy.

Show the damage from a blanket fillna call

To see what should not happen, we do a “negative control” by applying fillna(0) after reindex. In practice:

fill_every_null = fill_during_reindex.fillna(0)

The pandas Series.fillna documentation describes scalar replacement of missing values. Here, supplying scalar zero replaces every remaining NA in the Series. The result is:

  • A → 5

  • B → 0 (was <NA>, now 0)

  • C → 0

  • D → 0

  • E → 0

So we get [5, 0, 0, 0, 0]. At first glance, there are fewer missing values; only A’s position is nonzero. But this operation overwrote B’s original unknown (a grave error under our contract) and still kept the unauthorized zero at E. In the defect check, we see defects(fill_every_null) == ["B","E"].

In other words, the blanket fillna(0) created two “wrongs”: it introduced a zero at B where the source was NA, and it did nothing to remove the unauthorized zero at E (which was already 0 from reindex). A lower null count (zero missing) is not our goal; it’s exactly the opposite of our contract. This highlights that not all filling operations are equal. For example, forward-fill or group-based filling in BI tools (see our article on Power Query fill-down and customer boundaries) have different rules and scopes. Here, scalar fillna(0) is an unconditional fill that ignores policy. It is only used in our negative control to show what went wrong. In practice, we would not apply fillna(0) to the whole report because it erases provenance (as it did for B).

Thus the fillna(0) branch is invalid under the transformation contract. It actually hides the fact that B was unknown and E was unauthorized. A fix or acceptance cannot rely on this. Instead, we move to explicitly applying only the approved zero.

Apply the approved zero without rewriting observations

To produce the final, policy-compliant output, we start again from the plain reindex and then set only D to zero by policy. In code:

approved = plain.copy(deep=True)
approved.loc[list(ZERO_AUTHORITY.keys())] = 0

The only key in ZERO_AUTHORITY is D, so this assigns zero only to D. The resulting Series is:

  • A → 5

  • B → <NA>

  • C → 0

  • D → 0 (filled per policy)

  • E → <NA>

In values form: [5, NA, 0, 0, NA], exactly as EXPECTED. We explicitly check that this matches our independent expectation. Each label is now correct: original values stayed, B’s NA stayed NA, C’s 0 stayed 0, D became 0, and E stayed NA. We also verified raw is unmodified and the dtype stays Int64.

This shows the safe way to perform the transform: do not batch-fill anything except what’s explicitly allowed. Notice that we used a separate copy (approved = plain.copy()) to avoid touching the original or intermediate series. We then only touched label D by indexing with loc. This preserves all other data exactly.

Validate the approval scope before assignment

Before doing the assignment, we perform a sanity check on the policy mapping. We require every key in ZERO_AUTHORITY to meet two conditions: it must be in the target but not in the original source, and it must have a non-empty reason string. In code, our validate_contract raises an error if these are violated. For example:

  • If the source already contained label B but someone tried ZERO_AUTHORITY={"B": "..."}, we reject it as nonsensical.

  • If an approval was listed with an empty string, we reject that too (the reason must be documented).

  • A policy key outside the target is rejected. Duplicate or empty source and target labels also fail the input contract.

This validation happens before we create the source dictionary, because dictionary construction would otherwise collapse duplicate keys. We do not simply drop or repair bad inputs; we reject them for review of the input definitions. For example, if ZERO_AUTHORITY names B or supplies no reason, the validator raises an error. The next section executes eight rejection controls. Only a well-formed policy reaches the transformation.

Run the complete fixture and retain the JSON evidence

Below is the complete Python fixture, reindex_provenance.py. It contains the definitions, pandas operations, validation checks and JSON report construction. First it validates the contract, then it performs the four branches: plain, fill_value_zero, fillna_every_null and approved_only. It checks the values against independent literal expectations and raises an exception when a check fails.

Only after its checks pass does the script print JSON. The report records versions, the ledger, branch values and defects, rejected contracts and bounded library controls. Its branches field includes each branch’s values, dtype, missing count, sum and defect labels. The transformation_result field records the approved-only acceptance result.

The supplied preflight was a local synthetic run on October 7, 2026, not a production test. A discrepancy in a checked result must be retained and reviewed; do not rewrite the expected manifest to approve a changed implementation. Record the exact runtime versions and the current process exit status with each fresh result.

"""Offline research fixture: label expansion and missing-value provenance."""
import json
import platform
import pandas as pd
import numpy as np
ROWS = [("A", 5), ("B", None), ("C", 0)]
TARGET = ["A", "B", "C", "D", "E"]
ZERO_AUTHORITY = {"D": "fixture-policy-v1: structural zero approved for D"}
EXPECTED = [5, None, 0, 0, None]
EXPECTED_MANIFEST = [
    ("A", True, 5, "source_observed"),
    ("B", True, None, "source_unknown"),
    ("C", True, 0, "source_observed"),
    ("D", False, None, "introduced_approved_zero"),
    ("E", False, None, "introduced_unknown"),
]
def require(condition, message):
    if not condition:
        raise RuntimeError(message)
def validate_contract(rows, target, authority):
    labels = [key for key, in rows]
    for name, keys in (("source", labels), ("target", target)):
        if not keys:
            raise ValueError(f"{name} must be nonempty under this fixture contract")
        if any(not isinstance(key, str) or not key.strip() for key in keys):
            raise ValueError(f"invalid {name} label")
        if len(keys) != len(set(keys)):
            raise ValueError(f"duplicate {name} labels")
    if set(labels) - set(target):
        raise ValueError("target would omit source labels")
    for , value in rows:
        if value is not None and (type(value) is not int or value < 0):
            raise ValueError("expected nullable nonnegative integer measurement")
    for key, reason in authority.items():
        if key not in target or key in labels:
            raise ValueError("structural-zero approval must name an introduced label")
        if not isinstance(reason, str) or not reason.strip():
            raise ValueError("zero approval needs an independent documented reason")
def values(series):
    return [None if pd.isna(value) else int(value) for value in series]
validate_contract(ROWS, TARGET, ZERO_AUTHORITY)
source_lookup = dict(ROWS)  # Constructed before either pandas operation.
raw = pd.Series([v for , v in ROWS], index=[k for k, in ROWS],
                dtype="Int64", name="count")
raw_snapshot = raw.copy(deep=True)
plain = raw.reindex(TARGET)
fill_during_reindex = raw.reindex(TARGET, fill_value=0)
fill_every_null = fill_during_reindex.fillna(0)  # Deliberate negative control.
approved = plain.copy(deep=True)
approved.loc[list(ZERO_AUTHORITY)] = 0
require(values(plain) == [5, None, 0, None, None], "plain branch changed")
require(values(fill_during_reindex) == [5, None, 0, 0, 0], "reindex branch changed")
require(values(fill_every_null) == [5, 0, 0, 0, 0], "fillna branch changed")
require(values(approved) == EXPECTED, "approved branch differs from literal oracle")
require(raw.equals(raw_snapshot), "source was modified")
ledger = []
for label in TARGET:
    present = label in source_lookup
    original = source_lookup.get(label)
    if present:
        state = "source_unknown" if original is None else "source_observed"
    else:
        state = (
            "introduced_approved_zero"
            if label in ZERO_AUTHORITY
            else "introduced_unknown"
        )
    ledger.append({"label": label, "in_source": present,
                   "source_value": original, "state": state,
                   "zero_authority": ZERO_AUTHORITY.get(label)})
require([row["in_source"] for row in ledger] == [True, True, True, False, False],
        "membership ledger differs from independent literal expectation")
require([(row["label"], row["in_source"], row["source_value"], row["state"])
         for row in ledger] == EXPECTED_MANIFEST,
        "provenance ledger differs from the independent literal manifest")
def defects(candidate):
    issues = []
    if candidate.index.tolist() != TARGET or not candidate.index.is_unique:
        issues.append("index_contract")
        return issues
    if str(candidate.dtype) != "Int64":
        issues.append("dtype_contract")
    for label, actual in zip(TARGET, values(candidate)):
        if label in source_lookup:
            expected = source_lookup[label]
        else:
            expected = 0 if label in ZERO_AUTHORITY else None
        if actual != expected:
            issues.append(label)
    return issues
branches = {"plain": plain, "fill_value_zero": fill_during_reindex,
           "fillna_every_null": fill_every_null, "approved_only": approved}
require(defects(fill_during_reindex) == ["E"], "did not identify unapproved new zero")
require(defects(fill_every_null) == ["B", "E"], "did not identify both zero defects")
require(defects(approved) == [], "approved output failed the transformation contract")
require(all(int(s.sum()) == 5 for s in branches.values()), "comparison sums changed")
contract_rejections = {}
bad_cases = {
    "duplicate_source": ([("A", 1), ("A", 2)], TARGET, {}),
    "duplicate_target": (ROWS, TARGET + ["A"], ZERO_AUTHORITY),
    "unexpected_source": (ROWS + [("F", 7)], TARGET, ZERO_AUTHORITY),
    "empty_source": ([], TARGET, ZERO_AUTHORITY),
    "empty_target": (ROWS, [], {}),
    "approval_for_existing_unknown": (ROWS, TARGET, {"B": "cannot justify by absence"}),
    "approval_without_reason": (ROWS, TARGET, {"D": ""}),
    "missing_label": ([(None, 1)], TARGET, {}),
}
for name, args in bad_cases.items():
    try:
        validate_contract(*args)
    except ValueError as exc:
        contract_rejections[name] = str(exc)
    else:
        raise RuntimeError(f"invalid contract accepted: {name}")
try:
    pd.Series([1, 2], index=["A", "A"], dtype="Int64").reindex(["A", "B"])
except ValueError as exc:
    duplicate_source_library_result = str(exc)
else:
    raise RuntimeError("duplicate-source comparison did not raise")
library_controls = {
    "duplicate_source": duplicate_source_library_result,
    "duplicate_target_values": values(raw.reindex(["A", "A"])),
    "empty_source_fill_value_zero": values(
        pd.Series([], dtype="Int64").reindex(TARGET, fill_value=0)
    ),
    "empty_target_length": len(raw.reindex([])),
}
require(
    library_controls["duplicate_target_values"] == [5, 5],
    "target-duplicate behavior changed",
)
require(
    library_controls["empty_source_fill_value_zero"] == [0, 0, 0, 0, 0],
    "empty-source behavior changed",
)
require(
    library_controls["empty_target_length"] == 0, "empty-target behavior changed"
)
report = {
    "versions": {
        "python": platform.python_version(),
        "pandas": pd.__version__,
        "numpy": np.__version__,
    },
    "scope": (
        "synthetic local transformation evidence; "
        "zero authority is a stipulated fixture policy"
    ),
    "source_rows": ROWS, "target": TARGET, "zero_authority": ZERO_AUTHORITY,
    "ledger": ledger,
    "branches": {name: {"values": values(s), "dtype": str(s.dtype),
                       "missing_count": int(s.isna().sum()), "sum_skipna": int(s.sum()),
                       "defect_labels": defects(s)} for name, s in branches.items()},
    "contract_rejections": contract_rejections, "library_controls": library_controls,
    "transformation_result": (
        "APPROVED_ONLY_ACCEPTED_WITH_ORIGINAL_UNKNOWNS_PRESERVED"
    ),
    "remaining_unknown_labels": ["B", "E"],
}
print(json.dumps(report, indent=2, allow_nan=False))

The supplied fixture evidence reports approved_only values of [5, null, 0, 0, null] with no defect labels, and fill_value_zero values of [5, null, 0, 0, 0] with defect label E. The transformation_result value is APPROVED_ONLY_ACCEPTED_WITH_ORIGINAL_UNKNOWNS_PRESERVED. JSON null represents an unknown value; source membership remains a separate ledger field.

A synthetic script still has a process exit status. Immediately after invoking it in the shell, capture that status with echo $? and retain the freshly generated JSON alongside the result. Do not present a prior output file as evidence that a failed rerun passed.

Check invalid inputs before pandas can normalize them

Before building the pandas Series, we guard against bad inputs. Our code runs eight contract-rejection cases to ensure validate_contract catches them:

  • Duplicate source labels: Two rows with the same label (e.g. [("A",1),("A",2)]) should be rejected. If we slipped duplicate keys into the rows list, a raw dict(ROWS) would silently pick one, so we forbid it.

  • Duplicate target labels: A repeated label in TARGET (e.g. ["A","A","B"]) is also forbidden by our contract, even though pandas would just duplicate values. We catch it early.

  • Unexpected source label: A source label not present in TARGET, such as an added ("F", 7), is an error under this fixture contract because the target must include every source member.

  • Empty source: No rows at all is not allowed (our fixture requires at least something to transform).

  • Empty target: Having an empty target list is also disallowed (then there’d be nothing to output).

  • Approval for existing label: If ZERO_AUTHORITY={"B": ...} was given, we reject it (cannot justify a policy by an existing row).

  • Approval without reason: If someone put ZERO_AUTHORITY={"D": ""}, with an empty reason string, we reject because every structural-zero entry needs documentation.

  • Invalid source label: The executed case supplies None as a source label. The same validation also rejects empty or whitespace-only strings in the source and target lists.

The fixture tries each case in turn. Each must raise a ValueError in validate_contract, and the error message is collected in contract_rejections. An unexpectedly accepted invalid case raises a RuntimeError. The JSON report lists each case name with its message, showing that the checks are active. For example, duplicate_source produces the message “duplicate source labels.”

Treat duplicate targets and empty inputs as explicit contracts

The pandas duplicate-labels guide explains that pandas generally permits nonunique labels, while its changed-target reindex example with a duplicate source raises ValueError. The supplied fixture includes that bounded rejection control; it does not claim that every possible no-op reindex with duplicates fails. By contrast, a duplicate target can repeat a source value:

pd.Series([5, None, 0], index=["A","B","C"], dtype="Int64").reindex(["A","A"])

This produces [5, 5], duplicating A’s value. The fixture checks library_controls["duplicate_target_values"] == [5, 5]. Our application nevertheless disallows duplicate output labels by design: each target must be unique.

Similarly, an empty source with fill_value=0 lets pandas create five zeros for A–E, but this fixture rejects an empty source. An empty target produces a zero-length Series, so len(raw.reindex([])) is zero. Pandas allows that operation, while our contract requires at least one target label.

These controls distinguish successful pandas operations from inputs accepted by this application. Duplicates, omissions and empty inputs are rejected where the declared contract forbids them. The library_controls results record the bounded library behavior; the validator applies the stricter application rules. This is not a claim that those rules are appropriate for every valid pandas use case.

Compare provenance before comparing totals

At this point we have four series to compare (we use the branches dictionary from the report):

Branch

Values

Missing count

Defect labels

Plain reindex

[5, NA, 0, NA, NA]

3

D

fill_value=0

[5, NA, 0, 0, 0]

1

E

fillna(0)

[5, 0, 0, 0, 0]

0

B, E

Approved only

[5, NA, 0, 0, NA]

2

(none)

“Missing count” is the number of remaining <NA> values: 3, 1, 0 and 2 respectively. All four branches sum to five under the default skipna=True behavior documented for pandas Series.sum. Missing values are excluded from that calculation, and adding zeros does not change the total. The equal sums cannot authorize any synthetic value. We therefore inspect provenance and defect labels instead.

The fill_value=0 branch has defect label E, an unauthorized zero. The fillna(0) branch has defects at B and E. Against the final contract, the plain branch has defect label D because its approved zero has not yet been applied. Only the approved branch has no defects.

Matching totals do not resolve those differences. Our separate article on missing-key rows in pandas GroupBy addresses aggregation-specific reconciliation. Here, the source labels remain present and the decision concerns the provenance and treatment of the added slots.

Neither zero-filling negative-control branch is acceptable. Plain reindex is an intermediate baseline, not the final approved result. Only the approved-only branch yields zero defects and satisfies the final transformation checks, including require(defects(approved) == []).

Decide whether to accept, repair or hold

We distill all this into a decision table for the transformation step:

Condition

Evidence

Action

Owner

Accept: Target labels and dtype match contract, all source values unchanged (including original NAs), and introduced zeros match exactly the approved policy.

No defects; ledger matches EXPECTED_MANIFEST; missing counts as expected; only D’s zero added.

Accept the transformation. Finalize the report only under the report owner’s qualified reporting policy.

Analytics Engineer (with Data Owner sign-off)

Repair: Exactly the unauthorized cases above (zero at E or B changed).

Defect on E (unauthorized new zero) or B (original NA overwritten by zero). Source snapshot is intact.

Restore original source, rerun approved transform to correct it. Preserve evidence. Continue with final output from approved-only branch.

Analytics Engineer (with review by data owner)

Hold: Any other mismatch (missing ledger rows, changed source values, unexpected targets, unknown policy).

Unrecoverable errors: e.g. source or target changed since last check, missing payload, validate_contract failure.

Stop and escalate. Request updated source/target/policy or more info. Do not publish report.

Data Owner or Stakeholder

In words: we only accept when all contract conditions hold exactly (case 1). We repair when the only issues are those we can fix by re-running the approved logic (case 2: unauthorized/overfill zeros). We hold (neither accept nor repair) if we lack necessary provenance or if fundamental assumptions changed.

Notice in repair, we specifically allow fixing the output by reusing the untouched source. We do not recalc B/E from the broken output, because the contract forbids guessing. If B was inadvertently filled (like in fillna), we throw it away and go back to source. The defect evidence (unauthorized zero or missing flag) is kept for audit, but the final accepted output is rebuilt correctly.

If we can’t recover (e.g. the source file got altered, or the policy is unclear, or target list mismatch), then the data owner must intervene. The transformation is effectively incomplete. In that case, unknowns (like B or E) remain unresolved, and the report must be held.

Rebuild an affected report from the preserved source

Suppose we fell into the “Repair” path (for example, E got a zero). We keep the faulty output only as evidence. Then we go back to the original source raw (unchanged by pandas) and re-run our approved operation. This yields the same [5, NA, 0, 0, NA] that we already tested. We compare it to the expected ledger to confirm the fix.

We do not try to infer values from the rejected output. Turning E’s zero back into NA by intuition is unreliable without source evidence. The reviewed repair is to rebuild from the preserved source using reindex and selective assignment. With multiple unknowns, the final zeros alone cannot reveal which were observed and which were introduced. A learned imputation step is a different workflow and is outside this article’s scope. Our SimpleImputer feature-schema validation article examines fitted preprocessing and model-interface expectations, not this source-value preservation contract.

After repairing, we re-evaluate the defects(approved) which should now be empty, and the transformation_result is effectively "approved." We would log both the original failed state and the fixed result. Any publication or downstream use should proceed only with the approved output.

If instead we had a “Hold” case (like missing source info), we would not attempt to fix. The report stays incomplete; we rely on data owners to supply missing pieces (like an updated target list or clarification of policy) before proceeding.

Assign ownership and define the limits of acceptance

To ensure accountability, we assign roles for each part of this procedure:

  • Source snapshot ownership: The team that collected the data (e.g. ETL or data ingress) owns the original values. They preserve the snapshot and its timestamp. Any change to those values requires rerunning this validation.

  • Target list ownership: The business or report owner owns the list A–E. Any change (adding/removing labels) triggers a revalidation.

  • Zero-policy authority: The policy that “D is zero” must be signed off by the report owner or data steward. Changes to the policy (keys or reasons) require re-running this check. The policy reason string must remain visible for audit.

  • Transformation review: The data engineering team applies this script. They verify the ledger before publishing. The reviewer logs the missing counts (e.g. B and E) and the policy rationale.

  • Report interpretation: Finally, the analyst or decision-maker who reads the report interprets the values, acknowledging which labels are unresolved. They know that an “approved zero” is a placeholder, not a measured fact.

We keep all approval references and missing-value counts explicitly in the report metadata. If any component changes (e.g. the list of expected labels expands, or D’s policy is updated), we re-run the fixture from scratch.

For completeness, note that this check does not address upstream data delivery or freshness. It only deals with the transformation of a given snapshot. Upstream concerns (like missing partitions or incomplete loads) fall under other processes (see our discussion of dbt source freshness and missing partitions). Similarly, the quality or “truth” of the measurements themselves is outside this script’s scope; we assume the raw values and the zero-policy as given. We also don’t handle schema changes beyond this step. Each piece (source, target, policy) is managed by its owner, and our validation checks the declared alignment contract.

Practice reviewed data preparation with Refonte Learning

This workflow, preserving raw data, expanding an index explicitly and applying narrowly defined rules, gives you concrete preparation artifacts to review. For broader study, explore Refonte Learning’s Data Analytics program, described as a Data Analytics Virtual Internship and Training Program. Its three-month curriculum includes Python, R, SQL, exploratory data analysis and industry projects, with an educational mentor listed on the program page. It also covers Tableau and advanced Excel. Successful completion offers a Training Certificate and a Certificate of Internship. Applicants must be working toward a bachelor’s or higher degree.

The lab’s audit trail is the unchanged source, the literal expected ledger, the invalid-contract checks and an accept, repair or hold decision linked to evidence. Keep those artifacts together. A passing transformation preserves the original unknown at B, adds only the approved zero at D and leaves E unresolved. It approves application of the stated policy, not the completeness or truth of the underlying measurements.