In our application, the business rule requires every numeric output to be a multiple of 0.05 (or 0.25) with half-even rounding. A simple counterexample immediately illustrates the issue: using Python’s Decimal.quantize() on 0.03 with a step of 0.05 yields 0.03, which is not a multiple of 0.05. This happens because quantize only aligns exponents and does not enforce membership in the 0.05 grid. In other words, quantize returning the exact input value (0.03) violates the step-increment contract.
To ensure correctness, we must treat inputs as exact decimal strings under a restricted grammar, then explicitly check and round to the nearest allowed step. We will preserve the original string form for reference, identify off-grid outputs, apply the chosen half-even tie policy, and compare each result against an independent rational oracle. A single successful example (e.g. “1.20” → “1.20”) is not proof of correctness. In fact, even if the naive call succeeds on some values, it may still misbehave on others (as with 0.03). Consequently, our acceptance criteria will require passing a comprehensive test suite with exact expected values, not just one or two cases. We implement and document this end-to-end, ensuring that failing tests, not naive success, determine whether the code is ready for release.
For reference, the decimal specification and IEEE standards define “round half to even” (banker’s rounding) as rounding ties to the nearest even integer. However, choosing a tie-breaking rule is a business requirement, not something imposed by Decimal. Our role is to apply that rule exactly. We do not rely on floating-point behaviors or implicit choices. In the sections below, we first state the exact contract, then expose the quantize shortcoming, build a precise rational oracle, implement a restricted rounding function, and finally set strict release gates for acceptance. (The examples use Python 3.11+ syntax and decimal with ROUND_HALF_EVEN.)
Define the increment your output must obey
Contract specification: The function must round any input string to the nearest multiple of a step, where step is 0.05 or 0.25. Ties (exact half-way cases) must resolve to the multiple with an even integer step count (banker’s rounding). Inputs are explicit decimal strings matching the grammar -?(?:0|[1-9][0-9]{0,7})(?:\.[0-9]{1,6})?. This means up to 8 integer digits and at most 6 fractional digits, no leading zeros (except “0” itself), no exponent notation, and no non-numeric tokens. For example, “0001.20” or “01” are invalid, whereas “0.00” or “99999999.999999” are valid. The output must have exactly two fractional digits (quantize to “0.01”) and use canonical zero (“0.00”, never “-0.00”). Negative inputs are allowed under the same rules, and steps must appear exactly as “0.05” or “0.25” (string equality), not equivalent variants. In summary, the function’s signature is round_to_step(raw_str: str, step_str: str) -> Decimal, raising ValueError if either argument is out of the approved format or values.
Tie policy is chosen by the application: here we use nearest-even (“banker’s rounding”) as required. This is not chosen by the library but by the product owner. We enforce it explicitly. The following table gives examples:
Raw input | Step | Valid? (Reason) |
"123.45" | 0.05 | Yes (matches grammar) |
"12.3456789" | 0.05 | No (too many decimals) |
"01.23" | 0.05 | No (leading zero) |
"-0.10" | 0.05 | Yes (negative allowed) |
"0" | 0.05 | Yes (canonical zero) |
"1" | 0.03 | No (step not in {0.05,0.25}) |
We will explicitly check each condition in code. For example, the type check rejects non-string input, and the regex disallows "01", while step not in {"0.05","0.25"} rejects any other step. Only after validating the grammar and approved step do we proceed to arithmetic.
Separate exponent alignment from step membership
A key subtlety is that Decimal.quantize(other) only cares about matching the exponent of other, not about integer multiples of its value. In particular, 0.05 and 0.25 both have exponent -2 (two decimal places), but they represent different grids. Quantizing to exponent -2 will simply adjust “to two decimal places” with half-even rounding, irrespective of the step increment. For these finite inputs under the stated context, matching exponents leave the value unchanged. For example, both 0.05 and 0.01 have exponent -2, but quantizing 0.03 to exponent -2 leaves it at 0.03, not moving it to the nearest 0.05. As Python’s quantize documentation states: “Return a value equal to the first operand after rounding and having the exponent of the second operand”. Since 0.03 and 0.05 share exponent -2, no rounding is applied. In other words, quantize would leave 0.03 unchanged, even though 0.03 is not a multiple of 0.05.
It is therefore wrong to think that supplying a “more suggestive” quantize operand (or changing context precision) magically enforces our step rule. To obey the contract, we must check membership in the step grid explicitly (via rational arithmetic) and apply rounding if needed. We treat exponent alignment and multiple-of-step as separate concerns. As a sanity check, compare the numeric value modulo the step with zero, or use the rational equivalent rather than trusting quantize alone. This distinction underlies the entire problem: exponent alignment ≠ step alignment. (Even if we change precision or use a coefficient, it does not change that quantize only sees exponents.)
Inspect the result as a value, not just formatted text
The output of quantize must be evaluated as a numeric value, not just by its printed string. For example, Decimal('0.03') has two fractional digits, so when quantized to two decimals it remains '0.03'. But 0.03/0.05 = 0.6, which is not an integer, so 0.03 is not on the 0.05 grid. Similarly, consider Decimal('1.13').quantize(Decimal('0.25'), rounding=ROUND_HALF_EVEN). Both have exponent -2, so the result string is '1.13', but 1.13 is not a multiple of 0.25. In each case, the correct membership check is to convert to a rational step count and verify it’s an integer. We must not judge correctness by string patterns (e.g. “two decimals”) alone, since formatting may coincide with step-size decimals but still fail the true check. Instead, we preserve both the numeric value and its text, and use both in our evidence.
Freeze the input grammar and arithmetic environment
We explicitly freeze all assumptions about inputs and math to prevent inadvertent variation. We define the precise decimal text grammar above and enforce it via re.fullmatch. Float inputs, booleans, exponent notation (e.g. "1e2"), over-precision (more than 6 decimal places), out-of-bound magnitude (e.g. >8 integer digits) or any unsupported step string immediately raise ValueError. For example, 01 or 0.0000001 will be rejected by our regex, and "0.050" is rejected because the string must match exactly "0.05" or "0.25".
For the arithmetic, we use a fresh decimal.Context with fixed parameters. Specifically, we use precision 28 (enough for intermediate exact arithmetic on our input bounds), Emin=-999, Emax=999 to avoid exponent overflow, and rounding=ROUND_HALF_EVEN. We also set traps for invalid operations (InvalidOperation, DivisionByZero, Overflow, FloatOperation) and control the Inexact trap around each step. Here is the core context setup in code:
import decimal
from decimal import Context, localcontext, ROUND_HALF_EVEN, Inexactctx = Context(
prec=28, rounding=ROUND_HALF_EVEN, Emin=-999, Emax=999,
capitals=1, clamp=0, flags=[],
traps=[decimal.InvalidOperation, decimal.DivisionByZero,
decimal.Overflow, decimal.FloatOperation],
)Within a with localcontext(ctx) as ctx: block, we clear any flags and explicitly toggle ctx.traps[Inexact] around the division and multiplication steps. This lets us detect any unintended rounding or overflow. (The reference documentation shows using an Inexact trap to validate fixed-point input; we use the same principle to make rounding steps visible.) Every context field, including the complete trap set, is explicit. The function therefore does not inherit these settings from ambient application state.
The core and complete checks shown here were executed on October 7, 2026, with CPython 3.13.5, decimal 1.70 and libmpdec 2.5.1. The twelve expected-value cases passed, all twelve invalid input pairs were rejected, and the off-grid negative control was rejected. These observations apply to this fixture and runtime.
Reproduce the successful call that violates the policy
First, we demonstrate the failure of the naive approach. Using Python’s Decimal.quantize():
from decimal import Decimal, ROUND_HALF_EVEN
with localcontext(ctx) as diagnostic:
diagnostic.clear_flags()
result_naive = Decimal('0.03').quantize(
Decimal('0.05'), rounding=ROUND_HALF_EVEN
)
print(result_naive) # 0.03, off-gridThe output is 0.03, unchanged. This clearly violates our rule (expected 0.05). In contrast, a value already on the step behaves “correctly” in appearance:
with localcontext(ctx) as diagnostic:
diagnostic.clear_flags()
print(Decimal('1.20').quantize(
Decimal('0.05'), rounding=ROUND_HALF_EVEN
))
# 1.20Here 1.20 remains 1.20, which is on the 0.05 grid. But consider a different step:
with localcontext(ctx) as diagnostic:
diagnostic.clear_flags()
print(Decimal('1.13').quantize(
Decimal('0.25'), rounding=ROUND_HALF_EVEN
))
# 1.13Again it stays 1.13, despite the expected result being 1.25. In each case above, quantize left the value unchanged because the exponents already matched. Both failures (0.03 and 1.13) are off-step relative to the business rule. We must reject these outputs. For comparison, the declared step policy rounds 0.075 with step 0.05 to 0.10. Evaluating Decimal("0.075").quantize(Decimal("0.05"), rounding=ROUND_HALF_EVEN) instead produces 0.08: exponent alignment still does not implement that policy.
Crucially, no single passing example proves correctness. If we tested only 1.20, we might incorrectly conclude that quantize is sufficient. The 0.075 tie, like the 0.03 and 1.13 counterexamples, exposes the difference. We treat the naive off-step results as negative controls: if our checks ever fail to reject them, something is wrong.
Test what an Inexact trap actually detects
Let us confirm how the Inexact signal behaves. We run the quantize under a context that traps Inexact:
with localcontext(ctx) as diagnostic:
diagnostic.clear_flags()
diagnostic.traps[Inexact] = True
result = Decimal('0.03').quantize(
Decimal('0.05'), rounding=ROUND_HALF_EVEN
)
print(result, diagnostic.flags[Inexact]) # 0.03 FalseNo exception is raised. Quantize does not mark this result as inexact because no non-zero digits are discarded. The Inexact signal detects that particular arithmetic condition, not whether a value belongs to an application’s grid. Therefore, the fixed-point validation example in Python’s documentation does not turn this trap into a step-membership check. We must compute membership ourselves, and clear the flags before each diagnostic run rather than reading a flag left by an earlier operation.
Write the expected values before testing the repair
Before implementing the fix, we enumerate the exact expected outputs according to our chosen tie policy. We will never compute them from the code under test; they must come from our independent policy. Here is the manifest of test cases (using our policy of rounding to nearest, ties-even):
Case | Raw input | Step | Expected output |
Off-step value | 0.03 | 0.05 | 0.05 |
Already on step | 1.20 | 0.05 | 1.20 |
Positive even lower tie | 0.025 | 0.05 | 0.00 |
Positive odd lower tie | 0.075 | 0.05 | 0.10 |
Negative tie to zero | -0.025 | 0.05 | 0.00 |
Negative nonzero tie | -0.075 | 0.05 | -0.10 |
Quarter off-step value | 1.13 | 0.25 | 1.25 |
Quarter even lower tie | 1.125 | 0.25 | 1.00 |
Quarter odd lower tie | 1.375 | 0.25 | 1.50 |
Negative quarter tie | -1.375 | 0.25 | -1.50 |
Zero | 0 | 0.05 | 0.00 |
Upper-bound carry | 99999999.999999 | 0.05 | 100000000.00 |
These values are derived by treating the input and step as exact rationals and rounding to the nearest integer multiple of step (with ties breaking to even). We will use this table to validate our implementation. Crucially, we do not recalculate these from our code; doing so would not catch errors in logic. Instead, we will compare our code’s output against these independent expectations.
Build an independent integer and rational oracle
To compute expected outputs, we build an oracle using integer arithmetic. For each case, let
units = Fraction(raw) / Fraction(step)where Fraction(raw) and Fraction(step) create exact rationals from the strings. Then we split units into an integer quotient and remainder with lower, remainder = divmod(units.numerator, units.denominator). In Python, because the denominator is positive, remainder is non-negative even if the input is negative (floor division semantics). Now consider twice = 2 * remainder.
If twice < units.denominator, the value is closer to the lower multiple (round down).
if twice > units.denominator, round up to lower+1.
If twice == units.denominator, we have a tie: we round to the even integer, i.e. we round to lower if lower is even, otherwise to lower+1. In code, this is written as:
lower, remainder = divmod(units.numerator, units.denominator)
twice = 2 remainder
rounded = lower + (twice > units.denominator or
(twice == units.denominator and lower % 2 != 0))
expected = rounded Fraction(step)For the positive tie 0.025 with step 0.05, the exact step count is 1/2. Integer division gives lower=0 and remainder=1 with denominator 2, so the result is 0, the even step count. For -0.075 with step 0.05, the step count is -3/2: lower=-2 and remainder=1 with denominator 2. The lower count is even, giving -0.10. Distance alone cannot distinguish ties because both neighbors are equally near. Parity (lower % 2) supplies the tie-breaker, consistent with the decimal specification and IEEE standards.
Thus the oracle is simple integer math and parity checks. It shares the declared policy and input strings with the candidate, but it does not use Decimal.quantize() as its rounding algorithm. This gives each case an independent arithmetic reference.
Resolve positive and negative ties explicitly
For clarity, consider one positive and one negative tie scenario:
Positive tie: raw=0.075, step=0.05. Here units = 0.075/0.05 = 3/2. Floor division gives lower=1, remainder=1 (denominator 2). So twice=2. Because twice == denom (tie) and lower % 2 == 1 (odd), we round up to lower+1 = 2. Then 2 * step = 0.10. This matches the “positive odd lower tie” case.
Negative tie: raw=-0.075, step=0.05. Here units = -3/2. Python’s floor division yields lower=-2, remainder=1 (denominator 2, since -1.5 floors to -2 with remainder +1). Then twice=2. It’s a tie again, and lower % 2 == 0 (even), so we round to lower = -2. Then -2 * step = -0.10.
This example shows that the sign of the number only affected the quotient, not the tie-breaking rule: in both cases we computed “twice == denom” and used the even/odd rule on lower. Importantly, we see that a tie cannot be resolved by distance alone, so the parity rule is needed. In our tests, we will verify that these computed fractions exactly match the expected outputs above.
Implement a bounded step-rounding candidate
Now we implement the bounded candidate. The code below preserves the same input contract (string inputs → Decimal result) and applies two steps: an exact division in step units, and a nearest-integer quantization. A final multiplication and quantization restore two decimal places. We keep precision 28 to cover any intermediate size (up to 8 integer digits and 6 fraction digits divided by 0.05 or 0.25), ensuring exactness. Inexact traps are turned on only to detect unintended rounding; any true rounding error (beyond our intentional ties rule) would raise an exception.
import decimal
import re
from decimal import (
Decimal, Context, localcontext, ROUND_HALF_EVEN, Inexact,
FloatOperation,
)
from fractions import FractionGRAMMAR = re.compile(r"-?(?:0|[1-9][0-9]{0,7})(?:\.[0-9]{1,6})?", re.ASCII)
STEPS = {"0.05", "0.25"}def require(condition, message):
if not condition:
raise RuntimeError(message)def check_input(raw, step):
if type(raw) is not str or GRAMMAR.fullmatch(raw) is None:
raise ValueError("Input must be canonical bounded decimal text")
if type(step) is not str or step not in STEPS:
raise ValueError("Unapproved step")def round_to_step(raw, step):
check_input(raw, step)
with localcontext(Context(
prec=28, rounding=ROUND_HALF_EVEN, Emin=-999, Emax=999,
capitals=1, clamp=0, flags=[],
traps=[decimal.InvalidOperation, decimal.DivisionByZero,
decimal.Overflow, decimal.FloatOperation],
)) as ctx:
ctx.clear_flags()
# Divide exactly within the admitted input and step bounds
ctx.traps[Inexact] = True
units = Decimal(raw) / Decimal(step)
ctx.traps[Inexact] = False
# Round to nearest integer number of units
rounded_units = units.quantize(Decimal("1"), rounding=ROUND_HALF_EVEN)
ctx.traps[Inexact] = True
# Multiply back and fix to two decimals
result = (rounded_units * Decimal(step)).quantize(Decimal("0.01"))
# Normalize zero sign
return Decimal("0.00") if result.is_zero() else resultdef expected_fraction(raw, step):
units = Fraction(raw) / Fraction(step)
lower, remainder = divmod(units.numerator, units.denominator)
twice = 2 remainder
rounded = lower + (twice > units.denominator or
(twice == units.denominator and lower % 2 != 0))
return rounded Fraction(step)This code uses a fresh local context so it does not inherit any global state. The public result is returned as a Decimal with exactly two fractional digits. We do not broaden the STEPS set or the regex without simultaneously adjusting the documented contract; any change to allowed inputs would need the same rigorous re-validation.
Compare every result against the declared contract
Append the runner to the core above. It iterates through the complete manifest, requiring exactly twelve unique cases. For each case, it computes:
expected = expected_fraction(raw, step) from our rational oracle,
candidate = round_to_step(raw, step),
naive = Decimal(raw).quantize(Decimal(step), rounding=ROUND_HALF_EVEN) for comparison.
We then require (via explicit exceptions) that:
Fraction(candidate) == expected and Fraction(expected_str) == expected (the candidate and the independently written manifest both match the oracle); the naive result is recorded for comparison, not required to fail on every input.
(Fraction(candidate) / Fraction(step)).denominator == 1 (on-grid),
abs(Fraction(candidate) - Fraction(raw)) <= Fraction(step) / 2 (the exact distance from the original input is no greater than half a step),
format(candidate, 'f') equals the literal expected string.
The complete test loop uses all twelve rows from the expected table:
import json
import platformcases = [
('Off-step value', '0.03', '0.05', '0.05'),
('Already on step', '1.20', '0.05', '1.20'),
('Positive even lower tie', '0.025', '0.05', '0.00'),
('Positive odd lower tie', '0.075', '0.05', '0.10'),
('Negative tie to zero', '-0.025', '0.05', '0.00'),
('Negative nonzero tie', '-0.075', '0.05', '-0.10'),
('Quarter off-step value', '1.13', '0.25', '1.25'),
('Quarter even lower tie', '1.125', '0.25', '1.00'),
('Quarter odd lower tie', '1.375', '0.25', '1.50'),
('Negative quarter tie', '-1.375', '0.25', '-1.50'),
('Zero', '0', '0.05', '0.00'),
('Upper-bound carry', '99999999.999999', '0.05', '100000000.00'),
]
require(len(cases) == 12, "Expected exactly twelve cases")
case_ids = [case_id for case_id, raw, step, expected_str in cases]
require(len(set(case_ids)) == 12, "Duplicate or missing case ID")def contract_checks(raw, step, expected_str, result):
oracle = expected_fraction(raw, step)
value = Fraction(result)
return {
"literal_matches_oracle": Fraction(expected_str) == oracle,
"value_matches_oracle": value == oracle,
"on_grid": (value / Fraction(step)).denominator == 1,
"within_half_step": abs(value - Fraction(raw)) <= Fraction(step) / 2,
"text_matches": format(result, "f") == expected_str,
}evidence = []
for case_id, raw, step, expected_str in cases:
candidate = round_to_step(raw, step)
with localcontext(Context(
prec=28, rounding=ROUND_HALF_EVEN, Emin=-999, Emax=999,
capitals=1, clamp=0, flags=[],
traps=[decimal.InvalidOperation, decimal.DivisionByZero,
decimal.Overflow, decimal.FloatOperation],
)) as diagnostic:
diagnostic.clear_flags()
naive = Decimal(raw).quantize(
Decimal(step), rounding=ROUND_HALF_EVEN
)
checks = contract_checks(raw, step, expected_str, candidate)
for name, passed in checks.items():
require(passed, f"{case_id}: failed {name}")
evidence.append({
"case": case_id, "input": raw, "step": step,
"expected": expected_str,
"naive_result": format(naive, "f"),
"corrected_result": format(candidate, "f"),
"on_grid": checks["on_grid"], "checks": checks,
})
require(len(evidence) == 12, "Missing result")control = contract_checks("0.03", "0.05", "0.05", Decimal("0.03"))
require(not control["on_grid"], "Negative control passed membership")
require(not control["value_matches_oracle"], "Negative control matched")These checks use explicit exceptions rather than bare assert statements, so they are not removed by Python’s optimization mode. A failed comparison stops the run and reports the condition. Completion means all twelve supplied cases matched the declared contract; it is not a claim about every possible input or downstream calculation.
Record numeric evidence and output text separately
For auditability, we capture the raw inputs and results in a structured log (JSON). Each entry contains: case ID, original raw string, step, expected output (string), naive quantize result (string), corrected result (string), whether the corrected result is on the grid (True/False), and a note of the policy/context. This illustrative two-record excerpt separates numeric evidence from output text:
[
{
"case": "Off-step value",
"input": "0.03",
"step": "0.05",
"expected": "0.05",
"naive_result": "0.03",
"corrected_result": "0.05",
"on_grid": true
},
{
"case": "Already on step",
"input": "1.20",
"step": "0.05",
"expected": "1.20",
"naive_result": "1.20",
"corrected_result": "1.20",
"on_grid": true
}
]Use json.dumps(...) to serialize the completed evidence and the policy/context identity. The source strings and policy drive the output. A missing or mismatched result triggers an exception, not a rewrite of the expected data. Keep illustrative excerpts separate from the complete twelve-case log used for a release decision.
Make invalid and boundary inputs fail visibly
We also test that all known invalid inputs are rejected by our function (raising ValueError). The following pairs should each cause an error (note the variety: non-string types, out-of-grammar formats, or disallowed steps):
("NaN", "0.05"), “NaN” not digits
("Infinity", "0.05"), not numeric literal
(0.03, "0.05"), wrong type (float)
(True, "0.05"), wrong type (bool)
("1e2", "0.05"), scientific notation disallowed
("01", "0.05"), leading zero
("0.0000001", "0.05"), too many decimals
("100000000", "0.05"), 9 integer digits (beyond 8)
("1", "0"), step zero (invalid step)
("1", "-0.05"), negative step (not in STEPS)
("1", "0.03"), unapproved step value
("1", "0.050"), step format not exactly "0.05"
Each of these should produce a ValueError. Append this rejection check and serialize the full report only after it completes:
invalid_cases = [
('NaN', '0.05'),
('Infinity', '0.05'),
(0.03, '0.05'),
(True, '0.05'),
('1e2', '0.05'),
('01', '0.05'),
('0.0000001', '0.05'),
('100000000', '0.05'),
('1', '0'),
('1', '-0.05'),
('1', '0.03'),
('1', '0.050'),
]
require(len(invalid_cases) == 12, "Expected twelve invalid cases")
rejected = 0
for raw, step in invalid_cases:
try:
round_to_step(raw, step)
except ValueError:
rejected += 1
else:
raise RuntimeError(f"Invalid input accepted: {raw!r}, {step!r}")
require(rejected == 12, "Missing input rejection")report = {
"policy": {
"steps": sorted(STEPS), "rounding": "ROUND_HALF_EVEN",
"input_grammar": GRAMMAR.pattern,
"output_text": "two fractional digits; zero is 0.00",
},
"runtime": {
"implementation": platform.python_implementation(),
"python": platform.python_version(),
"decimal": decimal.__version__,
"libmpdec": decimal.__libmpdec_version__,
},
"context": {
"prec": 28, "rounding": "ROUND_HALF_EVEN",
"Emin": -999, "Emax": 999, "capitals": 1, "clamp": 0,
"initial_flags": [],
"initial_traps": ["InvalidOperation", "DivisionByZero",
"Overflow", "FloatOperation"],
"inexact_trapped_during": ["division", "reconstruction"],
},
"cases": evidence, "invalid_cases_rejected": rejected,
"negative_control_rejected": True,
}
print(json.dumps(report, indent=2))We keep the upper-bound rounding (“99999999.999999” → “100000000.00”) separate as a valid carry case, but "100000000" with no decimal point fails grammar. Note that "0.050" looks like 0.05, but it is not an exact member of STEPS. It would match the numeric input regex, but the step argument has its own stricter contract; we deliberately reject it. Any extension of the accepted grammar (e.g. allowing "0.050") would require a conscious policy and revised tests.
Diagnose precision and policy changes before widening support
Any change to the domain requires re-evaluation. For instance:
25. Wider input precision: Allowing more than 6 decimals or larger integers adds new combinations. We would need to revise the input bounds and independently update the manifest and oracle, and verify nothing breaks.
Huge exponents: Supporting scientific notation or very large/small numbers is a new design; we’d need to reevaluate rounding logic for those edge cases.
New step values: If step values beyond 0.05/0.25 are added (say 0.10), we must explicitly document and code the new step list and re-run all tests. The current qualification covers only 0.05 and 0.25.
Different tie rules: Switching to “ties away from zero” or another rule changes the oracle completely. That is a semantic change, requiring new expected values and tests (no longer just removing a trap flag).
Signed zero: Our code normalizes any zero to +0.00. If business rules required retaining a “-0.00” in some context, that would again be a policy change needing reevaluation.
In short, we do not casually disable traps or increase precision without a full review. Extending the API (more steps, new input forms) requires a reviewed contract and, where the policy changes, an updated oracle. By analogy, reviewing a JSON Schema extension means reviewing which new fields or patterns the declared schema will accept, rather than adding ad hoc exceptions. We take the same approach here: any widening of support is treated as a new versioned change, with its own tests and decision process.
Distinguish unsupported inputs from implementation defects
We classify errors clearly:
A caller passing an input not matching the contract (malformed number or step) should get a ValueError; this is by design (a contract violation), not a bug. Likewise, passing an unapproved step string is a policy rejection.
By contrast, a valid input that disagrees with our oracle indicates a bug in round_to_step. For instance, if round_to_step("0.075","0.05") returned 0.05 instead of the expected 0.10, that would be a logic error. In testing we catch this with an exception, not a silent fallback.
We never fall back to regular float or round() on invalid inputs; that would hide defects. Unapproved or malformed inputs must fail loudly. Any arithmetic signal (Overflow, DivisionByZero) that occurs with valid inputs is also treated as a test failure.
Recompute incorrect outputs from preserved inputs
If this rounding function has already been run in production on some data, we must preserve the original raw inputs to correct past errors. Ideally, the system’s audit log or database retains the raw source strings. After fixing round_to_step(), we would recompute any affected records by passing the original raw values again. We would then present the downstream owner (or data steward) with a report of “old vs new” values. This way, changes are visible and can be approved. We do not infer the raw input by looking at a rounded value (that’s lossy). If only the output is stored, we would have to mark it “inexact” and seek the true source data.
In general, when building data pipelines, keep a copy of the original values or a reproducible history. This supports controlled data cleaning: preserve the raw values, then transform them. A pandas DataFrame containing only rounded outputs cannot recover the discarded information. If raw inputs are missing, hold any exact-reconstruction claim and escalate to the data origin owner rather than guess.
Set a release gate that can reject the wrong implementation
Before approving this fix for production, we use a strict checklist as a gate. The implementation is acceptable only if all of the following hold:
Contract frozen: The input grammar, step list, and tie rule are fully documented and unchanged.
Test manifest complete: Our test cases (including ties and boundary values) cover the declared comparison, including positive and negative ties and the upper-bound carry. We have exactly 12 cases plus invalid inputs tested.
Code passes all tests: Running the code on the manifest produces exactly the expected results and rejects all invalid cases.
Independent verification: An outside reviewer runs the same manifest and independent oracle, confirming agreement. (Both the code and oracle should be provided as artifacts.)
Deterministic formatting: Outputs use consistent formatting (e.g. always “0.00” for zero, no extraneous signs or exponents).
Negative control check: The naive example (e.g. 0.03→0.03) is included in the tests and must fail our membership check. If our corrected code ever accepted 0.03 as valid for step 0.05, we would reject the release.
In practice, we implement all these checks as automated tests. For instance, we feed the naive off-step result into the same verification logic and explicitly require that it not match the expected output. If it did (or if our code accepted it as “on-grid”), the test fails and we do not proceed. Only when the code correctly flags the control case do we consider it correct. In other words, the test harness itself contains a built-in “sanity check” using the failing counterexample. An implementation that would accept the off-step value has not truly met the contract and cannot pass the gate.
Preserve context isolation when tests fail
The localcontext context manager restores the previous decimal context when its block exits, including when an exception is raised. It does not erase flags from an object retained after the block. The candidate and diagnostics therefore start from an explicitly specified context and clear flags before monitoring arithmetic. These calculations do not rely on external fixtures or files. Like pytest fixture cleanup, the purpose is to prevent one failed test from contaminating another. Keep a failed run’s evidence separate from subsequent runs and from the final approval; an earlier output file must never stand in for a freshly completed run.
Assign the rounding policy and the repair decision
Finally, we capture the decision. We use a simple decision matrix:
Decision | When | Required Artifacts | Owner/Approver |
ACCEPT | All tests pass; no oracle mismatch | Test manifest, code, JSON log | QA/Engineer Team |
REPAIR & RECOMPUTE | Tests found mismatches but raw inputs exist | Bug report, fixed code, raw inputs | Developer, Data Owner |
HOLD | Policy or provenance uncertain or missing | Document request, policy note | Product Manager |
REQUALIFY | Requirements expanded (new steps or rules) | Revised spec, new tests, code | Product & QA Managers |
The policy (step list, tie rule) is owned by the product/domain owner. They must formally approve any changes to it.
The implementation is owned by the software engineer. They provide the code, tests, and demonstration artifacts.
An independent reviewer (QA or peer) verifies that the implementation matches the policy, by reviewing the test evidence.
Passing this transformation (moving to nearest 0.05 with banker’s rounding) does not certify all downstream calculations. It qualifies the small, isolated function against the stated contract and evidence. Any new use cases (e.g. different currency conversions) would require re-running a similar review.
Practice reviewed implementations through software engineering projects
In summary, the result of this effort is a complete, auditable artifact: a clearly documented input contract, a rational oracle, well-isolated code, and a concise release record. The artifact shows how the declared inputs are handled, how the checks are performed, and who is responsible for each decision. This end-to-end review process is the kind of practical work that builds real expertise in software development.
If you found this approach helpful, consider exploring the Refonte Learning Software Engineering program. It covers the full software lifecycle in a hands-on way (system design, testing, cloud concepts, performance, etc.) and culminates in a capstone project. The same skills, precise implementation, testing against a specification, and careful documentation, are exactly what we practiced here.
