In Python testing, it’s all too easy for a mock to mask a bug. A green test (one that passes) can be misleading if the test double was more lenient than the actual dependency. For example, a mock might silently accept invalid arguments or return a placeholder result, while the real function would raise an error. This discrepancy creates risk: a passing test suite does not guarantee correct integration with the real code. We need a clear acceptance-and-repair playbook for such cases: whether to accept the verified double, repair the test contract, hold the integration claim, or run a real-boundary check.
A faithful test double must satisfy four contracts: it should only expose the same attributes and methods as the real dependency, allow exactly the same call shapes (positional/keyword parameters), return an object of the correct type with expected contents, and defer business-rule validation to real code (see below). Simply getting a green test result is not enough to trust this alignment. As one QA roadmap suggests, checking each layer and interface carefully is part of robust testing. Moreover, Python type hints are not enforced at runtime: “the Python runtime does not enforce function and variable type annotations”. In practice this means even if the mock’s return value is typed, Python won’t stop you from returning something else.
This guide will demonstrate a hands-on lab for one concrete synchronous interface, using only CPython standard library tools (unittest, unittest.mock, inspect, dataclasses) under Python 3.13.x. We’ll record exactly what happens when mocks and autospecced doubles handle valid and invalid calls, show how to patch the correct name, and keep business-rule tests separate. For consistency, we refer to CPython 3.13 documentation (autospec, spec_set rules, inspect signature binding, etc.), noting any behavior observed in our 3.13 test run. All test fixtures and commands are provided below so you can reproduce this lab without external dependencies or credentials.
Define what a faithful test double must prove
A correct test double must replicate its real counterpart in four ways:
Attribute contract: The mock should only expose the attributes, methods, and signature of the real object. A plain Mock() can create arbitrary attributes on access, but a well-designed double (using spec or spec_set) only allows those in the interface.
Call-signature contract: The double should accept exactly the same positional and keyword parameters (names and positions) as the real method. If the mock allows a call shape that the real function rejects, the test is unsound. Python’s create_autospec(..., spec_set=True) sets up a mock that mimics the real signature. A failing call (e.g. missing a keyword-only argument) should raise TypeError exactly as it would in production.
Result-object contract: The double should return a realistic object. In our example, the real interface returns a Reservation object (a dataclass instance). The mock should not simply return another mock or arbitrary value; we explicitly configure it to return a Reservation("MOCK"). Tests must then assert on that returned value. We will check that the consumer’s result is indeed a Reservation with the correct reservation_id. Otherwise a loosely-typed mock might “succeed” with a meaningless object.
Business contract: Business rules (e.g. “units must be positive”) are enforced by the real code and should be tested separately. For example, if the real method raises ValueError on negative units, the autospecced mock will not raise; it only checks the signature. We must keep such logic tests (“units=-1” should raise) outside the mock-based tests. Mock-based tests cover only the interface shape, not business logic validation.
A green test suite alone is insufficient evidence. If only the interface was checked, releasing code requires reviewing these four aspects. The end result is a decision matrix:
ACCEPT the double if all four contracts are satisfied (signature and attributes match, result checked).
REPAIR the test if the double is too permissive (e.g. accepted an invalid call) by tightening spec, switching to autospec, or patching the correct name.
HOLD the release if evidence is incomplete (e.g. only signature validated, not business rules).
RUN A REAL-BOUNDARY CHECK (integration test) before asserting true compatibility with the provider.
This approach avoids over-relying on mocks for rules they can’t enforce, and aligns tests with real behavior. (For more on QA layers and systematic testing, see the discussion of test layers and automation projects in a 2026 QA guide.)
Build a side-effect-free callable reference
First, define the real interface in lab_gateway.py. This is a simple class with no external I/O, just logic so we can trust its behavior. We give it a keyword-only parameter to test signature checking, and a frozen dataclass Reservation to avoid mock chaining. For example:
# lab_gateway.py
from dataclasses import dataclass
@dataclass(frozen=True, slots=True)
class Reservation:
reservation_id: str
class Gateway:
def reserve(self, order_id: str, *, units: int,
warehouse: str = "A") -> Reservation:
if type(units) is not int or units <= 0:
raise ValueError("units must be a positive integer")
return Reservation(f"{order_id}:{warehouse}:{units}")This Gateway().reserve("O1", units=2) will return Reservation("O1:A:2"). No real side effects occur (no network or database). We can safely call it in tests. In a separate file lab_consumer.py, we’ll import this and call it:
# lab_consumer.py
from lab_gateway import Gateway
def place_order():
return Gateway().reserve("O1", units=2)place_order() uses the imported Gateway name as patched in this module’s namespace.
Bind the method without inventing self. To get the independent oracle for call shapes, use inspect.signature on the bound method (so it doesn’t require providing self manually). For example, in Python shell:
>>> import inspect
>>> from lab_gateway import Gateway
>>> sig = inspect.signature(Gateway().reserve)
>>> sig
(S, order_id: str, *, units: int, warehouse: str = 'A') # pseudo-output
>>> bound = sig.bind("O1", units=2)
>>> bound.arguments
{'order_id': 'O1', 'units': 2, 'warehouse': 'A'}If we try sig.bind("O1", 2), Python raises a TypeError because reserve expects units as keyword-only. Using Signature.bind checks only the signature; it does not call the function or enforce business logic. We will rely on this for expected call-acceptance or TypeError outcomes.
Record the call cases before creating mocks
Next, enumerate the specific call cases we will test. These cases are defined by positional and keyword arguments and what the Gateway.reserve signature should allow or reject:
Case | Positional args | Keyword args | Expected (signature) |
Valid | ("O1",) | units=2 | Accepted |
Missing units | ("O1",) | (none) | TypeError |
Misspelled parameter | ("O1",) | quantity=2 | TypeError |
Keyword-only positionally | ("O1", 2) | (none) | TypeError |
Unexpected keyword | ("O1",) | units=2, region="X" | TypeError |
Each of the “TypeError” cases comes from Python’s argument binding rules: either a required keyword was not provided, or an unknown keyword was given, or a keyword-only argument was given positionally. These outcomes are what inspect.signature(...).bind(...) (and likewise calling the real method) enforce. We record these expected results independently before we even involve any mocks.
At this point, we haven’t generated any mocks; we’re just documenting what the true interface does. (As one QA article notes, understanding automation trends and test strategy means clearly defining expected outcomes for each test case.)
Expose the permissive green test
Now, create mocks and see what they do with those cases. First, the plain Mock with no spec:
from lab_gateway import Reservation, Gateway
from unittest.mock import Mock
plain_double = Mock()
plain_double.reserve.return_value = Reservation("MOCK")This configures plain_double so that any call to plain_double.reserve(...) returns Reservation("MOCK"), regardless of arguments. For example, if our consumer accidentally calls plain_double.reserve("O1", quantity=2) (the parameter name is wrong), the plain mock will happily return "MOCK". It does not enforce the real signature; it only cares that .reserve was configured to return something.
To illustrate:
result = plain_double.reserve("O1", quantity=2)
print(result.reservation_id) # Output: MOCKEven though quantity is not a parameter of reserve, no TypeError is raised. The plain mock simply ignored the fact that the argument name is wrong. This demonstrates the permissive behavior: a malformed call slips through. In a real test suite, that would appear as a passing test (the double returned something), but it would mask the real bug. The green test creates a false sense of success.
Compare spec and spec_set without conflating their roles
Next, compare Mock(spec=Gateway) and Mock(spec_set=Gateway). Both create mocks that know about the attributes of Gateway, but spec_set is stricter (as the docs say, “If used, attempting to set or get an attribute on the mock that isn’t on the object passed as spec_set will raise an AttributeError”). Set up fresh doubles similarly:
spec_double = Mock(spec=Gateway)
spec_set_double = Mock(spec_set=Gateway)
for d in (spec_double, spec_set_double):
d.reserve.return_value = Reservation("MOCK")When we make calls with the five test cases above, surprisingly both spec and spec_set also accept all of them without error in this particular example. That’s because spec=Gateway (and spec_set=Gateway) both allow calling the method; the difference shows up in attribute access, not in call invocation. (The docs specifically note that with a spec, the mock introspects the signature on method calls, so it can match arguments by position/name, but if the signature doesn’t match at all, a TypeError should arise. It can detect missing args, but in practice the plain Mock and Mock(spec=Gateway) both ignored our wrong-parameter names. This is due to how Mock vs autospec works in current CPython.)
For attribute checks:
Accessing an attribute that doesn’t exist on Gateway, e.g. mock.resreve (misspelling reserve), behaves differently. A plain Mock returns a new mock; Mock(spec=Gateway) will raise AttributeError; and Mock(spec_set=Gateway) will also raise AttributeError on reads. This is because spec_set forbids both get and set of unknown attributes, while spec by itself forbids only setting unknown attributes (though in practice our spec mock also disallowed the misspelled read).
Setting a new attribute mock.new_field = 1 raises for spec_set and for an autospec mock (when instance/spec_set used) but does not raise for spec or plain mocks.
Summarizing these behaviors:
Plain mock: allows any new attribute read or write.
Mock(spec=Gateway): forbids reading a nonexistent attribute, but allows writing one.
Mock(spec_set=Gateway): forbids both reading and writing any attribute not defined on Gateway.
Autospec (instance with spec_set=True): likewise forbids both.
These attribute checks answer a different question: they ensure we’re not accidentally using parts of the spec that aren’t real. They do not replace the call-shape checks below. Still, we see that using spec_set (or autospec with spec_set) is the stricter setting, as documented.
Apply autospec to the declared instance interface
To faithfully catch bad calls, we use create_autospec on our Gateway class. This creates a mock that has the same method signature as Gateway().reserve, and we specify instance=True so the mock represents an instance:
from unittest.mock import create_autospec
auto_double = create_autospec(Gateway, instance=True, spec_set=True)
auto_double.reserve.return_value = Reservation("MOCK")
The auto_double now enforces the real call signature. A correct call still returns our fake reservation:
valid_result = auto_double.reserve("O1", units=2) # OK
print(valid_result.reservation_id) # MOCKThis matches the binder and real result (aside from business logic; here 2 is valid). If we call auto_double.reserve("O1", units=2, warehouse="B"), it works too because the signature allows the optional warehouse. The return value is our configured Reservation("MOCK"), matching the type the consumer expects.
Make invalid calls fail at invocation
Crucially, with the autospecced mock, all the malformed call cases raise TypeError at the call boundary, not just later. This enforces our call-shape contract. For each of the four invalid cases above (missing keyword, wrong name, keyword-only positional, extra keyword), do:
with self.assertRaises(TypeError):
auto_double.reserve("O1", **{}) # missing unitswith self.assertRaises(TypeError):
auto_double.reserve("O1", quantity=2) # wrong kw namewith self.assertRaises(TypeError):
auto_double.reserve("O1", 2) # keyword-only used positionallywith self.assertRaises(TypeError):
auto_double.reserve("O1", units=2, region="X") # unexpected kwEach TypeError here is exactly what our signature oracle predicted. Note: this behavior comes from the autospec signature introspection; the mock itself does not check argument types or custom logic, only that the call signature matches the function’s signature. We use assertRaises(TypeError) (the standard unittest pattern) to assert that these invalid calls are correctly rejected, turning our expected-failure scenario into a passing test outcome (it is a success to have the error, because it indicates the mock caught the issue).
Turn the matrix into executable acceptance tests
Now we consolidate these checks into a unittest-based test suite. Below is the complete test_contract.py with one test method per case. Each test uses a fresh mocked Gateway, an explicit Reservation("MOCK") return, and asserts either success or failure as appropriate. We use deterministic identifiers for clarity.
# test_contract.py
impbort unittest
from lab_gateway import Gateway, Reservation
from lab_consumer import place_order
from unittest.mock import Mock, create_autospec, patch
class TestContract(unittest.TestCase):
def setUp(self):
# Real gateway for reference
self.real_gateway = Gateway()
# Base reservation object for wrong-call tests
self.mock_res = Reservation("MOCK")
def test_valid_call_autospec(self):
double = create_autospec(Gateway, instance=True, spec_set=True)
double.reserve.return_value = self.mock_res
result = double.reserve("O1", units=2)
# The mock should have returned our fake Reservation
self.assertIsInstance(result, Reservation)
self.assertEqual(result.reservation_id, "MOCK")
def test_missing_units_autospec(self):
double = create_autospec(Gateway, instance=True, spec_set=True)
double.reserve.return_value = self.mock_res
with self.assertRaises(TypeError):
double.reserve("O1")
def test_misspelled_param_autospec(self):
double = create_autospec(Gateway, instance=True, spec_set=True)
double.reserve.return_value = self.mock_res
with self.assertRaises(TypeError):
double.reserve("O1", quantity=2)
def test_keyword_positional_autospec(self):
double = create_autospec(Gateway, instance=True, spec_set=True)
double.reserve.return_value = self.mock_res
with self.assertRaises(TypeError):
double.reserve("O1", 2)
def test_unexpected_keyword_autospec(self):
double = create_autospec(Gateway, instance=True, spec_set=True)
double.reserve.return_value = self.mock_res
with self.assertRaises(TypeError):
double.reserve("O1", units=2, region="X")
def test_plain_mock_allows_invalid(self):
# Negative control: plain Mock should accept bad call shapes
plain = Mock()
plain.reserve.return_value = self.mock_res
result = plain.reserve("O1", quantity=2)
self.assertIsInstance(result, Reservation)
# It returns our fake reservation without error, which is a failing scenario
# if unchecked
def test_result_type_and_value(self):
# Check that place_order returns a Reservation with the expected ID
res = place_order()
self.assertIsInstance(res, Reservation)
self.assertEqual(res.reservation_id, "O1:A:2")
def test_invalid_business_rule(self):
# Negative units should be caught by real code, not by mock
with self.assertRaises(ValueError):
self.real_gateway.reserve("O1", units=-1)
def test_patch_gateway_not_used(self):
# Patching lab_gateway.Gateway should not affect lab_consumer (it uses its
# own import)
with patch('lab_gateway.Gateway') as mock_gateway:
# Configure patch (which will not actually be invoked)
mock_gateway.return_value.reserve.return_value = Reservation("PATCHED")
result = place_order()
# The patch should not be used, so result is the real one
self.assertEqual(result.reservation_id, "O1:A:2")
self.assertFalse(mock_gateway.called)
def test_patch_consumer_gateway(self):
# Correctly patching lab_consumer.Gateway
with patch('lab_consumer.Gateway', autospec=True) as mock_consumer_gateway:
mock_consumer_gateway.return_value.reserve.return_value = (
Reservation("PATCHED")
)
result = place_order()
self.assertEqual(result.reservation_id, "PATCHED")
mock_consumer_gateway.assert_called_once()A few notes on this test suite:
We use create_autospec for the call-signature matrix tests, and assertRaises(TypeError) for the invalid calls. These are positive tests: catching the expected error means the test passes. The suite will produce a nonzero exit code if any unexpected exception is thrown.
We explicitly check the result of a valid call in test_result_type_and_value, asserting both the type and the reservation_id string. This ensures we caught the mocked return (even though we configured "PATCHED" in the patch test below, here we used the real place_order() to get "O1:A:2").
The negative controls test_plain_mock_allows_invalid and test_invalid_business_rule show that a plain Mock does not reject the invalid call (thus if our test suite relied on it, it would incorrectly pass), whereas the real gateway instance raises ValueError for negative units.
The last two tests demonstrate patching pitfalls: patching lab_gateway.Gateway has no effect on the already-imported Gateway in lab_consumer, so mock_gateway should remain unused. Patching lab_consumer.Gateway is the correct fix; we verify it is called exactly once and returns a Reservation("PATCHED").
Run the tests with higher verbosity to list each case:
$ python3.13 -m unittest -vtest_invalid_business_rule (test_contract.TestContract) ... ok
test_keyword_positional_autospec (test_contract.TestContract) ... ok
test_missing_units_autospec (test_contract.TestContract) ... ok
test_misspelled_param_autospec (test_contract.TestContract) ... ok
test_patch_consumer_gateway (test_contract.TestContract) ... ok
test_patch_gateway_not_used (test_contract.TestContract) ... ok
test_plain_mock_allows_invalid (test_contract.TestContract) ... ok
test_result_type_and_value (test_contract.TestContract) ... ok
test_unexpected_keyword_autospec (test_contract.TestContract) ... ok
test_valid_call_autospec (test_contract.TestContract) ... ok
----------------------------------------------------------------------
Ran 10 tests in 0.005s
OKAnd here is a summary in JSON of the key cases, showing expected vs actual outcomes (example run on CPython 3.13.x, Linux):
{
"interpreter": "CPython 3.13.6 on linux (x86_64)",
"cases": [
{
"id": "Valid",
"patch_target": "None",
"mock_type": "autospec",
"args": ["O1"],
"kwargs": {"units": 2},
"expected_exception": null,
"actual_exception": null,
"expected_result": {"reservation_id": "O1:A:2"},
"actual_result": {"reservation_id": "MOCK"},
"passed": true
},
{
"id": "Missing units",
"patch_target": "None",
"mock_type": "autospec",
"args": ["O1"],
"kwargs": {},
"expected_exception": "TypeError",
"actual_exception": "TypeError",
"passed": true
},
{
"id": "Misspelled parameter",
"patch_target": "None",
"mock_type": "autospec",
"args": ["O1"],
"kwargs": {"quantity": 2},
"expected_exception": "TypeError",
"actual_exception": "TypeError",
"passed": true
},
{
"id": "Keyword-only positionally",
"patch_target": "None",
"mock_type": "autospec",
"args": ["O1", 2],
"kwargs": {},
"expected_exception": "TypeError",
"actual_exception": "TypeError",
"passed": true
},
{
"id": "Unexpected keyword",
"patch_target": "None",
"mock_type": "autospec",
"args": ["O1"],
"kwargs": {"units": 2, "region": "X"},
"expected_exception": "TypeError",
"actual_exception": "TypeError",
"passed": true
}
],
"exit_code": 0
}This JSON reports, for each case, whether the double (with autospec) raised the expected exception, and what (fake) result was returned. The "expected_result" for the valid case comes from the real interface (via Reservation("O1:A:2")), while "actual_result" is the mock’s configured output ("MOCK"). All expected TypeErrors were indeed raised, and no unexpected exceptions occurred. The exit code 0 indicates all tests passed as designed.
At this point, all declared call cases agree with the independent signature oracle, and we have proven the consumer uses the real lookup (patching tests). We explicitly checked return types/values. We also captured that business-rule failures (units≤0) are only caught by the real Gateway and must remain in a separate test.
Patch the name the consumer actually resolves
The lab_consumer.py imported Gateway at module load time. This means patches must target lab_consumer.Gateway, not lab_gateway.Gateway. We demonstrate both:
Wrong patch target:
We patch('lab_gateway.Gateway') in a context. This replaces Gateway in lab_gateway, but lab_consumer.place_order() still uses its own Gateway reference. In our test, we assert that mock_gateway was never called and the returned reservation ID remains "O1:A:2". This confirms the patch didn’t intercept the call (as expected).Correct patch target:
We use patch('lab_consumer.Gateway', autospec=True, spec_set=True). Now, when place_order() is called, it uses the patched Gateway. We configure the patch so that the mocked reserve returns Reservation("PATCHED"). Our test then verifies that place_order() returns "PATCHED" and that the mock was indeed called once.
Throughout, no real network calls occur; it’s all in-memory and isolated. (For patch examples, see Python’s mock docs on using patch as a context manager.)
Prove that the ineffective patch stayed uncalled
It’s important to note that simply entering a with patch(...) block does not guarantee it’s used. Only when the code under test actually references that patched name will the mock be invoked. In our negative-control test (test_patch_gateway_not_used), we showed exactly that: the patch object remained untouched. The test asserts mock_gateway.called is false, proving the patch had no effect. Only after switching the target to lab_consumer.Gateway did the mock get invoked.
By restoring the correct patch target (and using autospec to enforce the signature), we repair the test. The consumer module now correctly uses the mock, and the test’s expectation (“PATCHED” result) passes. This fixes the drift between where the test thought it was patching and where the code actually looked up Gateway.
Keep results from becoming unconfigured child mocks
Even with correct call-shape checking, we must ensure the returned object is meaningful. A common mistake is to rely on "truthiness" (e.g. assertTrue(res)) when a mock was returned. Instead, always assert on the properties of the object. In our case, the consumer’s return should be a Reservation instance with a specific reservation_id.
We explicitly checked this in test_result_type_and_value and in the patch test. By asserting isinstance(..., Reservation) and comparing the string field, we avoid a false pass where the mock returned some generic Mock or incorrect value. Python will happily return a different type if you misconfigure the mock, since it does not enforce annotations. For example, if we inadvertently set reserve.return_value = Reservation("WRONG"), a test that only checked assertTrue(res) would pass, but our explicit assertEqual(res.reservation_id, "O1:A:2") would catch the mistake. In short, do not confuse “it didn’t crash” with “it gave the right result.”
Retain business-rule tests outside the double
Finally, remember that autospec enforces signature but not logic. In our Gateway, the rule “units must be positive int” is enforced by code (ValueError). An autospecced mock will not check this; if we passed units=-1 to the mock, it would happily return "MOCK" because the signature is satisfied. Thus, tests for domain rules belong in real (or integration) tests. For example, we include:
with self.assertRaises(ValueError):
Gateway().reserve("O1", units=-1)This confirms the real method enforces the rule. We do not expect the mock to do it. As the typing docs remind us, annotations or signature checks are only for static analysis, not for validating values. We keep these business-rule checks separate so that an autospecced test cannot falsely claim the business logic was verified. (See also how API test responsibilities separate interface vs logic in API testing and debugging responsibilities.)
Name the validation autospec does not perform
In summary, autospec does validation of parameter names and positions, but does not validate parameter values or business rules. It won’t check types (it only checks syntax), or domain rules like “units > 0”. Any such “validation” in your test code actually comes from the real function’s logic. Do not rely on autospec=True to guarantee correct values. Always explicitly write separate tests for negative values or types (e.g. via assertRaises on the real object).
Handle introspection limits without production side effects
The above experiments assume an interface that’s easy to introspect: a plain class with a visible reserve method. In more complex cases, autospec can have blind spots. For example, if an attribute is added to instances at runtime (in init) rather than defined on the class, Mock(spec=Class) may not know about it. Or if methods use @property or other descriptors, creating a mock may not capture them correctly unless you explicitly use PropertyMock or allow calls through wraps.
In general, autospec builds its spec from the class (or instance) you give it. It does not execute init, so any attributes created there will be missing. It also won’t enforce any complex logic or decorators not visible at class definition time. The autospec documentation cautions that autospec “creates mock objects that have the same attributes and methods as the objects they are replacing”, but it cannot predict dynamically-added attributes or private API.
The safe decision in those edge cases is often to hold the mock-based claim and require an integration or more targeted test. If your interface is not introspectable, it may be better to rely on explicit integration tests. Do not try to hack around this by adding side effects to the real class just to make mocking work; keep production code clean. Instead, document the limitation and ensure tests reflect it. For most plain classes this lab covers, autospec works fine, but always review whether the mock’s spec truly matches what the consumer uses at runtime.
Capture a reviewable test-contract record
We have now a complete set of tests. The three files (lab_gateway.py, lab_consumer.py, and test_contract.py) together form the suite. Here are their contents for clarity:
# lab_gateway.py
from dataclasses import dataclass
@dataclass(frozen=True, slots=True)
class Reservation:
reservation_id: str
class Gateway:
def reserve(self, order_id: str, *, units: int,
warehouse: str = "A") -> Reservation:
if type(units) is not int or units <= 0:
raise ValueError("units must be a positive integer")
return Reservation(f"{order_id}:{warehouse}:{units}")
# lab_consumer.py
from lab_gateway import Gateway
def place_order():
return Gateway().reserve("O1", units=2)
# test_contract.py
import unittest
from lab_gateway import Gateway, Reservation
from lab_consumer import place_order
from unittest.mock import Mock, create_autospec, patch
class TestContract(unittest.TestCase):
def setUp(self):
self.real_gateway = Gateway()
self.mock_res = Reservation("MOCK")
def test_valid_call_autospec(self):
double = create_autospec(Gateway, instance=True, spec_set=True)
double.reserve.return_value = self.mock_res
result = double.reserve("O1", units=2)
self.assertIsInstance(result, Reservation)
self.assertEqual(result.reservation_id, "MOCK")
def test_missing_units_autospec(self):
double = create_autospec(Gateway, instance=True, spec_set=True)
double.reserve.return_value = self.mock_res
with self.assertRaises(TypeError):
double.reserve("O1")
def test_misspelled_param_autospec(self):
double = create_autospec(Gateway, instance=True, spec_set=True)
double.reserve.return_value = self.mock_res
with self.assertRaises(TypeError):
double.reserve("O1", quantity=2)
def test_keyword_positional_autospec(self):
double = create_autospec(Gateway, instance=True, spec_set=True)
double.reserve.return_value = self.mock_res
with self.assertRaises(TypeError):
double.reserve("O1", 2)
def test_unexpected_keyword_autospec(self):
double = create_autospec(Gateway, instance=True, spec_set=True)
double.reserve.return_value = self.mock_res
with self.assertRaises(TypeError):
double.reserve("O1", units=2, region="X")
def test_plain_mock_allows_invalid(self):
plain = Mock()
plain.reserve.return_value = self.mock_res
result = plain.reserve("O1", quantity=2)
self.assertIsInstance(result, Reservation)
def test_result_type_and_value(self):
res = place_order()
self.assertIsInstance(res, Reservation)
self.assertEqual(res.reservation_id, "O1:A:2")
def test_invalid_business_rule(self):
with self.assertRaises(ValueError):
self.real_gateway.reserve("O1", units=-1)
def test_patch_gateway_not_used(self):
with patch('lab_gateway.Gateway') as mock_gateway:
mock_gateway.return_value.reserve.return_value = Reservation("PATCHED")
result = place_order()
self.assertEqual(result.reservation_id, "O1:A:2")
self.assertFalse(mock_gateway.called)
def test_patch_consumer_gateway(self):
with patch('lab_consumer.Gateway', autospec=True) as mock_consumer_gateway:
mock_consumer_gateway.return_value.reserve.return_value = (
Reservation("PATCHED")
)
result = place_order()
self.assertEqual(result.reservation_id, "PATCHED")
mock_consumer_gateway.assert_called_once()Running python3.13 -m unittest -v gave the verbose output shown above, which we include. Then we captured and assembled the JSON summary of results (shown previously).
We disclaim that this is a single-version lab (CPython 3.13.x on Linux) and covers one interface example. The observed behavior matches CPython docs: autospec enforced call signature as promised, and specs limited attributes. We do not claim these results hold for every environment or for asynchronous or C-accelerated methods. We also do not assert that autospec guarantees any business logic; on the contrary, our lab shows it does not (see negative units=-1 test).
Repair drift and rerun the real-boundary control
Based on our findings, we apply repairs as needed and document decisions. In this case, our autospec usage was correct for checking call shapes, so no fix was needed there. The main repair was adjusting the patch target for Gateway. Once we patched lab_consumer.Gateway correctly, the tests passed and the consumer used the intended mock.
To generalize the decision-making:
ACCEPT: If the double and real interface fully agree on all checked contracts (attributes, call shape, return type), we can accept the double as verified for release, pending final integration tests. For example, the “Valid” case under autospec passes signature check.
REPAIR: If the double diverged (e.g. plain Mock accepted invalid calls), we repair. In our lab, using create_autospec(..., spec_set=True) replaced the plain mock to reject bad calls. Also patching lab_consumer.Gateway replaced the ineffective patch on lab_gateway.
HOLD: If only part of the contract was checked, we hold the claim. For instance, we did not verify attribute access contracts in detail; if that were a concern, we would hold for additional tests. Or if only the mock’s behavior was observed without verifying output type, that would be incomplete evidence.
RUN REAL-BOUNDARY CHECK: Even after all repairs, we require an integration test with the actual dependency (or a tightly controlled sandbox of it) to confirm real compatibility. The autospec and mock tests show interface alignment, but not runtime reliability beyond this lab.
Each test and outcome above maps to one of these actions. Because our interface here is introspectable and our patches now correct, we would accept the autospec double for release testing, while also scheduling actual integration testing of Gateway.reserve. If this code were production, we’d now run an end-to-end test against a staging or mock service to ensure no surprises. Any change to the Gateway.reserve signature or location (import path) should trigger a re-run of this contract test suite before trusting new releases.
Use the four release decisions consistently
We can tabulate responsibilities:
Interface Owner (API developer): Must document and maintain the method signature and expected exceptions. If the signature changes (e.g. a new parameter is added), this test suite must be updated and re-run.
Consumer/QA Maintainer: Writes and maintains these contract tests. If a test fails (e.g. a double rejects a call unexpectedly), the maintainer either accepts the interface change (if intentional) or repairs the test (if the interface shouldn’t have changed).
Reviewer: Ensures that a green test suite includes these contract checks. The reviewer should confirm that a passing autospec test means true compatibility, and that business-rule violations are tested separately.
In any case, the test failure modes lead directly to our four actions: accept, repair, hold, or further integration testing. We explicitly did not rely on undocumented behavior (like automatically copying arguments, which we saw a plain Mock won’t do). Each acceptance check has a reason: matching signature (autospec) ensures no extra keyword was allowed; explicit assertEqual on the result ensures correct return content; real ValueError tests ensure domain rules are enforced.
Whenever we see “expected_failure = True” in a test, remember that in a contract test suite it’s a desired outcome (the test passed by catching an error). That’s how the gate works: to admit a change, the test must fail on any invalid inputs.
Assign ownership of the dependency contract
This contract between producer and consumer has clear owners. The Gateway interface owner is responsible for the published signature and documented behavior of reserve(). If they change it, they should communicate that change (otherwise consumers will fail binding). The consumer (place_order) maintainer is responsible for updating the tests (and possibly the code) if the interface changed. The QA team ensures that after any change to the signature, parameter names, or import path, this contract test suite is re-validated.
Concretely, if the Gateway signature is modified or moved, the test suite should break, indicating either the contract evolved or the tests need patching. For example, if reserve dropped the warehouse argument or gained a new one, our autospec.bind tests would start failing, signaling that a decision is needed.
In summary: the API provider ensures the official interface, the consumer ensures their calls match it, and QA ties it together by running these tests. This clear division prevents a benign-looking test break from being ignored; instead, it triggers a decision to accept, repair, or further test.
Build broader QA engineering foundations
This walkthrough showcased a focused skill: using unittest.mock’s spec and autospec to validate call signatures against the real code. For a broader automation and test design foundation, consider structured learning. Refonte Learning’s Quality Assurance Engineering program offers a curriculum on QA processes, automation frameworks, unit testing, and CI/CD best practices. It can reinforce concepts like test-driven design, dependency injection, and test contracts, all of which underpin robust QA. For example, understanding when and how to isolate components (vs full integration tests) is a core topic there.
For a broader foundation in test design and automation workflows, review Refonte Learning’s Quality Assurance Automation Engineering / Quality Assurance Engineering program and its published prerequisites. (Completion of that program includes a certificate and often practical projects that would cover dependency mocking in more depth.)
This concludes the acceptance-and-repair playbook. By following the steps above (defining cases, using autospec, patching correctly, and keeping business checks separate), Python QA engineers and reviewers can catch interface drifts early. The key is evidence: each mock test documents a call scenario, the outcome, and the decision. Armed with this record, teams can confidently decide when to ship code or hold for fixes, ensuring that a seemingly green test suite truly reflects reality.
