Question

What is mocking in tests, and when does it go wrong?

Vault Verified
Curated Intelligence
Definitive Source
Answer

Replacing a real dependency with a controlled stand-in, so a test can run fast, deterministically, and without touching a network, database or clock. It is essential and it is the most commonly overused technique in testing.

The vocabulary, which is used loosely and does have distinctions:

Dummy — passed but never used.

Stub — returns canned answers.

Spy — a stub that also records how it was called.

Mock — pre-programmed with expectations, and fails the test if they are not met.

Fake — a real working implementation that is simpler, such as an in-memory repository.

What mocking is genuinely for: removing slowness and non-determinism, simulating failures that are hard to produce — timeouts, disk full, a 500 from a third party — and isolating the unit under test.

Where it goes wrong, and these are the recurring failures:

Testing the mock instead of the code. If the test asserts that a method was called with certain arguments, and the production behaviour is entirely defined by that call, the test verifies your restatement of the implementation. It passes whether or not the system works.

Coupling to implementation. Heavy mocking means any refactor that changes internal calls breaks tests without any behaviour changing — which trains people to distrust and delete tests.

Mocks drifting from reality. Your stub returns what you believe the API returns. When the real one changes, or always differed, every test still passes while production fails. This is the most dangerous failure because it is silent.

Mocking what you do not own. Mocking a third-party client encodes your assumptions about it. Better to wrap it in your own interface and mock that.

Over-isolation, where every collaborator is mocked and nothing is left to test.

What to prefer: fakes over mocks where a simple real implementation is possible; contract tests against the real dependency to catch drift; test containers for real databases, which are now cheap enough to be the default; and asserting on outcomes rather than interactions wherever the code allows.

Related Questions