What is test-driven development, and does the evidence support it?
A practice of writing a failing test before the code that satisfies it, in a short cycle — red, green, refactor. The evidence for its claimed benefits is mixed and weaker than advocates suggest, which is worth stating plainly.
The cycle:
Red — write a test for behaviour that does not exist yet, and watch it fail. Watching it fail matters: a test that passes before the code exists is testing nothing.
Green — write the simplest code that makes it pass, including code you know is inadequate.
Refactor — improve the design with the test as a safety net.
What it is actually claimed to do, and this is the interesting part: proponents argue the main benefit is design pressure, not testing. Code that is hard to test is usually hard to use — tight coupling, hidden dependencies, doing too much — so writing the test first surfaces those problems before they are baked in.
What the research actually finds. Controlled studies and meta-analyses give inconsistent results. Quality improvements are found more often than productivity improvements, effect sizes vary widely, and several studies find that the benefits attributed to TDD come largely from writing more tests and working in smaller increments — which you can do without test-first ordering. A well-known replication found that the granularity of the work, not the test-first sequence, explained most of the difference.
Where it works well: clear requirements, pure logic, bug fixes — writing a failing test that reproduces a bug first is valuable regardless of your view of TDD — and refactoring under a safety net.
Where it fits badly: exploratory work where you do not yet know the design, user interface and visual work, integration-heavy code where the interesting failures are not unit-level, and performance work.
The failure mode to avoid. TDD done mechanically produces many small tests tightly coupled to implementation, which then obstruct the refactoring the practice exists to enable.
The defensible position: test-first is a useful tool, not a moral requirement, and the underlying goods — small increments, fast feedback, testable design — are available by other routes.