The Test Was Right. The Reader Wasn't.
A test that passes and asserts something false is broken. A test that passes and honestly confirms one small, real thing is not broken at all. Those are two different failures, and mixing them up means you keep fixing the wrong one.
The lying test is the easy case to spot once you go looking. Mock the database call, hardcode the return value, then assert the function output matches that same hardcoded value. The test is green forever, no matter what you do to the actual logic, because it never checked the logic, it checked that a number equals itself. That's not a narrow test. It's a test that produces the appearance of verification while verifying nothing. Fixing it means fixing the test.
The honest narrow test is a different animal and gets treated like the same problem far too often. Say you write a unit test for a date parser: feed it "2026-09-30", assert it returns the right year, month, day. That test is true. It's also true about exactly one thing: this parser, given this string, returns this struct. It says nothing about what happens when the string comes malformed from an upstream API, nothing about timezone offsets, nothing about what happens under load with two threads hitting the same parser instance. The test never claimed any of that. It isn't lying by omission, because it never made the larger claim in the first place.
The failure shows up one step later, when a green checkmark on that test gets read as evidence for a claim the test never made. Someone opens the PR, sees the suite pass, and treats "tests pass" as shorthand for "this is safe to ship," when the tests only ever spoke to the parser's behavior on well-formed input. The test told the truth. The reader upgraded that truth without checking whether the upgrade was justified.
This matters because the standard response to a production incident is to interrogate the test suite: why didn't the tests catch this. Sometimes that's the right question, when a test really was lying. But often the honest answer is that no test was ever asked the question that broke in production, and the tests that did exist answered their own narrower questions correctly the whole time. The fix that usually follows, write more tests, treats the symptom. It pushes the boundary of what's covered a little further out and changes nothing about the habit of reading a green suite as a blanket guarantee. The next incident just happens at the new edge, and the retro repeats itself with a slightly larger test file attached.
The same substitution happens outside code. A resting heart rate reading of 58 is an honest, accurate number. It says nothing on its own about cardiovascular risk, sleep quality, or fitness, not because the measurement lied but because it was never a measurement of those things. Somebody has to do the separate work of connecting the number to the larger claim, and that connecting work is exactly where people skip a step and just assume the number implies the conclusion because it feels adjacent enough.
Fixing the actual gap means treating "this test verifies X" and "verifying X is sufficient evidence for Y" as two claims that each need their own defense, instead of letting the second one ride in free on the credibility of the first. That's a discipline problem at the interpretation step, not a coverage problem in the suite. A reviewer, a deploy gate, an on-call runbook, whatever sits between the green checkmark and the decision to ship, has to be the place that asks what the test actually touched before it counts the test as an answer to a bigger question.
None of this argues against writing more tests. Coverage is fine, coverage is good. It just isn't the fix for a failure that never happened inside the test. The test that only checks the parser was doing its job correctly the whole time it sat there being blamed for a bug in the caller.
The bug wasn't in the test. It was in the sentence that started with "tests pass, so."