The crash test dummy that's been checking car safety in America for over half a century represents a man. Not figuratively — literally: five foot nine, about 171 pounds, the proportions of an adult American male from the seventies. It was built in 1976 by a team led by engineer Harold Mertz at General Motors, drawing on data gathered earlier by Wayne State University professor Lawrence Patrick, who spent years hurling dummies — and before that, human bodies — at obstacles to understand exactly what happens on impact. They called it Hybrid III, and it's still, to this day, the basis for almost the entire automotive safety standard.
Behind the wheel and in the passenger seat, surprisingly enough, there happen to be quite a few people who aren't a man from the seventies. A 2019 University of Virginia study found that a female occupant has 73 percent higher odds of injury in a frontal crash than a male occupant in the same crash. According to Consumer Reports, a belted woman is 17 percent more likely to die in a crash than a belted man, all else equal. The reason is in the body itself — muscle mass is distributed differently, the neck holds differently, and a belt and airbag calibrated for the dummy respond to a different body differently.
For a while, the gap was patched by the "5th percentile female" dummy. Sounds like a fix. In reality it was the same scaled-down male dummy, same proportions, just smaller — engineering ingenuity at its finest — representing, per the documentation, only the smallest five percent of women by mid-1970s standards. The other ninety-five percent never showed up in the test at all.
The dummy, all this time, was doing exactly what it was built to do: producing the same measurable, repeatable response to impact, test after test, for decades. It passed. It was just answering a different question — not the one about who'd actually be in the car.
The More It Passes, the Less You Ask
There's a heavier version of the same story, and it explains something important about how confidence itself behaves — not the test, the confidence around it.
Years before the Challenger disaster, engineers at NASA and contractor Morton Thiokol had already seen O-ring erosion on the joints of the solid rocket boosters — on earlier flights. The erosion exceeded engineering tolerances. Formally, that meant: don't fly with this part. But flights with the erosion kept ending safely, and with each successful one, the decision to fly again felt a little less alarming. Sociologist Diane Vaughan, who reconstructed the entire chain of decisions from NASA's documents and hundreds of interviews, called this the normalization of deviance: the team wasn't consciously ignoring the rule — the rule gradually stopped feeling broken. By January 1986, the erosion wasn't a deviation to the engineers anymore. It had become a normal, expected part of flight, and on a cold launch morning, that exact part failed.
Notice where the confidence actually moved. It didn't sit still while the risk grew around it — it grew right alongside the number of successful launches, and it grew in exactly the wrong direction relative to reality. Every successful flight with the erosion looked like one more confirmation that the tolerance could be stretched. In fact, each one was just another time it got lucky, while the tolerance itself, still violated on paper, never moved an inch.
Here's the actual difference between "the test didn't check something" and what's really going on. A dummy and a test mode don't just stay silent about what's outside their scope. Run them long enough, successfully enough, and they start convincingly implying there's nothing to check out there. Fifty years of successful crash tests didn't add a single new fact about a woman's body in a car. But they seem to have added the industry's confidence that the question was closed — otherwise the dummy would hardly have stayed the same for this long.
Test Mode Works the Same Way
There's a similar blind spot in a far more modest thing — the test mode that comes with just about any payment provider today. You spin up a test store, plug in a test card, run the whole customer journey from the "pay" button to confirmation — and it's all green. I've built a flow like this from scratch myself, running it back and forth a few times before showing it to a client. And with every successful run, naturally, I felt more confident — that's exactly how trust in a tool is supposed to work. The problem is that confidence was expanding toward a question the test was never asking.
Then recently it turned out that, for one payment provider, a green checkmark guarantees nothing beyond the mechanism itself. A rule that let sellers of a certain category count on automatically issued receipts had simply vanished a few months earlier — with zero connection to whose code worked and whose didn't. Nobody sent a warning, of course — why would they, the test was still green. The feature just stopped existing one day. The payment mechanism still passed the test flawlessly. Permission for that particular way of doing business — no, and the test was never built to check that in the first place.
A test store won't tell you whether you're actually allowed to accept that money given your specific legal standing.
Governments Build the Same Rooms
Governments came up with a similar construction on their own — regulatory sandboxes for fintech startups, a controlled mode where you can trial a new financial product on real customers without the full weight of the usual regulatory requirements. The logic is sound: let people try it before demanding full compliance.
In write-ups on how these sandboxes actually play out, the same caveat keeps coming up: while a company is inside, it doesn't have to clear the same compliance bar waiting for it outside, and it's not always clear in advance which requirements will land the moment it exits. The sandbox honestly checks whether the product works. Whether it complies with everything else on the books is a separate question, settled somewhere else, later.
The first instinct, faced with a gap like this, is to run the test again, more carefully. Sometimes that genuinely helps — some problems really do hide in rare scenarios that only surface on the tenth run. But it won't help here. Slam the dummy into a wall a thousand times instead of ten, and it gets no closer to how a different body is built. Spend a month in the sandbox instead of an hour, and it — much to everyone's regret — still won't start checking your legal status, because that's simply not what it was built to do. The room was constructed so that nothing inside it could go wrong, and extra time in a room like that won't show you what it was built not to show.
What This Actually Changes
"Works" answers whether the mechanism behaves predictably under the conditions a test knows how to create. "Allowed" is a separate question: whether you're actually permitted to use that mechanism in your specific situation — your legal status, your paperwork, your country. These are two different questions that someone, at some point, decided to ask in the same place — and ever since, the answer to the first gets mistaken for the answer to the second out of sheer habit.
It isn't one. Worse: the more often the first question comes back "yes," the calmer people get about the second one too, even if nobody ever asked it out loud. NASA's engineers weren't stupid or reckless — their confidence grew exactly the way anyone's does on seeing the same green checkmark for the tenth time running. Most of us, thankfully, aren't measuring the stakes in human lives — more like a blown quarter and a couple of awkward calls to the accountant. The mechanism is identical all the same.
Which leads to a rule less convenient than "test it better." Growing calm after a run of successful passes doesn't mean the second question is closed. It's a signal — separately, out loud, loudly and clearly — to ask whoever's actually responsible for permission: a lawyer, an accountant, an inspector, anyone — what exactly about your situation this test could never have noticed in the first place. A test can grow your confidence. It can't grow its own coverage.
It's worth finishing the Challenger story properly, because it offers a prescription, not just a diagnosis.
Morton Thiokol engineer Roger Boisjoly had warned about the O-ring problem in cold weather long before the disaster, and on the night before launch, on learning about the record-low temperature at the pad, he and his colleagues flatly asked for a delay — charts and numbers in hand. The question got asked out loud, by a specific person, with specific data. Thiokol's management, after a long argument, withdrew the objection anyway — according to testimony before the commission, partly so as not to let down a major customer against a tight launch schedule. Boisjoly was right. It changed nothing.
Here's what that adds to the picture. A closed question does sometimes get reopened — but the longer the run of successful passes by that point, the more awkward the voice reopening it sounds. After two dozen successful flights with the erosion, the erosion itself starts to read as evidence in its own defense, and the person holding up the pad with data at the edge of tolerance ends up arguing not so much against the facts as against the accumulated success of the whole program. By that point, the burden of proof has quietly shifted onto the one with doubts.
Fortunately, someone figured out a countermeasure long before any particular launch. In 2007, psychologist Gary Klein proposed a simple exercise: before launching anything, gather the team and imagine it has already failed — then, from inside that imagined future, ask what exactly went wrong. It's called a premortem, and it works because it shifts the speaker's own position: hunting for reasons a failure happened in some hypothetical future is far more comfortable than holding up a real launch with doubt in the present. Daniel Kahneman, citing the method, noted that a single such conversation before launch raises the number of identified risks by roughly a third.
Applied to test mode, this needs no adaptation at all: before reading a green checkmark as permission, it helps to spend five minutes out loud imagining the whole scheme has already failed — not from a bug, but from something turning out to not be allowed — and ask exactly what. Asking yourself that question ahead of time is a lot cheaper than becoming the next Boisjoly — full set of data, zero chance of being heard in time.
Be the first to comment