← Back to blog
Data Literacy

The Plane That Never Made It Back

3 min read

During the Second World War, Allied bombers came home from missions over Europe covered in damage, and the military kept careful track of where. Engineers plotted every bullet and flak hole on diagrams of the aircraft, and a clear pattern emerged: the wings and the rear fuselage were peppered with hits, while the engines and cockpit area came back almost untouched. The obvious recommendation, the one engineers were ready to make, was to add armor plating to the wings and fuselage — reinforce where the damage was.

The military brought the problem to the Statistical Research Group at Columbia, and to a mathematician named Abraham Wald. Wald looked at the same diagrams and recommended the opposite: armor the engines, not the wings. His reasoning was simple once he said it out loud. Every plane in that dataset had survived the flight home. A bullet hole in the wing was a hole a plane could take and still fly back to be counted. The planes that took a hit to the engine weren't in the sample at all, because they went down over enemy territory. The blank space on the diagram wasn't the safe zone. It was the data the Air Force could never collect.

The mechanism: the sample was already filtered by the thing you're studying

This is survivorship bias, and once you see it, it's hard to unsee. It happens whenever the group you can observe has already been shaped by the outcome you're trying to measure — so the absence of a signal doesn't mean the signal isn't dangerous, it can mean the opposite: the cases that would have shown it never made it into the room. Wald's contribution wasn't a new formula. It was insisting that the Air Force's dataset had a structural hole in it exactly where the most important information would have been, and that no amount of careful analysis of the visible data could fill it in. You have to go looking for what's missing on purpose, because it will not announce itself.

What makes survivorship bias harder to catch than a simple missing-data problem is that the visible data doesn't just fail to help — it actively points the wrong way. Every engineer looking at those diagrams was doing careful, honest work. The pattern on the page was real. It just answered a different question than the one they thought they were answering: not "where do bombers get hit," but "where can a bomber get hit and still be the kind of bomber we get to study." Those are not the same question, and no amount of statistical sophistication applied to the wrong dataset gets you to the right answer.

Where else this shows up

The bombers are the vivid version. The same blind spot shows up constantly in far less dramatic settings:

  • Academic and grant-funded research. A literature review of "what works" is built almost entirely from studies that got results interesting enough to publish and researchers persistent enough to write up. The interventions that fizzled, the trials that were quietly discontinued, the null results that never became a paper — those don't show up in the review, and their absence looks like consensus instead of like a gap.
  • Non-profit program evaluation. A follow-up survey sent to program graduates measures the people who stayed engaged long enough to be reachable. Participants who dropped out, lost contact, or had the worst outcomes are frequently the hardest to survey — which means the follow-up data can look strong for reasons that have nothing to do with the program's actual effect.
  • AI and LLM evaluation. A model looks reliable if you're only reviewing the outputs a human bothered to check, or the transcripts that made it into a demo. The failures that got silently discarded, timed out, or never surfaced to a reviewer at all are exactly the ones you'd want to know about, and they're the ones missing from the sample by construction.

What to ask for instead

Before trusting a dataset that looks clean, ask what had to happen for a case to end up in it — and what would have had to happen for a case like it to be missing. One question does most of the work: what happened to the ones that aren't in front of me? If nobody can answer that, the analysis isn't wrong so much as incomplete in a spot nobody's looked at yet.

If you're building an evaluation, a follow-up study, or a review and you're not sure what's missing from your sample, talk to us before you draw conclusions from what's left.

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}