← Back to blog
Data Literacy

Part of our Statistical paradoxes series · The prosecutor's fallacy

The Number That Convicted an Innocent Mother

4 min read

In 1996, a twelve-week-old boy named Christopher Clark died suddenly at home in England. Two years later, in 1998, his baby brother Harry died too, at eight weeks old. Their mother, Sally Clark, was charged with murdering both of them. At her trial in Chester Crown Court in 1999, one of the most memorable pieces of evidence wasn't medical at all. It was a number.

Professor Sir Roy Meadow, an eminent paediatrician, told the jury that for an affluent, non-smoking family like the Clarks, the chance of a single cot death (sudden infant death syndrome, or SIDS) was about 1 in 8,543. The chance of two, he said, was that number squared: roughly 1 in 73 million. In November 1999, Sally Clark was convicted by a 10–2 majority and sentenced to life. It took until January 2003, a second appeal and the discovery of withheld microbiology results showing a serious bacterial infection in Harry for the convictions to be quashed. She was released, but never recovered, and she died in 2007.

The mechanism: the chance of the evidence is not the chance of guilt

The 1-in-73-million figure had two separate problems, and the second one is the one this post is about.

The first was arithmetic. You can only multiply two probabilities together when the events are independent, and cot deaths in the same family aren't. Siblings share genes, a home, sleeping arrangements and a set of parents. If one baby in a family dies of SIDS, the chance of a second is higher than for a random family, not the same. Squaring 1 in 8,543 threw that away and made the coincidence look far rarer than it was.

The second problem survives even if you fix the first. Suppose double SIDS really were extraordinarily rare. The number would tell you how likely two unexplained infant deaths are if the mother is innocent. What a jury needs is the reverse: how likely is she to be innocent, given two unexplained infant deaths? Treating the first number as if it answered the second question is called the prosecutor's fallacy, or the transposed conditional. In October 2001, the Royal Statistical Society issued a public statement saying the 1-in-73-million figure had "no statistical basis". It also warned that some press reports had presented it as the chance that the deaths were accidental, which it called a serious error of logic.

The fix is to compare explanations, not to judge one in isolation. Two babies in one family had died; that much was certain. The only question was which rare explanation was the more likely one. Double SIDS is very rare. But a mother murdering two of her infants is also very rare. Ray Hill, a mathematician at the University of Salford, later estimated that double SIDS was somewhere between 4.5 and 9 times more likely than double infant homicide. On the statistics alone, before any medical evidence, the odds favoured innocence. The jury heard a number that sounded like near-certain guilt, when the right comparison pointed the other way.

Where else this shows up

  • Academic and grant-funded research. A p-value of 0.01 is the chance of seeing data this extreme if the null hypothesis is true. It is not a 1% chance that the null is true, or a 99% chance that your effect is real. How likely the finding is to be real depends on how plausible the hypothesis was to begin with, and that's the same comparison the Clark jury never got to make.
  • Non-profit program evaluation. "Participants who dropped out almost always had attendance problems first" is not the same as "participants with attendance problems almost always drop out." If attendance problems are common and dropout is rare, most people you flag will finish anyway. Build an early-warning system on the first statement while believing the second, and you'll spend your outreach budget on the wrong people.
  • AI and LLM evaluation. A detector that flags 95% of AI-written essays tells you P(flag | AI-written). A teacher holding a flagged essay needs P(AI-written | flag), and that depends on how many essays were AI-written to begin with and how often human writing gets flagged. With a low base rate, a large share of flags can land on students who did nothing wrong. Moderation filters, fraud models and hallucination detectors all have the same gap.

What to ask for instead

Whenever someone hands you a tiny probability as proof, ask one question: tiny compared to what? A rare event happened; that part isn't in dispute. Every competing explanation for it has to be scored on the same evidence, and the rarity of one explanation means nothing until you've put it next to the rarity of the others. If the person quoting the number can't tell you what the alternative is, or how likely it is, the number isn't evidence yet. It's just a very convincing figure.

If a number in your study, evaluation or model report is carrying more weight than it should, or you're not sure which way round a probability actually runs, talk to us before it ends up in front of a reviewer, a board or a customer.

Related reading

Sources and further reading

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}