← Back to blog
Data LiteracyAI Evaluation

Part of our Statistical paradoxes series · Condorcet's jury theorem

The 27 Programs That Failed Together

5 min read

In the early 1980s, engineers building software that could not be allowed to fail were attracted to a simple idea. Don't write one program and hope it has no bugs. Have several teams write separate versions from the same specification, run them side by side, and let them vote. If one version hits a bug on some odd input, the other two outvote it. The approach is called N-version programming, and the arithmetic behind it is very appealing: if each version fails once in ten thousand runs, and the versions fail independently, two of them failing on the same input should happen about once in a hundred million.

John Knight at the University of Virginia and Nancy Leveson at the University of California, Irvine decided to test the "if." With NASA funding, they had 27 programmers, nine at Virginia and eighteen at Irvine, each write their own version of the same program, a launch-interceptor routine that took simulated radar readings and returned 241 yes/no decisions. Nobody compared notes. Every version was then run on one million randomly generated test cases and checked against a reference implementation. Individually, the programs were extremely reliable. But on 1,255 of those test cases, two or more versions failed on the same input, and on some inputs as many as eight failed at once. In their 1986 paper in IEEE Transactions on Software Engineering, Knight and Leveson rejected the independence assumption at the 99% confidence level. The programmers hadn't copied each other. They had just made their mistakes in the same places.

The mechanism: a vote is only as good as its independence

The math N-version programming relied on is two centuries older than software. In 1785 the Marquis de Condorcet proved what is now called Condorcet's jury theorem: if each juror is more likely to be right than wrong, and the jurors decide independently, then the chance that the majority is right rises toward certainty as the jury gets bigger. The numbers are striking. Say each juror gets the answer right 70% of the time. A single juror is right 70% of the time, a majority of three is right about 78% of the time, a majority of eleven about 92%, and a majority of fifty-one better than 99.8%.

The theorem is correct, but it only holds if every juror decides independently. Condorcet's guarantee needs each juror's mistakes to be unrelated to everyone else's. Real voters break that assumption all the time: they read the same newspaper, learned from the same textbook, or find the same part of the problem hard. Once errors are correlated, adding voters stops multiplying your reliability. Eleven jurors who all fall for the same misleading piece of evidence add up to roughly one juror, counted eleven times.

That's what the radar programs showed. Difficulty is a property of the problem, not only of the person solving it. Some inputs sat right on an awkward edge of the specification, so many independent programmers got them wrong, and the vote on exactly those inputs had several wrong answers in it. Independent people don't necessarily make independent mistakes.

The experiment has just been rerun with new programmers. In a 2026 preprint, Javier Ron, Benoit Baudry and Martin Monperrus had AI coding agents write 48 versions of the same launch-interceptor program and ran them on a million randomized inputs. Under independence, about 115 coincident failures would be expected. They found 429. Voting still helped: a three-version majority cut the average failure count from about 387 for a single version to about 131. That's the honest version of the lesson. Redundancy works. It just doesn't work as well as the independence math promises, and if you never measure the correlation, you can't tell how much you're actually getting.

Where else this shows up

  • Academic and grant-funded research. Three coders trained together on the same codebook, checking each other's ratings, can reach high agreement while making the same systematic mistake. And a meta-analysis that counts several papers drawn from one dataset as separate evidence ends up with a jury of one counted several times.
  • Non-profit program evaluation. Triangulation is the evaluator's version of a majority vote: if staff reports, participant surveys and administrative records all point the same way, the finding feels solid. But if program staff collected all three, encouraged participants to fill in the survey and entered the records, the three sources share one point of failure. Ask who produced each source before you count them as separate lines of evidence.
  • AI and LLM evaluation. A panel of LLM judges, majority voting across samples from one model, or an ensemble of models trained on largely the same web data is a jury. If the judges come from the same model family or the same training data, they tend to fail on the same items, and high agreement among them says little about whether any of them is right. The question isn't how many judges you have. It's how differently they get things wrong.

What to ask for instead

Whenever a result rests on several raters, models, versions or sources agreeing, ask one question: when one of them is wrong, how often are the others wrong too? Take the items where at least one voter erred and count how often a second voter erred on the same item, then compare that with what you'd expect if their errors were unrelated (roughly the product of their individual error rates). If the observed overlap is far higher, your panel is smaller than it looks. The fix is to build in real diversity: raters trained separately, models from different families, data sources collected by different people. Knight and Leveson's programmers really did work independently. Their mistakes still lined up, and nobody would have known if Knight and Leveson hadn't counted the overlaps.

If your study, evaluation or AI pipeline relies on agreement among raters, sources or models, and you want to know how much that agreement is actually worth, talk to us.

Related reading

Sources and further reading

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}