Part of our Statistical paradoxes series · Condorcet's jury theorem
In the early 1980s, engineers building software that could not be allowed to fail were attracted to a simple idea. Don't write one program and hope it has no bugs. Have several teams write separate versions from the same specification, run them side by side, and let them vote. If one version hits a bug on some odd input, the other two outvote it. The approach is called N-version programming, and the arithmetic behind it is very appealing: if each version fails once in ten thousand runs, and the versions fail independently, two of them failing on the same input should happen about once in a hundred million.
John Knight at the University of Virginia and Nancy Leveson at the University of California, Irvine decided to test the "if." With NASA funding, they had 27 programmers, nine at Virginia and eighteen at Irvine, each write their own version of the same program, a launch-interceptor routine that took simulated radar readings and returned 241 yes/no decisions. Nobody compared notes. Every version was then run on one million randomly generated test cases and checked against a reference implementation. Individually, the programs were extremely reliable. But on 1,255 of those test cases, two or more versions failed on the same input, and on some inputs as many as eight failed at once. In their 1986 paper in IEEE Transactions on Software Engineering, Knight and Leveson rejected the independence assumption at the 99% confidence level. The programmers hadn't copied each other. They had just made their mistakes in the same places.
The math N-version programming relied on is two centuries older than software. In 1785 the Marquis de Condorcet proved what is now called Condorcet's jury theorem: if each juror is more likely to be right than wrong, and the jurors decide independently, then the chance that the majority is right rises toward certainty as the jury gets bigger. The numbers are striking. Say each juror gets the answer right 70% of the time. A single juror is right 70% of the time, a majority of three is right about 78% of the time, a majority of eleven about 92%, and a majority of fifty-one better than 99.8%.
The theorem is correct, but it only holds if every juror decides independently. Condorcet's guarantee needs each juror's mistakes to be unrelated to everyone else's. Real voters break that assumption all the time: they read the same newspaper, learned from the same textbook, or find the same part of the problem hard. Once errors are correlated, adding voters stops multiplying your reliability. Eleven jurors who all fall for the same misleading piece of evidence add up to roughly one juror, counted eleven times.
That's what the radar programs showed. Difficulty is a property of the problem, not only of the person solving it. Some inputs sat right on an awkward edge of the specification, so many independent programmers got them wrong, and the vote on exactly those inputs had several wrong answers in it. Independent people don't necessarily make independent mistakes.
The experiment has just been rerun with new programmers. In a 2026 preprint, Javier Ron, Benoit Baudry and Martin Monperrus had AI coding agents write 48 versions of the same launch-interceptor program and ran them on a million randomized inputs. Under independence, about 115 coincident failures would be expected. They found 429. Voting still helped: a three-version majority cut the average failure count from about 387 for a single version to about 131. That's the honest version of the lesson. Redundancy works. It just doesn't work as well as the independence math promises, and if you never measure the correlation, you can't tell how much you're actually getting.
Whenever a result rests on several raters, models, versions or sources agreeing, ask one question: when one of them is wrong, how often are the others wrong too? Take the items where at least one voter erred and count how often a second voter erred on the same item, then compare that with what you'd expect if their errors were unrelated (roughly the product of their individual error rates). If the observed overlap is far higher, your panel is smaller than it looks. The fix is to build in real diversity: raters trained separately, models from different families, data sources collected by different people. Knight and Leveson's programmers really did work independently. Their mistakes still lined up, and nobody would have known if Knight and Leveson hadn't counted the overlaps.
If your study, evaluation or AI pipeline relies on agreement among raters, sources or models, and you want to know how much that agreement is actually worth, talk to us.
One email whenever we publish something new. No spam, unsubscribe anytime.
{{ subscribeErrorMsg }}