Part of our Statistical paradoxes series · Benford's law
In the early 1990s, a manager in the Arizona State Treasurer's office named Wayne James Nelson started writing checks to a vendor that didn't exist. In all there were 23 of them, adding up to nearly $2 million of the state's money. When the scheme came apart and the case reached court in 1993 as State of Arizona v. Wayne James Nelson, his explanation was that he'd diverted the funds on purpose, to show that a new computer system had no real safeguards. The court was not persuaded, and he was found guilty.
The checks themselves told a different story from his. Like a lot of embezzlers, Nelson started small and worked his way up. And like a lot of people making up numbers that need to look unremarkable, he kept them just under a threshold: most of the amounts landed just below $100,000. Nothing about any single check was absurd. But when the forensic accountant Mark Nigrini later laid all 23 out as a worked example in the Journal of Accountancy, one thing stood out that nobody had set out to create: the first digits. Far too many of the checks started with a 7, an 8 or a 9. Honest numbers almost never look like that.
Ask most people how often a number in a ledger should start with each digit, and they'll say about equally: one time in nine for each. That's wrong for a huge range of real-world data. In 1881, the astronomer Simon Newcomb noticed that the front pages of books of logarithm tables wore out faster than the back pages, as if people looked up numbers starting with 1 far more often than numbers starting with 9. In 1938, the physicist Frank Benford rediscovered the pattern, documented it across many different kinds of data, and it now carries his name.
Benford's law says that in many naturally occurring datasets, the leading digit is 1 about 30% of the time, 2 about 18%, 3 about 12%, and so on down to 9, which leads only about 5% of the time. The intuition is growth. To get from a leading 1 to a leading 2, a quantity has to double; to get from 9 back round to 1 (90 to 100), it only has to grow by about 11%. Numbers that grow, shrink and multiply across several orders of magnitude, like invoices, populations, river lengths and account balances, spend much more of their lives starting with small digits.
People inventing numbers don't know this, and even if they did, they couldn't easily fake it. They pick amounts that feel random, avoid round numbers, and stay under whatever limit they think triggers a review. The result is a first-digit distribution that looks nothing like Benford's. Nigrini had already shown in a 1996 study that the digit patterns in tax data could flag likely evasion, and digit analysis is now a standard first pass in forensic auditing.
Two caveats matter as much as the law itself. It only applies to data that spans several orders of magnitude and isn't assigned or capped: phone numbers, ZIP codes, prices set to end in .99, and survey scores from 1 to 5 won't follow it and aren't supposed to. And a deviation is a reason to look, not a verdict. Nelson was convicted on the evidence, not on a histogram.
Before you trust a column of numbers you didn't collect yourself, ask one question: does this look like it was measured, or like someone chose it? Plot the leading digits, the last digits, and how the values sit around any threshold that matters. Real data is lumpy in predictable ways. Invented data is usually smooth in the wrong places, and a two-minute check will often tell you which kind you're holding.
If you're sitting on data from many sources and aren't sure all of it was measured rather than filled in, or you want a systematic integrity check before the numbers go into a report or a model, talk to us.
One email whenever we publish something new. No spam, unsubscribe anytime.
{{ subscribeErrorMsg }}