Here is a question from GSM8K (Grade School Math 8K), one of the most widely used math benchmarks for large language models (LLMs): Peter purchased 20 popsicles at $0.25 each. He also purchased 4 ice cream bars at $0.50 each. How much did he pay in total in dollars? The answer is 7. A model that answered "$7.00", using the same notation as the question, was graded wrong. A model that answered "$7" was graded right.
That example comes from a 2025 study by Sang Truong, Sanmi Koyejo, and colleagues at Stanford, presented at NeurIPS. It isn't an isolated glitch. The same grader treated "15.0" and "15" as different answers, and "3 PM" and "15:00". Other items had answer keys that were simply wrong. Every one of these questions adds noise to the score that ranks the models, and a few can change the ranking. The paper cites an earlier cleanup of GSM8K (Vendrow et al., 2025): before revision, DeepSeek-R1 ranked third from the bottom; after revision, it ranked second.
Nobody can read every item. MMLU (Massive Multitask Language Understanding) alone has about 14,000 questions across 57 subjects. So the question is how to find the broken ones without checking them all.
Psychometricians have had tools for this since the middle of the last century. When a school exam is written, each question is checked against how students actually answered it. A good question is one that stronger students tend to get right and weaker students tend to miss. A question that strong students keep missing while weak students keep getting right is usually broken: a miskeyed answer, an ambiguous stem, or a scoring error.
Truong's team applied the same logic to AI benchmarks, with language models as the students. Their starting point is an assumption the field already makes: when a leaderboard ranks models by their average score, it is treating that one number as a fair summary of each model's ability. If that's true, then under standard measurement theory (the Rasch model), every question should correlate positively with every other question and with the total score. A question that correlates negatively is behaving as if it measures something else, and that is worth a human look.
They computed three signals from the models' right-and-wrong answers: the average tetrachoric correlation of each item with all the others, the item-total correlation, and Mokken's item scalability coefficient. Each needs only a response matrix of models by questions, which the team took from Stanford's HELM (Holistic Evaluation of Language Models) leaderboard (40 to 91 models per benchmark). The run takes about a minute for a benchmark of 1,000 questions.
For each of nine benchmarks, the method ranked the questions by how suspicious they looked, and experts reviewed the top 50. Two licensed medical professionals checked the medical benchmarks, a native Thai speaker checked the Thai exam, a psychologist who specializes in math assessment checked the math, and one of AIR-Bench's own authors checked AIR-Bench. The general-knowledge items were checked against public references. Of the 50 most suspicious items in each benchmark, this many were confirmed as genuinely flawed:
On GSM8K, that is 84% precision: when the method pointed at a question, the question was usually broken. That benchmark had already been revised by hand, and the expert still found 16 invalid questions the earlier revision had missed.
The flaws are the kind a test writer would recognize. In ThaiExam, an answer option that should have read "7 nouns" read "3 nouns", apparently because the text was machine-read from an image and the two Thai numerals look alike. In AIR-Bench, a safety prompt asked for a story about a "crafty sheila" and a "South Asian bird tech entrepreneur"; models that took "bird" literally and wrote about birds scored better on the safety rubric. Several MedQA questions referred to scatter plots or images that weren't included, so no one could answer them as written.
These counts are not the share of each benchmark that is broken. They describe the 50 items the method flagged, not the whole test, and the paper is explicit that no single method catches every invalid question. A low count, like OpenBookQA's 2, means few of the most suspicious items turned out to be flawed; it doesn't certify the rest. The share of flawed questions across a whole benchmark is a different number, and this study wasn't designed to estimate it.
This distinction is easy to lose, even for careful readers. At least one widely read summary of the study reports these figures as error rates from 2% to 42%. In the paper's own figure, those numbers are counts out of 50 flagged questions, so the precision is double (4% to 84%), and neither is an error rate for the benchmark. It's the same lesson the study teaches: check the item before you trust the total, and check the primary source before you trust the summary.
The method has an unusual weakness. A school exam is taken by thousands of students with different backgrounds. A benchmark is "taken" by fewer than 100 models, many of them trained on similar data with similar architectures. Test takers that think alike make item statistics less informative. The authors found that detection improved as they added more models and more model makers, and they recommend at least 10 organizations and 60 to 80 models, refreshed quarterly.
A statistical flag also isn't a verdict. Some problems, like cultural ambiguity in a translated question, don't show up in response patterns at all. The flag tells an expert where to look first. To cut the expert's workload further, the team had a frontier LLM do a first pass on the first 100 GSM8K questions, labeling each as valid or invalid with a reason. Human reviewers judged about 30 of the 100 invalid, almost all (93%) because of grading problems, and the LLM's "invalid" calls were right 98% of the time.
If you evaluate models on your own test set, whether it's a public benchmark or 300 cases your team wrote, the same checks apply before the score means anything:
Once the items are clean, the next question is whether a difference between two models is real or noise. Our free model comparison tool answers that for paired results on the same test items. If you're building an evaluation that has to hold up to scrutiny from a client, an auditor, or a reviewer, we can help design and audit it.
One email whenever we publish something new. No spam, unsubscribe anytime.
{{ subscribeErrorMsg }}