← Back to blog
AI Evaluation

How to Vet an AI Vendor's Accuracy Claims

3 min read

Ask for three things before you believe an AI vendor's accuracy number: the test set it was measured on, the exact scoring method used to grade it, and what it's being compared against. If a vendor can't produce clean answers to all three, treat the headline figure as marketing copy, not evidence.

What the number is actually hiding

"95% accurate" sounds precise, but accuracy is only as meaningful as the test it was measured against. A model can hit 95% on a benchmark built from easy, well-formed examples and fall apart on the messy, ambiguous inputs your actual users send it. The number itself tells you almost nothing until you know what it was measured on, who scored it, and what it's being measured against — a competitor, a prior version, a human baseline, or nothing at all.

The three things to ask for

  • The test set. Was it held out from training and tuning, or did the vendor peek at it while building the product? A test set the team has already seen isn't measuring performance — it's measuring memorization. Ask how the set was built and whether it's ever been used to make a decision about the model.
  • The scoring method. Was each output graded by a fixed rubric applied consistently, or by a person (or another model) making a judgment call case by case? A vague rubric or an unspecified grader means two people scoring the same outputs could land on very different numbers.
  • The comparison baseline. Accuracy without a baseline is a number in a vacuum. Better than what — a previous version, a competitor, a simple rule-based system, an unaided human? "Better than nothing" and "better than the thing we're replacing" are very different claims, and vendors are rarely eager to specify which one they mean.

Red flags in how a claim is reported

  • A single number with no spread. Real evaluations vary run to run and across subgroups of inputs. A report with one clean percentage and nothing about variance or worst-case performance usually hasn't been stress-tested, or the stress test isn't being shown to you.
  • A benchmark that doesn't resemble your use case. A model can score well on a public leaderboard built from short, generic queries and still struggle badly on your longer, domain-specific, or higher-stakes inputs. Ask what the test inputs actually looked like.
  • No mention of failure cases. Every real model fails sometimes. A vendor who can't show you a handful of representative failures, and what they did about them, likely hasn't looked closely enough to find them.

What a credible vendor can produce on request

A vendor confident in its numbers can usually hand over a short methodology write-up: how the test set was built and kept separate from training, how scoring was done and by whom, a breakdown by category or difficulty rather than one blended average, and a handful of documented failure cases. If a written methodology doesn't exist, the accuracy number wasn't built to survive scrutiny — it was built to be quoted.

If you're weighing whether to trust a vendor's claims before signing a contract, or want to build your own evaluation to hold whatever you buy accountable, our AI evaluation work covers exactly this ground.

Related reading

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}