← Back to blog
AI Evaluation

Part of our The Judgment Gap series · Access is not understanding

The AI Users Who Scored 13 and Guessed 17

6 min read

In a study published in Computers in Human Behavior, a team led by Daniela Fernandes at Aalto University gave 246 people twenty logical reasoning problems from the Law School Admission Test and a chat window running ChatGPT. They could prompt it as often as they liked. On average they solved about 13 of the 20. A comparison group of 3,543 people who had worked the same problems without any AI, in an earlier study, averaged about 9.5. So the AI helped.

Then everyone was asked how many they thought they had gotten right. The average answer was about 17. The AI users overestimated their own score by roughly four points, which is about a quarter more correct answers than they had actually produced. Two other details matter. First, ChatGPT working alone averaged about 13.7. A little over half the participants (55%) beat it, but the rest fell far enough behind that the group's average ended up slightly below what the AI managed on its own. Second, they barely talked to it: the mean was 1.15 prompts per question, and 46% of participants never sent more than one prompt on any question. The researchers classified most participants (59%) as copying the problem in and taking the answer back out without checking it.

The mechanism: having an answer feels like knowing it

This isn't new with AI. In 2015, Matthew Fisher, Mariel Goddu, and Frank Keil at Yale ran nine experiments on what searching the web does to people's sense of their own knowledge. In the first, one group looked up answers to everyday "why" questions online, and another answered from memory. Afterward, both groups rated how well they could explain answers to questions in completely different subjects that nobody had searched. The group that had used the internet rated themselves higher on those unrelated topics. The effect held when the researchers matched time spent and content seen, and it wasn't general overconfidence. Having access to explanations made people feel the knowledge was in their heads.

Generative AI is that effect with the friction removed. A search engine gives you ten sources you still have to read, compare, and put together. A language model gives you the finished explanation, in fluent prose, with the vocabulary of someone who knows the field. The work that used to force at least some understanding (choosing a source, noticing that two sources disagree, rewriting the idea in your own words) is exactly the work it takes off your hands.

Two more findings from the Fernandes study are uncomfortable. The usual Dunning–Kruger pattern, where weak performers overestimate themselves and strong performers underestimate, disappeared: with AI, low and high performers overestimated by similar amounts. And people who rated their own AI literacy higher were less accurate about their own performance, not more. A second study of 452 people, 245 of them using AI, replicated the pattern, even though participants were told they would be paid for estimating their score accurately.

Ten years of experience versus thirty seconds of prompting

This is the problem this series is about. Picture a meeting where someone with a decade of experience in program evaluation says a design won't support the conclusion the team wants. Someone else, who asked an AI system half a minute earlier, reads out a confident, well-organized explanation of why it will. Both sound informed. In the room, the two claims can feel like they deserve equal weight.

They don't, but not because experience is always right. The difference is what each person can do when the answer is pushed. Experience brings a model of why: the assumptions the answer depends on, the cases where it fails, and what evidence would change it. An AI answer can carry all of the vocabulary of that model without any of it being in the head of the person repeating it. As the Fernandes results show, that person usually can't tell the difference either.

Where else this shows up

  • Academic and grant-funded research. A graduate student asks an AI which model fits their nested data and gets a correct-sounding recommendation for a mixed model with random slopes. The recommendation may even be right. But if the student can't say what the random slopes assume or what to do when the model won't converge, the first reviewer question will show that the choice was borrowed, not made.
  • Non-profit program evaluation. A board member pastes an evaluation report into a chatbot and comes to the meeting with a crisp summary of what the program "proved." The summary drops the caveats the evaluator spent three pages on, and it sounds more certain than the report does.
  • AI and LLM evaluation. Teams that grade model outputs with AI assistance face the same trap one level up. A reviewer who reads an AI-written rationale for why an output is correct can come away confident about a judgment they never actually made.

Try this: the four-question check

Researchers who study the "illusion of explanatory depth" found a simple way to deflate it: ask people to explain, step by step, something they just rated themselves as understanding. Confidence drops once they try. The same test works on AI-assisted answers. Close the chat window, and without looking at it, answer four questions:

  1. Mechanism. Why is this the answer? What produces the result?
  2. Assumptions. What has to be true for it to hold?
  3. Boundaries. Under what conditions would the answer change?
  4. Disconfirmation. What evidence would show it's wrong?

If you can answer all four, you've learned something. If you can't, you have an answer but not the understanding, and that's fine as long as you treat it as a claim to check rather than something you know. The same questions work on a colleague who says "the AI says…": ask the four questions, and let the answers set how much weight the claim gets.

What to ask for instead

Don't ask whether someone used AI. Ask whether they can defend the answer without it. When a recommendation will shape a study design, an analysis plan, or a funding decision, the person presenting it should be able to name its assumptions and its failure conditions. AI doesn't remove the need for that judgment. It removes the signals that used to tell you whether the judgment was there.

If your team is adopting AI for analysis or evaluation work and wants a review process that catches borrowed answers before they become decisions, talk to us.

Next in The Judgment Gap: why a well-written AI answer feels truer than it is.

Related reading

Sources and further reading

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}