← Back to blog
AI Evaluation

Part of our The Judgment Gap series · Understanding can't be delegated

AI Can Think for You, But It Can't Understand for You

5 min read

In a field experiment published in PNAS, Hamsa Bastani and colleagues gave nearly 1,000 high-school math students in Turkey different kinds of AI help while they practiced. One group used a tutor that worked like ordinary ChatGPT. During practice, they did 48% better than students with no AI. Then the AI was taken away for the exam, and the same students scored 17% worse than the students who never had it.

A second group used a tutor built to give hints instead of answers, with input from teachers. Their practice scores were up 127%, and on the exam they did about the same as the no-AI group. The hint-giving tutor removed the harm. It did not create an advantage on the exam. Neither group could tell from inside the practice session how they would do without the tool, because the practice scores looked excellent.

Thinking done for you is not understanding gained

AI can now produce the chain of reasoning for a task: which model fits, why an assumption is acceptable, what the coefficients mean. What it can't hand over is your grip on that reasoning. In practice, understanding an analysis means three things. You can explain it without the screen in front of you. You can apply it to a case that differs from the one you saw. And you know which assumption, if it failed, would break the result.

People are poor judges of whether they have this. In a classic set of studies, Leonid Rozenblit and Frank Keil asked Yale graduate students to rate on a 7-point scale how well they understood eight everyday devices, including a zipper and a flush toilet. Then they asked them to write step-by-step explanations of how each one works, and to rate themselves again. The ratings dropped. Having the feeling of understanding, it turned out, is not the same as being able to explain. The gap between the two is what the first article in this series measured with AI in the room.

What workers say changed

A CHI 2025 survey by Lee and colleagues at Microsoft Research and Carnegie Mellon asked 319 knowledge workers to describe 936 examples of using generative AI at work. Workers who were more confident in the AI reported less critical thinking, while workers who were more confident in themselves reported more. The authors describe the work shifting toward verifying AI output, integrating it, and what they call task stewardship. This was a self-reported survey, so it shows an association and not a cause. But it fits the field experiment: the more the tool carries, the more the person's job becomes checking, and checking is only as good as what the person understands.

Why statistics is where this shows

An analysis is a series of choices: which model, which observations to exclude, which assumption to accept, which comparison to report. Each one can be questioned, and the questions come from people with no interest in whether an AI was involved. A reviewer asks why you used a random slope. A funder asks why the effect size is that large. A committee asks what happens if the two smallest sites are dropped. The person who signs the methods section answers, not the tool. A model can run the analysis. It can't stand in the room for it.

Try this: the transfer test

Before you accept an AI-assisted analysis, change one thing and predict the result. Drop the two smallest clusters. Add the covariate a reviewer will ask about. Halve the sample. Write down what you expect to happen to the estimate, the interval, and the conclusion, and only then rerun it. If your prediction is close, you understand the analysis. If it is far off, you have an answer without the understanding, and that is worth knowing before someone else finds out. This is our recommendation, not a finding from a study, and it is a different check from the four questions in the first article: those probe your reasons, and this one probes your model of how the result behaves.

Three habits make the test easier to pass:

  • Ask for hints, not answers. The tutor that gave hints did no harm in the PNAS study, and the one that gave answers did. Prompt your AI the same way: ask what to check next, not what the answer is.
  • Write the methods section from the code. Write it in your own words from the script and output, then compare it with what the AI suggested. Where the two differ, one of them is wrong.
  • Keep a short decision log. One line per choice: what you decided, why, and what would change your mind. It is the record you will reach for when a question comes.

What to ask for instead

The useful question about AI-assisted work is not how much AI was involved. It is whether the person presenting it can explain, extend, and defend it. If you want a second set of eyes on an analysis before it reaches a reviewer or a funder, talk to us.

Next in The Judgment Gap: why a well-written AI answer feels truer than it is.

Related reading

Sources and further reading

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}