← Back to blog
Program EvaluationAI Evaluation

An Active AI Tutor Isn't Yet an Effective One

4 min read

High tutor usage answers one question. It does not answer whether students learned. Logins, minutes and completed exercises tell you the tool was adopted and used. Whether students can do more without it is a separate question, and it needs its own measure, its own timing and its own comparison, all decided before anyone opens the usage dashboard.

That gap matters right now because the research hasn't caught up. A 2026 research update from the Institute of Education Sciences' Regional Educational Laboratory Northeast & Islands concluded that strong causal studies of how current AI tools help or hinder student learning are still extremely limited, and that results depend heavily on how a tool is designed. In other words, the published evidence can't tell a district whether its tutor, in its classrooms, is working. A local pilot has to answer that, and it can only answer it if it measures the right things.

Here is a hypothetical to make it concrete. Say a district pilots an AI math tutor in eight middle schools for a semester. By December the dashboard looks great: 85% of students logged in weekly, the average student spent 40 minutes a week with the tutor, and completed problem sets are up sharply. The board asks whether to expand. Every one of those numbers is real, and none of them answers the question.

1. Implementation: who used it, and how

Usage data is genuinely valuable, as long as it's labeled correctly. It describes adoption (who used the tool at all) and fidelity (whether they used it the way it was meant to be used). Both matter: a tutor that's never opened can't help anyone, and a tutor used mainly to fetch answers isn't the intervention the district thought it was buying.

  • Adoption: what share of students and teachers used it, and how that varies by school, grade and student group.
  • Fidelity: did students work through hints and attempts, or jump to solutions? Did teachers use it as planned, as a supplement rather than a substitute?
  • What it can't tell you: whether minutes caused learning. Students who log the most time are often the ones who were already more motivated or better supported. A correlation between usage and grades is a description of who chose to engage, not an effect of the tool.

2. An independent outcome: what can students do without it?

If the AI tutor disappears on test day, what can students still do? That's the outcome that matters, and it's easy to skip. Exercises completed inside the tutor measure performance with help. Learning is what's left when the help is gone.

  • Completed without the tutor: a unit assessment, a district benchmark, or a state test, taken without access to the tool.
  • Measured at a sensible follow-up: not just the day after the unit ends, but late enough to show whether the skill lasted, such as the next benchmark window.
  • Reliable enough to detect a change: a short quiz with a few items carries a lot of measurement error. Know how much a score bounces around on its own before you read a difference as a gain.

3. A comparison, decided in advance

A gain on the outcome still needs something to be compared against: students who didn't have the tutor, or the same schools' results in a prior year. Choose the comparison and the analysis before looking at results. Choosing them afterward is how a pilot ends up finding whatever it hoped to find.

  • Who's in each group: if schools volunteered for the pilot, they probably differ from schools that didn't. Compare baseline scores and demographics, and adjust for them, before attributing a gap to the tutor.
  • Who's missing: check who has outcome data. If the students who stopped using the tutor also skipped the benchmark, the pilot group's average is flattered by who's left.
  • What a comparison can't fix: a comparison group strengthens the case only when selection, baseline differences, missing outcomes and implementation have been dealt with. It isn't a box to tick.

Put three numbers on the dashboard

A pilot dashboard worth taking to a board shows three things side by side: who used the tool and how (implementation), what students could do afterward on their own (the independent outcome), and how that compares with similar students who didn't use it, including who dropped out of the data (the comparison). Usage alone is a progress report. All three together are an evaluation.

The practical step is small: write the learning outcome and comparison plan before opening the usage dashboard. If you're running or planning an AI tutor pilot and want that plan to hold up when the expansion decision comes, that's the kind of evaluation work we do.

Related reading

Sources and further reading

Get new posts by email

One email whenever we publish something new. No spam, unsubscribe anytime.

Check your inbox — click the confirmation link to finish subscribing.

{{ subscribeErrorMsg }}