A two year Tennessee trial of Khanmigo, Khan Academy's AI tutor, found math gains on the order of 1.3 national percentile ranks per term, but only for the students who actually used it.
Khanmigo, Khan Academy's AI tutor, was rolled out across 18 Tennessee middle schools for two school years, configured to coach students through problems rather than give them answers during daily remedial math sessions. Ninety-six percent of the students who got access tried it at least once. The median student messaged it on only a third of the days they practiced math, and substantive mathematical dialogue, the kind the model is designed to elicit, occurred in just 17 percent of exercise sessions where a student made a mistake. The students had access. Most of them did not use the tool in a way that would have changed the result.
In the first two-year cluster randomized trial of an AI tutor in U.S. public schools, released this fall by Philip Oreopoulos and Nina Low of the Annenberg Institute at Brown University (EdWorkingPaper 26-1551, DOI 10.26300/kner-hv33), Khanmigo raised math achievement by about 1.3 national percentile ranks per term of assignment, or 0.06 to 0.08 standard deviations over a school year. A full year of active participation topped out at roughly 0.14 standard deviations. The paper frames that 0.14-SD ceiling as the realistic upside of the current design.
The paper is also explicit about a second finding: the AI premium over plain Khan Academy practice was null. Students who used Khanmigo performed at about the same level as students who used the same practice platform without the AI coach. The model, in this configuration, was not the bottleneck on learning. The product was.
Khan Academy is reading the same numbers. In "Learning in the Open: What AI Is (and Isn't) Changing" and "Built in the Open: How Pilot Districts Shaped the Reimagined Khan Academy," the organization frames the next phase of work as a product and pedagogy problem, not a capability one. EdTech Innovation Hub's coverage of a Khan disclosure that only 15 percent of students with access to Khanmigo actually use it, alongside the company's annual report, are the public markers that Khan sees the same constraint the trial measured.
What changes the constraint is design. The trial points to specific levers: prompts that pull students into the conversation when they make a mistake rather than after; teacher workflows that surface Khanmigo as part of class, not as an extra; friction in the path from "ask for the answer" to "ask for a hint." These are the things the next deployment has to get right, and the kinds of things a smarter underlying model cannot substitute for. The 0.14-SD ceiling is the line a redesign will be measured against.
The Tennessee data points to a clean separation: giving every student an AI tutor is one problem; teaching every student to use it is another. The first is solved by an API key. The second is solved by a product, a teacher, and a student who shows up. The past two years of AI-tutoring announcements have mostly been about the first. The next two years will be about the second.
There is a clear falsifier for the engagement-as-binding-constraint read. A future experiment that holds the product constant and varies model capability, or a meaningful capability-tier gap inside this same trial, would push the explanation back toward model quality. The Tennessee result does not rule that out. It just sets the prior: until the engagement problem is solved, swapping in a smarter tutor moves the line only at the margin.
The Hacker News thread on the paper surfaced the student-side version of the engagement gap, including reports of tutors that hallucinate answers and students who type "just give me the answer" until the model complies. The next two years of district pilots will be the second RCT-scale test of whether the redesign, not the model, closes that gap.