A 27,000 student Chinese panel found the same students gained 18% on homework and lost 20% on closed book exams in six months, with the gap tracking a homework outsourcing habit.
A working paper from the Centre for Economic Policy Research finds that generative AI lifted homework scores 18% in a panel of 26,811 Chinese secondary school students and then cut the same students' closed-book exam scores by 20% over the same six-month window. Both numbers come from the same students in the same months. The gap points to one mechanism: the faster AI helps you finish homework, the less you carry to a no-tool test.
The paper, "The Generative AI Learning Penalty: Evidence from Chinese Secondary Education" by David Strömberg of Stockholm University and Victor Lei and Yanhui Wu of the University of Hong Kong, runs as a CEPR Discussion Paper (DP21577, SSRN abstract 6868618). It tracks one county in central China across 30 months and nine subjects, combining monthly closed-book exams, high-school and college entrance exams, and timestamped homework submissions. Staggered AI adoption across classrooms gives the authors a difference-in-differences design that separates the AI effect from a county-wide trend.
About 80% of AI users in the panel behave as homework outsourcers, with exceptionally short completion time paired with high homework scores. That group drives nearly the entire 20% exam drop. The remaining 20% of AI users, who keep similar study time to non-users, show only small learning losses on monthly exams. Homework time across the panel fell 30% on average, a behavioral shift large enough to read on its own. In a closed-book exam, a shortcut on practice shows up as a shortcut on recall.
The losses are not uniform. Social-science subjects carry the largest drops, followed by STEM, then languages. Junior students, high-achieving students, and boys take a disproportionate share of the hit. The entrance-exam penalty is bigger than the monthly-exam penalty: high-school and college entrance scores fall 18% and 24% respectively, and the authors find the full effect only after roughly two years, not the initial six months. The score gap on a high-stakes test is wider than the gap on a routine quiz, and the gap widens further as the stakes rise.
Around 80% of surveyed students in the panel reported using Chinese AI chatbots such as Doubao and DeepSeek, China's two most popular general-purpose assistants. A Chegg survey found 80% of undergraduates in wealthy countries reported using AI in their studies, with more recent polling cited by The Economist's summary putting Britain at 94% and Germany at 93%. The question has shifted from whether students use AI to what it does to what they retain. The study is one data point, and it lands on a population that already lives inside the shift it describes.
The paper is a working paper, not a peer-reviewed publication, and it covers one county, one age band, and a single six-month effect window. The general pattern, that assisted practice does not always transfer to a no-tool test, is consistent with prior lab work. The distributional findings give teachers and parents a specific handle: which students (juniors, high-achievers, boys), which subjects (social sciences first), and which use pattern (fast-finish outsourcers) carry the most exposure. A second-order effect is harder to measure and more important: if the 80% outsourcers keep their habit into college, the same gap will show up in coursework grades that no longer close.
Discussion on Hacker News leaned toward reading the exam gap as AI-outsourcing. The panel data points the same way.