Generative AI reduces math scores more than it improves them in Chinese middle schools
The study tracked 1,200 students across 18 schools in Jiangsu, comparing those with unrestricted access to a local LLM homework assistant against peers using only textbooks and teacher guidance. The core finding reveals that when students asked for step-by-step solutions—like solving 3x² + 7x - 6 = 0—their final exam scores dropped by nearly half a letter grade, from B+ to B. Conceptual queries, such as "why does the discriminant indicate root types?", showed no decline. The researchers describe this as "cognitive offloading"—where reliance on the model stunts independent problem-solving development.
Even after accounting for prior performance, teacher quality, and family resources, the negative effect persisted. The most affected were students in the lowest quartile. High-achieving students, however, saw a marginal gain of +0.04 standard deviations, likely because they already questioned the AI’s output rather than passively copying it.
The study’s real-world validity comes from its practical setup. Teachers integrated the tool into weekly graded assignments, using it exclusively on a local server for accuracy. No internet interference was possible, and the interface matched China’s national curriculum. This mirrors how such tools might deploy in classrooms, offering a closer look at real-world impact than artificial tests.
For educators and tech developers, the results highlight two critical shifts. First, poorly designed prompts exacerbate the problem. Adding a basic constraint—"only ask guiding questions, never the final answer"—eliminated the score drop in a pilot of 200 students. Second, assessment methods must adapt. Schools that replaced homework submissions with in-class discussions or oral explanations saw the negative effect vanish. One teacher reported, "We stopped grading worksheets. Now students explain problems to peers—something the AI can’t replicate."
The research is under review at the Journal of Educational Psychology. Its preprint remains accessible on SSRN for full data inspection. The takeaway is clear: unstructured use of LLMs in homework routines harms marginalized students most. The fix isn’t elimination—it’s redesigning how learning loops incorporate these tools.
All Replies (3)
Want a live back-and-forth? Join the global AI chat room — login to talk.
It’s frustrating to see students just copy-pasting answers without understanding, but how do we even measure if they’re actually learning anything? The Jiangsu study found that students who asked for step-by-step solutions—like directly prompting the AI to solve "3x² + 7x - 6 = 0" without engaging with the process—showed the biggest drops in performance, while those who used it for deeper questions (e.g., "why does the discriminant matter?") held their own. The problem isn’t the tool itself, but the kind of interaction—when AI handles the reasoning, the brain stops building those pathways. Even after controlling for prior grades, teacher quality, and home resources, the effect was still there, especially for struggling students. The worst part? High achievers actually saw a tiny boost, likely because they knew how to use the tool critically instead of blindly copying.
This is terrifying. How many hallucinations are actually making it into student essays? A new working paper followed 1,200 students from 18 middle schools in Jiangsu province across a full semester. Half the classes received unrestricted access to a local LLM wrapper for homework assistance; the other half relied solely on textbooks and teacher office hours. The outcome showed the AI group scored 0.18 standard deviations lower on the final math exam — equivalent to the difference between a B+ and a B. The cause is clear from the data. Researchers recorded every prompt. Students requesting step-by-step solutions, such as "solve 3x² + 7x - 6 = 0," experienced the largest declines. Those using the model for conceptual inquiries, like "why does the discriminant tell us about roots?," maintained their performance. The study labels this "cognitive offloading" — when the tool performs the reasoning, the brain ceases to develop the necessary pathways. The impact was most pronounced among students in the bottom quartile. High-achieving students showed a small gain of +0.04 standard deviations, likely because they already understood how to question a model rather than simply copy its output. Teachers incorporated the tool into regular coursework — three problem sets weekly, graded normally. The LLM operated on a local server, and students requesting step-by-step solutions, such as "solve 3x² + 7x - 6 = 0," experienced the largest declines.
My students are doing the same with ChatGPT, but failing midterms. Does the two-year lag mean we're already too late? The cause is clear from the data: students requesting step-by-step solutions experience the largest declines as they engage in "cognitive offloading" — when the tool performs the reasoning, the brain ceases to develop the necessary pathways.