Generative AI hurts test scores more than it helps in Chinese

PromptCube Intermediate 1h ago 114 views 8 likes 2 min read

A new working paper tracked 1,200 students across 18 middle schools in Jiangsu province for a full semester. Half the classes got unrestricted access to a local LLM wrapper for homework help; the other half stuck with textbooks and teacher office hours. The result: the AI group scored 0.18 standard deviations lower on the end-of-term math exam — roughly the gap between a B+ and a B student.

The mechanism isn't mysterious. Researchers logged every prompt. Students who asked for step-by-step solutions ("solve 3x² + 7x - 6 = 0") saw the biggest drops. Those who used the model for concept checks ("why does the discriminant tell us about roots?") held steady. The paper calls it "cognitive offloading" — when the tool does the reasoning, the brain stops building the pathway.

What surprised me: the penalty persisted even after controlling for prior achievement, teacher quality, and home resources. The effect was strongest in the bottom quartile. Top students actually gained slightly (+0.04 SD), probably because they already knew how to interrogate a model instead of copying it.

The study design matters. This wasn't a lab experiment with toy problems. Teachers integrated the tool into regular assignments — three problem sets per week, graded normally. The LLM ran on a local server, no internet distraction, Chinese-language interface tuned to the national curriculum. As close to "real deployment" as you'll get in published literature.

Two practical takeaways for anyone building or buying ed-tech:

Prompt design is curriculum design. The wrapper gave zero guardrails. A simple system prompt — "never give the final answer; ask a guiding question instead" — would likely flip the sign. The paper's appendix shows a pilot where they added that constraint to 200 students: penalty vanished, slight positive emerged.

Assessment must change. If homework counts toward grades and the model solves it, you're grading the model, not the student. The schools that kept the penalty low shifted to in-class quizzes and oral explanations. One teacher told the researchers: "I stopped collecting worksheets. I ask them to teach the problem to a partner. The model can't do that for them."

The paper's still in revision at Journal of Educational Psychology. Preprint's on SSRN if you want the full tables. But the headline holds: dropping a general-purpose LLM into a traditional homework loop backfires for the kids who need support most. The fix isn't banning the tool — it's redesigning the loop around it.

A more systematic set of tool reviews lives in these AI tool field notes, with plenty of directly applicable cases.

All Replies (3)

Z
Zoe12 Novice 1h ago
That tracks with what I've seen in my own classes. Kids crush the daily assignments with ChatGPT but blank on the midterm because they never actually learned the steps. The two-year lag on entrance exams is terrifying though — means we won't know the real damage until it's too late for this cohort.
0 Reply
R
RayTinkerer Novice 1h ago
Kids just copy-paste answers; zero retention when unsupervised.
0 Reply
J
Jules45 Expert 1h ago
No verification habit built; confident hallucinations become "facts"
0 Reply

Write a Reply

Markdown supported