Key Points
- Chinese study tracked 27,000 pupils and linked AI use to higher homework marks but 20% lower exam scores.
- Learning losses clustered among pupils who finished assignments faster, while longer-working AI users faced little penalty.
- A separate college experiment found chatbot-assisted learning gains persisted when students repeated the task one week later.
The latest
A six-month study of 27,000 Chinese pupils aged 12-18 found that students using artificial intelligence improved homework scores while performing substantially worse in exams than classmates who did not use it. Around 80% reported using models including Doubao and DeepSeek. Their average homework marks rose 18% across subjects and assignment time fell from 64 minutes to 45, but their exam scores were 20% lower than the control group’s, according to the reported findings.
Details
- Scale of use: The research comes amid rapid student adoption. A Chegg survey last year found 80% of undergraduates in wealthy countries used AI in their studies. More recent polls put usage at 94% in Britain and 93% in Germany, while teachers reported receiving formulaic essays they suspected were generated by ChatGPT.
- Homework disconnect: Before AI became widespread in the Chinese sample, homework performance predicted exam results. The relationship reversed among users: pupils with the highest homework scores became more likely to perform poorly under exam conditions, indicating that stronger submitted work did not necessarily reflect retained understanding.
- Time mattered: The exam-score decline was concentrated among pupils who rushed homework. AI users who spent roughly as long on assignments as non-users incurred little penalty. The report says stronger performers appeared to use chatbots to explain difficult concepts or address specific problems rather than copying answers, though that interpretation was not directly established.
- Researchers involved: David Stromberg of Stockholm University and Victor Lei and Wu Yanhui of the University of Hong Kong tracked pupils over six months. The material does not provide the study’s publication venue, sampling method, subject breakdown or controls, limiting assessment of causality and wider applicability.
- Separate experiment: Zara Contractor and Germán Reyes of Middlebury College asked undergraduates to study an unfamiliar topic with or without a chatbot. AI-assisted students scored higher in tests, and that advantage remained when they repeated the task a week later. No sample size or effect magnitude was supplied.
Between the lines
The findings distinguish answer generation from guided learning. Faster completion coincided with weaker exam performance in the school study, while sustained engagement appeared to reduce the downside. The college experiment points in the other direction, showing that chatbot support can accompany retention when a task is structured around learning rather than merely producing homework.
What’s next
The concrete test for schools is whether future assessments track both assignment time and closed-book exam performance among AI users and non-users. Publication of the Chinese study’s methods, controls and subject-level results would determine how confidently its 20% exam gap can be attributed to AI rather than differences between the groups.