HomeArticle

30-month follow-up of 26,000 Chinese students: AI boosts homework scores by 18% while lowering senior high school entrance examination scores by 24%

AI唱反调2026-08-24 08:23
The more beautifully polished your assignments are, the emptier your mind will be when you sit for the exams.

Children use AI to do their homework, their scores rise and their speed improves — two years later, their scores in the higher education entrance examination drop by 24%.

This is not clickbait from anxiety-driven marketing accounts, but a formal economics paper. In June, the Centre for Economic Policy Research (CEPR) published the discussion paper DP21577, which tracked 26,811 Chinese middle school students for a full 30 months, and its conclusion is strikingly stark: Generative AI improved homework performance by 18%, but reduced closed-book monthly exam scores by 20% within half a year.

The inherent correlation between homework performance and exam performance has been completely reversed by AI.

The CEPR paper tracked 26,800 Chinese middle school students for 30 months: AI increased homework scores by 18%, reduced monthly exam scores by 20%, and reduced high school entrance exam scores by 24%. 80% of students are practicing "cognitive outsourcing", and top students are the ones who suffer the most.

What exactly does this paper study

First, let's clarify the foundation of this paper. Its credibility comes from the high quality of its dataset.

Titled *The Learning Penalty of Generative AI: Evidence from Secondary Education in China*, the paper was written by David Strömberg from Stockholm University, and Victor Lei and Yanhui Wu from the University of Hong Kong. It was released on June 2 under the CEPR Discussion Paper No. DP21577.

The sample covers 26,811 students in Grades 7 to 12 from a county with a population of one million in China, spanning 30 months and covering 9 subjects. The data includes monthly closed-book monthly exams, homework scores and completion duration, as well as high-stakes exams such as the high school entrance examination and college entrance examination. The difference-in-differences method is adopted: a natural experiment is constructed using the time difference of different students' access to AI, to separate the causal effect of "using AI or not" from the correlation.

During the study period, the self-reported AI usage rate of students rose from nearly zero to about 80%, and the outbreak node was right around the release of DeepSeek V2.5 and R1. The most commonly used tools are Doubao, DeepSeek, ChatGLM, Wenxin and Tongyi. This dataset is basically a complete record of the entire process of AI popularization among Chinese middle school students.

Three sets of figures, one paradox

The core findings can be condensed into three sets of numbers:

On the homework side: after using AI, homework scores increased by 18%, the completion time was reduced from 64 minutes to 45 minutes, saving 30% of the time;

On the monthly exam side: within six months, the scores of closed-book monthly exams dropped by 20%;

On the entrance exam side: scores of the high school entrance exam dropped by 24%, scores of the college entrance exam dropped by 18%, and the full penalty will not be fully manifested until about two years later.

In the past, homework performance could basically predict exam performance. After the emergence of AI, this relationship is directly reversed — students with "excellent" homework performance are more likely to have questionable real abilities.

The harm is extremely unevenly distributed, and several details are worth highlighting separately:

In terms of subjects, social sciences suffered the sharpest drop (politics and geography down about 27%), followed by STEM subjects (down 22%), English down 17%, and Chinese the least (down 9%). In terms of groups, **top students are more harmed than underperforming students**: the top third of students saw their scores drop by 24%, while the bottom third only dropped by 16%; junior high school students are worse off than high school students, and male students are worse off than female students.

The most alarming finding is the dose effect: using AI no more than 1 hour per week results in about 5% loss of performance; using AI for more than 5 hours per week directly pushes the loss up to 30%.

Mechanism: 80% of students are practicing "cognitive outsourcing"

Why can high homework scores no longer translate to high exam scores? The paper divides students into two categories based on their behavior patterns.

Outsourcing type: the homework completion time is abnormally short, and the score is abnormally high — AI generates the answers, students copy and submit, and the brain is not involved at all. About 80% of AI users fall into this trap, and almost all the learning losses are concentrated on them. Among students who have used AI for five full months, 81% finish their homework faster than the fastest non-AI user.

Aided type: they use AI, but the homework duration is not significantly shortened — they use AI as an explanation tool, and the thinking process remains in their own minds. These students have very little, near-zero, exam performance loss.

The researchers speculate that top students suffer the most loss precisely because they originally had stronger self-learning abilities, and the intervention of AI abruptly interrupts the process of them building their own knowledge mental models.

The famous "teaching machine" incident was also brought up. Neuroscientist Howard noted in a Fortune China report that the conclusion is very piercing: the "teaching machine" experiment at Ohio State University in 1924 already proved that automation hinders learning — "Cognitive offloading tools that experts use to reduce workload should not be used by children to learn to become experts. What you learn is not a skill, but just dependence."

Why this finding is coming to light now

One key reason: the penalty is delayed.

The monthly exam performance loss manifests in half a year, while the entrance exam performance loss takes two years to show up. This means that all previous short-term studies spanning only a few weeks or months have systematically underestimated the harm of AI — by the time the "bomb" explodes, the remedy window has already closed. The current generation of students who are deeply using AI will take the entrance exams in batches from 2027 to 2032, and face the exams that AI cannot take for them for the first time.

Policy responses have already begun. The Norwegian Prime Minister announced a near-total ban on generative AI for students aged 6 to 13, and students aged 14 to 16 can only use it under teacher supervision. The UK Department for Education released more detailed guidelines in May: school AI tools are set to not give answers by default, require students to try to solve problems independently first, and record who is "offloading their thinking". The World Bank has included this paper in its 2026 Human Capital Blog as a warning to prevent the "AI learning penalty".

A comparison makes it more interesting: AI tutoring under the World Bank's controlled framework brings an effect of +0.31 standard deviations; under the "unregulated usage" scenario in this CEPR paper, the effect is -1.4 standard deviations. For the same technology, the usage determines the direction of its impact.

Some caveats

Of course, we also need to clarify the boundaries of this research.

This is a working paper that has not yet undergone formal peer review. The sample is from a single county, and the 30-month period is exactly the window when AI tools themselves are iterating at a high speed — the research captures the effect of early tools combined with "unregulated usage", while today's product forms and usage norms are constantly changing.

Moreover, the solution proposed by the paper itself is not "banning AI". The performance of aided-type users is almost undamaged, which shows that the problem never lies with AI, but with who does the thinking. A reasonable inference is that what schools really need to adjust is the assessment structure: shift the scoring weight to closed-book exams, monitor abnormal homework duration, and design assessment methods that AI cannot replace, such as oral defense and on-site problem derivation.

Conclusion

The Mark Twain-style irony is fully applicable here: AI makes homework more perfect than ever, but makes learning more hollow than ever.

The homework score, a "competency proxy indicator" that has been used for decades, has officially failed in the AI era. While parents comfort themselves with the high scores on the exercise books, the real knowledge accumulation may be collapsing quietly — and it will not show the bill until the exam paper two years later where no AI can be brought in.

The words of UCL professor Wayne Holmes are worth printing on the startup screen of every education product: "They get better grades, but in reality they learn worse."

This article is from the WeChat Official Account "AI Contrarian", written by Bai Ke, and published with authorization from 36Kr.