Generative artificial intelligence raised the homework scores of Chinese secondary students by 18 percent and cut the time they spent on assignments by roughly a third. Their closed-book exam scores fell by 20 percent. A working paper posted to the Social Science Research Network on Tuesday, June 23, 2026 tracked 26,811 pupils across 30 months and found that the two movements were caused by the same behaviour.

The paper, titled The Generative AI Learning Penalty: Evidence from Chinese Secondary Education, is written by David Stromberg of Stockholm University, Victor Lei of the University of Hong Kong, and Yanhui Wu of the University of Hong Kong and the Centre for Economic Policy Research. It carries a written date of June 2, 2026. The SSRN listing records a posting date of June 23, 2026 alongside a last-revised date of June 3, 2026, which precedes the posting date; the discrepancy appears in the source record and is not resolved there. At the time of the listing capture, the abstract had drawn 9,665 views and the full paper 5,217 downloads.

What the study measured

The research covers a single county in central China with a population above one million and GDP per capita slightly below 6,000 US dollars, which the authors describe as representative of most Chinese counties outside the developed coastal belt. The sample spans five junior high schools and four senior high schools, 90 percent of local secondary enrolment, with class sizes of 50 to 60 pupils across roughly 524 classes.

Three data streams were combined. The local education bureau supplied monthly closed-book exam results covering seven subjects in lower secondary and nine in upper secondary; weekly homework scores from a digital platform that timestamps both the retrieval of an assignment and the submission of answers, giving a direct measure of completion time; and scores for the two high-stakes gatekeeping exams, the Zhongkao, which sorts pupils into academic or vocational upper-secondary schools, and the Gaokao, which is the sole determinant of college admission.

Treatment timing came from a survey administered in late June 2025 through a class-based WeChat group, with teachers instructing pupils to check the registration dates on the apps they had used before answering. The valid response rate exceeded 96 percent. Reported adoption rose from close to zero in September 2022 to around 80 percent by June 2025, with two accelerations, in September 2024 and January 2025, coinciding with the releases of DeepSeek V2.5 and DeepSeek R1. Doubao, launched in August 2023, was upgraded in August 2024 and again in January 2025.

By June 2025, Doubao was the most reported tool at 47 percent of AI-using pupils, followed by DeepSeek at 36 percent, ChatGLM at 14 percent, Ernie Bot at 9 percent and Qwen at 8 percent. Reported use was heaviest in mathematics, at 66 percent, and English, at 55 percent.

The identification strategy

Because adoption was staggered rather than assigned, the authors use the difference-in-differences estimator of Callaway and Sant'Anna, comparing each adoption cohort's change in outcomes against pupils who had not adopted by the end of the window. Standard errors are clustered at class level. Balance tests show standardised differences between ever-adopters and never-adopters below 0.1 on every demographic variable except parental employment in knowledge-intensive work, which reaches 0.103, well under the 0.25 threshold conventionally treated as a warning sign. Pre-adoption trends are close to flat: ever-adopters recorded exam scores about 0.2 percent higher before adoption, a gap the authors call statistically detectable and substantively negligible.

Output up, capability down

Six months after first use, homework scores among adopters had risen 18 percent above the baseline mean, equal to 1.9 standard deviations, while never-adopters stayed flat. Average completion time per assignment fell from 64 minutes to around 45, a difference of roughly 19 minutes. Averaged across every post-adoption month, including the ramp-up, the homework score gain is 11 percent.

Over the same window, monthly closed-book exam scores fell 20 percent of the baseline mean, or 1.4 standard deviations. The overall average treatment effect across all post-adoption months is a fall of 13 percent. The decline builds over roughly six months and then holds steady.

That six-month build is itself informative. In the randomised trial by Bastani and colleagues, involving about 1,000 Turkish high-school pupils, the negative effect on closed-book mathematics performance appeared inside a single 90-minute session. Here the first month shows only a small effect, growing to 22 percent in mathematics after six months. The authors attribute the difference to tooling: the trial supplied a streamlined system pre-loaded with the homework questions, whereas these pupils had to work out for themselves how to apply general-purpose assistants.

A natural objection is that teachers might have simply written easier exams as class performance sagged. The paper tests this against county-wide exams, where every pupil in a grade sits identical papers set outside the individual school. The estimated effect there is a fall of about 23 percent, or 1.3 standard deviations, close to the school-level result. A separate test regressing never-adopters' scores on the share of adopters in their class and in their school-by-grade cell finds no significant spillover.

The two-year lag on entrance exams

Effects on the Zhongkao and Gaokao accumulate far more slowly. Regression estimates controlling for pre-adoption scores, class fixed effects and demographics put the gap between adopters and non-adopters at roughly 5 to 6 percent on both exams. The difference-in-differences estimate of the total average effect is about 7 percent for each.

Those averages understate the mature effect because most adopters in June 2025 had adopted recently. For pupils who had been using generative AI for two years or more, the estimated effects reach 24 percent for the Zhongkao, equal to 1.5 standard deviations, and 18 percent for the Gaokao, equal to 1.3 standard deviations. Those magnitudes sit alongside June 2025 regular-exam effects of roughly 21 percent for ninth-graders and 14 percent for twelfth-graders.

The proposed mechanism is arithmetic. Entrance exams test cumulative material across several years, so a pupil who adopted in January 2025 carries into the exam a body of pre-adoption learning that dilutes the measured effect. Regular exams test recently covered material, most of it processed after adoption. A placebo test using 2022 Zhongkao scores returns a precisely estimated null.

The 50-minute line

The paper's most operationally interesting section separates AI users by time spent on homework. Among pupils not using AI, homework almost always takes at least 50 minutes, and both homework and exam scores rise monotonically with time spent. Among AI users, more than half finish in 20 to 50 minutes, faster than even the quickest non-AI pupil, and that group records very high homework scores paired with very low exam scores.

The authors label this pattern homework outsourcing, as distinct from tutoring. Their evidence for the label is the homework score itself: the median in the fast group matches the documented accuracy of Chinese assistants on comparable problems. Benchmark work cited in the paper found Qwen-72B and Ernie Bot 4.0 answering roughly 90 percent of questions drawn from Chinese K-12 material in zero-shot settings, and GPT-4 answering 72 percent of Gaokao questions against 57 percent for Ernie Bot.

Applying a 50-minute cutoff, 58 percent of AI users qualify as outsourcing, and 34 percent as fully outsourcing under a 45-minute cutoff. After more than five months of use, those shares climb to 81 percent and 50 percent.

The counterpoint matters as much as the main result. In the 50-to-65-minute band, where AI users and non-users overlap on time spent, median exam scores and interquartile ranges are effectively identical. Those AI users score significantly higher on homework, which indicates they are using the tools, yet their exam performance tracks non-users, and their pre-adoption scores sit close to the baseline mean. In the paper's framing, generative AI reduced time spent learning for the majority without reducing learning efficiency for those who kept their hours.

The time arithmetic complicates a simple reading. A one-third cut in completion time amounts to 2.2 to 2.8 fewer hours a week, or 5 to 6 percent of a weekly study load of 42 to 44 hours. That is small against a 20 percent fall in exam scores, which the authors take as evidence that outsourcing extends beyond weekend assignments into weekday work and classroom engagement, and that homework time is disproportionately productive for learning.

Who loses most

Effects vary sharply, and the ordering is nearly identical across regular and entrance exams despite the two differing in scope, stakes, timing and grading authority. That consistency is offered as cross-validation for the weaker entrance-exam estimates.

By subject, social sciences take the largest hit at 27 percent on average, with Politics, History and Geography clustered around 27 to 31 percent. STEM subjects average 22 percent. English falls 17 percent and Chinese 9 percent. Mathematics, the subject that dominates the experimental literature, sits at 22 percent, equal to 0.84 within-subject standard deviations.

Lower-secondary pupils lose 24 percent against 17 percent for upper-secondary pupils, a difference the authors connect to interviews in which teachers described high-school pupils as more heavily supervised and more restricted in tool use. There is a clean dose response: pupils reporting 0 to 1 hours of weekly use lose 5 percent, while those reporting five hours or more lose about 30 percent.

Boys lose 21.6 percent against 18.4 percent for girls, a gap of 3.2 points, or 17 percent more damage. Adoption rates barely differ, at 82 percent for boys and 80 percent for girls. Weighting the estimated effects by reported weekly hours reproduces 2.7 of the 3.2-point gap, which locates the difference in intensity rather than in any distinct male pattern of use.

The distributional result runs against the grain of the workplace literature. Pupils in the top pre-adoption tercile lose 24 percent against 16 percent for the bottom tercile, an 8-point gap worth roughly half a standard deviation. Workplace studies have generally found generative AI raising task performance most for lower-skilled users. Here the distribution compresses because the strongest pupils fall furthest.

Why nobody noticed

The estimated penalty shrinks over calendar time, from about 25 percent in early 2023 to about 16 percent by June 2025, and the same trajectory holds within a fixed cohort of pupils who adopted in October 2022, ruling out sample composition as the driver. Adaptation is happening. It has not closed the gap.

Aggregate visibility arrived late. Multiplying the treatment effect by the treated share, the authors calculate that generative AI pulled the county-wide average exam score down 3.4 percent across the two school years from September 2022 to June 2024, then 9.9 percent by June 2025 as adoption reached 80 percent.

Individual teachers face a harder detection problem. A drop of the estimated size across all subjects is close to unheard of among non-users, occurring for only five pupils in the entire sample. Within a single subject, a decline that large occurs in roughly 4 percent of pupil-month observations among non-users, so from any one teacher's desk the signal is unremarkable. Pupils and parents see the full picture but, the authors note, tend not to connect it to the tool, partly because the deterioration is gradual and partly because the cognitive effort of unassisted work is commonly misread as evidence of poor learning.

The paper's policy discussion ends on monitoring inputs rather than outputs. Parents and teachers, the authors write, may be more effective if they "monitor inputs, such as homework time and study effort, rather than outputs".

Why this matters for marketing

The finding that transfers out of the classroom is not about students. It is that a productivity metric and a capability metric moved in opposite directions for two and a half years while routine reporting showed only the first.

Marketing organisations have been assembling exactly that measurement structure. IAB Europe reported in September 2025 that 85 percent of European digital advertising companies already deployed AI-based tools, with lack of internal expertise and training cited by 45 percent of respondents as the primary barrier to going further. A year earlier, IAB Europe and Microsoft found 91 percent adoption alongside only 38 percent of professionals confident they could define what artificial intelligence is. Adoption and comprehension were already decoupled before anyone started measuring output.

The Chinese data adds an uncomfortable point: the metric that improved most, the homework score, became actively misleading as a signal of what it was meant to proxy. Among AI users with above-average homework scores, higher scores predicted lower exam scores. The paper calls this a monitoring problem, and the analogue in a marketing department is close to exact. Asset volume, turnaround time and deliverable throughput are the homework score. Campaign outcomes measured months later, on work requiring judgment nobody rehearsed, are the entrance exam.

Evidence from professional settings points the same way. Anthropic research covered in April 2026 found that experienced users of its models succeed roughly 10 percent more often than newer ones, with the gap widening rather than closing, and acknowledged that separating learning-by-doing from survivorship effects would need longer panel data. The Chinese study supplies part of what that request is asking for. Its answer is that the direction of the effect depends on whether the tool substitutes for effort or supplements it, and the group that supplements is a minority.

The talent question follows directly. IAB Polska's Point of Youth report, covered in March 2026, documented artificial intelligence absorbing the repetitive operational tasks that historically served as entry points into digital marketing careers, naming campaign reconciliation, basic segmentation and performance reporting. That report warned of a competency gap emerging within a decade. The Chinese data describes the same mechanism at a shorter timescale and locates the loss precisely: not in the tool, but in the hours of effortful practice the tool removes. An earlier argument for deliberate slowness in marketing workflows made a version of that case without a controlled measurement behind it.

What the paper does not establish

The relationship between homework completion time and exam scores is explicitly described as non-causal. Completion time serves to characterise which pupils are affected, not to prove that adding minutes would restore learning. The authors are direct that mandating longer homework time might not work if the underlying constraint is that outsourcing pupils do not know how to use the tools productively.

Adoption timing is self-reported and retrospective. Teachers instructed pupils to verify registration dates, but compliance is unobserved. The authors argue that late reporting would have produced pre-trends, which are absent, while early reporting would flatten the initial slope without altering the mature effect. The sample is one county, entrance-exam identification is weaker because the outcome is observed once or twice per pupil, and the tools themselves changed substantially across the window.

Whether the pattern generalises from Chinese secondary schools to professional knowledge work is untested. What it does establish, on 26,811 pupils and 526,120 pupil-month exam observations, is that the gap between a fast, high-scoring output and the skill that output was meant to represent can persist for years without appearing in any routine measurement.

Timeline

  • September 2022 - Sample period opens; generative AI use among the county's secondary pupils is close to zero
  • October 2022 - Fixed cohort of early adopters used in the paper's adaptation analysis begins using generative AI
  • Early 2023 - Estimated five-month learning penalty stands at roughly 25 percent
  • August 2023 - Doubao launched
  • August 2024 - Doubao significantly upgraded
  • June 2024 - Cumulative county-wide exam score effect reaches 3.4 percent across two school years
  • September 2024 - First acceleration in adoption, coinciding with the release of DeepSeek V2.5
  • January 2025 - Second acceleration in adoption, coinciding with the release of DeepSeek R1; Doubao upgraded again
  • June 2025 - Survey of 26,811 pupils administered; reported adoption reaches about 80 percent; cumulative county-wide effect reaches 9.9 percent; five-month penalty has fallen to about 16 percent
  • June 2025 - Sample period closes after 30 months
  • September 2025 - IAB Europe reports 85 percent AI tool adoption among European digital advertising companies
  • December 2025 - Analysis argues that practitioner capabilities atrophy under wholesale marketing automation
  • February 2026 - Ahrefs research puts AI Overview click-through reduction at 58 percent for top-ranking pages
  • March 2026 - IAB Polska documents AI absorbing the junior tasks that formed entry points into marketing careers
  • April 2026 - Anthropic research finds experienced users succeeding about 10 percent more often, with the gap widening
  • June 2, 2026 - Paper written
  • June 3, 2026 - Last-revised date recorded on the SSRN listing, preceding the posting date
  • June 23, 2026 - Paper posted to the Social Science Research Network as abstract 6868618

Summary

Who: David Stromberg of Stockholm University, Victor Lei of the University of Hong Kong, and Yanhui Wu of the University of Hong Kong and the Centre for Economic Policy Research, using administrative data on 26,811 secondary pupils supplied by a county education bureau in central China.

What: A working paper finding that generative AI adoption raised homework scores 18 percent and cut homework completion time from 64 to 45 minutes, while lowering monthly closed-book exam scores 20 percent within six months and entrance-exam scores 18 to 24 percent after roughly two years. Losses concentrate among the 81 percent of established users whose completion times indicate homework outsourcing; users who keep their hours show small losses.

When: The panel covers 30 months from September 2022 to June 2025, with the adoption survey administered in late June 2025. The paper carries a written date of June 2, 2026 and was posted to SSRN on June 23, 2026, with the listing also showing a last-revised date of June 3, 2026 that precedes the posting date.

Where: One county in central China with a population above one million and GDP per capita slightly below 6,000 US dollars, covering five junior high schools and four senior high schools that together hold 90 percent of local secondary enrolment.

Why: The finding matters beyond education because it separates a productivity gain from a capability loss inside the same population, over a horizon long enough for the second effect to appear. Marketing organisations measuring generative AI through output volume and turnaround time are reading the equivalent of the homework score, a metric that in this dataset inverted its relationship with the outcome it was meant to predict.