The article covers a large study of roughly 27,000 students aged 12 to 18 in China. It found a clean split between performance on homework and performance on exams after generative AI adoption. Homework scores went up and time spent fell, but exam scores dropped by about 20% relative to non-users. The paper’s own breakdown pushed the conversation past the headline. Students who spent about the same time on homework as non-AI peers learned about the same. The losses showed up when AI reduced the time and effort spent wrestling with the material, which is exactly the part homework was supposed to force.
That led most people to a blunt conclusion. AI did not reveal a new way to learn faster. It broke a proxy. Homework used to stand in for practice and give teachers a rough read on understanding. Once students could outsource the work, higher homework marks stopped meaning much and in some cases inverted, with the best homework scorers doing worse on exams. Several people also pointed out a nastier detail from the paper. The effect was not limited to weak students. The paper says the negative learning effects were larger for higher-achieving students, which fits the idea that AI gradually crowds out effort even among capable kids.
The conversation mostly landed on pedagogy, not model quality. People compared AI to calculators, Wikipedia, tutors, and GPS, but the useful distinction was simpler. Tools help when they support practice and feedback. They hurt when they remove the struggle that builds fluency. That is why the strongest practical suggestions were not “ban AI” or “embrace AI,” but redesign assessment around closed-book exams, oral questioning, in-class work, short quizzes on pre-class reading, and other formats where you can still observe a student thinking. A second recurring point was that many schools were already leaning too hard on homework and grades as cheap measurement tools. AI just made that weakness impossible to ignore.
Treat AI in education and training as a workflow shortcut unless you have a way to preserve deliberate practice and independent assessment. If your team, school, or product measures progress through take-home work alone, expect the metric to get gamed and to lose its connection to real capability.
Mostly negative and unsurprised. People saw the study as evidence that generative AI is great at inflating visible output while eroding the practice, effort, and independent recall that exams and real work still demand.
Key insights
01
Same study time meant same learning
The paper’s most useful nuance is that AI did not make students who put in equal time learn more. It mostly let many students spend less time on homework and then know less later. That cuts against the popular claim that AI is a pure efficiency gain for learning. In this dataset, when practice time overlapped, outcomes also overlapped.
Do not assume faster completion means better learning. If you introduce AI into training, track time on task and independent performance, not just output quality.
Higher prior achievement did not shield students from the downside. A commenter pulled the paper’s estimate that the negative learning effects were largest in the highest achievement tercile. That matters because it weakens the comforting story that only disengaged students get hurt. Capable students can also let convenience displace effort.
Do not reserve concern for struggling learners. High performers also need guardrails if the tool makes it easy to skip the hard parts of practice.
Several comments pushed past AI and attacked the exams themselves. The sharper version was that most academic assessment is ad hoc, rarely psychometrically validated, and often used for screening, hiring, and status sorting far beyond what it was designed to measure. AI is making that long-standing mismatch harder to ignore because it severs the old link between coursework and whatever exams were actually capturing.
If you rely on assignments, take-homes, or grades as proxies for capability in hiring or internal training, revisit the assessment itself. The proxy may already be weak, and AI will weaken it further.
The most concrete classroom adaptation was not a grand AI policy. It was small, frequent, hard-to-outsource checks. One instructor used flipped teaching plus short closed-book quizzes on the assigned reading at the start of class. Another commenter made the broader point that homework used to help teachers see who was tracking the material, and AI destroys that signal unless understanding is checked live.
Prefer low-friction assessment you can observe directly. Short in-class quizzes, oral follow-ups, and supervised problem solving will tell you more than polished take-home work.
The tutoring analogy got dismantled in a useful way. Good tutors do not just supply answers. They impose pacing, push back, and build independence because they are scarce humans with judgment. A chatbot can simulate that only if the student deliberately asks for constraint, which is exactly what many students trying to avoid effort will not do.
If you want AI to behave like tutoring, the default product needs to withhold answers and force intermediate reasoning. Leaving that discipline to the user will fail for the users who most need help.
Some people argued the study overstates novelty because students have always shared answers, copied from peers, or pasted from Wikipedia. AI changes the ease and scale, but not the basic temptation. That framing matters because it suggests schools should avoid treating this as an isolated model problem and instead design for a world where assignment cheating is normal whenever the opportunity exists.
Do not wait for perfect AI detection or model restrictions. Build courses, interviews, and training around the assumption that unsupervised take-home work is vulnerable by default.
A minority view held that falling exam scores are not automatically alarming because schools often reward memorization and lag real work. In many jobs, the relevant skill is using tools well, not reproducing facts unaided. Even that camp still drew a line at fundamentals. Outsourcing too early can leave students without enough domain knowledge to judge the tool’s output.
Separate obsolete recall tasks from genuine foundations before rewriting policy. Some assessments should change for an AI-heavy workplace, but abandoning independent competence entirely is a mistake.
A few comments resisted the strong anti-AI reading by comparing this moment to earlier fights over calculators and the internet. The useful version of that argument was narrow. School often restricts tools until students first demonstrate baseline competence, then permits them in more advanced settings. That is a different posture from blanket acceptance or blanket bans.
Stage AI access by skill level. Require students or junior staff to first show they can do core tasks unaided, then let tools accelerate higher-order work.