HN Debrief

AI boosted homework scores, then exam scores dropped: study

  • AI
  • Education
  • Workplace
  • Productivity

The article covers a large study of roughly 27,000 students aged 12 to 18 in China. It found a clean split between performance on homework and performance on exams after generative AI adoption. Homework scores went up and time spent fell, but exam scores dropped by about 20% relative to non-users. The paper’s own breakdown pushed the conversation past the headline. Students who spent about the same time on homework as non-AI peers learned about the same. The losses showed up when AI reduced the time and effort spent wrestling with the material, which is exactly the part homework was supposed to force.

Treat AI in education and training as a workflow shortcut unless you have a way to preserve deliberate practice and independent assessment. If your team, school, or product measures progress through take-home work alone, expect the metric to get gamed and to lose its connection to real capability.

Discussion mood

Mostly negative and unsurprised. People saw the study as evidence that generative AI is great at inflating visible output while eroding the practice, effort, and independent recall that exams and real work still demand.

Key insights

  1. 01

    Same study time meant same learning

    The paper’s most useful nuance is that AI did not make students who put in equal time learn more. It mostly let many students spend less time on homework and then know less later. That cuts against the popular claim that AI is a pure efficiency gain for learning. In this dataset, when practice time overlapped, outcomes also overlapped.

    Do not assume faster completion means better learning. If you introduce AI into training, track time on task and independent performance, not just output quality.

      Attribution:
    • kzz102 #1
  2. 02

    Top students were not protected

    Higher prior achievement did not shield students from the downside. A commenter pulled the paper’s estimate that the negative learning effects were largest in the highest achievement tercile. That matters because it weakens the comforting story that only disengaged students get hurt. Capable students can also let convenience displace effort.

    Do not reserve concern for struggling learners. High performers also need guardrails if the tool makes it easy to skip the hard parts of practice.

      Attribution:
    • bonoboTP #1
    • kzz102 #1
  3. 03

    AI exposed broken assessment design

    Several comments pushed past AI and attacked the exams themselves. The sharper version was that most academic assessment is ad hoc, rarely psychometrically validated, and often used for screening, hiring, and status sorting far beyond what it was designed to measure. AI is making that long-standing mismatch harder to ignore because it severs the old link between coursework and whatever exams were actually capturing.

    If you rely on assignments, take-homes, or grades as proxies for capability in hiring or internal training, revisit the assessment itself. The proxy may already be weak, and AI will weaken it further.

      Attribution:
    • doctorpangloss #1
    • jkhdigital #1
  4. 04

    In-class checks restore the missing signal

    The most concrete classroom adaptation was not a grand AI policy. It was small, frequent, hard-to-outsource checks. One instructor used flipped teaching plus short closed-book quizzes on the assigned reading at the start of class. Another commenter made the broader point that homework used to help teachers see who was tracking the material, and AI destroys that signal unless understanding is checked live.

    Prefer low-friction assessment you can observe directly. Short in-class quizzes, oral follow-ups, and supervised problem solving will tell you more than polished take-home work.

      Attribution:
    • comicjk #1
    • yoyohello13 #1
    • ezst #1
  5. 05

    A chatbot is not a tutor by default

    The tutoring analogy got dismantled in a useful way. Good tutors do not just supply answers. They impose pacing, push back, and build independence because they are scarce humans with judgment. A chatbot can simulate that only if the student deliberately asks for constraint, which is exactly what many students trying to avoid effort will not do.

    If you want AI to behave like tutoring, the default product needs to withhold answers and force intermediate reasoning. Leaving that discipline to the user will fail for the users who most need help.

      Attribution:
    • Barrin92 #1
    • jasonfarnon #1

Against the grain

  1. 01

    Cheating is older than AI

    Some people argued the study overstates novelty because students have always shared answers, copied from peers, or pasted from Wikipedia. AI changes the ease and scale, but not the basic temptation. That framing matters because it suggests schools should avoid treating this as an isolated model problem and instead design for a world where assignment cheating is normal whenever the opportunity exists.

    Do not wait for perfect AI detection or model restrictions. Build courses, interviews, and training around the assumption that unsupervised take-home work is vulnerable by default.

      Attribution:
    • elictronic #1
    • shimman #1
  2. 02

    School may be testing the wrong skills

    A minority view held that falling exam scores are not automatically alarming because schools often reward memorization and lag real work. In many jobs, the relevant skill is using tools well, not reproducing facts unaided. Even that camp still drew a line at fundamentals. Outsourcing too early can leave students without enough domain knowledge to judge the tool’s output.

    Separate obsolete recall tasks from genuine foundations before rewriting policy. Some assessments should change for an AI-heavy workplace, but abandoning independent competence entirely is a mistake.

      Attribution:
    • grahammccain #1
    • iambenm #1
    • nick486 #1
  3. 03

    Calculator analogy is not fully dead

    A few comments resisted the strong anti-AI reading by comparing this moment to earlier fights over calculators and the internet. The useful version of that argument was narrow. School often restricts tools until students first demonstrate baseline competence, then permits them in more advanced settings. That is a different posture from blanket acceptance or blanket bans.

    Stage AI access by skill level. Require students or junior staff to first show they can do core tasks unaided, then let tools accelerate higher-order work.

      Attribution:
    • OroPla #1
    • graemep #1

In plain english

AI
Artificial intelligence, software systems that perform tasks such as analyzing code or generating text.
generative AI
AI tools that create new text, images, code, or other content in response to prompts.
tercile
One of three equal-sized groups in a ranked dataset, such as top, middle, and bottom third.

Reference links

Primary sources

Homework and pedagogy references

Related discussions and examples

Broader background links