Why AI help can raise homework marks but lower exam scores

A pattern reported in education research has drawn attention: students who use AI assistants score better on the work done with them, then worse on.

A pattern reported in education research has drawn attention: students who use AI assistants score better on the work done with them, then worse on unaided tests. It suggests help that improves output can weaken the learning the work was meant to produce.

Key takeaways

  • Reports of students improving on AI-assisted assignments while declining on later closed-book assessments describe a gap between performance and learning.
  • Performance during practice and retention afterwards are distinct measures, and educational research has long documented cases where the two diverge.
  • The effect is usually attributed to the assistant doing the cognitive work — retrieval, structuring, error correction — that the exercise was designed to force the student to do.
  • The size and reliability of any such effect depend heavily on subject, task design, student level and how the tool is used, none of which can be generalised from a single result.
  • The practical response being discussed is not banning tools outright but redesigning what is graded, so assessment measures unaided capability where that is the goal.

What is actually being described

The claim at the centre of the discussion is a divergence between two measurements taken from the same students. On assignments completed with access to a generative AI assistant, scores go up. On subsequent assessments completed without that access, scores go down relative to a comparison group that worked unaided throughout.

Both halves matter. If only the first were reported, the natural reading would be that the tool helps. If only the second, the natural reading would be that it harms. Together they describe something more specific: the assistant reliably improves the artefact a student hands in, while the internal change the assignment was supposed to cause does not occur, or occurs less.

It is worth being precise about what such a finding does and does not establish. It concerns a particular set of tasks, a particular subject area, a particular way of using the tool, and a particular kind of test. Without access to the underlying methodology, the exact conditions, sample and effect size cannot be verified here, and should not be assumed. What can be discussed is the mechanism the result points at, which is not new and is not specific to AI.

Why this is being discussed now

Generative AI assistants moved into student workflows faster than institutions could form policy about them. The first phase of the response was detection and prohibition: rules about what counts as misconduct, tools claiming to identify machine-written text, and a great deal of argument about whether either worked.

That phase produced few settled answers. Detection has proven unreliable enough that many institutions have stepped back from relying on it, and blanket bans are difficult to enforce on work done outside supervised settings. Attention has therefore shifted from policing use to measuring effects, which is where results of this kind land.

The timing also reflects a data lag. Assistants became widely available only recently, and studies that compare assisted practice against later unaided assessment take time to design, run and write up. Findings are now beginning to appear in enough volume to be argued over rather than treated as isolated curiosities. Interest is amplified by the fact that the result cuts against the most common institutional framing of these tools, which is that they function as tutors.

The background a newcomer needs

Education research has a long-standing distinction between performance and learning. Performance is how well someone does a task at the moment they do it. Learning is the durable change that lets them do it later, unaided, possibly in a different context. These are correlated but not identical, and conditions that raise one can lower the other.

The related idea is that of difficulty which is productive rather than merely obstructive. Effortful recall of information from memory tends to strengthen it more than reviewing that information does. Spacing practice out helps retention more than massing it together, even though massed practice feels more effective at the time. Struggling with a problem before being shown the method often produces better later transfer than being shown first.

Each of these makes practice feel harder and often look worse in the moment. A tool that removes friction — by supplying the next step, correcting an error before it is felt as an error, or producing structure the student would otherwise have to impose — is acting on exactly the mechanism that these findings identify as valuable. That is the theoretical reason a result of this shape is unsurprising to people in the field.

Who is affected and how

Students are affected unevenly. Someone using an assistant to check finished work, or to get an explanation of a concept they then practise on their own, is in a different position from someone who uses it to generate the work itself. The same tool supports both, and the distinction is invisible in the submitted artefact.

Teachers face a measurement problem. If assignments done outside supervision no longer indicate what a student can do, then a large part of the routine feedback loop stops functioning. Marks stay high while the signal they carry degrades, which is harder to notice than marks falling.

Institutions face a design problem, and it is expensive. Moving assessment weight towards supervised conditions means more invigilated exams, more oral assessment and more in-class work, all of which cost staff time. Doing nothing means credentials that certify less than they used to.

Employers and licensing bodies sit downstream of all of this. Their interest is in whether a qualification still predicts unaided capability, in fields where unaided capability continues to matter.

Where informed people disagree

There is real disagreement about how much unaided capability will continue to matter. One view holds that if assistants are permanently available in professional practice, then testing people without them measures an increasingly irrelevant skill, in the way that mental arithmetic became less central once calculators were universal. The counter-view is that competent use of an assistant depends on the underlying knowledge needed to judge its output, so hollowing out that knowledge undermines the assisted performance too.

There is disagreement about generalisability. Findings drawn from one subject, one age group and one style of task may not transfer. Well-designed instructional uses of these tools may produce different results from unstructured use, and studies of the former are not interchangeable with studies of the latter.

There is disagreement about what the exam measures. If later assessment is heavily weighted towards recall, a decline may indicate weaker memorisation rather than weaker understanding — a real effect, but a narrower one than the framing implies.

And there is disagreement about durability, since tools, and the habits people form around them, are both changing quickly.

What this implies in practice

The response most often proposed is not prohibition but a change in what carries assessment weight. If a task is meant to build a capability, and a tool can complete that task without the capability being exercised, then the task has stopped doing its job and needs redesigning rather than defending.

Concrete versions of this include shifting graded weight towards supervised conditions, using unassisted work primarily as low-stakes practice with feedback, and assessing process — drafts, reasoning, defence of choices — alongside product. Some courses make the assistant’s role explicit, requiring students to submit the exchange and critique it.

For individual learners, the operative distinction is between using a tool to produce an answer and using it to check one already attempted. The second preserves the effortful step; the first removes it. This is a claim about mechanism rather than a verified prescription, but it follows from well-established findings about retrieval and effort.

What to watch next

Watch for replication across subjects and levels, since a single result establishes far less than a consistent pattern. Watch whether studies begin to separate modes of use rather than treating any AI access as one condition, because that separation is where the practically useful answers are.

Watch institutional assessment policy, which is the clearest indicator of what decision-makers actually believe. A shift of grade weight towards supervised settings is a costly move and signals genuine concern; statements of principle without that shift signal less.

Watch how the tools themselves change, as some are being built deliberately to withhold answers and prompt the student instead. Whether that design produces measurably different outcomes is an open empirical question.

Finally, watch for longer-horizon evidence. Almost all current data covers short windows. What happens to capability over a full course, a degree or a career is not yet known.

Frequently asked questions

Does using AI for homework make students worse at exams?

Some reported studies describe exactly that pattern: higher scores on assisted work, lower scores on later unaided assessment. Whether it holds generally is not established. The result depends on subject, task type, how the tool is used and what the exam measures. Treat it as a documented risk in specific conditions rather than a settled rule that applies to all students and all forms of AI use.

Why would help during practice reduce later performance?

Because the effort a student expends is part of what produces learning. Retrieving information from memory, structuring an argument and working through an error all strengthen the capability being practised. A tool that supplies those steps removes the effort and the strengthening along with it. The work still gets done, and gets done better, but the internal change the exercise was designed to cause does not happen.

Is this specific to AI, or does it apply to other study aids?

The underlying mechanism is not specific to AI. Education research has long documented cases where conditions that make practice easier improve immediate performance and reduce retention. Re-reading notes rather than testing yourself is a familiar example. What is distinctive about generative assistants is scope and convenience: they can complete a far wider range of tasks, on demand, at no cost in effort.

Should schools ban AI assistants?

That is contested, and enforcement is a serious obstacle for work done outside supervision. The more commonly proposed response is to change what is graded rather than what is forbidden — shifting assessment weight towards supervised conditions, using unassisted work as low-stakes practice, and assessing reasoning as well as output. Which approach institutions adopt varies widely, and no consensus has formed.

Are there ways to use AI that support learning instead?

The distinction most often drawn is between generating an answer and checking one you have already attempted. The second preserves the effortful step; the first removes it. Using an assistant to explain a concept you then practise unaided, or to critique your reasoning after you have produced it, keeps the difficulty in place. This follows from established findings about effort and retrieval rather than from verified trial results.

Does this mean AI tutoring does not work?

No. Assisted homework use and designed instructional tools are different things, and a finding about one does not settle the other. Systems built to withhold answers, prompt recall and adapt difficulty are attempting to preserve the effort that unstructured assistance removes. Whether they succeed is an open question, and evidence on well-designed educational applications should be assessed separately from evidence on general-purpose assistants.

Sources and further reading

  • Peer-reviewed education and cognitive psychology journals, for the established literature on the performance–learning distinction, retrieval practice and spaced repetition.
  • University teaching and learning centres, which publish practical guidance on assessment redesign in response to generative AI.
  • Higher education trade press, for coverage of institutional policy changes on assessment and academic integrity.
  • Aggregator discussion threads such as Hacker News, useful as an indicator of which findings are circulating, not as evidence for the findings themselves.

Surfaced from the hackernews signal “AI use and exam performance”. AI-assisted draft, editorially reviewed.

Visited 2 times, 2 visit(s) today
share this recipe:
Facebook
X
WhatsApp
Telegram
Email
Reddit