BetaMaShop is in public beta. We improve it continuously, and your feedback shapes what comes next.
MaShop/Blog/Research/AI Grading Works for Comments, Not for Scores
ResearchAugust 6, 2026
Read · 5 min
ai grading · automated essay scoring

AI Grading Works for Comments, Not for Scores

Scoring and feedback are two products under one name. What the agreement research found, why the rubric decides more than the model, and the rules to set.

Key takeaways
  • Scoring and feedback are two different products sold under one name, and the evidence behind them is not remotely equal. That split runs through the rest of the week too, as a task by task look at where teacher time actually goes sets out.
  • On scoring, a study of five models across 67 university essays found agreement with human raters that was low and not statistically significant, with the models also disagreeing with themselves across repeat runs.
  • On feedback, a meta-analysis of 36 studies found a moderate positive effect on academic achievement, g = 0.61 with a confidence interval of 0.42 to 0.80.
  • Automated writing evaluation gains decay. One programme covering more than 130,000 students went from a 9.14 point gain in year one to 1.37 by year three.
  • The honest benchmark for a score is not correctness, it is how much two human markers agree with each other, and that figure is lower than most people assume.
  • Rubric quality decides more than model choice. A vague criterion produces unstable scores from humans and machines alike.

It is Sunday evening and there are 140 essays in the pile. The district has offered a tool that will mark them by Tuesday. The fear is not that it will be lazy, it is the parent meeting in three weeks where a grade has to be defended and the only honest answer is that software produced it.

That fear is well calibrated, and the way out of it is a distinction the vendors deliberately blur. Assigning a grade and producing comments are different tasks with different evidence and different consequences when they go wrong. Take them separately and most of the anxiety resolves into a policy decision you can actually make.

Diagram comparing automated scoring and automated feedback across what each produces, its stakes and its evidence base
Two jobs, one product name. The evidence behind them is not equal.

What does the evidence say about scoring?

That it is unreliable on open ended work, and that the models are not even consistent with themselves.

The clearest recent test is a study of five models scoring 67 university essays, using a four criterion rubric covering pertinence, coherence, originality and feasibility. Each model scored every essay three times so the researchers could measure stability, not just accuracy.

The results are worth reading slowly. Agreement between human raters and models was consistently low and non-significant. Within model reliability across the three repeated runs was similarly weak, with a median Kendall's W below 0.30, meaning a model asked the same question three times produced meaningfully different answers. The models tended to inflate coherence. Across models there was moderate convergence on coherence and originality, and negligible concordance on pertinence and feasibility.

That last split is the useful part. The dimensions where the models agreed with each other are the ones you can assess from the surface of the text. The dimensions where they did not are the ones requiring knowledge of the subject and the assignment. The models were consistent about form and inconsistent about substance.

Note

The study is small: 67 essays, one course, one language. Do not treat it as the last word. Treat it as a reason to demand agreement statistics from any vendor selling you scoring, measured on your rubric and your students rather than on a public dataset.

What does the evidence say about feedback?

Something quite different, and considerably more positive.

A meta-analysis of 36 experimental and quasi-experimental studies published between 2023 and 2025, yielding 72 effect sizes, found a moderate positive effect of generative AI feedback on academic achievement: g = 0.61 with a 95 percent confidence interval from 0.42 to 0.80. Cognitive outcomes came in at g = 0.60 and non-cognitive outcomes at g = 0.29. A metacognitive figure of g = 1.43 appears in the paper and rests on only four studies, so it should be read as a signal rather than a number.

The single significant moderator was teaching method, at p = 0.026. Feedback helped more where students were actively constructing something, with collaborative learning at g = 0.71 and self-directed learning at g = 0.68, and helped less in direct instruction. In plain terms: the tool works better when the student has to do something with the comment.

The authors state their own limits clearly, including a concentration of studies in Asia, the small overall sample, and heterogeneity that the subgroup analyses largely failed to explain. Read the effect as real and the precision as provisional.

Does the benefit last?

Often not, and this is the finding most likely to change how a school buys. ASCD's review of automated writing evaluation cites a meta-analysis of 20 studies putting the effect around a gain of 7 percentile points, with larger gains for multilingual learners. Then it reports a programme tracking more than 130,000 students where the state assessment gain fell from 9.14 points in year one to 4.38 in year two and 1.37 in year three.

Novelty is a plausible explanation, and so is the more uncomfortable one: students learn what the system rewards and produce it. Either way, a purchasing decision built on first year results is buying a number that has been observed to decay.

Which tasks hold up, and which do not?

TaskReliability evidenceDefensible to a parent?Sensible use
Multiple choiceDeterministic, no model neededYes, triviallyAutomate fully
Short factual answersGood where the answer set is boundedYes, with a spot checkAutomate with sampling
Rubric scored essaysWeak. Low, non-significant human agreement in the study aboveOnly if a human confirms the scoreMachine drafts, human decides
Open creative writingWeakest. The criteria that matter are the ones models handle worstNoComments only, never a grade
CodeStrong where tests exist, weak on style and designYes for the tested partAutomate the tests, mark the reasoning yourself

What is the honest benchmark for a score?

Agreement between two humans, not truth. This reframing does more work than any statistic in this article.

There is no correct grade sitting in the essay waiting to be discovered. Two experienced markers using the same rubric routinely differ, which is why serious assessment programmes use double marking and reconciliation. So the question is never whether the machine got it right. It is whether machine to human agreement is comparable to human to human agreement on the same task.

That standard is fair to the technology and brutal in practice, because it requires you to know your own inter-rater agreement first. Most schools do not. If you have never had two teachers mark the same twenty scripts blind and compared the results, you do not have the baseline that would let you evaluate any vendor claim, and the vendor knows it.

The measurement discipline here is identical to the one used for any model based feature, where the temptation is to trust a single number produced under favourable conditions. We set out the general method in how to tell whether a change actually helped, and the classroom version needs one extra ingredient: the human baseline you are comparing against.

Why does the rubric matter more than the model?

Because a vague criterion is unmeasurable by anything, and swapping models does not fix a definition problem.

Take a criterion written as: Argument is well developed and shows insight. Nothing in that sentence tells a marker what to look for. Two teachers reading the same paragraph will disagree about insight, and the same model asked three times will disagree with itself, which is precisely the instability the essay study measured.

Now the same criterion specified: States a claim in the opening paragraph. Supports it with at least two pieces of evidence drawn from the set texts. Names one objection and answers it. Every clause is checkable. A marker can point at the sentence that satisfies it or the absence of one. Scores tighten immediately, for humans and machines alike, and the appeal conversation becomes a matter of pointing at the page.

The order of operations follows: specify the rubric, measure your own agreement, then evaluate a tool. Doing it in the reverse order produces a purchase decision made on a metric nobody can interpret.

Card listing four governance rules a teacher should apply before any machine assigns a grade to student work

What is the failure mode nobody names?

That fluent structure gets rewarded over correct reasoning. The essay study found models inflating coherence while handling context dependent dimensions inconsistently, and that pattern has a predictable victim.

The student who writes in clean five paragraph form, signposts every transition and says very little scores well. The student who has an unusual idea, expresses it awkwardly and buries the good sentence in the middle scores badly. Human markers make this error too. A machine makes it consistently, at scale, and without the moment of hesitation where a teacher rereads a strange paragraph and realises it is right.

This is also the mechanism behind gaming. Once students work out that surface features move the score, they optimise for surface features, which is the most plausible explanation for effect sizes that shrink year after year. It is the same dynamic playing out wherever machine judgement meets people who benefit from a particular verdict, from content detection getting harder to fool to platforms rewriting their rules on generated posts. The measurement changes the behaviour it measures.

How do you test a tool before it touches a real grade?

With about two hours of work and a set of scripts you have already marked. This is the part vendors will not do for you and it is the only evidence that describes your students.

Take twenty pieces of work from a previous term, with your marks removed. Run them through the tool. Compare its scores to yours and count how often it lands on the same band, one band away, and two or more bands away. Then look only at the disagreements and ask which of you is right, because that is where you learn whether the tool is wrong or your rubric is ambiguous.

Two further checks cost almost nothing. Run five of the same scripts through the tool twice and see whether the scores move, since a tool that disagrees with itself cannot be defended regardless of how well it agrees with you. And include two pieces that you know are strong but oddly written, because that is the exact case where a machine is most likely to be harsh and where the consequence for a student is largest.

Write the result down, even informally. When a colleague or a parent asks how you know the tool is reliable, a page saying it matched your band on fourteen of twenty scripts and disagreed on two in a specific direction is a far better answer than a vendor brochure.

What should you ask a vendor?

Four questions, and the quality of the answers tells you more than the product demo does. What is the agreement statistic with human raters, on what task, and how many pieces of work? Was it measured on a public dataset or on work resembling ours? How stable is the score when the same piece is submitted twice? And is student work retained, and can it be used to train anything?

A vendor that answers all four precisely is worth taking seriously. A vendor that answers the first with an accuracy percentage and no denominator is describing a marketing claim, not a measurement.

What does this change about the workload problem?

It moves the saving rather than removing it. The reason the pile of 140 essays is exhausting is not arithmetic, it is the reading. A tool that produces a first pass of comments changes what the reading is for: instead of generating every observation from scratch, you are confirming, correcting and adding the things a machine cannot see.

That is a genuine change, and it is smaller than the sales pitch. The claim that marking time collapses assumes you stop reading carefully, and if you stop reading carefully then the human on every score commitment above is decoration. The realistic version is that the second half of the pile stops being worse than the first, which anyone who has marked at midnight will recognise as a real benefit even though it does not fit on a slide.

What governance does a teacher actually need?

Four commitments, short enough to put on a single page and give to students.

  1. Disclosure. Tell students when a tool is used, on what, and at which stage. A grade that arrives without this is the one that becomes a complaint.
  2. A human on every score that counts. The machine may draft. A person decides, and that person's name is on the mark.
  3. An appeal route. A stated way to ask for a remark by a human, with a deadline and a named person. Publishing it costs nothing and it is what turns a dispute into a procedure.
  4. A retention rule. Say where student work goes, how long it is kept, and whether it can be used to improve anybody's model. This is the question parents ask second and it should not need research to answer.

Those four sit inside a broader habit of writing down what tools may be used for before somebody improvises an answer under pressure. Ours is published on the MaShop AI policy page, and the format transfers: state the permitted use, the human checkpoint and the escalation route in language a non specialist can read.

Should students be allowed to see the machine's comments directly?

Usually yes, and this is where the evidence points. The meta-analysis found the strongest effects where students were actively working with the material, and feedback a student never reads cannot produce that. Routing comments through the teacher as a filter costs most of the benefit and keeps most of the risk.

The line to hold is between comments and consequences. A student reading detailed comments on a draft is the use case with the best evidence. A student receiving a final grade from the same system is the use case with the worst.

One more workload note. The saving is largest on the tasks where the evidence is also strongest, which is a rare and welcome alignment. Short factual answers and tested code are both fast to automate and safe to automate, and together they are a large share of what actually gets marked in a term.

The practical recommendation

Use ai grading for the comments and keep the number. That is not a compromise position, it is what the evidence separates into: moderate measured benefit on feedback, weak and unstable agreement on scores.

In practice, for that pile of 140 essays: run the tool to produce comments and a suggested band, read the comments as a first pass at speed, mark the score yourself with the rubric open, and pay attention to the essays where you disagree with the suggestion, because those are where the interesting work is. The time saving is real and comes from the reading order, not from delegating the decision.

Practitioner guides tend to arrive at the same place from the other direction. Edutopia's round up of tools that help teachers work more efficiently groups the useful applications into personalised learning, productivity and content creation, and its recurring instruction is to verify the output and check for bias rather than to trust it. That is the correct instinct, and this article is an attempt to say precisely which parts need verifying and why.

Comments 0

0 / 4000Your email stays private.
No comments yet. Be the first.

Keep reading picked for you.

Describe it. MaShop builds it.

Commerce apps and websites from one sentence. No card to start.

Start building