BetaMaShop is in public beta. We improve it continuously, and your feedback shapes what comes next.
MaShop/Blog/Tools/AI Moved the Work From Producing to Checking
ToolsAugust 10, 2026
Read · 5 min
ai for teachers · ai tools for teachers

AI Moved the Work From Producing to Checking

Weekly users report saving 5.9 hours. Here is where those hours come from, task by task, and the review step that has to survive or the saving is not real.

Key takeaways
  • The best sourced figure available: teachers who use AI weekly report saving 5.9 hours a week, from a Gallup survey of 2,232 US public school teachers run in spring 2025.
  • That is the number for weekly users, who were 32 percent of the sample. Forty percent were not using AI at all, so the headline is not a description of the profession.
  • The time does not disappear. It moves from producing to checking, and every hour saved comes with a review step that has to survive or the saving is fake.
  • The review step is different for every task. A reading list needs its citations verified, a differentiated text needs its reading level checked, a family email needs its tone read aloud.
  • Three things belong on a refusal list: a final grade with no human, identifiable student work in a consumer tool, and feedback you would not put your name to.

It is 7:40pm on a Tuesday. There are 26 exit tickets on the table, a differentiated version of tomorrow's reading that does not exist yet, an email to a parent that has been drafted three times, and a district form asking for evidence of intervention for four students. The district gave you an AI licence in September and an hour of training in which somebody demonstrated a lesson plan generator.

The question is not whether the tool can do any of this. It can do all of it, at a speed that is genuinely startling the first time. The question is which parts of that pile it gives back to you and which parts it quietly makes worse by handing you something plausible that you now have to check line by line.

What does the evidence actually say about time saved?

There is one figure worth quoting and it needs its conditions attached. Gallup, working with the Walton Family Foundation, surveyed 2,232 US public K-12 teachers between 18 March and 11 April 2025, recruited through the RAND American Teacher Panel, a nationally representative probability based panel, with a margin of error of 2.5 percentage points. The result, published as Three in 10 teachers use AI weekly, saving six weeks a year, is that weekly users save an average of 5.9 hours a week, which over a 37.4 week school year the report converts into roughly six weeks.

Now the conditions, because they change what the number means. Thirty two percent of teachers used AI at least weekly. Another 28 percent used it monthly or less. Forty percent were not using it at all. The 5.9 hours applies to the first group only, and people who adopt a tool weekly are not a random sample of teachers: they are the ones for whom it worked.

The task breakdown is more useful than the headline. Preparing to teach was the most common use at 37 percent monthly or more, making worksheets or activities at 33 percent, and modifying materials for particular student needs at 28 percent. Between 57 and 74 percent said quality improved across tasks, with fewer than 16 percent reporting a decline.

Read those three lines together and a shape appears that no summary of the survey states. Every one of the top uses is a production task. None of them is a judgement task. That distinction turns out to explain almost everything about where this works.

Diagram showing the five parts of a teaching week that AI touches, planning and differentiation and assessment and family communication and reporting

Planning: the clearest win, with one specific trap

Before. Ninety minutes to build a unit sequence: objectives, a hook, three activities with escalating difficulty, a formative check and a list of texts.

After. Around twenty five minutes. You describe the class, the standard and the constraints you actually have, meaning the length of the period, the equipment in the room and what the group did last week. You get a full sequence back and you rebuild the parts that do not fit.

The review step that has to survive: verify every text, link and citation before it reaches students. This is the failure mode with the highest cost and the lowest visibility, because a fabricated reference looks exactly like a real one. A title that sounds right, an author who exists, a page number, and no such book. If a resource list has six items, six searches is the price of using it. Not one of them is optional, and the moment you skip it once, you will skip it always.

Edutopia's roundup of tools that help teachers work more efficiently makes the same point in its own register: alongside the lesson and quiz generators, its guidance is to verify information and check for bias rather than to trust the output. What the tool changes is who does the drafting, not who is accountable for what ends up in front of a class.

Differentiation: the largest saving and the easiest to get wrong

Before. Rewriting one source text at three reading levels, keeping the content and losing the difficulty. An hour, and honestly, often skipped.

After. Fifteen minutes for all three versions.

The review step: check the reading level actually moved, and that the content survived. Two distinct failures hide here. The first is that a simplified version drifts up in complexity partway through, because the model relaxes back toward its default register after a few paragraphs. Read the last paragraph, not the first, and you will catch it. The second is subtler and worse: simplification quietly removes the hard idea rather than making it accessible. A text about causes of a war that becomes a text about events of a war has been made easier by deleting the thing you were teaching.

If you generate images or diagrams for these materials, there is a specific obligation that comes with them. The W3C's success criterion on non-text content is a Level A requirement, meaning the baseline rather than an aspiration: all non-text content presented to the user needs a text alternative serving an equivalent purpose, with narrow exceptions for controls, decoration and cases where a text version would invalidate a test. A generated diagram with no alt text is not differentiated material. It is material one of your students cannot use.

Assessment: help with feedback, not with the grade

Before. Two and a half hours to give 26 pieces of written work a comment each that says something specific.

After. Around an hour, if the model drafts comments against your rubric and you rewrite each one.

The review step: read every comment as the student who receives it, and rewrite the ones that could belong to anyone. Generic praise is worse than no comment, because it tells a student you did not read their work, and they can tell. The comments that need the most rewriting are the ones for the strongest and the weakest pieces, where the draft will reach for the same three encouragements.

The line worth holding is between feedback and scores. Feedback is a draft a human finishes. A score is a decision with consequences, and the reliability case for handing it over is much weaker than the vendor demonstration suggests. We went through the evidence on where that line sits in why AI grading works for comments and not for scores, and the short version is that agreement with human markers holds up on average and falls apart exactly on the unusual answers that most need a person.

Family communication: fastest to draft, most dangerous to send

Before. Twenty minutes on a difficult email you rewrite four times because the tone keeps landing wrong.

After. Five minutes to a draft that is clear, structured and slightly too formal.

The review step: read it out loud, and cut anything you would not say in the room. Drafted messages tend toward a register that reads as institutional distance, which in a difficult conversation is exactly the wrong signal. The specific things to check are that the child's name appears more than once, that the first sentence is not a complaint, and that you have not accidentally imported a fact about a different student from an earlier draft in the same session.

Two rules that are worth being absolute about. Never paste a student's identifiable work or personal circumstances into a consumer tool that is not covered by your district's agreement. And never send a message about a safeguarding or wellbeing concern that a model drafted, because the wording in those messages is a professional judgement and it may be read back to you months later. What data leaves your device and where it lands is a question worth asking of any tool before the first use rather than after the first incident, which is the reasoning behind how we set out our own data handling.

Administrative reporting: the quiet winner

Before. Forty five minutes converting your own notes into the district's required format for four students.

After. Fifteen minutes, because this is a reformatting task and reformatting is what these tools are genuinely best at.

The review step: check that nothing was added. When a form has a field your notes do not answer, a model will fill it plausibly rather than leave it blank. In an intervention record, a plausible invention is a false statement in a document that may be read in a meeting about a child. Compare field by field against your original notes, and leave blanks blank.

TaskRealistic time savedWhat to check every timeWhat it costs when you do not
Unit and lesson planningRoughly an hour per unitEvery text, link and citation existsA fabricated source handed to a class
Differentiating a textAround 45 minutes per text setReading level held to the end, hard idea intactAn easier text that no longer teaches the concept
Written feedbackRoughly 90 minutes per class setEvery comment is specific to that pieceStudents learn their teacher did not read it
Family emails15 minutes per difficult messageTone read aloud, no detail from another studentA relationship damaged by an institutional voice
Reporting and formsAround 30 minutes per batchNothing added that was not in your notesAn invented statement in a formal record

Add the checking column up and you get the honest version of the headline. The work does not vanish. It changes shape, from generating to verifying, and verifying is less tiring but requires a different and more suspicious kind of attention. Somebody who does the first half and skips the second has not saved five hours. They have moved a risk.

Card listing three uses a teacher should refuse outright, automated final grades and identifiable student work and unsigned feedback

What should a teacher refuse?

Three things, and the reasoning for each is different, which is why a blanket policy tends to be either useless or unfollowable.

A final grade with no human in it. Not because the marking is always wrong, but because a grade is a consequential decision about a person and the accountability for it cannot be delegated to a system that cannot explain itself to a parent. The NIST AI Risk Management Framework, released in January 2023 and explicitly voluntary, organises exactly this kind of thinking around four functions: govern, map, measure and manage. The map function is the one that applies here, and it asks you to establish the context and the stakes before deciding on a control. High stakes, irreversible for the student, low explainability. That is a map result, not a preference.

Identifiable student work in a tool your district has not approved. This one is not a judgement call. It is a question of what agreement covers the data, and the answer is usually written down somewhere you have not read.

Any output you would not sign. The simplest test in the whole subject. If you would be uncomfortable having your name on a comment, an email or a report as its author, do not send it. The tool does not change who wrote it in the way that matters.

Why does the checking cost more for some teachers than others?

Because verification cost scales with how specific your standards are, and that is not distributed evenly across subjects or stages.

A worksheet of arithmetic practice is cheap to check: the answers are either right or they are not, and you can verify thirty of them in two minutes. A set of comprehension questions on a text you know well is more expensive, because a question can be answerable, well phrased and still miss the thing the lesson was about. A source analysis task in history is the most expensive of all, because plausibility and accuracy come apart completely and only somebody who knows the period can tell.

This explains a pattern in the survey data that is otherwise puzzling. Adoption is not uniform, and the tasks with the highest reported use are the ones where checking is fastest. Making worksheets and activities sits at 33 percent partly because a worksheet declares its own errors, while a nuanced piece of feedback does not.

The practical consequence is that the honest estimate of time saved is task specific and cannot be borrowed from a colleague in another department. A maths teacher reporting that this saves four hours a week and an English teacher reporting that it saves none may both be describing their experience accurately.

Will students know, and does it matter?

Sometimes, and yes, but not in the direction people expect. Students are quick to identify generic feedback and quicker to conclude from it that their work went unread, which costs more than the time the comment saved. They are far less able to identify a well edited draft, and there is no reason they should: a teacher who used a tool to produce a first version of a worksheet and then reworked it has done what teachers have always done with textbooks and shared resources.

The detection question has its own literature and it is more unsettled than the marketing suggests, as we set out in what has changed in AI content detection. For classroom purposes the useful conclusion is narrow: detection tools should inform a conversation and never conclude one, because a false positive on a student's own writing is a serious harm and the tools do produce them.

How to get better output without a course

Most disappointing results come from prompts that describe the artefact rather than the situation. Compare asking for a lesson plan on fractions with asking for a 45 minute lesson for a mixed attainment year 5 group who can add fractions with the same denominator but not different ones, with no tablets in the room, ending in a five minute check you can mark at a glance.

The second gets something usable because it contains the constraints you would otherwise have to fix by hand. The pattern generalises: give the context, the constraint and the finished form you need, and you will spend your time editing rather than re specifying. Five patterns that carry across most drafting tasks, each with a before and an after, are set out in the prompt patterns that reliably change the output.

Where to start on Monday

Pick one task, not five. Choose the one that is highest volume and lowest stakes, which for most teachers is worksheets or differentiation rather than feedback or family communication. Run it for two weeks, and each time write down two things: the minutes it took including your review, and what you had to fix.

After two weeks the second list is the valuable one. It tells you what this particular tool gets wrong on your particular subject, and that is the review checklist you keep. It is also the honest test of whether you are in the 32 percent for whom this works or the 40 percent who tried it and found the checking cost more than the drafting saved. Both answers are legitimate, and only one of them appears in the headline.

Comments 0

0 / 4000Your email stays private.
No comments yet. Be the first.

Keep reading picked for you.

Describe it. MaShop builds it.

Commerce apps and websites from one sentence. No card to start.

Start building