How Should Schools Measure Whether AI Saves Teachers Time?
Find out whether AI is saving your team time once checking, rework and support are included.
In this playbook9 parts
- Prepare for the session
- 01 Agree what you are comparing
- 02 Find out how long the usual work takes
- 03 Keep the cost of getting started visible
- 04 Follow the AI-assisted task all the way through
- 05 Check whether the finished work is good enough
- 06 Add up the team's time, not just drafting time
- 07 Decide what happens next
- Completion check
Imagine a teacher who can now draft a lesson pack in half the time. That sounds promising. But the pack still needs checking and adapting, a colleague has to review it, and someone has to help when the tool gets things wrong. How much time has the team actually saved?
The useful comparison runs from starting the task to having something good enough to use. It includes everyone who helped, not just the person who generated the first draft.
This playbook gives you a way to make that comparison for one workflow at your school. It is a proposed local method, not a validated measurement instrument or a forecast of your results.
Agree what you are comparing
Start with what finished means. A generated worksheet and a checked, adapted lesson pack are different outputs. Agree what each pack must contain, which learner needs it must meet and how you will judge its quality.
Then describe what teachers actually do now. They may adapt last year's lesson, draw on department resources or already use AI. Compare AI-assisted work with that normal approach, rather than assuming every lesson starts from scratch. If you are changing an existing AI workflow, say so.
Choose tasks that are similar enough to compare. Note the phase, subject, curriculum, language, available resources and teacher experience, and agree which differences would rule a task out. There is no universal minimum number of tasks that makes a local result reliable.
Add these details to the first row of your record. The provider-question playbook can help if you are still unsure what a provider's workload claim covers.
Find out how long the usual work takes
Log a representative set of tasks before the change, including the awkward ones and attempts that do not produce a usable result. Separate drafting, checking, correction and adaptation, finalising and handing over, peer review and support. Use the same categories later.
Count each person's active working time. If two colleagues spend ten minutes reviewing a pack together, that is twenty person-minutes. Record waiting separately: it can delay the job, but someone doing other work while they wait is not necessarily spending that time on this task.
Give each stretch of active work one category so you do not count it twice. OECD's TALIS workload analysis explains why this matters: separately reported task hours can add up to more than a teacher's reported total. Count the creation of a shared template once, then each teacher's adaptation of it.
Note estimates and gaps rather than smoothing them away. Keep difficult cases in the record, and count both the outputs attempted and those finally accepted.
Keep the cost of getting started visible
Training, arranging access, integration and preparing shared templates all take staff time. Include the champion or technical colleague who helps behind the scenes.
Keep those costs separate from the work that repeats each time. You need to see both what the workflow takes now and what it cost to get there. Log the time spent measuring it too, from collecting diaries to checking records and discussing results, in both periods.
Describe the learning period your team actually used. Do not assume everyone became proficient on a particular date. If staff still need regular support, count it as part of the ongoing workload, not as a one-off setup cost.
EEF's implementation guidance distinguishes preparation, delivery and decisions about sustaining an approach. That is helpful context, although it does not prescribe this record or its calculations. Add these costs to the record before comparing the overall time spent.
Follow the AI-assisted task all the way through
Now log comparable tasks with AI, using the same categories. Include preparing inputs, trying and revising prompts, checking, correcting, adapting and handing over the final pack. Failed drafts and repeated attempts still took time.
Note changes to the tool or its version, the difficulty of the task, available resources and the number of outputs. Producing more packs is a change in volume; raw totals cannot tell you whether equivalent work became quicker.
Two studies show why it matters to be clear about what you are measuring.
A 2024 EEF trial involved 259 Year 7 and 8 science teachers in 68 English state secondary schools. Schools were randomly allocated to ChatGPT with a guide or no generative AI. Reported preparation time was 56.2 versus 81.5 minutes per week, a 31% reduction. Teachers kept weekly diaries during weeks 6 to 10, after a five-week learning period. That finding concerns a supported preparation task. It is not a measure of total weekly workload or a savings target for your school.
A 2025 Gallup/Walton survey, by contrast, asked 2,232 US public K-12 teachers about their experience. Weekly AI users estimated 5.9 hours saved per week across different tasks and tools. This was a retrospective, weighted survey, not a timed experiment. The two findings answer different questions and should not be used as interchangeable benchmarks.
Check whether the finished work is good enough
Apply the quality criteria you agreed at the start. Where possible, ask the quality lead to review the final outputs without knowing which method produced them. Note whether each first draft was accepted, needed revision or was rejected, as well as the final outcome and any missing review.
Put correction time in the work log. If a pack takes several rounds of rework before it passes, those rounds are part of the job. Do not drop essential checks to improve the timing.
For lesson materials, the keep, revise or reject review sets out the checks. An unresolved critical problem with accuracy, curriculum purpose, safety or accessibility means pausing that use, however quickly the draft arrived.
The NFER evaluation used a blinded panel to review 30 sets of final resources and found no evidence of a quality difference. That does not establish equivalent quality in every setting or improved pupil learning. If a pack hasn't been reviewed, don't record it as having passed.
Add up the team's time, not just drafting time
Before comparing totals, check that the task counts and conditions match, and flag missing entries. Add the work minutes for each role and for the whole team.
Subtract the after total from the baseline to find how much time was saved on the same amount of work. A negative answer means the work took longer.
To express the saving as a percentage, divide the difference by the baseline and multiply by 100. Only do this if you know the baseline and it is greater than zero. You can also show a median or range if your task-level records support it.
A fictional comparison
Every figure below is invented to show how the calculation works. This is not a school result or a benchmark.
F-01 compares four batches of five equivalent lesson-material packs. Each period attempts and finally accepts 20 packs. Four packs in the after period need revision, and their correction time is included. The example assumes complete records and no unresolved critical quality issue.
| Recurring person-minutes for four batches | Baseline | After | Baseline minus after |
|---|---|---|---|
| Teacher preparation/drafting | 120 | 32 | 88 |
| Teacher verification | 40 | 48 | -8 |
| Teacher correction/adaptation | 20 | 32 | -12 |
| Teacher finalisation/handoff | 0 | 8 | -8 |
| Peer review | 20 | 28 | -8 |
| Technical/champion support | 10 | 20 | -10 |
| Total recurring effort | 210 | 168 | 42 |
Drafting is 88 minutes quicker. Once the rest of the work is included, the team saves 42 minutes: 42 divided by 210, or 20%.
The split matters. Teacher effort falls from 180 to 120 minutes, while colleagues in other roles spend 48 minutes rather than 30. Some work has moved to them. Both changes belong in the result.
Now add the cost of getting started and doing the comparison.
| First-period cost, person-minutes | Baseline | After |
|---|---|---|
| Recurring delivery | 210 | 168 |
| Setup/training | 0 | 120 |
| Measurement/evaluation | 12 | 16 |
| Total first-period effort | 222 | 304 |
The after period costs 82 minutes more once setup and measurement are included. So the example has a lower repeat-use workload and a higher first-period cost. Both are useful to know.
There are no task-level figures or baseline first-pass decisions here, so the example cannot supply a median, range or baseline first-pass acceptance rate.
In this example, the team keeps using AI for the same task, then checks whether the saving lasts without adding to the workload of colleagues who provide support. They do not extend its use. Your school's decision would depend on its own quality results, costs and working conditions.
Decide what happens next
Choose the decision your evidence supports:
- Continue bounded use if quality is acceptable and the effort across all roles makes this limited use worthwhile.
- Revise if a workflow, training or support problem can be addressed, then compare again.
- Stop if an unresolved critical failure or unacceptable burden makes this use unsuitable.
- Record insufficient evidence if gaps or mismatched tasks prevent a sound judgement.
Write down why, what needs to happen, who will do it and when you will return to the decision.
A small before-and-after comparison can tell you what happened locally. It cannot, on its own, tell you how much of the difference AI caused. Experience, topic, time of year, reused resources and tool changes may all play a part. Matched tasks or a comparison running at the same time can help you interpret the result, without automatically establishing cause and effect.
If time has been freed up, note how it was actually used. If you do not know, write Unavailable. Spending less time on this task is different from working fewer hours or using the time for other work. It doesn't show a benefit for pupils, savings in staffing costs or extra capacity across the school over a year. The number of times staff used the tool doesn't tell you how many minutes were saved, either.
Use this comparison to decide whether to keep using AI for the task you tested. Before buying a tool or using it more widely, carry out a separate review.
If you'd like to discuss a pilot with TopSchool, get in touch. You don't need to share your school's working records to start that conversation.