AI and Assessment in Schools: What Counts as Evidence of Learning Now?
Generative AI changes what a finished assessment can prove. Explore how schools can use output, process and explanation to keep evidence of learning credible.
On this page9 sections
A student submits a thoughtful essay, a polished project or a convincing solution. The teacher knows that generative AI may have played a part in producing it.
The immediate questionDid the student use AI?
The educational questionWhat does this work tell us about what the student understands and can do?
Generative AI has not removed the purpose of assessment. Schools still need reliable ways to understand what students know, where they are struggling and what they are ready to do next.
What AI has changed is the evidence. When a tool can help plan, explain, draft, revise or solve parts of a task, a polished final submission may no longer tell the whole story.
AI has changed the evidence problem, not the purpose of assessment
Assessment has always involved inference. A teacher looks at an essay, calculation, presentation or project and uses it as evidence of a studentâs knowledge, reasoning or skill.
An UNESCO IdeasLAB article on the future of assessment argues that increasingly sophisticated AI-generated outputs create a reason to reconsider whether the finished product alone provides enough evidence of individual effort, reasoning or understanding.
That does not make essays, projects or coursework worthless. It means schools need to be more precise about what those activities are intended to measure.
AuthorshipWho produced the work?
UnderstandingWhat does the student understand?
Authorship matters where independent work is required. But establishing authorship does not automatically establish understanding, just as producing every word independently does not necessarily prove deep understanding. Good assessment is clear about which of those things it needs to establish.
Start with what the assessment is supposed to measure
Before deciding whether AI should be allowed, restricted or prohibited, schools need to return to the learning objective. A single school-wide rule such as âAI is allowedâ or âAI is bannedâ cannot reflect every assessment purpose.
| If we want evidence of... | The important AI question becomes... |
|---|---|
| Independent writing | Did AI perform the writing being assessed? |
| Factual recall | Did external assistance replace the recall required? |
| Conceptual understanding | Can the learner explain or apply the concept? |
| Critical evaluation | Can the learner judge the quality of the response? |
| Research judgement | Can the learner verify, select and justify information? |
| AI literacy | Can the learner use AI critically and responsibly? |
The OECDâs guidance on effective generative AI use in education makes a related point: educational value depends on how AI use relates to the learning goal, not simply whether the technology is present.
Curriculum, learning expectations and school policy all matter. It is one reason school-aware AI needs to fit existing curriculum and assessment expectations, rather than operating as a separate layer of classroom practice.
Not every assessment needs the same AI rule
The availability of AI does not mean every assessment should treat it in the same way. Schools need to distinguish three conditions.
Independent performance
Where the objective is a student's own writing, recall, calculation or subject knowledge, unrestricted AI support may obscure the capability being assessed.
Independent assessment therefore retains a clear purpose, particularly in regulated work where candidates must demonstrate their own knowledge, skills and understanding.
Limited AI assistance
Defined assistance may support early research, planning or feedback when the course and assessment conditions allow it.
Students still need to know what is acceptable, what must be acknowledged and what must remain their own. Cambridge International's coursework guidance provides an example of those explicit boundaries.
Deliberate AI use
AI can form part of the task when students are being assessed on judgement, verification or responsible use.
The output is not the evidence of learning. The student's judgement is. The International Baccalaureate's guidance similarly emphasises critical review and acknowledgement of AI-generated material.
The learning objective should determine the AI conditions, rather than the technology determining the assessment.
Build a stronger evidence picture
The finished artefact still matters, but it may no longer carry the full evidential burden when AI can contribute substantially to the same output. Disclosure helps, but it does not show which ideas a student understood, which recommendations they evaluated or whether they can reproduce the reasoning.
One useful way for schools to think about assessment evidence is across three layers. This is TopSchoolâs synthesis of the available evidence, not a model attributed to one external authority.
Output
What did the student produce?The final essay, project, solution, analysis or presentation remains valuable evidence, but may not always be sufficient on its own.
Process
How did the learner get there?Drafts, revisions, working notes or source choices can make enough thinking visible without creating a surveillance record of every interaction.
Explanation
Can the learner demonstrate understanding?A short explanation, a transfer question or recognition of an error can reveal whether the learner understands the work.
Different assessments need different combinations of evidence. The aim is not another mandatory framework. It is to make sure the evidence collected matches what the school wants to know.
Match AI conditions to the learning purpose
Schools need both independent assessment and thoughtfully designed AI-enabled assessment. Neither condition should become the default for every learning outcome.
Some assessment should remain independent of AI
Students may need to show that they can write coherently, recall essential knowledge, calculate, interpret evidence or reason through a problem without AI doing the core cognitive work.
That does not require schools to reject AI. It requires them to identify where independence is the capability being measured and make it visible in the assessment conditions.
Some assessment can deliberately include AI
Students can critique an AI response, identify errors, verify a claim, compare approaches or explain why they rejected a recommendation.
When the objective is to assess critical AI use, the system should be visible in the task and the student's judgement should remain visible in the evidence.
Avoid the wrong operational response
AI detection is not an assessment strategy
A detection system asksMight AI have contributed?
An educator needs to askWhat has the student demonstrated?
Current JCQ guidance on AI use in assessments recognises that detection tools vary in accuracy and should be considered alongside other evidence, including a teacher's knowledge of the student's normal work.
Authenticity still matters. But identifying inappropriate AI use and collecting credible evidence of learning are related, separate responsibilities. Better assessment design cannot be replaced by a detection score.
Do not solve the AI problem by creating a teacher workload problem
Requiring drafts, AI-use records, reflections and individual verification for every task could make assessment unmanageable.
What is the minimum additional evidence needed to make a sound judgement?
The answer should reflect the learning objective, assessment stakes and risk that AI could obscure what the learner can do. The same proportional approach matters when designing an AI pilot without overloading teachers.
Seven questions leadership teams should answer about AI and assessment
Assessment practices are often decided at subject or classroom level, but generative AI creates questions that school leadership cannot leave entirely to individual teachers.
- Which learning outcomes must students demonstrate independently?
Schools should know where independence is essential.
- Where can AI support learning without compromising those outcomes?
Guidance needs to go beyond a simple permitted or prohibited model.
- What AI use needs to be disclosed?
Expectations should be understandable for students and workable for teachers.
- Where is the final submission no longer enough evidence?
These are the places where process or explanation may add value.
- What additional evidence is proportionate?
The answer should reflect the stakes and the workload involved.
- Are expectations sufficiently consistent?
Unmanaged variation across subjects, departments or campuses can create conflicting assumptions.
- Do teachers have practical guidance for applying them?
A policy is not an implementation model if teachers must make every difficult judgement alone.
These questions sit alongside the wider governance, pedagogy and readiness issues that school leaders need to consider before scaling AI.
The question is what the assessment can still tell us
Generative AI has made it easier to produce work that looks capable. That makes it more important, not less important, for schools to understand what their assessments reveal.
Sometimes the right response is to protect independent performance. Sometimes it is to make more of the student's process visible. Sometimes AI can form part of the task because the skill being measured is the learner's ability to question, verify and use it responsibly.
Clearer answers can protect assessment quality without asking teachers to become AI investigators or treating every use of generative AI as the same educational problem.