Assessment after AI: the question schools can no longer defer
The tools that promised to police AI use are proving unreliable. That is uncomfortable—but it reopens a better question: what kind of assessment is actually worth setting?
For two years, the reassurance offered to worried teachers was simple: if pupils use AI to write their essays, a tool will catch it. That reassurance is falling apart. Detectors are probabilistic, and even a modest error rate becomes unacceptable when the result could be an accusation of misconduct.
The change is not simply technical. More leaders are deciding that a detector score cannot carry a disciplinary decision on its own. Once that is accepted, the ground shifts: the real question is no longer how to catch invisible tool use, but what kind of assessment is worth setting when a machine can complete the task unseen.
That is not a loss of control. Detection was doing the policing—imperfectly—while allowing assessment design to escape scrutiny. Removing that false certainty creates room for better evidence of learning: annotated drafts, short spoken defences, supervised checkpoints and honest conversations about how a pupil arrived at an answer.
The thread through this Brief is what can replace detection without creating a marking mountain. Begin with the evidence, pilot the ten-minute protocol, then use the policy checklist to make the change defensible.
The rest · curated items
4 itemsThe Library
Where the evidence on AI-writing detection actually stands
Independent evaluations find substantial reliability problems, while research has also identified a particular risk of misclassifying writing by non-native English speakers.
Why it matters for schools: It gives a school a defensible basis for treating detector output as a prompt for review—not proof of misconduct.
The Lab
A spoken-defence protocol you can run in ten minutes
Ask a pupil to explain one decision, challenge one claim and connect the work to something taught in class. Record a simple judgement: secure, partial or not yet evidenced.
Why it matters for schools: It turns a vague call for “more oral assessment” into a small pilot a department can run next week.
The Library
The clauses an AI-use policy needs once detection is off the table
State what evidence may trigger a review, what other evidence must be considered, how pupils can explain their process and how they can contest an automated judgement.
Why it matters for schools: Clear review and appeal language matters more than a list of tools when a decision is questioned by a pupil or family.
The Journal
What we lose when we stop policing prose
The useful question is whether a pupil is thinking with AI or using it to escape the work of thinking. That distinction belongs in pedagogy, not in a detector score.
Why it matters for schools: It offers constructive language for the staff conversation that follows when old certainties are removed.
From The Lab · use it tomorrow
The ten-minute spoken defence
- 01Which decision in this work was most important, and why?
- 02Which claim would you defend if I challenged it?
- 03Show me where this connects to something we studied in class.
Record only whether understanding is secure, partial or not yet evidenced. The conversation is corroborating evidence—not a second examination.