UK pilot swaps GPT 5 essays into the writing samples ~2,000 moderators use to standardise primary literacy grades. No impact assessment yet; spring 2027 decision.
The Department for Education is running a pilot in which GPT-5 produces the sample writing that about 2,000 moderators use to check whether teachers across England are grading pupils' work to the same standard at the end of primary school. The department says the swap cuts a moderation budget of roughly £100,000 (about $125,000 at recent exchange rates) by approximately 95 per cent. It is a trial, not a national rollout, but it reaches the documents that anchor how every primary teacher's writing grade gets compared against an agreed-upon reference point.
For 2025-26, GPT-5 is generating three collections of writing for one standardisation exercise, according to New Scientist's reporting on the policy. The two exercises that follow, in 2026-27 and 2027-28, will each still include scripts written by children. The change is incremental in volume but absolute in direction: for the first time, the benchmark texts that moderators calibrate against are not written by the population the test is designed for.
The DfE says AI replaces an external supplier's annual contract. The "95 per cent cost saving" is a department-supplied figure, and the Department for Education did not respond to New Scientist's request for comment. A search of the DfE's published transparency data lists no formal impact assessment for the system, a point the department has not publicly addressed.
About 20 experienced local authority moderation managers will review the AI-derived material for authenticity before it is used. They are the human check. The process guards against obvious tells, but it does not answer the harder question Rebecca Clarkson, who researches KS2 writing assessment and moderation at Anglia Ruskin University, raises in the same report. Clarkson argues that when the exemplars are no longer written by real children, the system is no longer comparing teachers to a child-authored standard. It is comparing them to a machine-authored one. That, she says, creates a "philosophical and ethical issue" and may shift what gets accepted as a standard piece of writing.
The standardisation exercise is the part of KS2 that sits behind the scenes. Roughly 2,000 moderators cross-check pupils' work across local authorities to make sure a teacher in one part of the country is not grading more harshly or leniently than a teacher in another. The samples they use, sometimes called exemplars, function like a reference scale: a piece of writing that meets the expected standard, a piece that exceeds it, a piece that falls short. Teachers are not told which of the three to imitate, but their grading is calibrated against the boundary the exemplars draw.
If the exemplars drift toward cleaner prose, more uniform structure, fewer of the messy hedges and run-on clauses that real eleven-year-olds produce, moderators would still be checking teachers against each other. The teachers would just be checking their students against a slightly different yardstick than the one that anchored the system before.
Clarkson's warning is not that AI will misread a child's handwriting or misjudge a paragraph. It is that the act of standardising moderation, by definition, assumes the reference is fixed. Replacing a child's writing with a model's may keep inter-teacher consistency intact while shifting the bar that determines what "meeting the expected standard" looks like. That second-order effect, on what gets accepted as competent writing at age eleven, is the part no published impact assessment has yet tested.
The pilot's calendar is short and public. The 2025-26 exercise is underway. A full mixed exercise runs in 2026-27. The department has set a decision window in spring 2027 on whether to keep producing the samples this way or return to procuring them from an external supplier.
For now, the trade is precise. About £95,000 of the original ~£100,000 annual moderation procurement budget is saved by the department's own arithmetic — a roughly 95 per cent reduction — leaving roughly £5,000 of the original budget unspent. In return, the reference texts that anchor how every Year 6 writing grade in England gets compared are partly written by the same kind of system that many of those Year 6 pupils have been told, by the same department, to treat with care. The spring 2027 decision is the public accountability window. No impact assessment, no public consultation document, and no DfE response to the original report have been published in the meantime.