The UK government is trialing AI-generated writing samples to standardize literacy assessments, aiming to reduce costs despite concerns from experts regarding ethics and potential bias.

Key facts
- •The DfE claims the use of AI-generated writing samples will cut annual assessment costs by 95 percent.
- •The current process of using real pupil work for moderation costs approximately £100,000 per year.
- •The DfE is using ChatGPT’s GPT-5 model to generate three collections of writing for standardisation exercises.
- •A panel of 20 local authority moderation managers is tasked with reviewing the AI-generated material for authenticity.
- •The DfE plans to make a final decision on whether to continue the program in spring 2027.
The UK Department for Education (DfE) has begun using AI to produce writing samples for assessing literacy standards in pupils transitioning from primary to secondary school. The department claims the initiative will reduce annual costs by 95 percent. Currently, about 2,000 moderators review pupil work at the end of key stage 2 (KS2) to ensure grading consistency, a process that typically costs £100,000 annually.
By the numbers
Implementation and Review Process
For the current and upcoming year, some moderators will evaluate AI-generated output created by ChatGPT’s GPT-5 model instead of samples from actual schoolchildren. The DfE is using the model to create three collections of writing for one standardisation exercise. A group of 20 experienced local authority moderation managers will review the AI-derived material for authenticity before it is used. The department plans to decide in spring 2027 whether to continue this method or return to procuring samples from external suppliers.
Expert Concerns and Potential Risks
Researchers have raised ethical and philosophical concerns regarding the use of synthetic writing samples. Rebecca Clarkson of Anglia Ruskin University noted that the practice may shift perspectives on what constitutes a good standard of writing. There are also concerns that AI-generated text, which can be flat or bland, may fail to reflect the vocabulary and sentence structures of neurodivergent or non-native English-speaking students. The DfE’s own risk assessment acknowledged that large language models often exclude atypical language, though the department plans to mitigate this through thorough review and editing of the outputs.
Impact on Educational Standards
While the KS2 assessments occur after pupils have been assigned to secondary schools, the alignment of children’s work with synthetic samples could influence the level of support they receive upon entry. Jo-Anne Baird of the University of Oxford warned that AI-generated materials could become the benchmark that teachers encourage pupils to emulate, potentially leading to unforeseen consequences in educational standards. The DfE has not conducted a formal impact assessment for the system.
Advertisement
This article was independently rewritten by ManyPress editorial AI from reporting originally published by New Scientist.


