AI in Education #1: Five Things I Learned from Daisy Christodoulou About AI and Assessment
This is the first in a series of AI in Education specials, where I speak to the world’s leading experts about how AI is changing teaching and learning.
We’re recruiting UK state-funded secondary schools with KS3 to take part in a major national RCT launching in September 2026, focused on closing the maths attainment gap. Learn more by clicking here, and submit your interest by clicking here.
A prediction that aged badly
In early November 2022, Daisy Christodoulou and her team at No More Marking sat down for a big strategy meeting. They spent the week looking at the state of AI, particularly handwriting recognition, and concluded it was still about 10 years away from being useful for their needs.
Three weeks later, ChatGPT launched.
As Daisy told me on the latest episode of my podcast: “We were like, you know, it’s still not there. There’s nothing there that’s really going to help us.”
Three years on, AI is helping them a great deal. But the journey from scepticism to cautious optimism has been fascinating, and full of lessons for anyone working in education. It was an absolute pleasure to have Daisy back on the show, and I wanted to share five practical takeaways from our conversation that I think every classroom teacher and headteacher will find interesting.
1. Good-looking stats can hide a terrible model
Here is something that blew my mind. The first AI essay marker was built in 1968 - a time when the England football team were World Champions! And on the surface, it looked impressive: it was consistent (unlike human markers), and it agreed well with human scores overall.
So why aren’t we all using it?
Because when you looked under the hood, it was essentially marking on the length of the essay. There is a correlation between length and quality, so the stats looked great. But as Daisy put it, a later study found you could write the same paragraph 37 times and get the top mark.
The lesson here goes way beyond essay marking. Whenever you see a headline claiming an AI tool is “better than humans,” ask yourself: what is it actually doing? Is it picking up on something meaningful, or has it found a shortcut that will fall apart the moment someone games it?
This is why Daisy insists on a qualitative review of the disagreements between humans and AI, not just the headline percentage agreement. As she said: “I would rather a model that had 80% agreement and wasn’t marking on length than one that had 95% and was marking on length, because that 95% one would degrade rapidly the minute people work out what it’s doing.”
2. AI is better at comparative judgement than absolute judgement
This one surprised me. No More Marking started by asking AI to do what everyone else in this space does: read an essay, consult a rubric, and give it a mark. This is absolute judgement, and the results were underwhelming.
Then they tried something different. They got the AI to do comparative judgement instead – looking at two pieces of writing and deciding which one is better. This is the same approach humans find easier, and the same approach used to train large language models in the first place.
The result? The AI makes fewer errors when judging comparatively. And here is the really remarkable bit: when they analysed the big disagreements between humans and AI across 400,000 pieces of student writing, Daisy said: “So far, I genuinely would say all of them are the result of human error. Not AI error.”
I did not see that twist coming.
A lot of the human errors turned out to be handwriting bias – teachers marking students down for messy handwriting. The AI, working from transcriptions, wasn’t affected by that.
3. Don’t take the human out of the loop
Given that the AI was outperforming humans in some areas, I asked Daisy the obvious question: Can we just remove the human?
Her answer was a firm “not yet, and maybe not ever.” She gave three reasons that I think apply far beyond assessment:
First, we don’t fully understand how these models work. Even their designers can’t explain why they make the decisions they do. Until we do, we need humans catching things we can’t predict.
Second, if students know a human is in the loop, they are far less likely to try and game the system. Remove the human, and even if the AI isn’t easy to game, students will try. And the act of trying is itself a problem.
Third – and this is the one that really made me think – teachers need to engage with student work to develop as professionals. Daisy drew a parallel with the legal profession: “Maybe a lot of the work of junior professionals can be done very well by AI. But that grunt work is what turns you into the senior professional.”
No More Marking’s model uses 10% human judgements and 90% AI. That 10% is enough for every piece of writing to be seen by a human twice, keeps teachers engaged with the work, and provides the oversight needed to catch problems. I think that ratio is worth thinking about for any AI tool we use in schools.
4. Personalised AI tutors still have three big problems to solve
We talked about AI tutors, and Daisy was generous about the research we did with Google DeepMind at Eedi. But she was also clear-eyed about the challenges that remain. She boiled it down to three:
Explanations are not enough. There is a naïve assumption that learning is about reading an explanation. As Daisy pointed out, we’ve had scalable, reproducible transmission of knowledge since the printing press. If explanations were all that were needed, everyone would understand thermodynamics. AI tutors need to be built around questions and practice, not just telling students things.
Hallucinations are not at zero. Our study got the error rate down to 0.14%. Daisy was complimentary about that, but then did some maths that made me wince: if every student in Britain used the system for two lessons a week, that would be around a million errors. She referenced Andrej Karpathy’s concept of “the march of the nines” – getting from 90% accuracy to 99% takes enormous effort, and each further nine takes the same effort again. We are not at 100% accuracy, and we may never be.
Are we asking kids to sit in front of screens all day? Good personalised tuition programs have existed for decades. They tend to get high scores in the lab but never scale to 8 million kids in 20,000 schools. Daisy’s instinct is that classrooms – especially in primary – should be relatively screen-free, with the technology doing its heavy lifting behind the scenes.
5. Always start with the problem you are trying to solve
When I asked Daisy what question I should have asked her, she didn’t hesitate: “What is the problem you are trying to solve?”
She described reading about a school where students used AI to generate an image of what Luton would look like if it were a car. And she just thought: why? What problem was that solving? Was there a world before AI where everyone was desperate for that image?
It is so easy – and I include myself in this – to get dazzled by what AI can do and forget to ask whether it should. Daisy’s north star at No More Marking has stayed constant from day one: improve marking reliability, improve marking efficiency, improve the impact on teaching and learning. Every technology decision gets tested against those three things.
That feels like a good test for any school leader considering an AI tool. Not “what can this do?” but “what problem does this solve for our students?”
Bonus: The best analogy I’ve heard for technology in schools
Finally, I asked Daisy to make a prediction about the future of schools. Her answer was a gem. She said: “Germany has no speed limits on its motorways and car-free medieval town centres. And that is my analogy for technology in schools.”
She wants a high-tech, high-speed infrastructure around the school – data whizzing around, AI solving the admin problems, teachers no longer wrestling with spreadsheets. And then classrooms that are quite old-fashioned, especially in primary. Handwriting on paper. Books. A teacher at the front. Protected spaces where screens don’t intrude.
I love that image. The autobahn and the medieval town centre. Technology serving the school, not dominating the classroom.
Over to you
This was a brilliant conversation, and I’ve only scratched the surface here. You can listen to the full episode here, check out the No More Marking Substack for all of Daisy’s latest research, and if you want to try their AI comparative judgement tool, you can get a free trial here.
This is the first in a series of conversations I’m having with leading thinkers about AI in education. I’d love to know: Who do you want me to speak to? What questions do you want me to ask? And how is your school navigating all of this? Let me know, and thanks for reading.
Craig



<< Third – and this is the one that really made me think – teachers need to engage with student work to develop as professionals. Daisy drew a parallel with the legal profession: “Maybe a lot of the work of junior professionals can be done very well by AI. But that grunt work is what turns you into the senior professional.” >>
So true, in a comment to someone's post on LinkedIn a few months back, I was suggesting the role of "AI Arbiter" to add that touch of 'humanity' to the process.
On a different, but related note, I'm trying an experiment at the moment with Claude to analyse revision areas based on the marks I gave them in a recent GCSE mock. This is anonymous, just based on the scores themselves being input, the question paper and the mark scheme. The output is useful and accurate at this point as it's effectively cross-referencing skills from the curriculum based on the mark scheme.
Next step is to extend this to a per student solution (whilst remaining GDPR compliant!), to help with a more personalised list of topics. This will be interesting as I will then see if what Claude comes up with matches my view of the students performance in class.