Solving GCSE mocks with Comparative Judgement
Double Marker - our newest project
Over the last ten years, we’ve worked with about 2,000 schools to assess approximately 3 million pieces of writing using Comparative Judgement.
Until 2025, we used human judges only. Since 2025, we’ve added the option of AI judges, which reduces the time it takes to judge and provides detailed feedback.
Whether you use human or AI judges, this approach works extremely well, and we have a lot of dedicated fans! In our ideal world, we’d love to see it adopted at a national system level.
However, that is unlikely to happen, and as long as the national system is different, we face a tension. Comparative Judgement relies on holistic judgements and does not require a mark scheme. Various studies - by us and others - show that it is more accurate than marking with a mark scheme, but in many circumstances that is not enough reassurance. In particular, when schools are marking GCSE mocks, they understandably want their teachers and students to engage with the detailed mark schemes produced by the exam boards.
We’ve devised a solution to this problem called “Double Marker.” The first marker is AI Comparative Judgement, which provides every script with a highly reliable mark. The second marker is the human teacher who uses the approved mark scheme on a carefully selected sample of questions. Those human marks are used to fine tune the AI Comparative Judgement marks and adjust them up and down as necessary.
Double Marker is very different from our standard Comparative Judgement assessments. It is optimised for GCSE mock marking, and is designed to resolve a number of trade-offs that we have run into over the last ten years or so. We know the mark scheme is important, even though it doesn’t always provide consistency! We know that marking English mocks is extraordinarily time consuming, but we also know that we don’t want to eliminate the time teachers spend reading student work. We know that AI can help, but we also know it needs carefully designed oversight.
In the summer term, we piloted Double Marker with 14 secondary schools on their Year 10 GCSE English Language and Literature mock examinations.
You can read the rest of the post for more details about how the system works, but the conclusion is that the pilots were very successful, and as a result we will run a larger project (that is still free!) in October-November 2026.
We still have a few spaces left – if you would like to learn more and take part, book a call with us next week.
Now read on for more detail about how Double Marker works.
How Double Marker Works
The logistics of barcodes and scanning
GCSE mock examinations in England are still mostly paper-based, with answers hand-written into paper based booklets. To replicate this experience, we asked the teachers to upload their normal booklets so we could add a name and a bar code onto every page. The teachers downloaded the modified booklets and printed them for completion by the pupils. After the examination, the teachers batch-scanned the booklets and uploaded them to us for processing. During the trial, one school told us that their reprographics department was able to manage this printing and scanning process within a couple of hours for a cohort of around 240 pupils.
Moderated judging
After checking all the scans at our end and ensuring there were no unexpected absentees we sent all the open-ended questions to our Comparative Judgement site for the first round of marking (judging). First, we judged each school independently so they were placed on their own scale. At this point, for a 40 mark question, every school would receive marks spread from 0 to 40.
Second, we took a 20% sample from every school and loaded those scripts into a separate moderation pot where they are directly judged against each other. The moderation sample acts as an anchor which allows us to place every school on a unified scale so their scores can be directly compared. At this point, the marks in one school may run from 4 to 39, while another school with relatively weaker performance may see their marks go from 0 to 36.
Making the scale meaningful
While the marks are now comparable between schools, they may still not reflect the mark scheme. By definition, the very best pupil in the cohort will receive the top mark, which is 40 in our example. What if the very best pupil does not deserve 40 marks according to the mark scheme?
It is at this point that we can introduce the lens of the national assessment system. Sampling from across their own pupils we asked schools to mark the work using the mark scheme they would normally use. To facilitate their marking, we gave the teachers an online marking platform where they view and mark their allocated responses, blind to the mark already given by the Comparative Judgement process.
Shifting our scale onto the mark scheme
Once we had collected the marks from the teachers, we were able to use them to ensure two things: first, that the AI wasn't universally too harsh or too generous compared to the teachers (bias); second, that individual AI scores closely matched the human marks (distance). We developed an optimisation routine that optimised against bias and distance and delivered a set of AI derived scores in line with the mark scheme.
Where individual teachers were wildly different from the AI scores, we were able to flag these marks for review by the senior marker. In cases of a discrepancy, the senior marker always has the final authority to override AI marks, assisted by our flagging system that highlights potential anomalies.
Returning the marks to pupils
Timelines were tight, but within two weeks of the examination date we had returned a full set of marks to teachers alongside comprehensive AI feedback for pupils and teachers separately at question level. The feedback was at pupil level, class level, school level and at cohort level we were able to deliver examiner reports.
Evaluation of the pilot
Overall, we were delighted with the way in which we were able to align the scales between Comparative Judgement and marking. Teachers enjoyed marking small samples (typically the equivalent of six to seven scripts per teacher) rather than the entire cohort which was a substantial time saver. Head of departments found the marks accurate and shared the reports on the precision of marking of their staff as professional development. Heads of Trusts appreciated the overview of marking across their schools.
Double Marker is a very powerful and full-featured assessment system that provides assessment co-ordinators with clarity about how every individual mark has been awarded. However, it is still extremely quick for teachers to use and lets you move quickly from the assessment to the final mark and feedback.
Our next pilot
If you would like to take part in our next pilot, which is entirely free, you can book a call with us next week to find out more. The next pilot will be for GCSE English Language and Literature and will take place after the October half-term.







