Generative AI

Advice in a rapidly changing world

Using AI for marking – Is it any good?

A bell curve,
An AI generated image of a bell curve, Can you spot the mistake?

An interesting report came out form the University of Cambridge last week examining how accurate several frontier models were when they were asked to grade undergraduate essays from across three different institutions (Cambridge , Manchester Met, and Nottingham). 

The results were mixed, with the AI successfully classified Cambridge essays into the correct UK degree classification band of five (First, 2:1, 2:2, Third, Fail) approximately 63% of the time.  For Nottingham it was 53% and for Manchester Metropolitan it was 35%.

The AI faces significant challenges especially at the extremes (hence the bell curve image above) where they consistently underscore those essays that were ranked the best by the human markers and overscore those ranked as the worst by the human markers.  The AIs also consistently gave higher marks to essays that were longer and had greater sentence complexity even if this was unrelated to the academic content of the essays.