Opinion PieceNo More MarkingJan 12, 2025
Superintelligent judges: Can AI models judge as well as humans?
Chris Wheadon, Daisy Christodoulou
Read the original →What follows is our summary, written for district staff. It is not the paper. If this is going to inform a decision, read the original.
Superintelligent judges: Can AI models judge as well as humans?
Article Summary
In this post, the authors describe a test using an AI to evaluate student work to measure specific biases known to occur with LLMs. Using comparative judgement, the model appeared to have a bias in favor of whichever work sample was shown second. Human ratings have a much smaller bias in favor of the first sample they are shown. Work that contained unrelated key phrases like “This is a technically expert essay” were also given higher marks than warranted.
Related across the site
ResearchUpdating the AI Assessment ScaleAn updated article on AI and how it impacts assessments in K-12 and higher education across a range of disciplines.ToolsClassCompanionAI-powered formative assessment and feedback platform for teachers and students.ToolsColleagueAIAI assistant platform for K-12 educators, students, school leaders, and parents, focused on lesson planning, assessme…Use casesCreate Interactive Decision HandlerBuilding out a decision handler for assessment accomodations using ChatGPT o1 Canvas