~/AI IN EDUCAT/study-finds-generative-ai-cannot-yet-reliably-grade-student-essays

Study Finds Generative AI Cannot Yet Reliably Grade Student Essays

A study by Cardiff and Melbourne universities found that ChatGPT cannot reliably grade undergraduate essays, showing significant discrepancies compared to human instructors. The AI tended to inflate low scores and deflate high scores, systematically regressing the grades toward the average. While educational institutions look to AI to reduce teacher workloads and increase grading efficiency, this study highlights that current LLMs are not yet ready to replace human judgment in evaluating subjective written work. Researchers tested two versions of ChatGPT on 50 bioscience essays using seven criteria and four prompting methods, finding a maximum score discrepancy of 40 points on a single essay compared to human grading.

## BACKGROUND

Automated Essay Scoring (AES) is an application of natural language processing that uses computer programs to evaluate and grade written essays. While traditional AES systems rely on statistical classification, the integration of Large Language Models (LLMs) introduces new capabilities and challenges in evaluating complex, subjective student work.

## REFERENCES

## KEYWORDS

#AI in Education#LLM Evaluation#Natural Language Processing#EdTech

$ subscribe --daily

Study Finds Generative AI Cannot Yet Reliably Grade Student Essays | Daily News