Automatic Short Answer Grading for Computer Science Placement Exam with Meta Llama
August 14, 2025
Authors
Anh Nguyen
Mt. Holyoke College
As part of a CAFÉCS summer fellowship program, funded by Mt. Holyoke, this project explores the use of large language models (LLMs), specifically Meta’s open-source Llama family, for automatic short answer grading (ASAG) in a computer science placement exam. The goal is to determine whether incoming students should begin with the introductory Exploring Computer Science (ECS) course or the more rigorous AP Computer Science Principles (CSP). Traditional human scoring of the ECS-based exam is labor-intensive, prompting the investigation of automated grading solutions. Initial experiments with BERT models showed promise but required extensive scored datasets. This project leverages Llama’s free, open-source nature to address privacy concerns and reduce dependency on large datasets. The study involves prompt engineering and iterative tuning to achieve high agreement with human scores. Advanced techniques like chain-of-thought prompting and multi-agent ensemble methods were employed to enhance grading consistency. Results indicate that while Llama’s performance still falls short of human graders, it shows potential for practical reliability with further refinement. The findings suggest that LLMs can offer rapid formative feedback, making them a viable option for large-scale automated grading in educational settings.
Suggested Citation
Nguyen, A. (2025, August). Automatic Short Answer Grading for Computer Science Placement Exam with Meta Llama. Chicago, IL: The Learning Partnership. https://doi.org/10.51420/Report.2025.4
Add your information to our mailing list for updates about what’s new at The Learning Partnership.