Multi-trait Automated Essay Scoring on ASAP Prompt 7 (test)
69.5Ideas ScoreHuman Rater 1 - Human Rater 2
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Human Rater 1 - Human Rater 2type=Human baseline2026.02 | 69.5 | 57.6 | 56.7 | 54.4 | 62 | |
| Llama 4 (multi-agent prompting framework)Prompting Strategy=Few-shot (Base), Rubric Detail=Detailed2026.02 | 69.3 | 64.7 | 61.5 | 59.7 | 63.8 | |
| Llama 4 (No Examples)Prompting Strategy=Zero-shot, Rubric Detail=Detailed2026.02 | 60.1 | 49.8 | 48.8 | 44.5 | 50.8 | |
| Llama 4 (Reduced Rubric)Prompting Strategy=Few-shot, Rubric Detail=Reduced2026.02 | 57.5 | 54.1 | 63.9 | 49 | 56.1 | |
| Llama 2Prompting Strategy=1 shot2026.02 | 9.1 | 2.3 | 32.7 | 15.4 | 15.1 | |
| GPT 3.5Prompting Strategy=1 shot2026.02 | 4.5 | 6.8 | 9.7 | 7.9 | 7.3 |