Structural curriculum understanding on K12-Bench
63.4Grounding EMGemini-1.5-Pro
Evaluation Results
| Method | Links | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Gemini-1.5-ProEvaluation Protocol=zero-shot, Response Format=answer-only, Model Category=Proprietary Models2026.05 | 63.4 | 83.3 | 34.8 | 58.2 | 33.4 | 63.5 | 47.4 | 72.4 | 81.7 | 82.6 | 57.1 | 73 | |
| Gemini-1.5-FlashEvaluation Protocol=zero-shot, Response Format=answer-only, Model Category=Proprietary Models2026.05 | 58 | 75.7 | 29.7 | 56.3 | 15.4 | 56 | 47.1 | 72 | 72.8 | 74 | 48.3 | 66.7 | |
| GPT-4 TurboEvaluation Protocol=zero-shot, Response Format=answer-only, Model Category=Proprietary Models2026.05 | 50.8 | 79.2 | 17.9 | 59.5 | 13.1 | 60 | 41.6 | 73.9 | 70.1 | 71.5 | 42.8 | 68 | |
| Gemma-2-9B-ITEvaluation Protocol=zero-shot, Response Format=answer-only, Model Category=Open Source Models2026.05 | 50.6 | 79 | 28.3 | 62.6 | 15 | 60.7 | 43.3 | 73.9 | 73.4 | 73.5 | 46.4 | 69.5 | |
| Qwen2-72B-InstructEvaluation Protocol=zero-shot, Response Format=answer-only, Model Category=Open Source Models2026.05 | 46.7 | 77.2 | 16.9 | 60.4 | 14.6 | 60.3 | 41.7 | 75.1 | 72.1 | 75.9 | 42.6 | 69.5 | |
| Mistral-7B-v0.3-InstructEvaluation Protocol=zero-shot, Response Format=answer-only, Model Category=Open Source Models2026.05 | 43 | 75.1 | 18.4 | 59.4 | 14.5 | 59.8 | 38.3 | 73.2 | 59.3 | 72.2 | 37.5 | 67.4 | |
| GLM-4-9B-ChatEvaluation Protocol=zero-shot, Response Format=answer-only, Model Category=Open Source Models2026.05 | 36.4 | 70.9 | 13.4 | 56.6 | 15.2 | 59.6 | 39.3 | 72.8 | 48.9 | 66.3 | 31.7 | 63.9 | |
| GPT-4o-miniEvaluation Protocol=zero-shot, Response Format=answer-only, Model Category=Proprietary Models2026.05 | 30.4 | 70.4 | 9.9 | 57.6 | 12.5 | 59 | 32.9 | 72.3 | 55.5 | 72.8 | 31.7 | 66.4 | |
| GPT-4oEvaluation Protocol=zero-shot, Response Format=answer-only, Model Category=Proprietary Models2026.05 | 30.3 | 70.6 | 9.3 | 57.5 | 10.1 | 57.9 | 32.4 | 71.8 | 55.6 | 72.3 | 31.1 | 65.9 | |
| Random guessEvaluation Protocol=zero-shot, Response Format=answer-only, Model Category=Random Baseline2026.05 | 6.7 | 36.2 | 6.7 | 37.9 | 6.7 | 41.3 | 6.7 | 37.7 | 6.7 | 32.9 | 6.7 | 36.4 | |
| Meta-LLaMA-3-8B-InstructEvaluation Protocol=zero-shot, Response Format=answer-only, Model Category=Open Source Models2026.05 | 6.2 | 54.9 | 4.3 | 47.9 | 3.8 | 53.4 | 5.2 | 55.4 | 11.5 | 53.9 | 7.2 | 52.6 |