Language Understanding on MMLU CoT
87.2AccuracyGPT-5 nano
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5 nanoevaluation setting=llama custom, match type=flexible2026.05 | 87.2 | |
| gpt-oss-20bevaluation setting=llama custom, match type=flexible2026.05 | 83.7 | |
| Qwen3-8Bevaluation setting=llama custom, match type=flexible2026.05 | 83 | |
| Ministral-3-8Bevaluation setting=llama custom, match type=flexible2026.05 | 79.7 | |
| Qwen3-4Bevaluation setting=llama custom, match type=flexible2026.05 | 79.2 | |
| gemma-3-12b-itevaluation setting=llama custom, match type=flexible2026.05 | 76.9 | |
| Llama-3.1-8B-Instructevaluation setting=llama custom, match type=flexible2026.05 | 74 | |
| EngGPT2-16B-A3Bevaluation setting=llama custom, match type=flexible2026.05 | 71.1 | |
| Moonlight-16B-A3B-Instructevaluation setting=llama custom, match type=flexible2026.05 | 68.9 | |
| Llama-3.2-3B-Instructevaluation setting=llama custom, match type=flexible2026.05 | 65.9 | |
| gemma-3-4b-itevaluation setting=llama custom, match type=flexible2026.05 | 63.2 | |
| Velvet-14Bevaluation setting=llama custom, match type=flexible2026.05 | 52.6 | |
| FastwebMIIA-7Bevaluation setting=llama custom, match type=flexible2026.05 | 47.8 | |
| deepseek-moe-16b-chatevaluation setting=llama custom, match type=flexible2026.05 | 45.3 | |
| Minerva-7B-instruct-v1.0evaluation setting=llama custom, match type=flexible2026.05 | 27.9 | |
| LLaMAntino-3evaluation setting=llama custom, match type=flexible2026.05 | 21.2 |