Reasoning on ARC Challenge (exact_match, flexible_extract)
95.1Exact Match AccuracyGPT-5 nano
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-5 nanoEvaluation Setting=chat2026.05 | 95.1 | — | |
| Qwen3-8BEvaluation Setting=chat2026.05 | 90.5 | — | |
| gemma-3-12b-itEvaluation Setting=chat2026.05 | 90 | — | |
| Ministral-3-8BEvaluation Setting=chat2026.05 | 89.6 | — | |
| gpt-oss-20bEvaluation Setting=chat2026.05 | 88.1 | — | |
| Qwen3-4BEvaluation Setting=chat2026.05 | 88 | — | |
| Llama-3.1-8B-InstructEvaluation Setting=chat2026.05 | 81.7 | — | |
| Moonlight-16B-A3B-InstructEvaluation Setting=chat2026.05 | 80.9 | — | |
| gemma-3-4b-itEvaluation Setting=chat2026.05 | 75.7 | — | |
| Llama-3.2-3B-InstructEvaluation Setting=chat2026.05 | 75.6 | — | |
| EngGPT2-16B-A3BEvaluation Setting=chat2026.05 | 71 | — | |
| Velvet-14BEvaluation Setting=chat2026.05 | 65.9 | — | |
| LLaMAntino-3-ANITA-8BEvaluation Setting=chat2026.05 | 45.7 | — | |
| FastwebMIIA-7BEvaluation Setting=chat2026.05 | 45.5 | — | |
| deepseek-moe-16b-chatEvaluation Setting=chat2026.05 | 39.8 | — | |
| Minerva-7B-instruct-v1.0Evaluation Setting=chat2026.05 | 26.2 | — | |
| deepseek-moe-16b-chatEvaluation Setting=cot_custom2026.05 | — | 54.9 | |
| EngGPT2-16B-A3BEvaluation Setting=cot_custom2026.05 | — | 87.2 | |
| FastwebMIIA-7BEvaluation Setting=cot_custom2026.05 | — | 56.2 | |
| gemma-3-12b-itEvaluation Setting=cot_custom2026.05 | — | 92.3 | |
| gemma-3-4b-itEvaluation Setting=cot_custom2026.05 | — | 77.8 | |
| GPT-5 nanoEvaluation Setting=cot_custom2026.05 | — | 94.4 | |
| gpt-oss-20bEvaluation Setting=cot_custom2026.05 | — | 95.4 | |
| Llama-3.1-8B-InstructEvaluation Setting=cot_custom2026.05 | — | 80.5 | |
| Llama-3.2-3B-InstructEvaluation Setting=cot_custom2026.05 | — | 73.7 | |
| LLaMAntino-3-ANITA-8BEvaluation Setting=cot_custom2026.05 | — | 75.4 | |
| Minerva-7B-instruct-v1.0Evaluation Setting=cot_custom2026.05 | — | 25.4 | |
| Ministral-3-8BEvaluation Setting=cot_custom2026.05 | — | 94.2 | |
| Moonlight-16B-A3B-InstructEvaluation Setting=cot_custom2026.05 | — | 78.9 | |
| Qwen3-4BEvaluation Setting=cot_custom2026.05 | — | 93 | |
| Qwen3-8BEvaluation Setting=cot_custom2026.05 | — | 95.1 | |
| Velvet-14BEvaluation Setting=cot_custom2026.05 | — | 74.1 |