Multi-task Language Understanding on MMLU (Normalized Score)
92Normalized ScoreGPT-o1
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-o1Model Category=Reasoning Models, Assessed Version Date=2024-092025.11 | 92 | |
| DeepSeek–R1Model Category=Reasoning Models, Assessed Version Date=2025-052025.11 | 91 | |
| GPT-4oModel Category=Closed-Source Models (non-reasoning), Assessed Version Date=2025-032025.11 | 86 | |
| GPT-4o-MiniModel Category=Closed-Source Models (non-reasoning), Assessed Version Date=2024-112025.11 | 82 | |
| GPT-3.5-TurboModel Category=Closed-Source Models (non-reasoning), Assessed Version Date=2023-112025.11 | 70 | |
| Qwen2.5-3B-InstructModel Category=Open-Source Models (non-reasoning)2025.11 | 65 | |
| LLaMA2-13B-ChatModel Category=Open-Source Models (non-reasoning)2025.11 | 51 | |
| Falcon-7B-InstructModel Category=Open-Source Models (non-reasoning)2025.11 | 28 |