General Knowledge Question Answering on MMLU-PRO
81.61AccuracySMCS
Evaluation Results
| Method | Links | |
|---|---|---|
| SMCSMaximum output tokens=32,7682025.07 | 81.61 | |
| GPT-4.1(2025-04-14)Maximum output tokens=32,7682025.07 | 80.43 | |
| Claude-3.5-Sonnet(2024-06-20)Maximum output tokens=32,7682025.07 | 78.34 | |
| QwQ-32BMaximum output tokens=32,7682025.07 | 78.34 | |
| GLM-Z1-32B-0414Maximum output tokens=32,7682025.07 | 77.84 | |
| Qwen3-32BMaximum output tokens=32,7682025.07 | 77.76 | |
| DeepSeek-R1-Distill-Llama-70BMaximum output tokens=32,7682025.07 | 76.09 | |
| DeepSeek-R1-Distill-Qwen-32BMaximum output tokens=32,7682025.07 | 75.17 | |
| HuatuoGPT-o1-72BMaximum output tokens=32,7682025.07 | 74.16 | |
| GPT-o3-mini(2025-01-31)Maximum output tokens=32,7682025.07 | 74 | |
| GPT-4o(2024-08-06)Maximum output tokens=32,7682025.07 | 73.83 | |
| Qwen-2.5-72B-InstructMaximum output tokens=32,7682025.07 | 72.16 | |
| EXAONE-Deep-32BMaximum output tokens=32,7682025.07 | 70.48 | |
| Llama-3.3-70B-InstructMaximum output tokens=32,7682025.07 | 69.87 | |
| Claude-3.7-Sonnet(2025-02-19)Maximum output tokens=32,7682025.07 | 69.43 | |
| Qwen2.5-32b-InstructMaximum output tokens=32,7682025.07 | 69.15 | |
| TeleChat2-35B-32KMaximum output tokens=32,7682025.07 | 67.98 | |
| Llama-3.3-Nemotron-Super-49B-v1Maximum output tokens=32,7682025.07 | 67.47 | |
| Gemma-3-27b-itMaximum output tokens=32,7682025.07 | 65.47 | |
| Qwen2.5-Coder-32B-InstructMaximum output tokens=32,7682025.07 | 61.79 | |
| DarkForestModels=Qwen2.5-7B-Instruct, Qwen2.5-Coder-7B-Instruct, Mathstral-7B-v0.12026.05 | 58.38 | |
| DebateModels=Qwen2.5-7B-Instruct, Qwen2.5-Coder-7B-Instruct, Mathstral-7B-v0.12026.05 | 55.86 | |
| Mixture-of-AgentModels=Qwen2.5-7B-Instruct, Qwen2.5-Coder-7B-Instruct, Mathstral-7B-v0.12026.05 | 53.38 | |
| Self-ConsistencyModels=Qwen2.5-7B-Instruct, Qwen2.5-Coder-7B-Instruct, Mathstral-7B-v0.12026.05 | 53 | |
| RefineModels=Qwen2.5-7B-Instruct, Qwen2.5-Coder-7B-Instruct, Mathstral-7B-v0.12026.05 | 52.62 | |
| ReConcileModels=Qwen2.5-7B-Instruct, Qwen2.5-Coder-7B-Instruct, Mathstral-7B-v0.12026.05 | 45.71 | |
| InternLM2.5-20B-ChatMaximum output tokens=32,7682025.07 | 44.23 | |
| Graph-of-Agent (Max)Models=Qwen2.5-7B-Instruct, Qwen2.5-Coder-7B-Instruct, Mathstral-7B-v0.12026.05 | 41.86 | |
| Graph-of-Agent (Mean)Models=Qwen2.5-7B-Instruct, Qwen2.5-Coder-7B-Instruct, Mathstral-7B-v0.12026.05 | 40.71 |