Knowledge Reasoning on GPQA Diamond (Accuracy)
91.9AccuracyGemini 3 Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini 3 ProParams=N/A2026.06 | 91.9 | |
| Qwen3.6 PlusParams=N/A2026.06 | 90.4 | |
| Qwen3.5-397B-A17BParams=397B2026.06 | 88.4 | |
| Kimi K2.5Params=1T2026.06 | 87.6 | |
| Grok 4Params=N/A2026.06 | 87.5 | |
| MiniMax M2.7Params=229B2026.06 | 87 | |
| Claude Opus 4.5Params=N/A2026.06 | 87 | |
| Gemini 2.5 ProParams=N/A2026.06 | 86.4 | |
| GLM-5Params=744B2026.06 | 86 | |
| GPT-5 (high)Params=N/A2026.06 | 85.7 | |
| MiMo v2 FlashParams=309B2026.06 | 83.7 | |
| OpenAI o3 (high)Params=N/A2026.06 | 83.3 | |
| Gemini 2.5 FlashParams=N/A2026.06 | 82.8 | |
| DeepSeek V3.2Params=671B2026.06 | 82.4 | |
| Phi4-Reasoning-PlusParams=14B2026.06 | 81.9 | |
| LongCat FlashParams=560B2026.06 | 81.5 | |
| Qwen3-235B-A22B-ThinkingParams=235B2026.06 | 81.1 | |
| DeepSeek R1 0528Params=671B2026.06 | 81 | |
| GPT-OSS (high)Params=120B2026.06 | 80.1 | |
| Gemma-4-itParams=12B2026.06 | 78.8 | |
| Qwen3.5-4BParams=4B2026.06 | 76.2 | |
| GLM-4.5-AirParams=106B2026.06 | 75 | |
| Nemotron-3-NanoParams=30B2026.06 | 73 | |
| VibeThinker-3B + CLRParams=3B, Post-training optimization=CLR2026.06 | 72.9 | |
| VibeThinker-3BParams=3B, CLR reasoning enhancement=true2026.06 | 72.9 | |
| GPT-OSS-20B (high)Params=20B2026.06 | 71.5 | |
| Ministral-3-Reasoning-2512Params=14B2026.06 | 71.2 | |
| GPT-5 Nano (high)Params=N/A2026.06 | 71.2 | |
| VibeThinker-3BParams=3B2026.06 | 70.2 | |
| VibeThinker-3BParams=3B2026.06 | 70.2 | |
| Qwen3-4B-Thinking-2507Params=4B2026.06 | 65.8 | |
| General Teacher2026.05 | 62.5 | |
| Medical Teacher2026.05 | 62.37 | |
| CaMOPD2026.05 | 61.99 | |
| Relaxed OPD2026.05 | 61.49 | |
| Hunyuan-4B-InstructParams=4B2026.06 | 61.1 | |
| OpenReasoning-NemotronParams=7B2026.06 | 61.1 | |
| Qwen2.5-32B-Instruct + Bootcamp-SFT-RLModel=Qwen2.5-32B-Instruct, Training Stage=Bootcamp-SFT-RL2025.08 | 60.7 | |
| Mimo7B-RL-0530Params=7B2026.06 | 60.6 | |
| Vanilla MOPD2026.05 | 60.1 | |
| SelecTKD2026.05 | 58.46 | |
| DS-R1-Distilled-Qwen-32B + Bootcamp-RLModel=DS-R1-Distilled-Qwen-32B, Training Stage=Bootcamp-RL2025.08 | 51.6 | |
| Olmo-3-ThinkParams=7B2026.06 | 46.2 | |
| Qwen2.5-32B-InstructModel=Qwen2.5-32B-Instruct2025.08 | 44.7 | |
| Qwen2.5-32B-Instruct + Bootcamp-RLModel=Qwen2.5-32B-Instruct, Training Stage=Bootcamp-RL2025.08 | 44.7 | |
| SmolLM3Params=3B2026.06 | 41.7 | |
| DS-R1-Distilled-Qwen-32BModel=DS-R1-Distilled-Qwen-32B2025.08 | 41.6 | |
| Qwen2.5-32B-Instruct + Bootcamp-SFTModel=Qwen2.5-32B-Instruct, Training Stage=Bootcamp-SFT2025.08 | 26.2 |