Mathematical Reasoning on Beyond AIME (Accuracy (%))
58.8AccuracyGemini 2.5 pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini 2.5 proEvaluation Protocol=pass@12025.09 | 58.8 | |
| FIGR2025.12 | 54 | |
| Qwen3-235B-A22BModel Category=Large Language Models, Model Scale=235B-A22B, Reasoning Strategy=Thinking2025.12 | 52 | |
| GLM-4.5VModel Category=Large Vision-Language Models, Model Scale=108B2025.12 | 47 | |
| Text-only RLMode=Text-only, Reasoning Strategy=RL2025.12 | 46 | |
| Qwen3-VL-32B-InstructModel Category=Large Vision-Language Models, Model Scale=32B, Reasoning Strategy=Instruct2025.12 | 43 | |
| Qwen3-32BModel Category=Large Language Models, Model Scale=32B, Reasoning Strategy=Thinking2025.12 | 40 | |
| Qwen3-VL-8B-InstructModel Category=Large Vision-Language Models, Model Scale=8B, Reasoning Strategy=Instruct2025.12 | 30 | |
| PACS-8BEvaluation Protocol=pass@1, Backbone=Qwen3-8B2025.09 | 28.86 | |
| PACS-4BEvaluation Protocol=pass@1, Backbone=Qwen3-4B2025.09 | 27.16 | |
| Qwen3-14BEvaluation Protocol=pass@12025.09 | 18.59 | |
| Qwen3-4BEvaluation Protocol=pass@12025.09 | 17.33 | |
| Qwen3-8BEvaluation Protocol=pass@12025.09 | 16.28 | |
| Qwen3-32BModel Category=Large Language Models, Model Scale=32B, Reasoning Strategy=Non-Thinking2025.12 | 16 | |
| Ouro2.6B-Thinking + RLTT2026.02 | 16 | |
| LIMOEvaluation Protocol=pass@12025.09 | 15.98 | |
| DyJRBackbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@256, alpha_JS=0.05, M=82026.03 | 12.7 | |
| OpenThinker3-7BEvaluation Protocol=pass@12025.09 | 12.33 | |
| Ex-GRPOBackbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@2562026.03 | 11.6 | |
| DyJR (alpha_JS = 0.01)Backbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@256, alpha_JS=0.012026.03 | 11.5 | |
| DyJR (Forward KL)Backbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@256, Divergence Type=Forward KL2026.03 | 11.3 | |
| DyJR (M = 16)Backbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@256, M=162026.03 | 11.3 | |
| DyJR (M = 32)Backbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@256, M=322026.03 | 10.9 | |
| DPH-RLBackbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@2562026.03 | 10.8 | |
| GRPOBackbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@2562026.03 | 10.7 | |
| DAPOBackbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@2562026.03 | 10.6 | |
| RLEPBackbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@2562026.03 | 10.6 | |
| DyJR (M = 64)Backbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@256, M=642026.03 | 10.5 | |
| DyJR (alpha_JS = 0.2)Backbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@256, alpha_JS=0.22026.03 | 10.1 | |
| Sky-T1-7BEvaluation Protocol=pass@12025.09 | 8.54 | |
| Eurus-2-7B-PRIMEEvaluation Protocol=pass@12025.09 | 6.67 | |
| Ouro2.6B-Thinking + SFT2026.02 | 6 | |
| Ouro2.6B-Thinking + GRPO2026.02 | 6 | |
| Ouro2.6B-Thinking2026.02 | 5 | |
| Base ModelBackbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@2562026.03 | 4.3 | |
| DeepSeekR1 – 7B2026.02 | 3 | |
| Marco-o1Evaluation Protocol=pass@12025.09 | 2.77 | |
| DeepSeekR1 – 1.5B2026.02 | 2 | |
| Bagel-7B-MoTModel Category=Unified Multimodal Models, Model Scale=7B, Reasoning Strategy=MoT2025.12 | 1 | |
| Deepthought-8BEvaluation Protocol=pass@12025.09 | 0.76 | |
| Bagel-Zebra-CoTModel Category=Unified Multimodal Models, Model Scale=7B, Reasoning Strategy=CoT2025.12 | 0 | |
| DeepEyesModel Category=Tool-Augmented Vision-Language Models2025.12 | 0 | |
| Chain-of-FocusModel Category=Tool-Augmented Vision-Language Models2025.12 | 0 | |
| Qwen3 – 1.7B2026.02 | 0 | |
| Qwen3 – 4B2026.02 | 0 |