Mathematical Reasoning on AIME 24
93.3AIME 24 AccuracyGPT-5
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5Evaluation Protocol=Closed-Source2026.01 | 93.3 | |
| DeepConfModel=DeepSeek-8B, Sample Count=@5122026.02 | 92 | |
| Gemini2.5-ProEvaluation Protocol=Closed-Source2026.01 | 92 | |
| DeepConfModel=DeepSeek-8B, Sample Count=@202026.02 | 91.3 | |
| CoRefine TreeModel=Qwen3-32B2026.02 | 90.7 | |
| DeepConfModel=PaCoRe-8B, Sample Count=@5122026.02 | 90.7 | |
| CoRefine TreeModel=PaCoRe-8B2026.02 | 90.7 | |
| MajorityModel=DeepSeek-8B, Sampling Mode=Parallel, Sample Count=@202026.02 | 90.6 | |
| CoRefine TreeModel=DeepSeek-8B2026.02 | 90.6 | |
| MajorityModel=DeepSeek-8B, Sampling Mode=Sequential, Sample Count=@202026.02 | 90 | |
| CoRefineModel=DeepSeek-8B2026.02 | 90 | |
| DeepConfModel=Qwen3-32B, Sample Count=@202026.02 | 90 | |
| MajorityModel=PaCoRe-8B, Sampling Mode=Parallel, Sample Count=@202026.02 | 90 | |
| DeepConfModel=Qwen3-32B, Sample Count=@5122026.02 | 89.3 | |
| CoRefineModel=Qwen3-32B2026.02 | 89.3 | |
| DeepConfModel=PaCoRe-8B, Sample Count=@202026.02 | 87.4 | |
| CoRefineModel=PaCoRe-8B2026.02 | 86.7 | |
| MajorityModel=Qwen3-32B, Sampling Mode=Parallel, Sample Count=@202026.02 | 86 | |
| MajorityModel=PaCoRe-8B, Sampling Mode=Parallel, Sample Count=@5122026.02 | 86 | |
| MajorityModel=Qwen3-32B, Sampling Mode=Parallel, Sample Count=@5122026.02 | 85.4 | |
| MajorityModel=Qwen3-32B, Sampling Mode=Sequential, Sample Count=@202026.02 | 85.4 | |
| MajorityModel=PaCoRe-8B, Sampling Mode=Sequential, Sample Count=@202026.02 | 85.4 | |
| MajorityModel=DeepSeek-8B, Sampling Mode=Parallel, Sample Count=@5122026.02 | 85.3 | |
| DeepSeek-8BStrategy=Pass, Sample Count=@12026.02 | 82 | |
| Qwen3-32BStrategy=Pass, Sample Count=@12026.02 | 80.7 | |
| PaCoRe-8BStrategy=Pass, Sample Count=@12026.02 | 76.7 | |
| DS-R1-Distill-Qwen-7B-ScaleQuestModel=DS-R1-Distill-Qwen-7B-ScaleQuest, Data=ScaleQuest2024.10 | 53.3 | |
| LycheeDecodeBase Model=DeepSeek-R1-Distill-Qwen-7B, Cache Correction=True2026.02 | 46.7 | |
| GPT-4.1Evaluation Protocol=Closed-Source2026.01 | 46.7 | |
| DS-R1-Distill-Qwen-7BModel=DS-R1-Distill-Qwen-7B2024.10 | 43.3 | |
| LycheeDecodeBase Model=DeepSeek-R1-Distill-Qwen-7B, Cache Correction=False2026.02 | 43.3 | |
| ATLAS (cluster)Evaluation Protocol=In-Distribution2026.01 | 43.3 | |
| ATLAS (RL)Evaluation Protocol=Out-of-Distribution2026.01 | 43.3 | |
| LycheeDecodeBase Model=DeepSeek-R1-Distill-Llama-8B, Cache Correction=True2026.02 | 40 | |
| Full AttentionBase Model=DeepSeek-R1-Distill-Qwen-7B, Cache Correction=False2026.02 | 40 | |
| RouterDCEvaluation Protocol=In-Distribution2026.01 | 40 | |
| SAGEBase Model=Qwen2.5-7B-Instruct2026.02 | 33.3 | |
| TidalDecodeBase Model=DeepSeek-R1-Distill-Llama-8B, Cache Correction=True2026.02 | 33.3 | |
| BertRouterEvaluation Protocol=In-Distribution2026.01 | 30 | |
| DPO (Full)Base Model=Qwen2.5-7B-Instruct2026.02 | 26.7 | |
| LycheeDecodeBase Model=DeepSeek-R1-Distill-Llama-8B, Cache Correction=False2026.02 | 26.7 | |
| TidalDecodeBase Model=DeepSeek-R1-Distill-Qwen-7B, Cache Correction=True2026.02 | 26.7 | |
| MLPRouterEvaluation Protocol=In-Distribution2026.01 | 26.7 | |
| ToRLBackbone=Qwen2.5-7B, Training Paradigm=RL-only2026.01 | 23.33 | |
| AutoTrajBackbone=Qwen2.5-7B, Training Paradigm=SFT–RL TIR, SFT Trajectory Count=13K2026.01 | 23.33 | |
| EASDTM=32B, DM=7B2025.12 | 23.33 | |
| VanillaBase Model=Qwen2.5-7B-Instruct2026.02 | 23.3 | |
| Full AttentionBase Model=DeepSeek-R1-Distill-Llama-8B, Cache Correction=False2026.02 | 23.3 | |
| FS RouterEvaluation Protocol=Training-free, Prompting Strategy=Few-shot2026.01 | 23.3 | |
| DPO (Random)Base Model=Qwen2.5-7B-Instruct2026.02 | 20 | |
| EASDTM=72B, DM=7B2025.12 | 20 | |
| TidalDecodeBase Model=DeepSeek-R1-Distill-Qwen-7B, Cache Correction=False2026.02 | 16.7 | |
| Tool-StarBackbone=Qwen2.5-7B, Training Paradigm=SFT–RL TIR, SFT Trajectory Count=54K2026.01 | 16.67 | |
| Single ModelTM=32B2025.12 | 16.67 | |
| Majority Voting (N=16)DM=7B2025.12 | 16.67 | |
| RSDTM=32B, DM=7B, PRM=1.5B2025.12 | 16.67 | |
| Single ModelTM=72B2025.12 | 16.67 | |
| RSDTM=72B, DM=7B, PRM=1.5B2025.12 | 16.67 | |
| SDTM=32B, DM=7B2025.12 | 13.33 | |
| SDTM=72B, DM=7B2025.12 | 13.33 | |
| SAGEBase Model=Qwen2.5-3B-Instruct2026.02 | 13.3 | |
| TidalDecodeBase Model=DeepSeek-R1-Distill-Llama-8B, Cache Correction=False2026.02 | 13.3 | |
| GPT-4oEvaluation Protocol=Closed-Source2026.01 | 13.3 | |
| ZS RouterEvaluation Protocol=Training-free, Prompting Strategy=Zero-shot2026.01 | 13.3 | |
| RouterDCEvaluation Protocol=Out-of-Distribution2026.01 | 13.3 | |
| MLPRouterEvaluation Protocol=Out-of-Distribution2026.01 | 13.3 | |
| ATLAS (cluster)Evaluation Protocol=Out-of-Distribution2026.01 | 13.3 | |
| DPO (Full)Base Model=Qwen2.5-1.5B-Instruct2026.02 | 10 | |
| SAGEBase Model=Qwen2.5-1.5B-Instruct2026.02 | 10 | |
| DPO (Full)Base Model=Qwen2.5-3B-Instruct2026.02 | 10 | |
| Tool-Star-SFTBackbone=Qwen2.5-7B, Training Paradigm=SFT-only, SFT Trajectory Count=54K2026.01 | 10 | |
| Single ModelDM=7B2025.12 | 10 | |
| Beam Search (N=16)DM=7B2025.12 | 10 | |
| VanillaBase Model=Qwen2.5-1.5B-Instruct2026.02 | 6.7 | |
| VanillaBase Model=Qwen2.5-3B-Instruct2026.02 | 6.7 | |
| Random RouterEvaluation Protocol=Training-free2026.01 | 6.7 | |
| BertRouterEvaluation Protocol=Out-of-Distribution2026.01 | 6.7 | |
| Qwen2.5-7B-InstructBackbone=Qwen2.5-7B, Training Paradigm=Instruct2026.01 | 6.67 | |
| AutoTIRBackbone=Qwen2.5-7B, Training Paradigm=RL-only2026.01 | 6.67 | |
| R1-SearcherBackbone=Qwen2.5-7B, Training Paradigm=RL-only2026.01 | 3.33 | |
| Vanilla SFT-RL TIRBackbone=Qwen2.5-7B, Training Paradigm=SFT–RL TIR2026.01 | 3.33 | |
| DPO (Random)Base Model=Qwen2.5-1.5B-Instruct2026.02 | 3.3 | |
| DPO (Random)Base Model=Qwen2.5-3B-Instruct2026.02 | 0 | |
| ReSearchBackbone=Qwen2.5-7B, Training Paradigm=RL-only2026.01 | 0 |