Mathematical Reasoning on MATH
95.63AccuracyCoD
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| CoDRisk tolerance (epsilon)=0.08, Confidence level (alpha)=0.12026.01 | 95.63 | — | 12.04 | |
| VanillaModel=Qwen3-8B, Budget=24002026.03 | 92.6 | — | — | |
| VanillaModel=Qwen3-8B, Budget=32002026.03 | 92.6 | — | — | |
| Pass@8 (Upper Bound)Decoding Strategy=Pass@82025.05 | 91.8 | — | — | |
| R-KVModel=Qwen3-8B, Budget=32002026.03 | 91.6 | — | — | |
| VATPModel=Qwen3-8B, Budget=32002026.03 | 91 | — | — | |
| H2OModel=Qwen3-8B, Budget=32002026.03 | 90.4 | — | — | |
| GHG-TDABase Model=GPT-4o2026.02 | 90 | — | — | |
| LongFlowModel=Qwen3-8B, Budget=32002026.03 | 89.8 | — | — | |
| Qwen3-VL-8BCategory=Visual-enhanced LLM2026.01 | 89.24 | — | — | |
| theory-guided context selection strategyModel=Qwen3-8B, Selection strategy=top-12026.02 | 89.2 | — | — | |
| AoTBase Model=GPT-4o2026.02 | 89.1 | — | — | |
| GHG-TDABase Model=Claude 3.5 Sonnet2026.02 | 89.1 | — | — | |
| R-KVModel=Qwen3-8B, Budget=24002026.03 | 89 | — | — | |
| Qwen2.5-7B + Reinforce ++Base Model=Qwen2.5-7B, Algorithm=Reinforce ++2025.12 | 88.8 | — | — | |
| Qwen2.5-7B + DAPOBase Model=Qwen2.5-7B, Algorithm=DAPO2025.12 | 88.8 | — | — | |
| Qwen2.5-7B + ARPOBase Model=Qwen2.5-7B, Algorithm=ARPO2025.12 | 88.8 | — | — | |
| GoTBase Model=GPT-4o2026.02 | 88.4 | — | — | |
| GenICLModel=Qwen3-8B, Selection strategy=top-12026.02 | 88.4 | — | — | |
| AoTBase Model=Claude 3.5 Sonnet2026.02 | 88 | — | — | |
| BM25Model=Qwen3-8B, Selection strategy=top-12026.02 | 88 | — | — | |
| DICLModel=Qwen3-8B, Selection strategy=top-12026.02 | 88 | — | — | |
| LongFlowModel=Qwen3-8B, Budget=24002026.03 | 88 | — | — | |
| Qwen2.5-7B + GRPOBase Model=Qwen2.5-7B, Algorithm=GRPO2025.12 | 87.8 | — | — | |
| SCOPEDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 87.7 | — | — | |
| SCOPE (Step Skipping)Decoding Strategy=Best-of-8 strategy, Ablation=Step Skipping2025.05 | 87.7 | — | — | |
| VATPModel=Qwen3-8B, Budget=24002026.03 | 87.6 | — | — | |
| SCOPE (w/o AST normalization)Decoding Strategy=Best-of-8 strategy, Ablation=w/o AST normalization2025.05 | 87.3 | — | — | |
| SCOPE (Step Replacement)Decoding Strategy=Best-of-8 strategy, Ablation=Step Replacement2025.05 | 87.3 | — | — | |
| GoTBase Model=Claude 3.5 Sonnet2026.02 | 87.3 | — | — | |
| SCOPE (w/o code translation)Decoding Strategy=Best-of-8 strategy, Ablation=w/o code translation2025.05 | 87 | — | — | |
| ZeroModel=Qwen3-8B, Selection strategy=top-1, contextual_information=none2026.02 | 87 | — | — | |
| Majority @8Decoding Strategy=Majority @82025.05 | 86.9 | — | — | |
| ToTBase Model=GPT-4o2026.02 | 86.9 | — | — | |
| Qwen2.5-Math-7B-PRM800KDecoding Strategy=Best-of-8 strategy, Training Data=manually annotated (PRM800K)2025.05 | 86.5 | — | — | |
| Skywork-PRM-1.5BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 86.3 | — | — | |
| TopicKModel=Qwen3-8B, Selection strategy=top-12026.02 | 86.2 | — | — | |
| Skywork-PRM-7BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 86.1 | — | — | |
| ToTBase Model=Claude 3.5 Sonnet2026.02 | 86 | — | — | |
| Qwen3-VL-4BCategory=Visual-enhanced LLM2026.01 | 85.83 | — | — | |
| RLHFlow-PRM-Mistral-8BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 85.7 | — | — | |
| RLHFlow-PRM-Deepseek-8BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 85.6 | — | — | |
| CoTBase Model=GPT-4o2026.02 | 85.4 | — | — | |
| H2OModel=Qwen3-8B, Budget=24002026.03 | 85.2 | — | — | |
| Gemma 3 12BParameters=12B2026.02 | 84.25 | — | — | |
| CoTBase Model=Claude 3.5 Sonnet2026.02 | 84.1 | — | — | |
| NBDiff-7B-INSTRUCTParameters=7B, Training Protocol=Instruct, Sampling Strategy=Greedy2025.12 | 84 | — | — | |
| EurusPRM-Stage1Decoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 83.4 | — | — | |
| VanillaModel=DeepSeek-R1-Distill-Llama-8B, Budget=24002026.03 | 83.4 | — | — | |
| VanillaModel=DeepSeek-R1-Distill-Llama-8B, Budget=32002026.03 | 83.4 | — | — | |
| EurusPRM-Stage2Decoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 83.1 | — | — | |
| GreedyDecoding Strategy=Greedy2025.05 | 83 | — | — | |
| Qwen2.5-3B + ARPOBase Model=Qwen2.5-3B, Algorithm=ARPO2025.12 | 82.5 | — | — | |
| R-KVModel=DeepSeek-R1-Distill-Llama-8B, Budget=32002026.03 | 82.4 | — | — | |
| Qwen2.5-7BBase Model=Qwen2.5-7B, Algorithm=Direct Reasoning2025.12 | 82 | — | — | |
| Math-Shepherd-PRM-7BDecoding Strategy=Best-of-8 strategy, Training Data=automatically constructed2025.05 | 81.7 | — | — | |
| H2OModel=DeepSeek-R1-Distill-Llama-8B, Budget=32002026.03 | 81.6 | — | — | |
| VATPModel=DeepSeek-R1-Distill-Llama-8B, Budget=32002026.03 | 81.6 | — | — | |
| Qwen2.5-3B + DAPOBase Model=Qwen2.5-3B, Algorithm=DAPO2025.12 | 81.2 | — | — | |
| GPT-4o-2024-08-06Synthesis Model=-, Chain-of-Thought (CoT) reasoning=true, Greedy decoding=true2024.10 | 81.1 | — | — | |
| Qwen2.5-3B + GRPOBase Model=Qwen2.5-3B, Algorithm=GRPO2025.12 | 81 | — | — | |
| R-KVModel=DeepSeek-R1-Distill-Llama-8B, Budget=24002026.03 | 80.8 | — | — | |
| SKPO (Ours)Base Model=Qwen2.5-Math-7B2026.04 | 80.8 | — | — | |
| Qwen2.5-3B + Reinforce ++Base Model=Qwen2.5-3B, Algorithm=Reinforce ++2025.12 | 80.2 | — | — | |
| Llama3.1-8B + ARPOBase Model=Llama3.1-8B, Algorithm=ARPO2025.12 | 80.2 | — | — | |
| H2OModel=DeepSeek-R1-Distill-Llama-8B, Budget=24002026.03 | 79.8 | — | — | |
| VATPModel=DeepSeek-R1-Distill-Llama-8B, Budget=24002026.03 | 79.8 | — | — | |
| LongFlowModel=DeepSeek-R1-Distill-Llama-8B, Budget=24002026.03 | 79.8 | — | — | |
| LongFlowModel=DeepSeek-R1-Distill-Llama-8B, Budget=32002026.03 | 79.8 | — | — | |
| CISPOBase Model=Qwen2.5-Math-7B2026.04 | 79.4 | — | — | |
| Llama3.1-8B + GRPOBase Model=Llama3.1-8B, Algorithm=GRPO2025.12 | 79.2 | — | — | |
| SDAR 8BParameters=8B, Sampling Strategy=Greedy2025.12 | 78.6 | — | — | |
| rStarMath PolicyParameter Count=7B2025.02 | 78.4 | — | — | |
| Qwen2.5-7B + TIR PromptingBase Model=Qwen2.5-7B, Algorithm=TIR Prompting2025.12 | 78.2 | — | — | |
| NoThinkingRisk tolerance (epsilon)=0.08, Confidence level (alpha)=0.12026.01 | 77.26 | — | 15.77 | |
| Llama3.1-8B + Reinforce ++Base Model=Llama3.1-8B, Algorithm=Reinforce ++2025.12 | 77.2 | — | — | |
| GPT-4o2025.02 | 76.6 | — | — | |
| GPT-4ochain-of-thought=true2024.10 | 76.6 | — | — | |
| Llama3.1-8B + DAPOBase Model=Llama3.1-8B, Algorithm=DAPO2025.12 | 76.4 | — | — | |
| SAPOBase Model=Qwen2.5-Math-7B2026.04 | 76.1 | — | — | |
| DAPOBase Model=Qwen2.5-Math-7B2026.04 | 75.8 | — | — | |
| FastMCTSParameter Count=7B2025.02 | 75.4 | — | — | |
| DictaLM 3.0 12B-InstParameters=12B, Variant=Instruct2026.02 | 74.99 | — | — | |
| MCNIGModel Size=8B2026.03 | 74.8 | — | — | |
| M2CLBackbone=Qwen-72B, Number of LLMs=322026.02 | 74.7 | — | — | |
| QwenPRMModel Size=7B2026.03 | 74.3 | — | — | |
| GSPOBase Model=Qwen2.5-Math-7B2026.04 | 74.2 | — | — | |
| ORMModel Size=8B2026.03 | 73.8 | — | — | |
| LLaDA2.0-mini preview 16BA1BParameters=16BA1B, Sampling Strategy=Greedy2025.12 | 73.5 | — | — | |
| GPT-4-Turbo-24-04-09Synthesis Model=-, Chain-of-Thought (CoT) reasoning=true, Greedy decoding=true2024.10 | 73.4 | — | — | |
| Qwen2-Math-7B-ScaleQuestBase Model=Qwen2-Math-7B, Synthesis Model=Qwen2-Math-7B-Instruct2024.10 | 73.4 | — | — | |
| Qwen2-Math-7B-InstructSynthesis Model=-, Chain-of-Thought (CoT) reasoning=true, Greedy decoding=true2024.10 | 73.1 | — | — | |
| IGModel Size=8B2026.03 | 73.1 | — | — | |
| M2CLBackbone Model=Qwen-72B, Number of LLMs=42026.02 | 72.5 | — | — | |
| Primitives-based MAS2026.02 | 72.4 | — | — | |
| MacNetBackbone=Qwen-72B, Number of LLMs=322026.02 | 72.1 | — | — | |
| Qwen2.5-3BBase Model=Qwen2.5-3B, Algorithm=Direct Reasoning2025.12 | 71.6 | — | — | |
| PRIMEBase Model=Qwen2.5-Math-7B2026.04 | 71.3 | — | — | |
| MiMo-V2 FlashModel Variant=Base, # Shots=4-shot, # Activated Params=15B, # Total Params=309B2026.02 | 71 | — | — | |
| MiMo-V2-Flash Base# Shots=4-shot, # Activated Params=15B, # Total Params=309B2026.01 | 71 | — | — |