Mathematical Reasoning on GSM8K (test)
98.8AccuracyReProbe
Evaluation Results
| Method | Links | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ReProbe# Sample=32K, Input Features=Hidden States, Annotation Source=DeepSeek-anno, Decoding Method=Beam search2025.11 | 98.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Claude 3.5 Sonnet2024.06 | 97.72 | — | 1.35 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ReProbe# Sample=32K, Input Features=Hidden States, Annotation Source=Self-anno, Decoding Method=Beam search2025.11 | 97.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4 Code2023.10 | 97 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4-0613Answer format=code, Eval method=pass@1, Self-measured=true2023.12 | 97 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4 Code+Self-Verification (K=5)GPT-3.5/4 usage=true, Data Augmentation=false, Code Execution=true, Self-Verification=true, K=52023.11 | 97 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4 Code InterpreterSize=-, Reasoning Mode=Tool-Integrated, Source Type=Closed-Source, Evaluation Protocol=Top12024.02 | 97 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4o2024.06 | 96.43 | — | 1.71 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SC-MASLLM backbone=LLM Pool, MAS=true, Routing=true2026.01 | 96.09 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MasRouterLLM backbone=LLM Pool, MAS=true, Routing=true2026.01 | 95.45 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Self-ContrastLLM=GPT4, #Call Avg.=7.82024.01 | 95.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-Math-PRM-7B# Sample=860K, Decoding Method=Beam search2025.11 | 95.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ReProbe# Sample=32K, Input Features=Attn+Logit, Annotation Source=Self-anno, Decoding Method=Beam search2025.11 | 95.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CoT + Skill-Based (maj@5)Base Model=GPT-4-0613, Prompting=CoT + Skill-Based (maj@5), self-consistency=maj@52024.05 | 95.38 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-Math-7B-InstructReasoning Strategy=short-CoT, Parameters=7B2025.03 | 95.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Self-ReflectionLLM=GPT4, #Call Avg.=32024.01 | 95.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Claude-3.5-Sonnet-1022Code Integration=No2025.02 | 95 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Claude-3-OpusSampling=N/A2024.06 | 95 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AFlowLLM backbone=gemini-1.5-flash, MAS=true, Routing=false2026.01 | 94.91 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AgentPruneLLM backbone=gpt-4o-mini, MAS=true, Routing=false2026.01 | 94.89 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4 (0314)Sampling=N/A2024.06 | 94.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPTSwarmLLM backbone=gpt-4o-mini, MAS=true, Routing=false2026.01 | 94.66 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4-Turbo (24-04-09)Sampling=N/A2024.06 | 94.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Gemini UltraAnswer format=nlp, Eval method=maj1@322023.12 | 94.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Gemini UltraSize=-, Reasoning Mode=Chain-of-Thought, Source Type=Closed-Source, Evaluation Protocol=Top12024.02 | 94.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CoT + Skill-BasedBase Model=GPT-4-0613, Prompting=CoT + Skill-Based2024.05 | 94.31 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SC-VoteLLM=GPT4, #Call Avg.=82024.01 | 94.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPTSwarmLLM backbone=gemini-1.5-flash, MAS=true, Routing=false2026.01 | 93.98 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PromptLLMLLM backbone=LLM Pool, MAS=false, Routing=true2026.01 | 93.92 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CoT PromptLLM=GPT4, #Call Avg.=12024.01 | 93.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Math-PromptLLM=GPT4, #Call Avg.=4.52024.01 | 93.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AgentPruneLLM backbone=gemini-1.5-flash, MAS=true, Routing=false2026.01 | 93.88 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ExpertPromptLLM=GPT4, #Call Avg.=22024.01 | 93.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Hint-PromptLLM=GPT4, #Call Avg.=6.72024.01 | 93.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| RouterDCLLM backbone=LLM Pool, MAS=false, Routing=true2026.01 | 93.68 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Multi-AgentLLM=GPT4, #Call Avg.=92024.01 | 93.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| RouteLLMLLM backbone=LLM Pool, MAS=false, Routing=true2026.01 | 93.42 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SC-ReflectLLM=GPT4, #Call Avg.=92024.01 | 93.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PROMPTCOT-Qwen-7BBase Model=Qwen2.5-Math-7B, Reasoning Strategy=short-CoT, Parameters=7B2025.03 | 93.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ReProbe# Sample=32K, Input Features=Attn+Logit, Annotation Source=DeepSeek-anno, Decoding Method=Beam search2025.11 | 93.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VanillaLLM backbone=gpt-4o-mini, MAS=false, Routing=false2026.01 | 93.17 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SC-SelectLLM=GPT4, #Call Avg.=92024.01 | 93.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CoTBase Model=GPT-4-0613, Prompting=CoT2024.05 | 93 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4oCode Integration=No2025.02 | 92.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| NuminaMath-7BBase Model=Qwen2.5-Math-7B, Reasoning Strategy=short-CoT, Parameters=7B2025.03 | 92.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CoT + RandomBase Model=GPT-4-0613, Prompting=CoT + Random2024.05 | 92.87 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VanillaLLM backbone=llama-3.1-70b, MAS=false, Routing=false2026.01 | 92.68 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VanillaLLM backbone=gemini-1.5-flash, MAS=false, Routing=false2026.01 | 92.67 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PROMPTCOT-DS-7BBase Model=DeepSeek-R1-Distill-Qwen-7B, Reasoning Strategy=long-CoT, Parameters=7B2025.03 | 92.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AFlowLLM backbone=gpt-4o-mini, MAS=true, Routing=false2026.01 | 92.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VanillaLLM backbone=claude-3.5-haiku, MAS=false, Routing=false2026.01 | 92.16 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4Training=5-shot ICL2023.08 | 92 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-42023.10 | 92 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4number_of_shots=5-shot, chain-of-thought=true2023.03 | 92 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4Size=-, Reasoning Mode=Chain-of-Thought, Source Type=Closed-Source, Evaluation Protocol=Top12024.02 | 92 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-42023.09 | 92 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Openmathinstruct-7BBase Model=Qwen2.5-Math-7B, Reasoning Strategy=short-CoT, Parameters=7B2025.03 | 92 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DeepSeek-R1-Distill-Qwen-7BReasoning Strategy=long-CoT, Parameters=7B, Reproduced with Paper Prompt=true2025.03 | 91.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BiRMBackbone=Qwen2.5-7B, K=1002025.03 | 91.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-Math-7Bshot=8-shot, prompting=few-shot chain-of-thought2024.09 | 91.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| NuminaMath-72BCode Integration=Yes2025.02 | 91.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Gemma-2 9B2024.06 | 91.28 | — | 2.17 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ORMBackbone=Qwen2.5-7B, K=1002025.03 | 91.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-Math-72Bshot=8-shot, prompting=few-shot chain-of-thought2024.09 | 90.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| FrugalGPTLLM backbone=LLM Pool, MAS=false, Routing=true2026.01 | 90.76 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BiRMBackbone=Qwen2.5-7B, K=202025.03 | 90.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DART-Math-Llama3-70B (Uniform)number of samples=0.59M, Base Model=Llama3-70B, Sampling Strategy=Uniform2024.06 | 90.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ORMBackbone=Qwen2.5-7B, K=202025.03 | 90.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Llama3-70B-VRTnumber of samples=0.59M, Base Model=Llama3-70B2024.06 | 90.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Llama2-70B-Xwin-Math-V1.1+number of samples=1.4M, Base Model=Llama2-70B2024.06 | 90.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ToTModel=GPT-4, Prompting Strategy=Tree-of-Thought2023.05 | 90 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Llama 3Model Size=405B, Few-shot settings=16-shot2024.07 | 90 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| KPDDS-7BBase Model=Qwen2.5-Math-7B, Reasoning Strategy=short-CoT, Parameters=7B2025.03 | 89.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DART-Math-Llama3-70B (Prop2Diff)number of samples=0.59M, Base Model=Llama3-70B, Sampling Strategy=Prop2Diff2024.06 | 89.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2-72Bshot=8-shot, prompting=few-shot chain-of-thought2024.09 | 89.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BiRMBackbone=Qwen2.5-7B, K=82025.03 | 89.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Llama3-70B-MMIQCnumber of samples=2.3M, Base Model=Llama3-70B2024.06 | 89.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AutoCode4Math-DeepSeekCode Integration=Autonomous2025.02 | 89.26 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PRMBackbone=Qwen2.5-7B, K=202025.03 | 89.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AutoCode4Math-Qwen2.5Code Integration=Autonomous2025.02 | 89.12 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2-Math-72Bshot=8-shot, prompting=few-shot chain-of-thought2024.09 | 89.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Llama-3.1-405Bshot=8-shot, prompting=few-shot chain-of-thought2024.09 | 89 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ORMBackbone=Qwen2.5-7B, K=82025.03 | 88.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Evol-Instruct-7BBase Model=Qwen2.5-Math-7B, Reasoning Strategy=short-CoT, Parameters=7B2025.03 | 88.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PRMBackbone=Qwen2.5-7B, K=1002025.03 | 88.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ToRA (SC, K=50)Size=70B, GPT-3.5/4 usage=true, Data Augmentation=true, Code Execution=true, Self-Consistency=true, K=502023.11 | 88.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Baichuan-3Size=-, Reasoning Mode=Chain-of-Thought, Source Type=Closed-Source, Evaluation Protocol=Top12024.02 | 88.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DeepSeekMath-RLSize=7B, Reasoning Mode=Chain-of-Thought, Source Type=Open-Source, Evaluation Protocol=Top12024.02 | 88.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Self-Consistency + MATH-Minos (ORM)Base Model=Mistral-7B: MetaMATH, Verifier=Self-Consistency + MATH-Minos (ORM), Verification outputs=2562024.06 | 88.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DeepSeekMath-7B-VRTnumber of samples=0.59M, Base Model=DeepSeekMath-7B2024.06 | 88.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DART-Math-DSMath-7B (Uniform)number of samples=0.59M, Base Model=DeepSeekMath-7B, Sampling Strategy=Uniform2024.06 | 88.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AutoCode4Math-Qwen2Code Integration=Autonomous2025.02 | 88.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PRMBackbone=Qwen2.5-7B, K=82025.03 | 88.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Llama3-70B-MetaMathnumber of samples=0.40M, Base Model=Llama3-70B2024.06 | 88 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Self-Consistency + MATH-Minos (PRM)Base Model=Mistral-7B: MetaMATH, Verifier=Self-Consistency + MATH-Minos (PRM), Verification outputs=2562024.06 | 87.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Dart-Math-DeepSeek-7BCode Integration=No2025.02 | 87.64 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLM-4Size=-, Reasoning Mode=Chain-of-Thought, Source Type=Closed-Source, Evaluation Protocol=Top12024.02 | 87.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MATH-Minos (PRM)Base Model=Mistral-7B: MetaMATH, Verifier=MATH-Minos (PRM), Verification outputs=2562024.06 | 87.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BiRMBackbone=Qwen2.5-3B, K=1002025.03 | 87.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ORMBackbone=Qwen2.5-3B, K=1002025.03 | 87.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — |