Mathematical Reasoning on MATH (test) (Pass@1)
94.8Pass@1GPT-o1
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-o1Inference Mode=Chain of Thought (CoT), Use External Tools=False2023.08 | 94.8 | |
| GPT-o1Decoding Strategy=Greedy, Evaluation Protocol=CoT without python tool2023.08 | 94.8 | |
| GPT-o1-miniInference Mode=Chain of Thought (CoT), Use External Tools=False2023.08 | 90 | |
| GPT-o1-miniDecoding Strategy=Greedy, Evaluation Protocol=CoT without python tool2023.08 | 90 | |
| Gemini-1.5 002Inference Mode=Chain of Thought (CoT), Use External Tools=False2023.08 | 86.5 | |
| Gemini-1.5 002Decoding Strategy=Greedy, Evaluation Protocol=CoT without python tool2023.08 | 86.5 | |
| WizardMath-QwenBase=Qwen2.5-Math, Params=7B2023.08 | 77.8 | |
| GPT-4o-2024-0513Inference Mode=Chain of Thought (CoT), Use External Tools=False2023.08 | 76.6 | |
| GPT-4o-2024-0513Decoding Strategy=Greedy, Evaluation Protocol=CoT without python tool2023.08 | 76.6 | |
| WizardMath-QwenBase=Qwen2.5, Params=7B2023.08 | 74.5 | |
| TRE-KBackbone=Qwen2.5-7B-Instruct2026.02 | 74.29 | |
| TRE-PBackbone=Qwen2.5-7B-Instruct2026.02 | 74.27 | |
| KL-CovBackbone=Qwen2.5-7B-Instruct2026.02 | 73.79 | |
| EntBackbone=Qwen2.5-7B-Instruct2026.02 | 73.66 | |
| Forking-TokensBackbone=Qwen2.5-7B-Instruct2026.02 | 73.59 | |
| Vanilla (PPO)Backbone=Qwen2.5-7B-Instruct2026.02 | 73.52 | |
| FullBackbone=Q-7B, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 73.2 | |
| FullBackbone=Q-7B, Setting=Online, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 73.2 | |
| r_pb_onlineBackbone=Q-7B, Setting=Online, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 72.9 | |
| PPLBackbone=Q-7B, Setting=Online, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 72.1 | |
| r_pbBackbone=Q-7B, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 72.1 | |
| ENTBackbone=Q-7B, Setting=Online, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 71.8 | |
| BaseBackbone=Qwen2.5-7B-Instruct2026.02 | 71.42 | |
| Claude 3.5 SonnetInference Mode=Chain of Thought (CoT), Use External Tools=False2023.08 | 71.1 | |
| Claude 3.5 SonnetDecoding Strategy=Greedy, Evaluation Protocol=CoT without python tool2023.08 | 71.1 | |
| PPLBackbone=Q-7B, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 71 | |
| WizardMath-MathstralBase=Mathstral-v0.1, Params=7B2023.08 | 70.9 | |
| RandomBackbone=Q-7B, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 70.8 | |
| K-centerBackbone=Q-7B, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 70.5 | |
| ENTBackbone=Q-7B, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 70.3 | |
| AskLLMBackbone=Q-7B, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 69.8 | |
| WizardMath-Qwen-RLBase=Qwen-Math-2.5, Parameters=1.5B2023.08 | 68.6 | |
| WizardMath-QwenBase=Qwen-Math-2.5, Params=1.5B2023.08 | 68.6 | |
| RandomBackbone=Q-7B, Setting=Online, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 68.2 | |
| Active PromptBackbone=Q-7B, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 65.1 | |
| WizardMath-DeepSeekBase=DeepSeekMath, Params=7B2023.08 | 64.6 | |
| GPT-4-turbo-0125Inference Mode=Chain of Thought (CoT), Use External Tools=False2023.08 | 64.5 | |
| GPT-4-turbo-0125Decoding Strategy=Greedy, Evaluation Protocol=CoT without python tool2023.08 | 64.5 | |
| ZS-CALBackbone=Q-7B, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 64 | |
| r_pb_onlineBackbone=Q-3B, Setting=Online, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 64 | |
| FullBackbone=Q-3B, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 63.8 | |
| FullBackbone=Q-3B, Setting=Online, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 63.8 | |
| r_pbBackbone=Q-3B, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 63.3 | |
| ENTBackbone=Q-3B, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 62.9 | |
| ENTBackbone=Q-3B, Setting=Online, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 62.8 | |
| PPLBackbone=Q-3B, Setting=Online, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 62.5 | |
| K-centerBackbone=Q-3B, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 62.5 | |
| RandomBackbone=Q-3B, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 62.2 | |
| WizardMath-Qwen-SFTBase=Qwen-Math-2.5, Parameters=1.5B2023.08 | 62.1 | |
| AskLLMBackbone=Q-3B, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 61.9 | |
| PPLBackbone=Q-3B, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 61.6 | |
| WizardMath-LlamaBase=Llama 3, Params=8B2023.08 | 58.8 | |
| RandomBackbone=Q-3B, Setting=Online, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 58.8 | |
| WizardMath-LlamaBase=Llama-2, Params=70B2023.08 | 58.6 | |
| TRE-PBackbone=Qwen2.5-1.5B-Instruct2026.02 | 58.28 | |
| TRE-KBackbone=Qwen2.5-1.5B-Instruct2026.02 | 58.26 | |
| KL-CovBackbone=Qwen2.5-1.5B-Instruct2026.02 | 58.23 | |
| Forking-TokensBackbone=Qwen2.5-1.5B-Instruct2026.02 | 57.16 | |
| Vanilla (PPO)Backbone=Qwen2.5-1.5B-Instruct2026.02 | 57.04 | |
| EntBackbone=Qwen2.5-1.5B-Instruct2026.02 | 56.64 | |
| WizardMath-MistralBase=Mistral-v0.3, Params=7B2023.08 | 55.6 | |
| WizardMath-MistralBase=Mistral-v0.1, Params=7B2023.08 | 55.4 | |
| ZS-CALBackbone=Q-3B, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 54.5 | |
| Active PromptBackbone=Q-3B, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 54.1 | |
| GPT-4-0314Inference Mode=Chain of Thought (CoT), Use External Tools=False2023.08 | 52.6 | |
| GPT-4-0314Decoding Strategy=Greedy, Evaluation Protocol=CoT without python tool2023.08 | 52.6 | |
| r_pb_onlineBackbone=L-8B-I, Setting=Online, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 52.5 | |
| FullBackbone=L-8B-I, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 52 | |
| FullBackbone=L-8B-I, Setting=Online, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 52 | |
| LEMMA (w/ MetaMath)Backbone=DeepSeekMath-7B, # Samples=403.59k2025.03 | 51.7 | |
| PPLBackbone=L-8B-I, Setting=Online, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 51.6 | |
| r_pbBackbone=L-8B-I, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 51.5 | |
| ENTBackbone=L-8B-I, Setting=Online, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 51.3 | |
| BaseBackbone=Qwen2.5-1.5B-Instruct2026.02 | 51.06 | |
| K-centerBackbone=L-8B-I, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 51 | |
| RandomBackbone=L-8B-I, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 50.9 | |
| ENTBackbone=L-8B-I, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 50.7 | |
| LEMMABackbone=DeepSeekMath-7B, # Samples=88.90k2025.03 | 50.6 | |
| WizardMath-LlamaBase=Llama 2, Params=13B2023.08 | 50.6 | |
| RandomBackbone=L-8B-I, Setting=Online, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 50.6 | |
| AskLLMBackbone=L-8B-I, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 50.5 | |
| WizardMath-Llama-RLBase=Llama 3.2, Parameters=3B2023.08 | 49.9 | |
| WizardMath-LlamaBase=Llama 3.2, Params=3B2023.08 | 49.9 | |
| PPLBackbone=L-8B-I, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 49.8 | |
| Active PromptBackbone=L-8B-I, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 49.5 | |
| Baichuan-3Inference Mode=Chain of Thought (CoT), Use External Tools=False2023.08 | 49.2 | |
| Baichuan-3Decoding Strategy=Greedy, Evaluation Protocol=CoT without python tool2023.08 | 49.2 | |
| ZS-CALBackbone=L-8B-I, Setting=Offline, Sample size (p)=30, Decoding Strategy=greedy2026.01 | 48.6 | |
| Llama-3.2-InstructBase=Llama 3.2, Parameters=3B2023.08 | 48 | |
| Llama-3.2-InstructBase=Llama 3.2, Params=3B2023.08 | 48 | |
| GLM-4Inference Mode=Chain of Thought (CoT), Use External Tools=False2023.08 | 47.9 | |
| GLM-4Decoding Strategy=Greedy, Evaluation Protocol=CoT without python tool2023.08 | 47.9 | |
| GPTAugBackbone=DeepSeekMath-7B, # Samples=88.62k2025.03 | 45.5 | |
| WizardMath-Llama-SFTBase=Llama 3.2, Parameters=3B2023.08 | 45.2 | |
| WizardMath-LlamaBase=Llama-2, Params=7B2023.08 | 43.5 | |
| GPT-3.5-TurboInference Mode=Chain of Thought (CoT), Use External Tools=False2023.08 | 43.1 | |
| GPT-3.5-TurboDecoding Strategy=Greedy, Evaluation Protocol=CoT without python tool2023.08 | 43.1 | |
| RefAug-90kBackbone=DeepSeekMath-7B, # Samples=89.92k2025.03 | 42.5 | |
| GPT-4 (original version)Inference Mode=Chain of Thought (CoT), Use External Tools=False2023.08 | 42.5 | |
| GPT-4 (original version)Decoding Strategy=Greedy, Evaluation Protocol=CoT without python tool2023.08 | 42.5 |