Mathematical Reasoning on MATH (test)
94.13Overall AccuracyIIPC
Evaluation Results
| Method | Links | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| IIPCLLM=Gemini 2.0 Flash2026.02 | 94.13 | — | — | — | — | — | — | — | — | — | — | — | |
| PoTLLM=Gemini 2.0 Flash2026.02 | 92.58 | — | — | — | — | — | — | — | — | — | — | — | |
| IIPCLLM=Llama 4 Maverick2026.02 | 91.23 | — | — | — | — | — | — | — | — | — | — | — | |
| IIPCLLM=Mistral 3.2 24B2026.02 | 90.83 | — | — | — | — | — | — | — | — | — | — | — | |
| IIPCLLM=Gemma 3 27B2026.02 | 90.56 | — | — | — | — | — | — | — | — | — | — | — | |
| CRLLM=Gemini 2.0 Flash2026.02 | 90.09 | — | — | — | — | — | — | — | — | — | — | — | |
| MACMLLM=Gemini 2.0 Flash2026.02 | 90.09 | — | — | — | — | — | — | — | — | — | — | — | |
| CRLLM=Llama 4 Maverick2026.02 | 89.94 | — | — | — | — | — | — | — | — | — | — | — | |
| PoTLLM=Mistral 3.2 24B2026.02 | 89.62 | — | — | — | — | — | — | — | — | — | — | — | |
| PoTLLM=Gemma 3 27B2026.02 | 89.01 | — | — | — | — | — | — | — | — | — | — | — | |
| PoTLLM=Llama 4 Maverick2026.02 | 88.94 | — | — | — | — | — | — | — | — | — | — | — | |
| MACMLLM=Llama 4 Maverick2026.02 | 88.67 | — | — | — | — | — | — | — | — | — | — | — | |
| Ours (theory-guided context selection strategy)Model=Qwen3-8B, Context Selection=Theory-guided2026.02 | 88.4 | — | — | — | — | — | — | — | — | — | — | — | |
| ReMemModel=Qwen3-8B, Context Selection=ReMem2026.02 | 88.2 | — | — | — | — | — | — | — | — | — | — | — | |
| DCModel=Qwen3-8B, Context Selection=DC2026.02 | 88 | — | — | — | — | — | — | — | — | — | — | — | |
| ExpRAGModel=Qwen3-8B, Context Selection=ExpRAG2026.02 | 88 | — | — | — | — | — | — | — | — | — | — | — | |
| MACMModel=GPT-4 Turbo, Prompting Method=Multi-Agent System for Condition Mining2024.04 | 87.92 | 88.46 | 62.74 | 98.04 | — | 94.11 | 96.07 | 78.43 | 97.95 | — | — | — | |
| BM25Model=Qwen3-8B, Context Selection=BM252026.02 | 87.2 | — | — | — | — | — | — | — | — | — | — | — | |
| CRLLM=Gemma 3 27B2026.02 | 87.05 | — | — | — | — | — | — | — | — | — | — | — | |
| ZeroModel=Qwen3-8B, Context Selection=None2026.02 | 87 | — | — | — | — | — | — | — | — | — | — | — | |
| MACMLLM=Gemma 3 27B2026.02 | 86.72 | — | — | — | — | — | — | — | — | — | — | — | |
| CRLLM=Mistral 3.2 24B2026.02 | 83.61 | — | — | — | — | — | — | — | — | — | — | — | |
| MACMLLM=Mistral 3.2 24B2026.02 | 82.13 | — | — | — | — | — | — | — | — | — | — | — | |
| PoTLLM=GPT-4o-mini2026.02 | 81.19 | — | — | — | — | — | — | — | — | — | — | — | |
| IIPCLLM=GPT-4o-mini2026.02 | 80.98 | — | — | — | — | — | — | — | — | — | — | — | |
| SC-CoTModel=GPT-4 Turbo, Prompting Method=Self-Consistency Chain-of-Thought, voters=52024.04 | 80.12 | 79.67 | 50.14 | 89.91 | — | 86.75 | 94.96 | 71.99 | 87.17 | — | — | — | |
| QuaSARModel=GPT-4o2025.02 | 79.5 | — | — | — | — | — | — | — | — | — | — | — | |
| FUNCODERModel=GPT-4, Reasoning Mode=Program-aided2024.05 | 78.2 | 56.4 | 59.5 | 82.2 | — | 89 | 92.8 | 63.3 | 83 | — | — | — | |
| CoTModel=GPT-4o2025.02 | 76.8 | — | — | — | — | — | — | — | — | — | — | — | |
| CRLLM=GPT-4o-mini2026.02 | 76.53 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5 7BPost-train=SFT+RL, Model Type=Autoregressive, Number of shots=4, Evaluation Protocol=RL alignment (†)2026.01 | 75.5 | — | — | — | — | — | — | — | — | — | — | — | |
| CoTModel=GPT-4 Turbo, Prompting Method=Chain-of-Thought2024.04 | 74.36 | 74.18 | 42.02 | 77.31 | — | 82.07 | 92.99 | 68.07 | 83.67 | — | — | — | |
| GPT-4-Turbo (24-04-09)Sampling=N/A2024.06 | 73.4 | — | — | — | — | — | — | — | — | — | — | — | |
| I-OModel=GPT-4 Turbo, Prompting Method=Input-Output2024.04 | 72.78 | 71.15 | 45.11 | 74.51 | — | 81.82 | 88.24 | 66.67 | 81.63 | — | — | — | |
| MACMLLM=GPT-4o-mini2026.02 | 72.62 | — | — | — | — | — | — | — | — | — | — | — | |
| Cumulative Reasoning (CR) w/ codeBackbone=GPT-4, Code Environment=Python, Decoding=Greedy (t=0.0), Prompting=2-shot2023.08 | 72.2 | 51.8 | 53.7 | 88.7 | 71.1 | 86.6 | 86.3 | — | — | — | — | — | |
| Self-RefineModel=GPT-4, Reasoning Mode=Program-aided2024.05 | 72.2 | 63.6 | 54.8 | 77.8 | — | 82.9 | 82 | 55.6 | 76.6 | — | — | — | |
| CRModel=GPT-4, Reasoning Mode=Program-aided2024.05 | 72.2 | 51.8 | 53.7 | 88.7 | — | 86.6 | 86.3 | 51.5 | 71.1 | — | — | — | |
| baselineModel=GPT-4o2025.02 | 70.4 | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4 Code InterpreterSize=-, Reasoning Mode=Tool-Integrated, Source Type=Closed-Source, Evaluation Protocol=Top12024.02 | 69.7 | — | — | — | — | — | — | — | — | — | — | — | |
| BayesFlowOptimizer=Claude-3.5-Sonnet, Executor=Claude-3.7-Sonnet2026.01 | 69.4 | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4-Turbo-0409Few-shot CoT=4-shot2024.05 | 69.2 | — | — | — | — | — | — | — | — | — | — | — | |
| CoTModel=GPT-4, Reasoning Mode=Program-aided2024.05 | 68.6 | 54.5 | 45.2 | 62.2 | — | 84.1 | 87.1 | 48.9 | 68.1 | — | — | — | |
| StandardModel=GPT-4, Reasoning Mode=Text-based2024.05 | 68.2 | 47.3 | 59.5 | 71.1 | — | 81.7 | 82.7 | 46.7 | 72.3 | — | — | — | |
| PoTModel=GPT-4, Reasoning Mode=Program-aided2024.05 | 68.2 | 58.2 | 50 | 75.6 | — | 79.3 | 80.6 | 47.8 | 72.3 | — | — | — | |
| Gemini 1.5 ProSampling=N/A2024.06 | 67.7 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-Math-72Bshot=4-shot, prompting=few-shot chain-of-thought2024.09 | 66.8 | — | — | — | — | — | — | — | — | — | — | — | |
| LEMMABackbone=Qwen2-Math-7B, # Samples=88.90k2025.03 | 62.9 | — | — | — | — | — | — | — | — | — | — | — | |
| VDLMPost-train=SFT+RL, Model Type=Diffusion, Number of shots=4, Evaluation Protocol=same protocol (*)2026.01 | 62.4 | — | — | — | — | — | — | — | — | — | — | — | |
| ToRABackbone=GPT-4, Code Environment=Python, Decoding=Greedy (t=0.0)2023.08 | 61.6 | 37.2 | 44.1 | 68.9 | 67.3 | 82.2 | 75.8 | — | — | — | — | — | |
| ToRABackbone=GPT-4, Code Environment=Python, Decoding=Greedy (t=0.0), Setting=reproduction, Prompting=4-shot2023.08 | 60.8 | 44.6 | 48.8 | 49.5 | 66.1 | 67.1 | 71.8 | — | — | — | — | — | |
| Qwen2-Math-72Bshot=4-shot, prompting=few-shot chain-of-thought2024.09 | 60.5 | — | — | — | — | — | — | — | — | — | — | — | |
| MemGenModel Backbone=Qwen 3 4B Instruct, Inference Strategy=MemGen2026.01 | 60.23 | — | — | — | — | — | — | — | — | — | — | — | |
| Claude-3-OpusSampling=N/A2024.06 | 60.1 | — | — | — | — | — | — | — | — | — | — | — | |
| AFlowOptimizer=Claude-3.5-Sonnet, Executor=Claude-3.7-Sonnet2026.01 | 60.1 | — | — | — | — | — | — | — | — | — | — | — | |
| DeepSeekMath-RLSize=7B, Reasoning Mode=Tool-Integrated, Source Type=Open-Source, Evaluation Protocol=Top12024.02 | 58.8 | — | — | — | — | — | — | — | — | — | — | — | |
| FlashMemModel Backbone=Qwen 3 4B Instruct, Inference Strategy=FlashMem2026.01 | 58.36 | — | — | — | — | — | — | — | — | — | — | — | |
| DeepSeekMath-InstructSize=7B, Reasoning Mode=Tool-Integrated, Source Type=Open-Source, Evaluation Protocol=Top12024.02 | 57.4 | — | — | — | — | — | — | — | — | — | — | — | |
| RefAug-90kBackbone=Qwen2-Math-7B, # Samples=89.92k2025.03 | 56.4 | — | — | — | — | — | — | — | — | — | — | — | |
| DART-Math-Llama3-70B (Prop2Diff)number of samples=0.59M, Base Model=Llama3-70B, Sampling Strategy=Prop2Diff2024.06 | 56.1 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-Math-7Bshot=4-shot, prompting=few-shot chain-of-thought2024.09 | 55.4 | — | — | — | — | — | — | — | — | — | — | — | |
| CoT-SCModel Backbone=Qwen 3 4B Instruct, Inference Strategy=CoT-SC2026.01 | 54.96 | — | — | — | — | — | — | — | — | — | — | — | |
| DART-Math-Llama3-70B (Uniform)number of samples=0.59M, Base Model=Llama3-70B, Sampling Strategy=Uniform2024.06 | 54.9 | — | — | — | — | — | — | — | — | — | — | — | |
| RFTBackbone=Qwen2-Math-7B, # Samples=86.52k2025.03 | 54.4 | — | — | — | — | — | — | — | — | — | — | — | |
| InternLM2-MathSize=20B, Reasoning Mode=Tool-Integrated, Source Type=Open-Source, Evaluation Protocol=Top12024.02 | 54.3 | — | — | — | — | — | — | — | — | — | — | — | |
| VanillaModel Backbone=Qwen 3 4B Instruct, Inference Strategy=Vanilla2026.01 | 54.22 | — | — | — | — | — | — | — | — | — | — | — | |
| FUNCODERModel=GPT-3.5, Reasoning Mode=Program-aided2024.05 | 54 | 41.8 | 34.1 | 55.6 | — | 76.8 | 61.2 | 36 | 59.6 | — | — | — | |
| GPT-4 Complex CoTPHP=true, Decoding Strategy=greedy2023.04 | 53.9 | 29.8 | 41.9 | 55.7 | 56.3 | 73.8 | 74.3 | 26.3 | — | — | — | 2.8494 | |
| CoT + Skill-BasedLLM Backbone=GPT-4-0613, Prompting Strategy=Skill-aligned example selection2024.05 | 53.88 | 33.7 | 41.75 | 51.1 | 58.01 | 74.28 | 73.12 | 27.02 | — | — | — | — | |
| Llama-3.1-405Bshot=4-shot, prompting=few-shot chain-of-thought2024.09 | 53.8 | — | — | — | — | — | — | — | — | — | — | — | |
| DART-Math-DSMath-7B (Prop2Diff)number of samples=0.59M, Base Model=DeepSeekMath-7B, Sampling Strategy=Prop2Diff2024.06 | 53.6 | — | — | — | — | — | — | — | — | — | — | — | |
| GPTAugBackbone=Qwen2-Math-7B, # Samples=88.62k2025.03 | 53.6 | — | — | — | — | — | — | — | — | — | — | — | |
| RefAugBackbone=Qwen2-Math-7B, # Samples=29.94k2025.03 | 53.5 | — | — | — | — | — | — | — | — | — | — | — | |
| Gemini UltraSize=-, Reasoning Mode=Chain-of-Thought, Source Type=Closed-Source, Evaluation Protocol=Top12024.02 | 53.2 | — | — | — | — | — | — | — | — | — | — | — | |
| Llama3-70B-VRTnumber of samples=0.59M, Base Model=Llama3-70B2024.06 | 53.1 | — | — | — | — | — | — | — | — | — | — | — | |
| DeepSeekMath-7B-VRTnumber of samples=0.59M, Base Model=DeepSeekMath-7B2024.06 | 53 | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4Size=-, Reasoning Mode=Chain-of-Thought, Source Type=Closed-Source, Evaluation Protocol=Top12024.02 | 52.9 | — | — | — | — | — | — | — | — | — | — | — | |
| DART-Math-DSMath-7B (Uniform)number of samples=0.59M, Base Model=DeepSeekMath-7B, Sampling Strategy=Uniform2024.06 | 52.9 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2 7BPost-train=SFT+RL, Model Type=Autoregressive, Number of shots=4, Evaluation Protocol=RL alignment (†)2026.01 | 52.9 | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4 (0314)Sampling=N/A2024.06 | 52.6 | — | — | — | — | — | — | — | — | — | — | — | |
| Llama2-70B-Xwin-Math-V1.1+number of samples=1.4M, Base Model=Llama2-70B2024.06 | 52.5 | — | — | — | — | — | — | — | — | — | — | — | |
| PALBackbone=GPT-4, Code Environment=Python, Decoding=Greedy (t=0.0), Setting=reproduction, Prompting=4-shot2023.08 | 52 | 23.2 | 31.7 | 66.1 | 57.9 | 73.2 | 65.3 | — | — | — | — | — | |
| PALBackbone=GPT-4, Code Environment=Python, Decoding=Greedy (t=0.0)2023.08 | 51.8 | 29.3 | 38 | 58.7 | 61 | 73.9 | 59.1 | — | — | — | — | — | |
| MetaMathBackbone=Qwen2-Math-7B, # Samples=394.99k2025.03 | 51.8 | — | — | — | — | — | — | — | — | — | — | — | |
| DeepSeekMath-RLSize=7B, Reasoning Mode=Chain-of-Thought, Source Type=Open-Source, Evaluation Protocol=Top12024.02 | 51.7 | — | — | — | — | — | — | — | — | — | — | — | |
| S3C-MathBackbone=Qwen2-Math-7B, # Samples=927k2025.03 | 51.7 | — | — | — | — | — | — | — | — | — | — | — | |
| DeepSeek-LLM-ChatSize=67B, Reasoning Mode=Tool-Integrated, Source Type=Open-Source, Evaluation Protocol=Top12024.02 | 51.1 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2-72Bshot=4-shot, prompting=few-shot chain-of-thought2024.09 | 51.1 | — | — | — | — | — | — | — | — | — | — | — | |
| SFTBackbone=Qwen2-Math-7B, # Samples=14.97k2025.03 | 50.9 | — | — | — | — | — | — | — | — | — | — | — | |
| ToRASize=34B, Reasoning Mode=Tool-Integrated, Source Type=Open-Source, Evaluation Protocol=Top12024.02 | 50.8 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2-Math-7Bshot=4-shot, prompting=few-shot chain-of-thought2024.09 | 50.4 | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4 Complex CoTPHP=false, Decoding Strategy=greedy2023.04 | 50.36 | 26.7 | 36.5 | 49.6 | 53.1 | 71.6 | 70.8 | 23.4 | — | — | — | — | |
| CoT + Topic-BasedLLM Backbone=GPT-4-0613, Prompting Strategy=Topic-based retrieval2024.05 | 50.31 | 31.13 | 39.45 | 47.03 | 54.64 | 71.16 | 67.9 | 24.14 | — | — | — | — | |
| MinervaModel Parameters=540B, Majority Voting (maj1@k)=true, k samples=642022.06 | 50.3 | — | — | — | — | — | — | — | — | — | — | — | |
| Complex CoTLLM Backbone=GPT-4-0613, Prompting Strategy=Complex Chain-of-Thought2024.05 | 50.3 | 26.7 | 36.5 | 49.6 | 53.1 | 71.6 | 70.8 | 23.4 | — | — | — | — | |
| Previous SOTAPHP=false2023.04 | 50.3 | — | — | — | — | — | — | — | — | — | — | — | |
| FlashMemModel Backbone=Qwen 2.5 1.5B Instruct, Inference Strategy=FlashMem2026.01 | 50.16 | — | — | — | — | — | — | — | — | — | — | — | |
| SnapKVModel Backbone=Qwen 3 4B Instruct, Inference Strategy=SnapKV2026.01 | 50.12 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-Math-1.5Bshot=4-shot, prompting=few-shot chain-of-thought2024.09 | 49.8 | — | — | — | — | — | — | — | — | — | — | — | |
| Llama3-70B-MMIQCnumber of samples=2.3M, Base Model=Llama3-70B2024.06 | 49.4 | — | — | — | — | — | — | — | — | — | — | — |