Mathematical Problem Solving on MATH
97.6AccuracyDeepSeek-R1
Evaluation Results
| Method | Links | |
|---|---|---|
| DeepSeek-R1Model Category=Closed-source Models2026.01 | 97.6 | |
| LoGos(32B)Model Category=Open-source Models (32B), Parameter Count=32B2026.01 | 96.5 | |
| Qwen2.5-32B-InstructModel Category=Open-source Models (32B), Parameter Count=32B, Model Variant=Instruct2026.01 | 95.2 | |
| o1-miniModel Category=Closed-source Models2026.01 | 95 | |
| DeepSeek-R1-Distill-Qwen-32BModel Category=Open-source Models (32B), Parameter Count=32B, Model Variant=Distill2026.01 | 94.5 | |
| LoGos(7B)Model Category=Open-source Models (7B), Parameter Count=7B2026.01 | 93.2 | |
| Qwen2.5-7B-InstructModel Category=Open-source Models (7B), Parameter Count=7B, Model Variant=Instruct2026.01 | 92.6 | |
| Qwen2.5-32B-BaseModel Category=Open-source Models (32B), Parameter Count=32B, Model Variant=Base2026.01 | 90.7 | |
| DeepSeek-R1-Distill-Qwen-7BModel Category=Open-source Models (7B), Parameter Count=7B, Model Variant=Distill2026.01 | 88.2 | |
| BaseBackbone=Qwen3-4B2025.05 | 84 | |
| Qwen3-8BEvaluation Protocol=Zero-shot, TPS=30.112025.12 | 83.28 | |
| Qwen2.5-7B-BaseModel Category=Open-source Models (7B), Parameter Count=7B, Model Variant=Base2026.01 | 83.2 | |
| Qwen2.5-72BModel Type=Instruct, Parameters=72B2025.02 | 83.1 | |
| Qwen2.5-VL-72BModel Type=Instruct, Parameters=72B2025.02 | 83 | |
| Claude3.7-SonnetModel Category=Closed-source Models2026.01 | 79.8 | |
| Primitives-based MASModel=DeepSeek-R1-Distill Qwen-32B2026.02 | 79.8 | |
| M2CLModel=Qwen-72B, Number of LLMs=642026.02 | 79.7 | |
| Primitives-based MASModel=DeepSeek-R1-Distill Llama-70B2026.02 | 79.3 | |
| SHE-LoRABackbone=Qwen3-4B2025.05 | 78.86 | |
| LatentMASModel=DeepSeek-R1-Distill Llama-70B2026.02 | 78.6 | |
| LatentMASModel=DeepSeek-R1-Distill Qwen-32B2026.02 | 78.2 | |
| Vanilla LoRABackbone=Qwen3-4B2025.05 | 77.63 | |
| VotingModel=DeepSeek-R1-Distill Llama-70B2026.02 | 77.2 | |
| ReviewModel=DeepSeek-R1-Distill Llama-70B2026.02 | 75.9 | |
| PlanningModel=DeepSeek-R1-Distill Llama-70B2026.02 | 75.6 | |
| VotingModel=DeepSeek-R1-Distill Qwen-32B2026.02 | 74.7 | |
| PlanningModel=DeepSeek-R1-Distill Qwen-32B2026.02 | 74.6 | |
| MacNetModel=Qwen-72B, Number of LLMs=642026.02 | 73.9 | |
| Llama-3.1-405BModel Type=Instruct, Parameters=405B2025.02 | 73.8 | |
| TextMASModel=DeepSeek-R1-Distill Llama-70B2026.02 | 72.8 | |
| ReviewModel=DeepSeek-R1-Distill Qwen-32B2026.02 | 71.3 | |
| SingleModel=DeepSeek-R1-Distill Llama-70B2026.02 | 69.6 | |
| Qwen2-72BModel Type=Instruct, Parameters=72B2025.02 | 69 | |
| GPTSwarmModel=Qwen-72B, Number of LLMs=642026.02 | 68.8 | |
| Llama-3.1-70BModel Type=Instruct, Parameters=70B2025.02 | 68 | |
| TextMASModel=DeepSeek-R1-Distill Qwen-32B2026.02 | 68 | |
| DyLANModel=Qwen-72B, Number of LLMs=642026.02 | 66.8 | |
| GUIDEDSAMPLINGModel=Qwen2.5-3B-Instruct, Decoding=Majority Voting2025.10 | 64.2 | |
| Primitives-based MASModel=Qwen3-8B2026.02 | 63.7 | |
| SingleModel=DeepSeek-R1-Distill Qwen-32B2026.02 | 63.4 | |
| LatentMASModel=Qwen3-8B2026.02 | 62.6 | |
| DebateModel=Qwen-72B, Number of LLMs=642026.02 | 62.3 | |
| M2CLModel=Qwen-14B, Number of LLMs=642026.02 | 61.9 | |
| TextMASModel=Qwen3-8B2026.02 | 61.4 | |
| VotingModel=Qwen3-8B2026.02 | 61.4 | |
| ReviewModel=Qwen3-8B2026.02 | 61 | |
| SingleModel=Qwen3-8B2026.02 | 60.8 | |
| PlanningModel=Qwen3-8B2026.02 | 60.8 | |
| M2CLModel=Llama-70B, Number of LLMs=642026.02 | 58 | |
| CoFiCotBase Model=GPT-3.5-Turbo2026.03 | 57.7 | |
| Qwen-2.5-72B-Instruct-DistilledBase Model=Qwen-2.5-7B, Seed Dataset=WizardLM2025.04 | 56.3 | |
| Qwen-2.5-32B-Instruct-DistilledBase Model=Qwen-2.5-7B, Seed Dataset=Alpaca2025.04 | 56.26 | |
| Qwen-2.5-32B-Instruct-DistilledBase Model=Qwen-2.5-7B, Seed Dataset=Condor2025.04 | 56 | |
| Qwen-2.5-72B-Instruct-DistilledBase Model=Qwen-2.5-7B, Seed Dataset=Alpaca2025.04 | 55.8 | |
| Qwen-2.5-32B-Instruct-DistilledBase Model=Qwen-2.5-7B, Seed Dataset=WizardLM2025.04 | 54.96 | |
| Qwen-2.5-72B-Instruct-DistilledBase Model=Qwen-2.5-7B, Seed Dataset=Condor2025.04 | 54.46 | |
| REFUSIONEvaluation Protocol=Zero-shot, TPS=81.772025.12 | 54.22 | |
| DART-Math-DSMath-7BTraining Method=SFT, Data Selection Strategy=Prop2Diff, Backbone=DeepSeekMath-7B2024.06 | 53.6 | |
| DeepSeekMath-7B-RLTraining Method=RL, Backbone=DeepSeekMath-7B2024.06 | 53.1 | |
| DART-Math-DSMath-7BTraining Method=SFT, Data Selection Strategy=Uniform, Backbone=DeepSeekMath-7B2024.06 | 52.9 | |
| GRPO (RLPR)#Data=10K2026.01 | 52.7 | |
| GRPO (RLVRR)#Data=10K2026.01 | 52.6 | |
| Qwen3 8BCompression Ratio=dense2026.02 | 52.57 | |
| GRPO (RM)#Data=10K2026.01 | 52.4 | |
| Instruct#Data=n/a2026.01 | 51.9 | |
| SFT#Data=100K2026.01 | 51.7 | |
| GRPO (GRM)#Data=10K2026.01 | 51.2 | |
| Repeated SamplingModel=Qwen2.5-3B-Instruct, Decoding=Majority Voting2025.10 | 51.2 | |
| k-way SCBase Model=GPT-3.5-Turbo, k=1202026.03 | 51.2 | |
| BoNModel=Qwen-72B, Number of LLMs=642026.02 | 51 | |
| DPO#Data=10K2026.01 | 51 | |
| M2CLModel=Qwen-7B, Number of LLMs=642026.02 | 50.7 | |
| SFT#Data=10K2026.01 | 50.6 | |
| Best-of-kBase Model=GPT-3.5-Turbo, k=1202026.03 | 50.6 | |
| Repeated SamplingModel=Llama-3.2-3B-Instruct, Decoding=Majority Voting2025.10 | 50.4 | |
| GRPO (BLEU)#Data=10K2026.01 | 50.2 | |
| SDAR-4B-ChatModel=SDAR-4B-Chat, Decoding Strategy=FourierSampler2026.01 | 50 | |
| Self-Refine + k-way SCBase Model=GPT-3.5-Turbo, k=1202026.03 | 49.8 | |
| SDAR-4B-ChatModel=SDAR-4B-Chat, Decoding Strategy=RWS2026.01 | 49.2 | |
| CondorBase Model=Qwen-2.5-7B, Seed Dataset=Condor2025.04 | 48.6 | |
| SDAR-4B-ChatModel=SDAR-4B-Chat, Decoding Strategy=Vanilla2026.01 | 48.2 | |
| CoFiCotBase Model=Llama3-8B-Instruct2026.03 | 47.9 | |
| GRABase Model=Qwen-2.5-7B, Seed Dataset=WizardLM2025.04 | 47.84 | |
| Dream-7B-InstructEvaluation Protocol=Zero-shot, TPS=18.992025.12 | 46.6 | |
| Tree-of-thoughtModel=Llama-3.2-3B-Instruct, Decoding=Majority Voting2025.10 | 45.8 | |
| Tree-of-thoughtModel=Qwen2.5-3B-Instruct, Decoding=Majority Voting2025.10 | 45.4 | |
| MacNetModel=Llama-70B, Number of LLMs=642026.02 | 44.2 | |
| GUIDEDSAMPLINGModel=Llama-3.2-3B-Instruct, Decoding=Majority Voting2025.10 | 43.4 | |
| GRABase Model=Qwen-2.5-7B, Seed Dataset=Condor2025.04 | 42.82 | |
| Best-of-kBase Model=Llama3-8B-Instruct, k=1202026.03 | 41.4 | |
| WavefrontDiffusionModel=LLaDA-8B-Instruct, Denoising steps (T)=10242025.11 | 41.04 | |
| SDAR-1.7B-ChatModel=SDAR-1.7B-Chat, Decoding Strategy=RWS2026.01 | 41 | |
| Truncated-BlockDiffusion (TBD)Model=LLaDA-8B-Instruct, Denoising steps (T)=10242025.11 | 40.84 | |
| GRPO (Random)#Data=10K2026.01 | 40.7 | |
| BlockDiffusionModel=LLaDA-8B-Instruct, Denoising steps (T)=10242025.11 | 40.62 | |
| k-way SCBase Model=Llama3-8B-Instruct, k=1202026.03 | 40.6 | |
| Self-Refine + k-way SCBase Model=Llama3-8B-Instruct, k=1202026.03 | 40.4 | |
| VanillaModel=Dream-Base, Block Size (B0)=322025.09 | 40.2 | |
| VanillaModel=Dream-Base, Block Size (B0)=642025.09 | 40.1 | |
| Running Confidence Remasking (RCR)Model=LLaDA-8B-Instruct, Denoising steps (T)=10242025.11 | 40.02 |