Mathematical Reasoning on MATH 500 (Accuracy and Length)
91.2AccuracyVanilla
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| VanillaModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 91.2 | 3,673 | |
| O1-PrunerModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 91.2 | 3,422 | |
| TLMREModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 91.2 | 2,419 | |
| AdaptThinkModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 91 | 1,862 | |
| LASERModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 91 | 1,871 | |
| ARLCPModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 91 | 2,086 | |
| PIRBackbone=DeepSeek-R1-32B2026.02 | 91 | 1,549.96 | |
| DPOShortestModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 90.8 | 3,340 | |
| PIRBackbone=QwQ-32B2026.02 | 90 | 1,557.84 | |
| SFTShortestModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 89 | 3,575 | |
| PIRBackbone=DeepSeek-R1-7B2026.02 | 88.8 | 2,055.53 | |
| PIRBackbone=DeepSeek-R1-14B2026.02 | 88.4 | 1,842.08 | |
| SOUPratioBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Truncation=Length-ratio-based2026.01 | 88.25 | — | |
| On-policyBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.01 | 88.08 | — | |
| QwQ-32BBackbone=QwQ-32B2026.02 | 88 | 3,318.15 | |
| M2POBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.01 | 87.28 | — | |
| DeepSeek-R1-7BBackbone=DeepSeek-R1-7B2026.02 | 87 | 2,576.49 | |
| DeepSeek-R1-32BBackbone=DeepSeek-R1-32B2026.02 | 87 | 2,021.68 | |
| DeepSeek-R1-14BBackbone=DeepSeek-R1-14B2026.02 | 86 | 2,408.36 | |
| LASERModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 84.6 | 2,365 | |
| ARLCPModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 84.6 | 1,812 | |
| DPOShortestModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 84.2 | 4,185 | |
| SOUPentropyBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Truncation=Entropy-based2026.01 | 83.59 | — | |
| LUFFYBackbone=Qwen2.5-Math-7B2026.01 | 83.48 | — | |
| On-policyBackbone=Qwen2.5-Math-7B2026.01 | 83.33 | — | |
| SOUPratioBackbone=Qwen2.5-Math-7B, Truncation=Length-ratio-based2026.01 | 83.25 | — | |
| AdaptThinkModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 83.2 | 1,869 | |
| SOUPentropyBackbone=Qwen2.5-Math-7B, Truncation=Entropy-based2026.01 | 82.89 | — | |
| O1-PrunerModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 82.6 | 4,375 | |
| LUFFYBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.01 | 82.38 | — | |
| TLMREModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 82.1 | 1,800 | |
| SFTShortestModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 81 | 4,531 | |
| NoThinkingModel Backbone=DeepSeek-R1-Distill-Qwen-7B2026.02 | 80.6 | 706 | |
| VanillaModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 80.2 | 4,582 | |
| M2POBackbone=Qwen2.5-Math-7B2026.01 | 79.23 | — | |
| NoThinkingModel Backbone=DeepSeek-R1-Distill-Qwen-1.5B2026.02 | 69.2 | 908 | |
| Initial policyBackbone=Qwen2.5-Math-7B2026.01 | 60.92 | — | |
| Info-GainModel=SDAR-8B-Chat, K (Decoding Rate)=12026.02 | 59.6 | — | |
| Info-GainModel=SDAR-8B-Chat, K (Decoding Rate)=22026.02 | 54.6 | — | |
| Info-GainModel=TraDo-8B-Instruct, K (Decoding Rate)=12026.02 | 45.9 | — | |
| ConfidenceModel=TraDo-8B-Instruct, K (Decoding Rate)=12026.02 | 44.2 | — | |
| Info-GainModel=TraDo-8B-Instruct, K (Decoding Rate)=22026.02 | 40.9 | — | |
| ConfidenceModel=SDAR-8B-Chat, K (Decoding Rate)=12026.02 | 40.6 | — | |
| LLaDA-8B-Instruct + DiffuGRPOModel Category=Language-Only dLLMs, per-dataset finetuning=true2026.02 | 40.2 | — | |
| Dream-7BModel Category=Language-Only dLLMs2026.02 | 39.6 | — | |
| ConfidenceModel=TraDo-8B-Instruct, K (Decoding Rate)=22026.02 | 39.2 | — | |
| LaViDa-R1Model Category=Unified-Understanding-and-Generation dLLMs2026.02 | 38.6 | — | |
| KLASSModel=SDAR-8B-Chat, K (Decoding Rate)=12026.02 | 38.3 | — | |
| KLASSModel=TraDo-8B-Instruct, K (Decoding Rate)=12026.02 | 37.3 | — | |
| LLaDA-8B-Instruct w/ D1Seq Len=2562026.02 | 37.2 | — | |
| ConfidenceModel=SDAR-8B-Chat, K (Decoding Rate)=22026.02 | 36.6 | — | |
| LLaDA-8B-Instruct w/ DAMSeq Len=2562026.02 | 36.49 | — | |
| LLaDA-8B-InstructModel Category=Language-Only dLLMs2026.02 | 36.2 | — | |
| MMaDa-8B-Base +UniGRPOModel Category=Unified-Understanding-and-Generation dLLMs, RL checkpoint not open-sourced=true2026.02 | 36 | — | |
| EntropyModel=SDAR-8B-Chat, K (Decoding Rate)=12026.02 | 35.4 | — | |
| LLaDA-8B-Instruct w/ DAMSeq Len=1282026.02 | 32.6 | — | |
| MarginModel=SDAR-8B-Chat, K (Decoding Rate)=12026.02 | 32.4 | — | |
| KLASSModel=SDAR-8B-Chat, K (Decoding Rate)=22026.02 | 32.3 | — | |
| KLASSModel=TraDo-8B-Instruct, K (Decoding Rate)=22026.02 | 32.3 | — | |
| LLaDA-8B-Instruct w/ D1Seq Len=1282026.02 | 31.2 | — | |
| LaViDa-O +SFTModel Category=Unified-Understanding-and-Generation dLLMs2026.02 | 31 | — | |
| LLaDA-8B-InstructSeq Len=2562026.02 | 30.8 | — | |
| LLaDA-8B-InstructSeq Len=1282026.02 | 28.8 | — | |
| Initial policyBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.01 | 27.56 | — | |
| MMaDa-8B-Base +CoT SFTModel Category=Unified-Understanding-and-Generation dLLMs2026.02 | 26.5 | — | |
| EntropyModel=SDAR-8B-Chat, K (Decoding Rate)=22026.02 | 24.4 | — | |
| LaViDa-OModel Category=Unified-Understanding-and-Generation dLLMs2026.02 | 23.4 | — | |
| MarginModel=SDAR-8B-Chat, K (Decoding Rate)=22026.02 | 22.4 | — | |
| EntropyModel=TraDo-8B-Instruct, K (Decoding Rate)=12026.02 | 22 | — | |
| MarginModel=TraDo-8B-Instruct, K (Decoding Rate)=12026.02 | 22 | — | |
| EntropyModel=TraDo-8B-Instruct, K (Decoding Rate)=22026.02 | 17 | — | |
| MarginModel=TraDo-8B-Instruct, K (Decoding Rate)=22026.02 | 17 | — | |
| MMaDa-8B-BaseModel Category=Unified-Understanding-and-Generation dLLMs2026.02 | 4.2 | — |