Mathematical Reasoning on College
48.1AccuracyDPO-R1 (ZHANG ET AL., 2025)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| DPO-R1 (ZHANG ET AL., 2025)Backbone=Qwen3-4B-Instruct, Compute Cost (Sampling Budget N)=N = 8, Trainable Pairs=8,403, Processing Time=2.82026.02 | 48.1 | — | — | |
| IDPOModel=Qwen2.5-7B-Math-SFT2025.05 | 47.9 | — | — | |
| DPO-R1 (MIDDLE)Backbone=Qwen3-4B-Instruct, Compute Cost (Sampling Budget N)=N = 4, Trainable Pairs=5,127, Processing Time=1.22026.02 | 47.6 | — | — | |
| DPO-R1 (HIGH)Backbone=Qwen3-4B-Instruct, Compute Cost (Sampling Budget N)=N = 16, Trainable Pairs=8,964, Processing Time=5.22026.02 | 47.6 | — | — | |
| PACEBackbone=Qwen3-4B-Instruct, Compute Cost (Sampling Budget N)=2 < N < 3, Trainable Pairs=10,717, Processing Time=1.02026.02 | 47.4 | — | — | |
| Base ModelModel=Qwen2.5-7B-Math-SFT2025.05 | 47.4 | — | — | |
| SAI-DPOModel=Qwen2.5-7B-Math-SFT2025.05 | 47.3 | — | — | |
| DPO-R1 (ZHANG ET AL., 2025)Backbone=Qwen3-8B-Instruct, Compute Cost (Sampling Budget N)=N = 8, Trainable Pairs=10,211, Processing Time=4.82026.02 | 47.1 | — | — | |
| DPO-R1 (LOW)Backbone=Qwen3-4B-Instruct, Compute Cost (Sampling Budget N)=N = 2, Trainable Pairs=3,479, Processing Time=0.92026.02 | 46.8 | — | — | |
| PACEBackbone=Qwen3-8B-Instruct, Compute Cost (Sampling Budget N)=2 < N < 3, Trainable Pairs=9,028, Processing Time=1.72026.02 | 46.8 | — | — | |
| DPO-R1 (HIGH)Backbone=Qwen3-8B-Instruct, Compute Cost (Sampling Budget N)=N = 16, Trainable Pairs=10,543, Processing Time=7.22026.02 | 46.7 | — | — | |
| DPO-R1 (MIDDLE)Backbone=Qwen3-8B-Instruct, Compute Cost (Sampling Budget N)=N = 4, Trainable Pairs=6,073, Processing Time=2.22026.02 | 46.5 | — | — | |
| SAI-DPOModel=Qwen2.5-7B-Math-Base2025.05 | 46.3 | — | — | |
| IDPOModel=Qwen3-8B2025.05 | 46.1 | — | — | |
| SAI-DPOModel=Qwen3-8B2025.05 | 46.1 | — | — | |
| IDPOModel=Qwen2.5-7B-Math-Base2025.05 | 45.9 | — | — | |
| DPO-R1 (LOW)Backbone=Qwen3-8B-Instruct, Compute Cost (Sampling Budget N)=N = 2, Trainable Pairs=3,403, Processing Time=1.42026.02 | 45.8 | — | — | |
| SAGEBase Model=Qwen2.5-3B-Instruct2026.02 | 45.14 | — | — | |
| DPO (Random)Base Model=Qwen2.5-3B-Instruct2026.02 | 45 | — | — | |
| DPO (Full)Base Model=Qwen2.5-3B-Instruct2026.02 | 44.9 | — | — | |
| INSTRUCT (BASE)Backbone=Qwen3-4B-Instruct, Compute Cost (Sampling Budget N)=N/A, Trainable Pairs=N/A, Processing Time=N/A2026.02 | 44.9 | — | — | |
| Base ModelModel=Qwen3-8B2025.05 | 44.9 | — | — | |
| INSTRUCT (BASE)Backbone=Qwen3-8B-Instruct, Compute Cost (Sampling Budget N)=N/A, Trainable Pairs=N/A, Processing Time=N/A2026.02 | 44.6 | — | — | |
| VanillaBase Model=Qwen2.5-3B-Instruct2026.02 | 44.5 | — | — | |
| SAGEBase Model=Qwen2.5-7B-Instruct2026.02 | 43.1 | — | — | |
| DPO (Full)Base Model=Qwen2.5-7B-Instruct2026.02 | 42.7 | — | — | |
| DPO (Random)Base Model=Qwen2.5-7B-Instruct2026.02 | 42.7 | — | — | |
| VanillaBase Model=Qwen2.5-7B-Instruct2026.02 | 42.4 | — | — | |
| Base ModelModel=Qwen2.5-7B-Math-Base2025.05 | 41.3 | — | — | |
| VanillaBase Model=Qwen2.5-1.5B-Instruct2026.02 | 38.4 | — | — | |
| SAGEBase Model=Qwen2.5-1.5B-Instruct2026.02 | 38.1 | — | — | |
| DPO (Full)Base Model=Qwen2.5-1.5B-Instruct2026.02 | 38 | — | — | |
| DPO (Random)Base Model=Qwen2.5-1.5B-Instruct2026.02 | 37.9 | — | — | |
| MMIQCBase Model=LLaMA3-8B, #Samples=2.3M2026.04 | 29.5 | — | — | |
| CRPSBase Model=LLaMA3-8B, #Samples=60K2026.04 | 29.4 | — | — | |
| DARTBase Model=Mistral-7B-v0.1, #Samples=590K2026.04 | 29.4 | — | — | |
| DARTBase Model=LLaMA3-8B, #Samples=590K2026.04 | 28.8 | — | — | |
| SIGMABase Model=LLaMA3-8B, #Samples=60K2026.04 | 28.1 | — | — | |
| CRPSBase Model=LLaMA3-8B, #Samples=30K2026.04 | 27.9 | — | — | |
| DARTBase Model=LLaMA3-8B, #Samples=60K2026.04 | 27.9 | — | — | |
| MathFusionBase Model=LLaMA3-8B, #Samples=60K2026.04 | 27.9 | — | — | |
| SIGMABase Model=LLaMA3-8B, #Samples=30K2026.04 | 26.3 | — | — | |
| CRPSBase Model=Mistral-7B-v0.1, #Samples=60K2026.04 | 26.1 | — | — | |
| CRPSBase Model=LLaMA3-8B, #Samples=15K2026.04 | 25.8 | — | — | |
| MathFusion (Sequential)Base Model=LLaMA3-8B, #Samples=30K2026.04 | 25.1 | — | — | |
| SAI-DPOModel=Llama3.1-8B-Instruct2025.05 | 25 | — | — | |
| PACEBackbone=Llama-3.1-8B-Instruct, Compute Cost (Sampling Budget N)=2 < N < 3, Trainable Pairs=6,197, Processing Time=0.82026.02 | 24.9 | — | — | |
| DPO-R1 (LOW)Backbone=Llama-3.1-8B-Instruct, Compute Cost (Sampling Budget N)=N = 2, Trainable Pairs=5,246, Processing Time=0.82026.02 | 24.7 | — | — | |
| CRPSBase Model=Mistral-7B-v0.1, #Samples=30K2026.04 | 24.3 | — | — | |
| MathFusionBase Model=Mistral-7B-v0.1, #Samples=60K2026.04 | 24.3 | — | — | |
| RFTBase Model=Mistral-7B-v0.1, #Samples=590K2026.04 | 24.2 | — | — | |
| SIGMABase Model=Mistral-7B-v0.1, #Samples=60K2026.04 | 24.1 | — | — | |
| RFTBase Model=LLaMA3-8B, #Samples=590K2026.04 | 23.9 | — | — | |
| IDPOModel=Llama3.1-8B-Instruct2025.05 | 23.8 | — | — | |
| DPO-R1 (MIDDLE)Backbone=Llama-3.1-8B-Instruct, Compute Cost (Sampling Budget N)=N = 4, Trainable Pairs=9,692, Processing Time=1.62026.02 | 23.7 | — | — | |
| DARTBase Model=Mistral-7B-v0.1, #Samples=60K2026.04 | 23.4 | — | — | |
| WizardMathBase Model=Mistral-7B-v0.1, #Samples=418K2026.04 | 23.1 | — | — | |
| DPO-R1 (HIGH)Backbone=Llama-3.1-8B-Instruct, Compute Cost (Sampling Budget N)=N = 16, Trainable Pairs=16,241, Processing Time=4.02026.02 | 23 | — | — | |
| CRPSBase Model=Mistral-7B-v0.1, #Samples=15K2026.04 | 22.4 | — | — | |
| DPO-R1 (ZHANG ET AL., 2025)Backbone=Llama-3.1-8B-Instruct, Compute Cost (Sampling Budget N)=N = 8, Trainable Pairs=13,224, Processing Time=2.42026.02 | 22.2 | — | — | |
| SIGMABase Model=Mistral-7B-v0.1, #Samples=30K2026.04 | 22.1 | — | — | |
| INSTRUCT (BASE)Backbone=Llama-3.1-8B-Instruct, Compute Cost (Sampling Budget N)=N/A, Trainable Pairs=N/A, Processing Time=N/A2026.02 | 21.3 | — | — | |
| Base ModelModel=Llama3.1-8B-Instruct2025.05 | 21.3 | — | — | |
| MetaMathBase Model=LLaMA3-8B, #Samples=400K2026.04 | 20.6 | — | — | |
| MetaMathBase Model=Mistral-7B-v0.1, #Samples=400K2026.04 | 19.3 | — | — | |
| MathFusion (Sequential)Base Model=Mistral-7B-v0.1, #Samples=30K2026.04 | 18.9 | — | — | |
| MetaMathBase Model=Mistral-7B-v0.1, #Samples=60K2026.04 | 14.1 | — | — | |
| Base ModelBackbone=LLADA 8B2025.10 | — | 30.2 | 49.1 | |
| CoconutBackbone=LLAMA 3.1 8B2025.10 | — | 40.2 | 42.9 | |
| CODIBackbone=LLAMA 3.1 8B2025.10 | — | 43.2 | 49 | |
| CoT SFTBackbone=LLADA 8B2025.10 | — | 38.9 | 48.3 | |
| Discrete LatentBackbone=LLAMA 3.1 8B2025.10 | — | 47.1 | 53.7 | |
| iCoTBackbone=LLAMA 3.1 8B2025.10 | — | 32.6 | 34.8 | |
| LaDiRBackbone=LLAMA 3.1 8B2025.10 | — | 48.6 | 60.3 | |
| LaDiR –w/o Stage 2Backbone=LLAMA 3.1 8B, Training Stage=without Stage 22025.10 | — | 32.8 | 38 | |
| LD4LGBackbone=LLAMA 3.1 8B2025.10 | — | 17.9 | 24.3 | |
| Pause TokenBackbone=LLAMA 3.1 8B2025.10 | — | 27.2 | 32.1 | |
| PLANNERBackbone=LLAMA 3.1 8B2025.10 | — | 23.6 | 29.1 | |
| SFTBackbone=LLAMA 3.1 8B, Decoding temperature (alpha)=0.12025.10 | — | 46.5 | 51 | |
| SFTBackbone=LLAMA 3.1 8B, Decoding temperature (alpha)=0.72025.10 | — | 45.3 | 54.3 | |
| SFTBackbone=LLAMA 3.1 8B, Decoding temperature (alpha)=12025.10 | — | 44.6 | 56 | |
| SFTBackbone=LLAMA 3.1 8B, Decoding temperature (alpha)=1.22025.10 | — | 43.6 | 57.8 | |
| Soft ThinkBackbone=LLAMA 3.1 8B2025.10 | — | 45.9 | 48 | |
| Sol-Only SFTBackbone=LLAMA 3.1 8B2025.10 | — | 15.9 | 20.6 | |
| TaH+Backbone=LLAMA 3.1 8B2025.10 | — | 46.7 | 50.1 |