Mathematical Reasoning on MATH (Accuracy, Delta Avg)
92.8AccuracyCoT2-Meta
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| CoT2-MetaBackbone=Claude-4.5, Strategy=Ours (CoT2-Meta)2026.03 | 92.8 | 14.5 | |
| Vanilla ToTBackbone=Claude-4.5, Strategy=Vanilla ToT2026.03 | 87.4 | 8.3 | |
| Best-of-16Backbone=Claude-4.5, Strategy=Best-of-162026.03 | 84.2 | 4.8 | |
| CoT2-MetaBackbone=DeepSeek-V3.2, Strategy=Ours (CoT2-Meta)2026.03 | 84.2 | 10.5 | |
| Vanilla ToTBackbone=DeepSeek-V3.2, Strategy=Vanilla ToT2026.03 | 78.6 | 5.7 | |
| Greedy CoTBackbone=Claude-4.5, Strategy=Greedy CoT2026.03 | 78.5 | — | |
| Best-of-16Backbone=DeepSeek-V3.2, Strategy=Best-of-162026.03 | 75.3 | 3 | |
| Greedy CoTBackbone=DeepSeek-V3.2, Strategy=Greedy CoT2026.03 | 70.8 | — | |
| Naive Avg.Parent Model Pair=RL2 & RL3, Parent Model Details=RLinf-math-1.5B & E1-Math-1.5B, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 65.23 | — | |
| CoT2-MetaBackbone=Qwen2.5-VL-7B, Strategy=Ours (CoT2-Meta)2026.03 | 64.2 | 12.2 | |
| Naive Avg.Parent Model Pair=RL1 & RL2, Parent Model Details=DeepScaleR-1.5B-Preview & RLinf-math-1.5B, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 60.17 | — | |
| SAR-MergingParent Model Pair=RL2 & RL3, Parent Model Details=RLinf-math-1.5B & E1-Math-1.5B, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 59.52 | — | |
| SAR-MergingParent Model Pair=RL1 & RL3, Parent Model Details=DeepScaleR-1.5B-Preview & E1-Math-1.5B, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 59.32 | — | |
| Vanilla ToTBackbone=Qwen2.5-VL-7B, Strategy=Vanilla ToT2026.03 | 59.1 | 6.4 | |
| SAR-MergingParent Model Pair=RL1 & RL2, Parent Model Details=DeepScaleR-1.5B-Preview & RLinf-math-1.5B, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 58.7 | — | |
| Naive Avg.Parent Model Pair=RL1 & RL3, Parent Model Details=DeepScaleR-1.5B-Preview & E1-Math-1.5B, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 57.9 | — | |
| Best-of-16Backbone=Qwen2.5-VL-7B, Strategy=Best-of-162026.03 | 55.8 | 3.4 | |
| TIES-MergingParent Model Pair=RL1 & RL3, Parent Model Details=DeepScaleR-1.5B-Preview & E1-Math-1.5B, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 52.92 | — | |
| Task ArithmeticParent Model Pair=RL2 & RL3, Parent Model Details=RLinf-math-1.5B & E1-Math-1.5B, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 51.8 | — | |
| RAM-MergingParent Model Pair=RL1 & RL3, Parent Model Details=DeepScaleR-1.5B-Preview & E1-Math-1.5B, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 51.56 | — | |
| DARE-MergingParent Model Pair=RL1 & RL3, Parent Model Details=DeepScaleR-1.5B-Preview & E1-Math-1.5B, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 51.42 | — | |
| RAM-MergingParent Model Pair=RL2 & RL3, Parent Model Details=RLinf-math-1.5B & E1-Math-1.5B, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 50.7 | — | |
| Task ArithmeticParent Model Pair=RL1 & RL2, Parent Model Details=DeepScaleR-1.5B-Preview & RLinf-math-1.5B, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 50.52 | — | |
| Greedy CoTBackbone=Qwen2.5-VL-7B, Strategy=Greedy CoT2026.03 | 50.4 | — | |
| DARE-MergingParent Model Pair=RL2 & RL3, Parent Model Details=RLinf-math-1.5B & E1-Math-1.5B, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 50.26 | — | |
| Task ArithmeticParent Model Pair=RL1 & RL3, Parent Model Details=DeepScaleR-1.5B-Preview & E1-Math-1.5B, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 50.08 | — | |
| RAM-MergingParent Model Pair=RL1 & RL2, Parent Model Details=DeepScaleR-1.5B-Preview & RLinf-math-1.5B, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 49.64 | — | |
| TIES-MergingParent Model Pair=RL2 & RL3, Parent Model Details=RLinf-math-1.5B & E1-Math-1.5B, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 49.64 | — | |
| Naive Avg.Parent Model Pair=SFT1 & SFT2, Parent Model Details=lul-sft & MiniMath, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 49.01 | — | |
| Linear-MergingParent Model Pair=RL2 & RL3, Parent Model Details=RLinf-math-1.5B & E1-Math-1.5B, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 48.34 | — | |
| DARE-MergingParent Model Pair=RL1 & RL2, Parent Model Details=DeepScaleR-1.5B-Preview & RLinf-math-1.5B, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 48.28 | — | |
| TIES-MergingParent Model Pair=RL1 & RL2, Parent Model Details=DeepScaleR-1.5B-Preview & RLinf-math-1.5B, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 48.1 | — | |
| Linear-MergingParent Model Pair=RL1 & RL2, Parent Model Details=DeepScaleR-1.5B-Preview & RLinf-math-1.5B, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 47.34 | — | |
| Linear-MergingParent Model Pair=RL1 & RL3, Parent Model Details=DeepScaleR-1.5B-Preview & E1-Math-1.5B, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 47.34 | — | |
| Linear-MergingParent Model Pair=SFT1 & SFT2, Parent Model Details=lul-sft & MiniMath, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 45.34 | — | |
| TIES-MergingParent Model Pair=SFT1 & SFT2, Parent Model Details=lul-sft & MiniMath, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 42.08 | — | |
| DARE-MergingParent Model Pair=SFT1 & SFT2, Parent Model Details=lul-sft & MiniMath, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 41.98 | — | |
| Task ArithmeticParent Model Pair=SFT1 & SFT2, Parent Model Details=lul-sft & MiniMath, Model Scale=1.5B, Evaluation Protocol=zero-shot2026.06 | 34.6 | — | |
| 4-bit-QLoRA 13BModel Size (GB)=6.632025.07 | 14.3 | — | |
| 4-bit-QLoRA 7BModel Size (GB)=3.452025.07 | 12.6 | — | |
| 4-bit-quantized 13BModel Size (GB)=6.512025.07 | 12.5 | — | |
| Full Fine-TuningRank (r)=N/A, Backbone Model=Llama 2-7B2026.06 | 11.74 | — | |
| 4-bit-quantized 7BModel Size (GB)=3.372025.07 | 8.8 | — | |
| LoRA-αRank (r)=128, Backbone Model=Llama 2-7B2026.06 | 8.3 | — | |
| RsLoRARank (r)=128, Backbone Model=Llama 2-7B2026.06 | 7.32 | — | |
| LoRAMRank (r)=128, Backbone Model=Llama 2-7B2026.06 | 7.25 | — | |
| PiSSARank (r)=128, Backbone Model=Llama 2-7B2026.06 | 7.04 | — | |
| LoRA-αRank (r)=16, Backbone Model=Llama 2-7B2026.06 | 6.84 | — | |
| LoRAMRank (r)=16, Backbone Model=Llama 2-7B2026.06 | 5.3 | — | |
| LoRA+Rank (r)=128, Backbone Model=Llama 2-7B2026.06 | 5.28 | — | |
| PiSSARank (r)=16, Backbone Model=Llama 2-7B2026.06 | 5.16 | — | |
| RsLoRARank (r)=16, Backbone Model=Llama 2-7B2026.06 | 4.94 | — | |
| LoRARank (r)=128, Backbone Model=Llama 2-7B2026.06 | 4.72 | — | |
| LoRARank (r)=16, Backbone Model=Llama 2-7B2026.06 | 4.16 | — | |
| LoRA+Rank (r)=16, Backbone Model=Llama 2-7B2026.06 | 3.98 | — |