Mathematical Reasoning on Math Reasoning Suite Average
63Average AccuracyReasoning Palette (Linear Decay)
Evaluation Results
| Method | Links | |
|---|---|---|
| Reasoning Palette (Linear Decay)Backbone=Qwen3-8B-Base, RL Algorithm=RLOO2025.12 | 63 | |
| Reasoning Palette (Two-Phase)Backbone=Qwen3-8B-Base, RL Algorithm=RLOO2025.12 | 62.28 | |
| Reasoning Palette (Linear Decay)Backbone=Qwen3-4B-Base, RL Algorithm=RLOO2025.12 | 61.28 | |
| Reasoning Palette (Linear Decay)Backbone=Qwen3-8B-Base, RL Algorithm=GRPO2025.12 | 61.17 | |
| Reasoning Palette (Two-Phase)Backbone=Qwen3-4B-Base, RL Algorithm=GRPO2025.12 | 60.77 | |
| Reasoning Palette (Two-Phase)Backbone=Qwen3-4B-Base, RL Algorithm=RLOO2025.12 | 60.69 | |
| Reasoning Palette (Two-Phase)Backbone=Qwen3-8B-Base, RL Algorithm=GRPO2025.12 | 60.57 | |
| Reasoning Palette (Linear Decay)Backbone=Qwen3-4B-Base, RL Algorithm=GRPO2025.12 | 60.12 | |
| GRPOBackbone=Qwen3-8B-Base, RL Algorithm=GRPO2025.12 | 60.05 | |
| RLOOBackbone=Qwen3-8B-Base, RL Algorithm=RLOO2025.12 | 59.91 | |
| RLOOBackbone=Qwen3-4B-Base, RL Algorithm=RLOO2025.12 | 59.39 | |
| GRPOBackbone=Qwen3-4B-Base, RL Algorithm=GRPO2025.12 | 58.27 | |
| Qwen2.5-7B-Math + SFT + RL-TCERBase Model=Qwen2.5-7B-Math, Training Strategy=SFT + RL-TCER2026.04 | 50.2 | |
| Qwen2.5-7B-Math + SFT + RL-EndoRBase Model=Qwen2.5-7B-Math, Training Strategy=SFT + RL-EndoR2026.04 | 48.9 | |
| ExOPDDistillation Setting=Single-Teacher Distillation2026.02 | 48 | |
| ExOPDDistillation Setting=Multi-Teacher Distillation2026.02 | 47.7 | |
| OPDDistillation Setting=Single-Teacher Distillation2026.02 | 46.5 | |
| OPDDistillation Setting=Multi-Teacher Distillation2026.02 | 46.4 | |
| TeacherDistillation Setting=Baseline2026.02 | 46 | |
| ExPODistillation Setting=Single-Teacher Distillation2026.02 | 45.8 | |
| ExPODistillation Setting=Multi-Teacher Distillation2026.02 | 45 | |
| SFTDistillation Setting=Multi-Teacher Distillation2026.02 | 44.3 | |
| Qwen2.5-7B-Math + SFTBase Model=Qwen2.5-7B-Math, Training Strategy=SFT2026.04 | 44.1 | |
| Llama3.1-8B-Instruct + SFT + RL-TCERBase Model=Llama3.1-8B-Instruct, Training Strategy=SFT + RL-TCER2026.04 | 33.6 | |
| Llama3.1-8B-Instruct + SFT + RL-EndoRBase Model=Llama3.1-8B-Instruct, Training Strategy=SFT + RL-EndoR2026.04 | 31.9 | |
| Llama3.1-8B-Instruct + SFTBase Model=Llama3.1-8B-Instruct, Training Strategy=SFT2026.04 | 30.9 | |
| DAPO++Backbone=Qwen3-8B, Sampling Strategy=Mixed2026.05 | 29.6 | |
| DAPO-Math-17kBackbone=Qwen3-8B, Sampling Strategy=Mixed2026.05 | 29.3 | |
| DeepScaleRBackbone=Qwen3-8B, Sampling Strategy=Mixed2026.05 | 26.1 | |
| DeepMath-103KBackbone=Qwen3-8B, Sampling Strategy=Mixed2026.05 | 25.1 | |
| Skywork-OR1-RL-DataBackbone=Qwen3-8B, Sampling Strategy=Mixed2026.05 | 25.1 | |
| OpenR1-Math-220kBackbone=Qwen3-8B, Sampling Strategy=Mixed2026.05 | 25 | |
| Reasoning Palette (Linear Decay)Backbone=Qwen3-1.7B-Base, RL Algorithm=GRPO2025.12 | 24.05 | |
| Reasoning Palette (Linear Decay)Backbone=Qwen3-1.7B-Base, RL Algorithm=RLOO2025.12 | 23.49 | |
| Reasoning Palette (Two-Phase)Backbone=Qwen3-1.7B-Base, RL Algorithm=GRPO2025.12 | 22.91 | |
| Reasoning Palette (Two-Phase)Backbone=Qwen3-1.7B-Base, RL Algorithm=RLOO2025.12 | 22.8 | |
| RLOOBackbone=Qwen3-1.7B-Base, RL Algorithm=RLOO2025.12 | 21.77 | |
| GRPOBackbone=Qwen3-1.7B-Base, RL Algorithm=GRPO2025.12 | 21.18 | |
| Qwen2.5-7B-MathBase Model=Qwen2.5-7B-Math, Training Strategy=None2026.04 | 19.1 | |
| Qwen3-8B-BaseBackbone=Qwen3-8B, Sampling Strategy=Mixed2026.05 | 18.6 | |
| Llama3.1-8B-InstructBase Model=Llama3.1-8B-Instruct, Training Strategy=None2026.04 | 17.6 | |
| DAPO++Backbone=Qwen3-1.7B, Sampling Strategy=Mixed2026.05 | 15.7 | |
| StudentDistillation Setting=Baseline2026.02 | 15.4 | |
| DeepMath-103KBackbone=Qwen3-1.7B, Sampling Strategy=Mixed2026.05 | 15.4 | |
| Skywork-OR1-RL-DataBackbone=Qwen3-1.7B, Sampling Strategy=Mixed2026.05 | 15.1 | |
| DAPO-Math-17kBackbone=Qwen3-1.7B, Sampling Strategy=Mixed2026.05 | 15 | |
| DeepScaleRBackbone=Qwen3-1.7B, Sampling Strategy=Mixed2026.05 | 14.7 | |
| OpenR1-Math-220kBackbone=Qwen3-1.7B, Sampling Strategy=Mixed2026.05 | 14 | |
| Qwen3-1.7B-BaseBackbone=Qwen3-1.7B, Sampling Strategy=Mixed2026.05 | 10.8 |