Mathematical Reasoning on MATH500 (Accuracy and Average Score)
96.2Accuracyw/o RLVR
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| w/o RLVRBase model=Qwen3-30B-A3B-Instruct2026.02 | 96.2 | — | |
| DeepSeek-R1-Distill-Qwen-7B + ERC-DAPOModel Backbone=Qwen-7B, Training Algorithm=ERC-DAPO2025.12 | 95.1 | — | |
| DeepSeek-R1-Distill-Qwen-7B + DAPOModel Backbone=Qwen-7B, Training Algorithm=DAPO2025.12 | 94.1 | — | |
| DeepSeek-R1-Distill-Qwen-7B + GRPOModel Backbone=Qwen-7B, Training Algorithm=GRPO2025.12 | 93.7 | — | |
| DeepSeek-R1-Distill-Qwen-7BModel Backbone=Qwen-7B, Training Algorithm=Baseline2025.12 | 93.6 | — | |
| DeepSeek-R1-Distill-Qwen-1.5B + ERC-DAPOModel Backbone=Qwen-1.5B, Training Algorithm=ERC-DAPO2025.12 | 90 | — | |
| DeepSeek-R1-Distill-Qwen-1.5B + DAPOModel Backbone=Qwen-1.5B, Training Algorithm=DAPO2025.12 | 89.4 | — | |
| DeepSeek-R1-Distill-Qwen-1.5B + GRPOModel Backbone=Qwen-1.5B, Training Algorithm=GRPO2025.12 | 88.3 | — | |
| DeepSeek-R1-Distill-Qwen-1.5BModel Backbone=Qwen-1.5B, Training Algorithm=Baseline2025.12 | 86 | — | |
| LUSPOBase model=Qwen2.5-7B-Base2026.02 | 78.4 | — | |
| MAmmoTH2-8BModel Size=>= 7B, Reasoning Protocol=few-shot CoT2025.12 | 73.2 | 65.2 | |
| GSPOBase model=Qwen2.5-7B-Base2026.02 | 71 | — | |
| M3POBackbone=Qwen2.5-3B-Instruct2025.12 | 63 | 70.5 | |
| w/o RLVRBase model=Qwen2.5-7B-Base2026.02 | 60.8 | — | |
| PPOBackbone=Qwen2.5-3B-Instruct2025.12 | 60.4 | 68.2 | |
| GRPOBackbone=Qwen2.5-3B-Instruct2025.12 | 60.4 | 69.1 | |
| HRPOBackbone=Qwen2.5-3B-Instruct2025.12 | 60.2 | 68.7 | |
| PPCVBackbone=Llama-3.1-8B-Instruct2026.02 | 50 | — | |
| M3POBackbone=Qwen2.5-1.5B-Instruct2025.12 | 48 | 60.3 | |
| Qwen2.5-7BModel Size=>= 7B, Reasoning Protocol=few-shot CoT2025.12 | 46.4 | 63.5 | |
| HRPOBackbone=Qwen2.5-1.5B-Instruct2025.12 | 45.8 | 58.1 | |
| GRPOBackbone=Qwen2.5-1.5B-Instruct2025.12 | 45.2 | 57.9 | |
| PPOBackbone=Qwen2.5-1.5B-Instruct2025.12 | 44.8 | 57.2 | |
| MAmmoTH2-7BModel Size=>= 7B, Reasoning Protocol=few-shot CoT2025.12 | 39.6 | 57.8 | |
| Phi-DecodingBackbone=Llama-3.1-8B-Instruct2026.02 | 38.2 | — | |
| Self-ConsistencyBackbone=Llama-3.1-8B-Instruct2026.02 | 37.8 | — | |
| Gemma-2-9BModel Size=>= 7B, Reasoning Protocol=few-shot CoT2025.12 | 36.4 | 55.6 | |
| SFTBackbone=Qwen2.5-3B-Instruct2025.12 | 36 | 46.1 | |
| DeepSeekMath-7BModel Size=>= 7B, Reasoning Protocol=few-shot CoT2025.12 | 34.6 | 51.9 | |
| Predictive DecodingBackbone=Llama-3.1-8B-Instruct2026.02 | 34 | — | |
| Tree-of-ThoughtBackbone=Llama-3.1-8B-Instruct2026.02 | 31.6 | — | |
| Guided DecodingBackbone=Llama-3.1-8B-Instruct2026.02 | 31.2 | — | |
| Chain-of-ThoughtBackbone=Llama-3.1-8B-Instruct2026.02 | 31 | — | |
| SFTBackbone=Qwen2.5-1.5B-Instruct2025.12 | 30.2 | 43.3 | |
| PPCVBackbone=Mistral-7B-Instruct-v0.22026.02 | 14.6 | — | |
| Predictive DecodingBackbone=Mistral-7B-Instruct-v0.22026.02 | 14.4 | — | |
| Self-ConsistencyBackbone=Mistral-7B-Instruct-v0.22026.02 | 14.2 | — | |
| Guided DecodingBackbone=Mistral-7B-Instruct-v0.22026.02 | 14 | — | |
| Phi-DecodingBackbone=Mistral-7B-Instruct-v0.22026.02 | 13.4 | — | |
| Chain-of-ThoughtBackbone=Mistral-7B-Instruct-v0.22026.02 | 12.2 | — | |
| Tree-of-ThoughtBackbone=Mistral-7B-Instruct-v0.22026.02 | 11.4 | — |