Math reasoning on MATH-500 (in-distribution)
62.4Pass@1 AccuracyG2D
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| G2DModel=Qwen2.5-7B, K (warm-up steps)=150, GPU-hrs=3.8, Decoding strategy=greedy decoding2026.05 | 62.4 | — | |
| G2DModel=Qwen2.5-7B, K (warm-up steps)=300, GPU-hrs=7.0, Decoding strategy=greedy decoding2026.05 | 59.2 | — | |
| G2DModel=Qwen2.5-7B, K (warm-up steps)=500, GPU-hrs=10.4, Decoding strategy=greedy decoding2026.05 | 58.5 | — | |
| G2DModel=Qwen2.5-7B, K (warm-up steps)=1000, GPU-hrs=20.8, Decoding strategy=greedy decoding2026.05 | 57.6 | — | |
| G2DModel=Qwen2.5-7B, K (warm-up steps)=700, GPU-hrs=14.5, Decoding strategy=greedy decoding2026.05 | 56.4 | — | |
| DPOModel=Qwen2.5-7B, K (warm-up steps)=0, GPU-hrs=2.2, Decoding strategy=greedy decoding2026.05 | 56.2 | — | |
| GRPO (G = 4)Model=Qwen2.5-7B, K (warm-up steps)=1000, GPU-hrs=52.2, Decoding strategy=greedy decoding2026.05 | 53.2 | — | |
| GRPO (G = 2)Model=Qwen2.5-7B, K (warm-up steps)=1000, GPU-hrs=26.6, Decoding strategy=greedy decoding2026.05 | 51.6 | — | |
| G2DModel=Llama-3.1-8B, K (warm-up steps)=500, GPU-hrs=11.4, Decoding strategy=greedy decoding2026.05 | 49.4 | — | |
| G2DModel=Llama-3.1-8B, K (warm-up steps)=700, GPU-hrs=16.8, Decoding strategy=greedy decoding2026.05 | 49.2 | — | |
| SFTModel=Qwen2.5-7B, Decoding strategy=greedy decoding2026.05 | 49.1 | — | |
| G2DModel=Llama-3.1-8B, K (warm-up steps)=1000, GPU-hrs=21.8, Decoding strategy=greedy decoding2026.05 | 47.9 | — | |
| GRPO (G = 4)Model=Llama-3.1-8B, K (warm-up steps)=1000, GPU-hrs=54.2, Decoding strategy=greedy decoding2026.05 | 47.5 | — | |
| G2DModel=Llama-3.1-8B, K (warm-up steps)=300, GPU-hrs=7.8, Decoding strategy=greedy decoding2026.05 | 47.4 | — | |
| DPOModel=Llama-3.1-8B, K (warm-up steps)=0, GPU-hrs=2.6, Decoding strategy=greedy decoding2026.05 | 47.1 | — | |
| G2DModel=Llama-3.1-8B, K (warm-up steps)=150, GPU-hrs=4.5, Decoding strategy=greedy decoding2026.05 | 46.4 | — | |
| GRPO (G = 2)Model=Llama-3.1-8B, K (warm-up steps)=1000, GPU-hrs=27.2, Decoding strategy=greedy decoding2026.05 | 46.1 | — | |
| SFTModel=Llama-3.1-8B, Decoding strategy=greedy decoding2026.05 | 45.8 | — | |
| CalibRLTraining Data=9k-sample subset of MATH dataset2026.02 | — | 70.2 | |
| GRPOTraining Data=9k-sample subset of MATH dataset2026.02 | — | 68.5 | |
| Qwen2.5-VL-7BTraining Data=9k-sample subset of MATH dataset2026.02 | — | 64.8 |