Mathematical Reasoning on MetaMath Insufficient
72.5Success Rate (SR)GRIL
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| GRILModel Scale=Qwen2.5-3B-Instruct2026.04 | 72.5 | 86.6 | 2.473 | 624 | |
| GRILModel Scale=Qwen3-1.7B2026.04 | 65.8 | 95.2 | 2.488 | 502 | |
| w/ PromptModel Scale=Qwen3-1.7B2026.04 | 64.2 | 86.1 | 2.609 | 966 | |
| w/ SFTModel Scale=Qwen3-1.7B2026.04 | 63 | 90.3 | 2.531 | 874 | |
| Base ModelModel Scale=Qwen3-1.7B2026.04 | 62 | 90.5 | 2.585 | 929 | |
| GRILModel Scale=Qwen2.5-1.5B-Instruct2026.04 | 58.4 | 88.2 | 2.941 | 581 | |
| w/ SFTModel Scale=Qwen2.5-3B-Instruct2026.04 | 56.1 | 83.9 | 2.8 | 750 | |
| w/ PromptModel Scale=Qwen2.5-3B-Instruct2026.04 | 51.5 | 77.1 | 2.985 | 824 | |
| GRILModel Scale=Qwen3-0.6B2026.04 | 45.2 | 79.2 | 3.025 | 1,616 | |
| w/ Multi-turn RLtraining=Multi-turn Reinforcement Learning2026.04 | 44.2 | 79.6 | — | — | |
| w/ Promptstrategy=prompting2026.04 | 42.9 | 73.6 | — | — | |
| w/ SFTModel Scale=Qwen3-0.6B2026.04 | 41.3 | 81.8 | 3.077 | 1,003 | |
| w/ PromptModel Scale=Qwen3-0.6B2026.04 | 40.6 | 75.7 | 2.869 | 1,345 | |
| w/ SFTModel Scale=Qwen2.5-1.5B-Instruct2026.04 | 31.1 | 60.4 | 3.248 | 999 | |
| w/ PromptModel Scale=Qwen2.5-1.5B-Instruct2026.04 | 19.3 | 48.1 | 3.426 | 743 | |
| Base ModelModel Scale=Qwen2.5-3B-Instruct2026.04 | 15.5 | 20.9 | 3.722 | 1,091 | |
| Base Model2026.04 | 15 | 64 | — | — | |
| Base ModelModel Scale=Qwen3-0.6B2026.04 | 14.7 | 25.6 | 3.793 | 1,907 | |
| Base ModelModel Scale=Qwen2.5-1.5B-Instruct2026.04 | 2.2 | 4.2 | 3.792 | 937 |