Mathematical Reasoning on GSM8K Insufficient
49.3Success Rate (SR)w/ Multi-turn RL
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| w/ Multi-turn RLtraining=Multi-turn Reinforcement Learning2026.04 | 49.3 | 80.4 | |
| w/ Promptstrategy=prompting2026.04 | 43.8 | 80.8 | |
| Base Model2026.04 | 17 | 65 |