Multi-turn response refinement on MT-Refine
89.8Math ScoreRLSTA
Evaluation Results
| Method | Links | |
|---|---|---|
| RLSTABackbone=Qwen3-4B-Instruct-25072026.03 | 89.8 | |
| GRPOBackbone=Qwen3-4B-Instruct-25072026.03 | 88.2 | |
| GRPOBackbone=Qwen2.5-7B-Instruct2026.03 | 83.6 | |
| RLSTABackbone=Qwen2.5-7B-Instruct2026.03 | 82.2 | |
| BaseBackbone=Qwen3-4B-Instruct-25072026.03 | 81.4 | |
| SFTBackbone=Qwen3-4B-Instruct-25072026.03 | 80.9 | |
| DPOBackbone=Qwen3-4B-Instruct-25072026.03 | 80.1 | |
| RLSTABackbone=Qwen2.5-3B-Instruct2026.03 | 74.5 | |
| GRPOBackbone=Qwen2.5-3B-Instruct2026.03 | 73.4 | |
| SFTBackbone=Qwen2.5-7B-Instruct2026.03 | 69.4 | |
| BaseBackbone=Qwen2.5-7B-Instruct2026.03 | 66.9 | |
| RLSTABackbone=Llama-3.2-3B-Instruct2026.03 | 64 | |
| GRPOBackbone=Llama-3.2-3B-Instruct2026.03 | 62 | |
| BaseBackbone=Qwen2.5-3B-Instruct2026.03 | 60.3 | |
| BaseBackbone=Llama-3.2-3B-Instruct2026.03 | 58.5 | |
| DPOBackbone=Qwen2.5-3B-Instruct2026.03 | 56.8 | |
| SFTBackbone=Qwen2.5-3B-Instruct2026.03 | 53.3 | |
| DPOBackbone=Qwen2.5-7B-Instruct2026.03 | 52.2 | |
| SFTBackbone=Llama-3.2-3B-Instruct2026.03 | 51.9 | |
| DPOBackbone=Llama-3.2-3B-Instruct2026.03 | 36.8 |