Tool Use Accuracy on tau2-Bench
80.3AccuracyMOPD
Evaluation Results
| Method | Links | |
|---|---|---|
| MOPDModel=MiMo-V2-Flash2026.06 | 80.3 | |
| TeacherModel=MiMo-V2-Flash, Status=RL-trained2026.06 | 79.6 | |
| StudentModel=MiMo-V2-Flash2026.06 | 75.9 | |
| Qwen2.5-32B-CodeGymCoT Pattern=Short-CoT, Training Strategy=CodeGym, Model Series=Qwen2.5, Model Size=32B, T=0.7, top-p=0.95, Inference Averaging=5 runs2025.09 | 30.7 | |
| QwQ-32B-CodeGymCoT Pattern=Long-CoT, Training Strategy=CodeGym, Model Series=QwQ, Model Size=32B, T=0.7, top-p=0.95, Inference Averaging=5 runs2025.09 | 30.7 | |
| QwQ-32BCoT Pattern=Long-CoT, Training Strategy=Base, Model Series=QwQ, Model Size=32B, T=0.7, top-p=0.95, Inference Averaging=5 runs2025.09 | 26.1 | |
| Qwen2.5-72B-CodeGymCoT Pattern=Short-CoT, Training Strategy=CodeGym, Model Series=Qwen2.5, Model Size=72B, T=0.7, top-p=0.95, Inference Averaging=5 runs2025.09 | 25.8 | |
| Qwen2.5-32B-InstructCoT Pattern=Short-CoT, Training Strategy=Instruct, Model Series=Qwen2.5, Model Size=32B, T=0.7, top-p=0.95, Inference Averaging=5 runs2025.09 | 24.7 | |
| Qwen2.5-72B-InstructCoT Pattern=Short-CoT, Training Strategy=Instruct, Model Series=Qwen2.5, Model Size=72B, T=0.7, top-p=0.95, Inference Averaging=5 runs2025.09 | 22.6 | |
| Qwen2.5-14B-InstructCoT Pattern=Short-CoT, Training Strategy=Instruct, Model Series=Qwen2.5, Model Size=14B, T=0.7, top-p=0.95, Inference Averaging=5 runs2025.09 | 20.9 | |
| Qwen2.5-14B-CodeGymCoT Pattern=Short-CoT, Training Strategy=CodeGym, Model Series=Qwen2.5, Model Size=14B, T=0.7, top-p=0.95, Inference Averaging=5 runs2025.09 | 19.9 | |
| Qwen2.5-7B-CodeGymCoT Pattern=Short-CoT, Training Strategy=CodeGym, Model Series=Qwen2.5, Model Size=7B, T=0.7, top-p=0.95, Inference Averaging=5 runs2025.09 | 15.5 | |
| Youtu-LLM 2BNon-thinking Mode=true, temperature=1, top-p=0.95, top-k=20, presence penalty=1.52025.12 | 15 | |
| Qwen2.5-7B-InstructCoT Pattern=Short-CoT, Training Strategy=Instruct, Model Series=Qwen2.5, Model Size=7B, T=0.7, top-p=0.95, Inference Averaging=5 runs2025.09 | 14.9 | |
| Qwen3 4BNon-thinking Mode=true, temperature=1, top-p=0.95, top-k=20, presence penalty=1.52025.12 | 10.9 | |
| SmolLM3 3BNon-thinking Mode=true, temperature=1, top-p=0.95, top-k=20, presence penalty=1.52025.12 | 9.7 | |
| Qwen3 1.7BNon-thinking Mode=true, temperature=1, top-p=0.95, top-k=20, presence penalty=1.52025.12 | 2.6 |