Multi-turn task completion on TauGym
21Task Completion ScoreGRPO + Qwen2.5-8B
Evaluation Results
| Method | Links | |
|---|---|---|
| GRPO + Qwen2.5-8BCategory=RL (GRPO) Trained Models2026.04 | 21 | |
| GPT-4o-miniCategory=Closed-Source LLM2026.04 | 20.6 | |
| GRPO + Qwen2.5-4BCategory=RL (GRPO) Trained Models2026.04 | 20 | |
| Gemini-1.5-ProCategory=Closed-Source LLM2026.04 | 19.4 | |
| Qwen2.5-14B + PRIMECategory=PRIME Models2026.04 | 12.4 | |
| Gemini-1.5-FlashCategory=Closed-Source LLM2026.04 | 12.1 | |
| Qwen2.5-14BCategory=Open-Source LLM2026.04 | 10.3 | |
| Qwen2.5-8B + PRIMECategory=PRIME Models2026.04 | 7.2 | |
| Qwen2.5-4B + PRIMECategory=PRIME Models2026.04 | 5.8 | |
| Qwen2.5-8BCategory=Open-Source LLM2026.04 | 4.8 | |
| Qwen2.5-4BCategory=Open-Source LLM2026.04 | 3.6 | |
| GPT-4oCategory=Closed-Source LLM2026.04 | 3 | |
| Qwen2.5-32BCategory=Open-Source LLM2026.04 | 0 |