Multi-turn task completion on TravelGym
57.3Task Completion ScoreGRPO + Qwen2.5-8B
Evaluation Results
| Method | Links | |
|---|---|---|
| GRPO + Qwen2.5-8BCategory=RL (GRPO) Trained Models2026.04 | 57.3 | |
| GRPO + Qwen2.5-4BCategory=RL (GRPO) Trained Models2026.04 | 50.9 | |
| GPT-4oCategory=Closed-Source LLM2026.04 | 36.4 | |
| Gemini-1.5-ProCategory=Closed-Source LLM2026.04 | 34.7 | |
| Gemini-1.5-FlashCategory=Closed-Source LLM2026.04 | 25.5 | |
| Qwen2.5-14B + PRIMECategory=PRIME Models2026.04 | 24.8 | |
| Qwen2.5-8B + PRIMECategory=PRIME Models2026.04 | 21.5 | |
| Qwen2.5-14BCategory=Open-Source LLM2026.04 | 19.2 | |
| Qwen2.5-4B + PRIMECategory=PRIME Models2026.04 | 18.5 | |
| Qwen2.5-32BCategory=Open-Source LLM2026.04 | 17.2 | |
| Qwen2.5-8BCategory=Open-Source LLM2026.04 | 15.8 | |
| Qwen2.5-4BCategory=Open-Source LLM2026.04 | 14.1 | |
| GPT-4o-miniCategory=Closed-Source LLM2026.04 | 9.8 |