Multi-turn task completion on PersuadeGym
57.9Task Completion ScoreGRPO + Qwen2.5-4B
Evaluation Results
| Method | Links | |
|---|---|---|
| GRPO + Qwen2.5-4BCategory=RL (GRPO) Trained Models2026.04 | 57.9 | |
| Qwen2.5-14B + PRIMECategory=PRIME Models2026.04 | 56.8 | |
| GPT-4o-miniCategory=Closed-Source LLM2026.04 | 53.2 | |
| Qwen2.5-14BCategory=Open-Source LLM2026.04 | 53.2 | |
| GRPO + Qwen2.5-8BCategory=RL (GRPO) Trained Models2026.04 | 53.2 | |
| Qwen2.5-4B + PRIMECategory=PRIME Models2026.04 | 49.8 | |
| Qwen2.5-32BCategory=Open-Source LLM2026.04 | 48.4 | |
| Qwen2.5-8B + PRIMECategory=PRIME Models2026.04 | 47.2 | |
| Qwen2.5-8BCategory=Open-Source LLM2026.04 | 44.1 | |
| Gemini-1.5-ProCategory=Closed-Source LLM2026.04 | 42.5 | |
| Gemini-1.5-FlashCategory=Closed-Source LLM2026.04 | 40.9 | |
| Qwen2.5-4BCategory=Open-Source LLM2026.04 | 40.5 | |
| GPT-4oCategory=Closed-Source LLM2026.04 | 37.7 |