Multi-turn task completion on FunctionGym
42.3Task Completion ScoreGRPO + Qwen2.5-8B
Evaluation Results
| Method | Links | |
|---|---|---|
| GRPO + Qwen2.5-8BCategory=RL (GRPO) Trained Models2026.04 | 42.3 | |
| Gemini-1.5-ProCategory=Closed-Source LLM2026.04 | 41 | |
| GRPO + Qwen2.5-4BCategory=RL (GRPO) Trained Models2026.04 | 39.7 | |
| Gemini-1.5-FlashCategory=Closed-Source LLM2026.04 | 32.1 | |
| Qwen2.5-14B + PRIMECategory=PRIME Models2026.04 | 29.8 | |
| GPT-4oCategory=Closed-Source LLM2026.04 | 28.2 | |
| Qwen2.5-8B + PRIMECategory=PRIME Models2026.04 | 27.8 | |
| Qwen2.5-4B + PRIMECategory=PRIME Models2026.04 | 24.5 | |
| Qwen2.5-14BCategory=Open-Source LLM2026.04 | 16.7 | |
| GPT-4o-miniCategory=Closed-Source LLM2026.04 | 15.4 | |
| Qwen2.5-32BCategory=Open-Source LLM2026.04 | 15.4 | |
| Qwen2.5-8BCategory=Open-Source LLM2026.04 | 9.5 | |
| Qwen2.5-4BCategory=Open-Source LLM2026.04 | 7.7 |