API-based agent interaction on AppWorld
27.6Success RateG2PO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| G2POType=RL Training, Base model=Qwen2.5-14B-Instruct2026.06 | 27.6 | 21.7 | |
| GiGPOType=RL Training, Base model=Qwen2.5-14B-Instruct2026.06 | 25.7 | 19.2 | |
| RLOOType=RL Training, Base model=Qwen2.5-14B-Instruct2026.06 | 24.8 | 19.5 | |
| GRPOType=RL Training, Base model=Qwen2.5-14B-Instruct2026.06 | 24.8 | 20.7 | |
| PPO (with critic)Type=RL Training, Base model=Qwen2.5-14B-Instruct2026.06 | 19.1 | 14.7 | |
| ReActType=Prompting, Base model=Qwen2.5-14B-Instruct2026.06 | 10.5 | 8 | |
| Qwen2.5Type=Prompting, Base model=Qwen2.5-14B-Instruct2026.06 | 0 | 0 |