Multi-turn Agent Interaction on ALFWorld (test)
100Success Rate (Pick)GRPO+SDPO(Advantage)
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| GRPO+SDPO(Advantage)Type=Hybrid2026.05 | 100 | 91.7 | 88.9 | 53.3 | 81 | 21.7 | 72.8 | |
| GraphGPOType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.05 | 100 | 92.86 | 100 | 94.44 | 91.4 | 92.06 | 95.31 | |
| GiGPOType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 98.81 | 78.57 | 95.16 | 94.44 | 81.46 | 93.65 | 90.88 | |
| GiGPOType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.05 | 97.53 | 90.48 | 100 | 94.44 | 83.98 | 100 | 94.27 | |
| GRPO+SDPO(Loss)Type=Hybrid2026.05 | 97.4 | 100 | 88.9 | 100 | 71.4 | 34.8 | 82.1 | |
| RLSDType=Hybrid2026.05 | 97.4 | 75 | 88.9 | 100 | 61.9 | 73.9 | 82.9 | |
| GraphGPOType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 95.15 | 85.71 | 100 | 96.3 | 85.26 | 93.65 | 92.71 | |
| GIGPOType=RL Training2026.05 | 93.5 | 83.3 | 78.9 | 86.7 | 76.2 | 85 | 83.9 | |
| Gemini-2.5-ProType=Prompting, Base Model=Closed-Source2026.05 | 92.8 | 69 | 63.3 | 26.6 | 62.1 | 58.7 | 60.3 | |
| PPO (with critic)Type=RL Training2026.05 | 92.3 | 64 | 92.5 | 89.5 | 80.3 | 68.8 | 80.4 | |
| HGPOType=RL Training2026.05 | 92.3 | 91.7 | 77.8 | 93.3 | 85.7 | 73.9 | 85.8 | |
| SERLType=Hybrid2026.05 | 92.3 | 100 | 88.9 | 100 | 76.2 | 82.6 | 90 | |
| PPOType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.05 | 92.3 | 89.5 | 64 | 80.3 | 92.5 | 68.8 | 80.4 | |
| GRPOType=RL Training2026.05 | 90.3 | 83.3 | 84.2 | 70 | 69.2 | 55 | 75.3 | |
| GRPOType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.05 | 88.98 | 78.57 | 91.98 | 90.74 | 77.89 | 71.43 | 83.33 | |
| RLOOType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 88.3 | 62.8 | 52.8 | 66.4 | 71 | 56.9 | 69.7 | |
| RLOOType=RL Training2026.05 | 87.6 | 78.2 | 87.3 | 81.3 | 71.9 | 48.9 | 75.5 | |
| RLOOType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.05 | 87.6 | 81.3 | 78.2 | 71.9 | 87.3 | 48.9 | 75.5 | |
| GRPOType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 82.89 | 78.57 | 82.14 | 77.78 | 73.86 | 71.43 | 77.86 | |
| GPT-4oType=Prompting, Base Model=Closed-Source2026.05 | 75.3 | 56.7 | 60.8 | 21.6 | 31.2 | 49.8 | 48 | |
| SDPOType=Self-Distillation2026.05 | 66.7 | 66.7 | 29.6 | 16.7 | 20 | 0 | 33.3 | |
| PPOType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 64.8 | 60.6 | 40.5 | 46.4 | 57.1 | 47.4 | 54.4 | |
| ReflexionType=Prompting2026.05 | 62 | 41.6 | 44.9 | 30.9 | 36.3 | 23.8 | 42.7 | |
| ReflexionType=Prompting, Base Model=Qwen2.5-7B-Instruct2026.05 | 62 | 30.9 | 41.6 | 36.3 | 44.9 | 23.8 | 42.7 | |
| ReActType=Prompting2026.05 | 48.5 | 35.4 | 34.3 | 13.2 | 18.2 | 17.6 | 31.2 | |
| ReActType=Prompting, Base Model=Qwen2.5-7B-Instruct2026.05 | 48.5 | 13.2 | 35.4 | 18.2 | 34.3 | 17.6 | 31.2 | |
| ReflexionType=Prompting, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 35.3 | 13.6 | 22.2 | 19.4 | 21.7 | 3.7 | 21.8 | |
| Qwen2.5-7B-InstructType=Prompting2026.05 | 33.4 | 21.6 | 19.3 | 6.9 | 2.8 | 3.2 | 14.8 | |
| Qwen2.5Type=Prompting, Base Model=Qwen2.5-7B-Instruct2026.05 | 33.4 | 6.9 | 21.6 | 2.8 | 19.3 | 3.2 | 14.8 | |
| ReActType=Prompting, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 17.4 | 6.2 | 20.5 | 7.7 | 15.7 | 2 | 12.8 | |
| Qwen2.5Type=Prompting, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 5.9 | 9.7 | 5.5 | 4.2 | 3.3 | 0 | 4.1 |