Interactive Task Execution on ALFWorld (in-distribution)
96.81Success Rate3SPO
Evaluation Results
| Method | Links | |
|---|---|---|
| 3SPOModel=Qwen2.5-7B-Instruct, Type=RL Training2026.06 | 96.81 | |
| HGPOModel=Qwen2.5-7B-Instruct, Type=RL Training2026.06 | 95.44 | |
| 3SPOModel=Qwen2.5-1.5B-Instruct, Type=RL Training2026.06 | 95.42 | |
| GiGPOModel=Qwen2.5-7B-Instruct, Type=RL Training2026.06 | 93.29 | |
| HGPOModel=Qwen2.5-1.5B-Instruct, Type=RL Training2026.06 | 92.77 | |
| GiGPOModel=Qwen2.5-1.5B-Instruct, Type=RL Training2026.06 | 90.16 | |
| GRPOModel=Qwen2.5-7B-Instruct, Type=RL Training2026.06 | 78.64 | |
| RLOOModel=Qwen2.5-7B-Instruct, Type=RL Training2026.06 | 77.86 | |
| PPO (with critic)Model=Qwen2.5-7B-Instruct, Type=RL Training2026.06 | 77.08 | |
| GRPOModel=Qwen2.5-1.5B-Instruct, Type=RL Training2026.06 | 72.8 | |
| RLOOModel=Qwen2.5-1.5B-Instruct, Type=RL Training2026.06 | 69.7 | |
| Gemini-2.5-ProModel=Closed, Type=Prompting2026.06 | 60.3 | |
| PPO (with critic)Model=Qwen2.5-1.5B-Instruct, Type=RL Training2026.06 | 54.4 | |
| GPT-4oModel=Closed, Type=Prompting2026.06 | 48 | |
| ReflexionModel=Qwen2.5-7B-Instruct, Type=Prompting2026.06 | 42.7 | |
| ReActModel=Qwen2.5-7B-Instruct, Type=Prompting2026.06 | 31.2 | |
| ReflexionModel=Qwen2.5-1.5B-Instruct, Type=Prompting2026.06 | 21.8 | |
| Qwen2.5Model=Qwen2.5-7B-Instruct, Type=Prompting2026.06 | 14.8 | |
| ReActModel=Qwen2.5-1.5B-Instruct, Type=Prompting2026.06 | 12.8 | |
| Qwen2.5Model=Qwen2.5-1.5B-Instruct, Type=Prompting2026.06 | 4.1 |