Embodied Physical Environment Decision-Making on ALFWorld
98.7Pick Success RateG2PO
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| G2POType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.06 | 98.7 | 90.6 | 97.6 | 100 | 98.8 | 93.7 | 96.9 | |
| GiGPOType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.06 | 97.7 | 82.7 | 98.8 | 83.7 | 89.3 | 79.2 | 90.8 | |
| G2POType=RL Training, Backbone=Qwen2.5-1.5B-Instruct2026.06 | 97.2 | 96.3 | 95.1 | 94.9 | 97.1 | 91.4 | 95 | |
| Qwen3.5-397BType=Prompting, Backbone=Off-the-shelf2026.06 | 96.6 | 72.7 | 70.4 | 0 | 70.4 | 85.7 | 71.9 | |
| GiGPOType=RL Training, Backbone=Qwen2.5-1.5B-Instruct2026.06 | 94.4 | 67.5 | 94.8 | 94.4 | 79.8 | 76.4 | 86.7 | |
| Gemini-2.5-ProType=Prompting, Backbone=Off-the-shelf2026.06 | 92.8 | 63.3 | 62.1 | 69 | 26.6 | 58.7 | 60.3 | |
| PPO (with critic)Type=RL Training, Backbone=Qwen2.5-7B-Instruct2026.06 | 92.3 | 64 | 92.5 | 89.5 | 80.3 | 68.8 | 80.4 | |
| GRPOType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.06 | 90.8 | 66.1 | 89.3 | 74.7 | 72.5 | 64.7 | 77.6 | |
| RLOOType=RL Training, Backbone=Qwen2.5-1.5B-Instruct2026.06 | 88.3 | 52.8 | 71 | 62.8 | 66.4 | 56.9 | 69.7 | |
| RLOOType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.06 | 87.6 | 78.2 | 87.3 | 81.3 | 71.9 | 48.9 | 75.5 | |
| GRPOType=RL Training, Backbone=Qwen2.5-1.5B-Instruct2026.06 | 85.3 | 53.7 | 84.5 | 78.2 | 59.7 | 53.5 | 72.8 | |
| DeepSeek-V3.2Type=Prompting, Backbone=Off-the-shelf2026.06 | 79.3 | 54.6 | 48.2 | 46.2 | 29.6 | 57.1 | 53.1 | |
| PPO (with critic)Type=RL Training, Backbone=Qwen2.5-1.5B-Instruct2026.06 | 64.8 | 40.5 | 57.1 | 60.6 | 46.4 | 47.4 | 54.4 | |
| ReflexionType=Prompting, Backbone=Qwen2.5-7B-Instruct2026.06 | 62 | 41.6 | 44.9 | 30.9 | 36.3 | 23.8 | 42.7 | |
| ReActType=Prompting, Backbone=Qwen2.5-7B-Instruct2026.06 | 48.5 | 35.4 | 34.3 | 13.2 | 18.2 | 17.6 | 31.2 | |
| ReflexionType=Prompting, Backbone=Qwen2.5-1.5B-Instruct2026.06 | 35.3 | 22.2 | 21.7 | 13.6 | 19.4 | 3.7 | 21.8 | |
| Qwen2.5Type=Prompting, Backbone=Qwen2.5-7B-Instruct2026.06 | 33.4 | 21.6 | 19.3 | 6.9 | 2.8 | 3.2 | 14.8 | |
| ReActType=Prompting, Backbone=Qwen2.5-1.5B-Instruct2026.06 | 17.4 | 20.5 | 15.7 | 6.2 | 7.7 | 2 | 12.8 | |
| Qwen2.5Type=Prompting, Backbone=Qwen2.5-1.5B-Instruct2026.06 | 5.9 | 5.5 | 3.3 | 9.7 | 4.2 | 0 | 4.1 |