Household Agent Interaction on ALFWorld
99.1Pick Success RateHCAPO
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| HCAPOType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.03 | 99.1 | 90.3 | 97.3 | 81.8 | 90.8 | 81.9 | 91.4 | |
| GiGPOType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.03 | 97.7 | 82.7 | 98.8 | 83.7 | 89.3 | 79.2 | 90.8 | |
| GiGPOType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.03 | 94.4 | 67.5 | 94.8 | 94.4 | 79.8 | 76.4 | 86.7 | |
| EMPGType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.03 | 92.9 | 75.2 | 74.8 | 86.3 | 73.7 | 65.3 | 78.5 | |
| Gemini-2.5-ProType=Prompting, Base Model=Closed-Source Model2026.03 | 92.8 | 63.3 | 62.1 | 69 | 26.6 | 58.7 | 60.3 | |
| PPO (with critic)Type=RL Training, Base Model=Qwen2.5-7B-Instruct2026.03 | 92.3 | 64 | 92.5 | 89.5 | 80.3 | 68.8 | 80.4 | |
| GRPOType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.03 | 90.8 | 66.1 | 89.3 | 74.7 | 72.5 | 64.7 | 77.6 | |
| HCAPOType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.03 | 88.6 | 75 | 97.6 | 90.7 | 84.2 | 74.2 | 87 | |
| RLOOType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.03 | 88.3 | 52.8 | 71 | 62.8 | 66.4 | 56.9 | 69.7 | |
| RLOOType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.03 | 87.6 | 78.2 | 87.3 | 81.3 | 71.9 | 48.9 | 75.5 | |
| EMPGType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.03 | 85.5 | 33.5 | 78.9 | 76.2 | 74.7 | 69.1 | 73.7 | |
| GRPOType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.03 | 85.3 | 53.7 | 84.5 | 78.2 | 59.7 | 53.5 | 72.8 | |
| GPT-4oType=Prompting, Base Model=Closed-Source Model2026.03 | 75.3 | 60.8 | 31.2 | 56.7 | 21.6 | 49.8 | 48 | |
| PPO (with critic)Type=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.03 | 64.8 | 40.5 | 57.1 | 60.6 | 46.4 | 47.4 | 54.4 | |
| ReflexionType=Prompting, Base Model=Qwen2.5-7B-Instruct2026.03 | 62 | 41.6 | 44.9 | 30.9 | 36.3 | 23.8 | 42.7 | |
| ReActType=Prompting, Base Model=Qwen2.5-7B-Instruct2026.03 | 48.5 | 35.4 | 34.3 | 13.2 | 18.2 | 17.6 | 31.2 | |
| ReflexionType=Prompting, Base Model=Qwen2.5-1.5B-Instruct2026.03 | 35.3 | 22.2 | 21.7 | 13.6 | 19.4 | 3.7 | 21.8 | |
| Qwen2.5Type=Prompting, Base Model=Qwen2.5-7B-Instruct2026.03 | 33.4 | 21.6 | 19.3 | 6.9 | 2.8 | 3.2 | 14.8 | |
| ReActType=Prompting, Base Model=Qwen2.5-1.5B-Instruct2026.03 | 17.4 | 20.5 | 15.7 | 6.2 | 7.7 | 2 | 12.8 | |
| Qwen2.5Type=Prompting, Base Model=Qwen2.5-1.5B-Instruct2026.03 | 5.9 | 5.5 | 3.3 | 9.7 | 4.2 | 0 | 4.1 |