Text-based embodied task completion on ALFWorld
96.8Task Completion Rate (Short)BEACON
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| BEACONType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 96.8 | 87 | 92.9 | 91.4 | |
| BEACONType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.05 | 95.1 | 94.9 | 90 | 94.5 | |
| GiGPOType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.05 | 93.6 | 91.8 | 79.2 | 90.8 | |
| GiGPOType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 90.7 | 84.3 | 79.5 | 86.1 | |
| RLOOType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.05 | 85.1 | 80.2 | 48.9 | 75.5 | |
| Gemini-2.5-Pro (ReAct)Type=Prompting, Base Model=Closed-Source Models2026.05 | 84.8 | 50.7 | 58.7 | 60.3 | |
| PPOType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.05 | 84.6 | 87.3 | 68.8 | 80.4 | |
| GRPOType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.05 | 84.1 | 79.7 | 64.7 | 77.6 | |
| RLOOType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 78.7 | 67.4 | 56.9 | 69.7 | |
| GRPOType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 76.7 | 73.9 | 53.5 | 72.8 | |
| GPT-4o (ReAct)Type=Prompting, Base Model=Closed-Source Models2026.05 | 71.4 | 33.7 | 49.8 | 48 | |
| PPOType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 58.2 | 54 | 47.4 | 54.4 | |
| ReflexionType=Prompting, Base Model=Qwen2.5-7B-Instruct2026.05 | 56.5 | 38.4 | 23.8 | 42.7 | |
| ReActType=Prompting, Base Model=Qwen2.5-7B-Instruct2026.05 | 45 | 23.4 | 17.6 | 31.2 | |
| ReflexionType=Prompting, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 31.8 | 18.9 | 3.7 | 21.8 | |
| Direct PromptType=Prompting, Base Model=Qwen2.5-7B-Instruct2026.05 | 30.2 | 10.3 | 3.2 | 14.8 | |
| ReActType=Prompting, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 18.2 | 10.5 | 2 | 12.8 | |
| Direct PromptType=Prompting, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 5.8 | 5.1 | 0 | 4.1 |