Embodied Task Completion on ALFWorld out-of-distribution (test)
95.93Task Success Rate3SPO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| 3SPOModel=Qwen2.5-7B-Instruct, Type=RL Training2026.06 | 95.93 | — | — | |
| 3SPOModel=Qwen2.5-1.5B-Instruct, Type=RL Training2026.06 | 93.18 | — | — | |
| GiGPOModel=Qwen2.5-7B-Instruct, Type=RL Training2026.06 | 92.18 | — | — | |
| HGPOModel=Qwen2.5-7B-Instruct, Type=RL Training2026.06 | 92.05 | — | — | |
| HGPOModel=Qwen2.5-1.5B-Instruct, Type=RL Training2026.06 | 90.16 | — | — | |
| GiGPOModel=Qwen2.5-1.5B-Instruct, Type=RL Training2026.06 | 84.76 | — | — | |
| GRPOModel=Qwen2.5-7B-Instruct, Type=RL Training2026.06 | 76.82 | — | — | |
| PPO (with critic)Model=Qwen2.5-7B-Instruct, Type=RL Training2026.06 | 76.23 | — | — | |
| RLOOModel=Qwen2.5-7B-Instruct, Type=RL Training2026.06 | 73.95 | — | — | |
| GRPOModel=Qwen2.5-1.5B-Instruct, Type=RL Training2026.06 | 70.1 | — | — | |
| RLOOModel=Qwen2.5-1.5B-Instruct, Type=RL Training2026.06 | 68.7 | — | — | |
| Gemini-2.5-ProModel=Closed, Type=Prompting2026.06 | 50.5 | — | — | |
| GPT-4oModel=Closed, Type=Prompting2026.06 | 46 | — | — | |
| Meta-RLFramework=LAMER, Training Mode=Meta Reinforcement Learning2025.12 | — | 0.81 | 0.502 | |
| PromptingType=Prompting-based2025.12 | — | 0.428 | 0.212 | |
| RLFramework=LAMER, Training Mode=Reinforcement Learning2025.12 | — | 0.581 | 0.36 |