Interactive Environment Task Completion on ALFWorld (Unseen)
91.8Average RewardEAGLET + GiGPO
Evaluation Results
| Method | Links | |
|---|---|---|
| EAGLET + GiGPOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (EAGLET)2025.10 | 91.8 | |
| EAGLET + GPT-5Type=Executor Agents w/o Training, Planning=Explicit Planning (EAGLET)2025.10 | 90.7 | |
| BPOApproach=System-2, Evaluation Protocol=Fine-Tuned, Base Model=Llama-3.1-8B-Instruct2025.08 | 89.6 | |
| GiGPOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 88.6 | |
| MPO + GiGPOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (MPO)2025.10 | 88.1 | |
| GPT-5Type=Executor Agents w/o Training2025.10 | 83.6 | |
| MPO + GPT-5Type=Executor Agents w/o Training, Planning=Explicit Planning (MPO)2025.10 | 83.6 | |
| EAGLET + GPT-4.1Type=Executor Agents w/o Training, Planning=Explicit Planning (EAGLET)2025.10 | 83.6 | |
| EAGLET + AgentTuningType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (EAGLET)2025.10 | 83.4 | |
| EAGLET + ETOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (EAGLET)2025.10 | 83.2 | |
| MPO + AgentTuningType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (MPO)2025.10 | 81.3 | |
| MPO + GPT-4.1Type=Executor Agents w/o Training, Planning=Explicit Planning (MPO)2025.10 | 79.1 | |
| MPO + ETOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (MPO)2025.10 | 79.1 | |
| MPOApproach=System-2, Evaluation Protocol=Fine-Tuned, Base Model=Llama-3.1-8B-Instruct2025.08 | 78.4 | |
| WKMType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 78.2 | |
| ETOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 76.4 | |
| Qwen-2.5-7B-InstructApproach=System-1, Evaluation Protocol=Zero-shot2025.08 | 76.1 | |
| KnowAgentType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 74.9 | |
| GPT-4.1Type=Executor Agents w/o Training2025.10 | 72.4 | |
| AgentTuningType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 71.6 | |
| SFTApproach=System-2, Evaluation Protocol=Fine-Tuned, Base Model=Llama-3.1-8B-Instruct2025.08 | 71.6 | |
| ETOApproach=System-2, Evaluation Protocol=Fine-Tuned, Base Model=Llama-3.1-8B-Instruct2025.08 | 71.6 | |
| DeepSeek-V3.1-ThinkType=Executor Agents w/o Training2025.10 | 69.4 | |
| Qwen-3-ThinkingApproach=System-2, Evaluation Protocol=Zero-shot2025.08 | 67.2 | |
| o3-miniApproach=System-2, Evaluation Protocol=Zero-shot2025.08 | 62.7 | |
| EAGLET + Llama-3.1-8B-InstructType=Executor Agents w/o Training, Planning=Explicit Planning (EAGLET)2025.10 | 55.2 | |
| Deepseek-R1Approach=System-2, Evaluation Protocol=Zero-shot2025.08 | 53.7 | |
| MPO + Llama-3.1-8B-InstructType=Executor Agents w/o Training, Planning=Explicit Planning (MPO)2025.10 | 52.2 | |
| Llama-3.1-8B-InstructApproach=System-1, Evaluation Protocol=Zero-shot2025.08 | 40.3 | |
| DeepSeek-V3.1-Non-ThinkType=Executor Agents w/o Training2025.10 | 37.3 | |
| Llama-3.1-8B-InstructType=Executor Agents w/o Training2025.10 | 28.4 |