Interactive Environment Task Completion on ALFWorld (Seen)
90.2Average RewardEAGLET + GPT-5
Evaluation Results
| Method | Links | |
|---|---|---|
| EAGLET + GPT-5Type=Executor Agents w/o Training, Planning=Explicit Planning (EAGLET)2025.10 | 90.2 | |
| EAGLET + GiGPOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (EAGLET)2025.10 | 88.6 | |
| MPO + GPT-5Type=Executor Agents w/o Training, Planning=Explicit Planning (MPO)2025.10 | 88.2 | |
| GPT-5Type=Executor Agents w/o Training2025.10 | 87.9 | |
| BPOApproach=System-2, Evaluation Protocol=Fine-Tuned, Base Model=Llama-3.1-8B-Instruct2025.08 | 87.9 | |
| EAGLET + ETOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (EAGLET)2025.10 | 87.3 | |
| MPO + GiGPOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (MPO)2025.10 | 86.6 | |
| GiGPOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 85.2 | |
| MPO + ETOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (MPO)2025.10 | 85 | |
| EAGLET + GPT-4.1Type=Executor Agents w/o Training, Planning=Explicit Planning (EAGLET)2025.10 | 84.3 | |
| MPOApproach=System-2, Evaluation Protocol=Fine-Tuned, Base Model=Llama-3.1-8B-Instruct2025.08 | 82.9 | |
| EAGLET + AgentTuningType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (EAGLET)2025.10 | 82.3 | |
| DeepSeek-V3.1-ThinkType=Executor Agents w/o Training2025.10 | 81.4 | |
| MPO + GPT-4.1Type=Executor Agents w/o Training, Planning=Explicit Planning (MPO)2025.10 | 81.4 | |
| MPO + AgentTuningType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (MPO)2025.10 | 80.7 | |
| KnowAgentType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 80 | |
| SFTApproach=System-2, Evaluation Protocol=Fine-Tuned, Base Model=Llama-3.1-8B-Instruct2025.08 | 80 | |
| AgentTuningType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 79.3 | |
| GPT-4.1Type=Executor Agents w/o Training2025.10 | 78.6 | |
| ETOApproach=System-2, Evaluation Protocol=Fine-Tuned, Base Model=Llama-3.1-8B-Instruct2025.08 | 78.6 | |
| WKMType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 77.5 | |
| ETOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 77.1 | |
| Qwen-2.5-7B-InstructApproach=System-1, Evaluation Protocol=Zero-shot2025.08 | 72.1 | |
| Deepseek-R1Approach=System-2, Evaluation Protocol=Zero-shot2025.08 | 61.4 | |
| o3-miniApproach=System-2, Evaluation Protocol=Zero-shot2025.08 | 58.9 | |
| EAGLET + Llama-3.1-8B-InstructType=Executor Agents w/o Training, Planning=Explicit Planning (EAGLET)2025.10 | 54.3 | |
| DeepSeek-V3.1-Non-ThinkType=Executor Agents w/o Training2025.10 | 50 | |
| MPO + Llama-3.1-8B-InstructType=Executor Agents w/o Training, Planning=Explicit Planning (MPO)2025.10 | 50 | |
| Qwen-3-ThinkingApproach=System-2, Evaluation Protocol=Zero-shot2025.08 | 49.3 | |
| Llama-3.1-8B-InstructApproach=System-1, Evaluation Protocol=Zero-shot2025.08 | 32.9 | |
| Llama-3.1-8B-InstructType=Executor Agents w/o Training2025.10 | 22.9 |