Interactive Environment Task Completion on ScienceWorld (Seen)
89.5Average RewardEAGLET + GPT-5
Evaluation Results
| Method | Links | |
|---|---|---|
| EAGLET + GPT-5Type=Executor Agents w/o Training, Planning=Explicit Planning (EAGLET)2025.10 | 89.5 | |
| MPO + GPT-5Type=Executor Agents w/o Training, Planning=Explicit Planning (MPO)2025.10 | 87.8 | |
| EAGLET + GiGPOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (EAGLET)2025.10 | 87.7 | |
| GPT-5Type=Executor Agents w/o Training2025.10 | 87.6 | |
| EAGLET + ETOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (EAGLET)2025.10 | 84.7 | |
| MPO + GiGPOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (MPO)2025.10 | 84.6 | |
| MPO + ETOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (MPO)2025.10 | 83.4 | |
| GiGPOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 83.3 | |
| EAGLET + GPT-4.1Type=Executor Agents w/o Training, Planning=Explicit Planning (EAGLET)2025.10 | 82.6 | |
| WKMType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 82.1 | |
| KnowAgentType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 81.7 | |
| ETOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 81.3 | |
| MPO + GPT-4.1Type=Executor Agents w/o Training, Planning=Explicit Planning (MPO)2025.10 | 80.4 | |
| DeepSeek-V3.1-ThinkType=Executor Agents w/o Training2025.10 | 78.7 | |
| GPT-4.1Type=Executor Agents w/o Training2025.10 | 76.2 | |
| EAGLET + AgentTuningType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (EAGLET)2025.10 | 74.3 | |
| MPO + AgentTuningType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (MPO)2025.10 | 70.2 | |
| AgentTuningType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 65.3 | |
| EAGLET + Llama-3.1-8B-InstructType=Executor Agents w/o Training, Planning=Explicit Planning (EAGLET)2025.10 | 59.3 | |
| DeepSeek-V3.1-Non-ThinkType=Executor Agents w/o Training2025.10 | 57.4 | |
| MPO + Llama-3.1-8B-InstructType=Executor Agents w/o Training, Planning=Explicit Planning (MPO)2025.10 | 56.5 | |
| Llama-3.1-8B-InstructType=Executor Agents w/o Training2025.10 | 47.7 |