Interactive Environment Task Completion on ScienceWorld (Unseen)
90.1Average RewardEAGLET + GPT-5
Evaluation Results
| Method | Links | |
|---|---|---|
| EAGLET + GPT-5Type=Executor Agents w/o Training, Planning=Explicit Planning (EAGLET)2025.10 | 90.1 | |
| MPO + GPT-5Type=Executor Agents w/o Training, Planning=Explicit Planning (MPO)2025.10 | 89 | |
| GPT-5Type=Executor Agents w/o Training2025.10 | 88.2 | |
| EAGLET + GPT-4.1Type=Executor Agents w/o Training, Planning=Explicit Planning (EAGLET)2025.10 | 85.6 | |
| MPO + GPT-4.1Type=Executor Agents w/o Training, Planning=Explicit Planning (MPO)2025.10 | 83.8 | |
| EAGLET + GiGPOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (EAGLET)2025.10 | 83.6 | |
| EAGLET + ETOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (EAGLET)2025.10 | 82.5 | |
| MPO + ETOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (MPO)2025.10 | 80.8 | |
| GPT-4.1Type=Executor Agents w/o Training2025.10 | 79.9 | |
| MPO + GiGPOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (MPO)2025.10 | 78.2 | |
| WKMType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 76.5 | |
| DeepSeek-V3.1-ThinkType=Executor Agents w/o Training2025.10 | 76.2 | |
| GiGPOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 74.5 | |
| ETOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 74.1 | |
| Q-EvolveBackbone=Llama-2-7B-Chat2026.06 | 69.7 | |
| KnowAgentType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 69.6 | |
| EAGLET + AgentTuningType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (EAGLET)2025.10 | 68.4 | |
| QLASSBackbone=Llama-2-7B-Chat2026.06 | 66.4 | |
| MPO + AgentTuningType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (MPO)2025.10 | 65.9 | |
| ETOBackbone=Llama-2-7B-Chat2026.06 | 65 | |
| GPT-4Prompting=ReAct2026.06 | 64.4 | |
| ReflexionPrompting=ReAct2026.06 | 64.4 | |
| DMPOBackbone=Llama-2-7B-Chat2026.06 | 61.7 | |
| EAGLET + Llama-3.1-8B-InstructType=Executor Agents w/o Training, Planning=Explicit Planning (EAGLET)2025.10 | 61.6 | |
| DeepSeek-V3.1-Non-ThinkType=Executor Agents w/o Training2025.10 | 58.1 | |
| Best-of-NBackbone=Llama-2-7B-Chat, N=62026.06 | 57.6 | |
| AgentTuningType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 57 | |
| MPO + Llama-3.1-8B-InstructType=Executor Agents w/o Training, Planning=Explicit Planning (MPO)2025.10 | 55.5 | |
| RFTBackbone=Llama-2-7B-Chat2026.06 | 54.3 | |
| SFTBackbone=Llama-2-7B-Chat2026.06 | 53 | |
| PPOBackbone=Llama-2-7B-Chat2026.06 | 51.7 | |
| Llama-3.1-8B-InstructType=Executor Agents w/o Training2025.10 | 42.2 | |
| GPT-3.5-TurboPrompting=ReAct2026.06 | 13 | |
| Base AgentBackbone=Llama-2-7B-Chat2026.06 | 3.1 |