Interactive Environment Task Completion on WebShop (Seen)
86.2Average RewardEAGLET + GiGPO
Evaluation Results
| Method | Links | |
|---|---|---|
| EAGLET + GiGPOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (EAGLET)2025.10 | 86.2 | |
| MPO + GiGPOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (MPO)2025.10 | 83.5 | |
| GiGPOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 82.5 | |
| EAGLET + GPT-5Type=Executor Agents w/o Training, Planning=Explicit Planning (EAGLET)2025.10 | 80.1 | |
| MPO + GPT-5Type=Executor Agents w/o Training, Planning=Explicit Planning (MPO)2025.10 | 76.6 | |
| GPT-5Type=Executor Agents w/o Training2025.10 | 75.3 | |
| EAGLET + GPT-4.1Type=Executor Agents w/o Training, Planning=Explicit Planning (EAGLET)2025.10 | 74.7 | |
| EAGLET + ETOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (EAGLET)2025.10 | 72.6 | |
| MPO + GPT-4.1Type=Executor Agents w/o Training, Planning=Explicit Planning (MPO)2025.10 | 72.5 | |
| DeepSeek-V3.1-ThinkType=Executor Agents w/o Training2025.10 | 70.8 | |
| GPT-4.1Type=Executor Agents w/o Training2025.10 | 70.2 | |
| MPO + ETOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (MPO)2025.10 | 70.2 | |
| ETOType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 68.4 | |
| WKMType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 66.9 | |
| EAGLET + Llama-3.1-8B-InstructType=Executor Agents w/o Training, Planning=Explicit Planning (EAGLET)2025.10 | 66.7 | |
| EAGLET + AgentTuningType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (EAGLET)2025.10 | 66.7 | |
| MPO + AgentTuningType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct, Planning=Explicit Planning (MPO)2025.10 | 65.5 | |
| KnowAgentType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 64.8 | |
| AgentTuningType=Executor Agents w/ Training, Base Model=Llama-3.1-8B-Instruct2025.10 | 63.3 | |
| MPO + Llama-3.1-8B-InstructType=Executor Agents w/o Training, Planning=Explicit Planning (MPO)2025.10 | 63.2 | |
| DeepSeek-V3.1-Non-ThinkType=Executor Agents w/o Training2025.10 | 58.6 | |
| Llama-3.1-8B-InstructType=Executor Agents w/o Training2025.10 | 56.3 |