Interactive Decision Making on WebShop (test)
97Success RateBPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| BPOApproach=System-2, Evaluation Protocol=Fine-Tuned, Base Model=Llama-3.1-8B-Instruct2025.08 | 97 | — | |
| Qwen-2.5-7B-InstructApproach=System-1, Evaluation Protocol=Zero-shot, Model Category=Base Models2025.08 | 91.5 | — | |
| ETOApproach=System-2, Evaluation Protocol=Fine-Tuned, Base Model=Llama-3.1-8B-Instruct2025.08 | 90 | — | |
| SFTApproach=System-2, Evaluation Protocol=Fine-Tuned, Base Model=Llama-3.1-8B-Instruct2025.08 | 88 | — | |
| MPOApproach=System-2, Evaluation Protocol=Fine-Tuned, Base Model=Llama-3.1-8B-Instruct2025.08 | 87.5 | — | |
| GiGPO + PA-MoEType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 82.8 | 93.1 | |
| GiGPO + PA-MoEType=RL Training, Backbone=Qwen2.5-1.5B-Instruct2026.02 | 82.3 | 91 | |
| GEPOType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 80.5 | 91 | |
| o3-miniApproach=System-2, Evaluation Protocol=Zero-shot, Model Category=Reasoning LLMs2025.08 | 76.5 | — | |
| GEPOType=RL Training, Backbone=Qwen2.5-1.5B-Instruct2026.02 | 75.6 | 89.9 | |
| GiGPO w/o stdType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 75.2 | 86.2 | |
| GiGPOType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 72.8 | 84.4 | |
| Qwen-3-8B-ThinkingApproach=System-2, Evaluation Protocol=Zero-shot, Model Category=Reasoning LLMs2025.08 | 71.5 | — | |
| PPO + PA-MoEType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 70.3 | 82.8 | |
| PPOType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 68.7 | 81.4 | |
| GRPO + PA-MoEType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 68.5 | 81.5 | |
| RLOO + PA-MoEType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 67.8 | 82.1 | |
| GiGPO w/o stdType=RL Training, Backbone=Qwen2.5-1.5B-Instruct2026.02 | 67.4 | 83.5 | |
| GRPOType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 66.1 | 79.3 | |
| RLOOType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.02 | 65.7 | 80.3 | |
| GiGPOType=RL Training, Backbone=Qwen2.5-1.5B-Instruct2026.02 | 65 | 83.1 | |
| Llama-3.1-8B-InstructApproach=System-1, Evaluation Protocol=Zero-shot, Model Category=Base Models2025.08 | 60 | — | |
| RLOO + PA-MoEType=RL Training, Backbone=Qwen2.5-1.5B-Instruct2026.02 | 58.6 | 73.9 | |
| Deepseek-R1Approach=System-2, Evaluation Protocol=Zero-shot, Model Category=Reasoning LLMs2025.08 | 58.5 | — | |
| GRPO + PA-MoEType=RL Training, Backbone=Qwen2.5-1.5B-Instruct2026.02 | 57 | 73.9 | |
| GRPOType=RL Training, Backbone=Qwen2.5-1.5B-Instruct2026.02 | 56.8 | 75.8 | |
| PPO + PA-MoEType=RL Training, Backbone=Qwen2.5-1.5B-Instruct2026.02 | 52.3 | 73.9 | |
| RLOOType=RL Training, Backbone=Qwen2.5-1.5B-Instruct2026.02 | 52.1 | 73.9 | |
| PPOType=RL Training, Backbone=Qwen2.5-1.5B-Instruct2026.02 | 51.5 | 73.8 | |
| Gemini-2.5-ProType=Prompting, Backbone=Closed-Source Model2026.02 | 35.9 | 42.5 | |
| ReflexionType=Prompting, Backbone=Qwen2.5-7B-Instruct2026.02 | 28.8 | 58.1 | |
| GPT-4oType=Prompting, Backbone=Closed-Source Model2026.02 | 23.7 | 31.8 | |
| ReflexionType=Prompting, Backbone=Qwen2.5-1.5B-Instruct2026.02 | 21.9 | 55.8 | |
| ReActType=Prompting, Backbone=Qwen2.5-7B-Instruct2026.02 | 19.5 | 46.2 | |
| ReActType=Prompting, Backbone=Qwen2.5-1.5B-Instruct2026.02 | 11.3 | 40.1 | |
| Qwen2.5Type=Prompting, Backbone=Qwen2.5-7B-Instruct2026.02 | 7.8 | 26.4 | |
| Qwen2.5Type=Prompting, Backbone=Qwen2.5-1.5B-Instruct2026.02 | 5.2 | 23.1 |