Online Shopping on WebShop (test)
95ScoreSkillMaster
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SkillMasterCategory=Memory-Augmented RL-based Methods2026.05 | 95 | 82 | |
| SERLType=Hybrid2026.05 | 89.5 | 80.1 | |
| GraphGPOType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 89.29 | 78.65 | |
| HGPOType=RL Training2026.05 | 88.4 | 77.8 | |
| GRPO+SDPO(Loss)Type=Hybrid2026.05 | 88.4 | 73 | |
| GiGPOType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 87.94 | 73.83 | |
| GraphGPOType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.05 | 86.94 | 80.31 | |
| GiGPOType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.05 | 86.72 | 78.38 | |
| SkillRLCategory=Memory-Augmented RL-based Methods2026.05 | 85.2 | 72.7 | |
| GRPO+SDPO(Advantage)Type=Hybrid2026.05 | 84.8 | 75.4 | |
| GRPOType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 84.73 | 71.35 | |
| GRPOType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.05 | 84.31 | 75 | |
| RLSDType=Hybrid2026.05 | 83.6 | 75.8 | |
| GIGPOType=RL Training2026.05 | 83.5 | 75.8 | |
| PPO (with critic)Type=RL Training2026.05 | 81.4 | 68.7 | |
| PPOType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.05 | 81.4 | 68.7 | |
| RLOOCategory=RL-based Methods2026.05 | 80.3 | 65.7 | |
| RLOOType=RL Training2026.05 | 80.3 | 65.7 | |
| RLOOType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.05 | 80.3 | 65.7 | |
| GRPOCategory=RL-based Methods2026.05 | 79.3 | 66.1 | |
| RLOOType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 73.9 | 52.1 | |
| PPOType=RL Training, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 73.8 | 51.5 | |
| GRPOType=RL Training2026.05 | 73.1 | 64.1 | |
| SimpleMem+GRPOCategory=Memory-Augmented RL-based Methods2026.05 | 67.8 | 46.9 | |
| SkillBrewBackbone=Qwen2.5-7B-Instruct2026.05 | 59.3 | 38.4 | |
| ReflexionCategory=Prompt-based Agentic or Memory-based Methods2026.05 | 58.1 | 28.8 | |
| Mem0+GRPOCategory=Memory-Augmented RL-based Methods2026.05 | 58.1 | 37.5 | |
| ReflexionType=Prompting2026.05 | 58.1 | 28.8 | |
| ReflexionType=Prompting, Base Model=Qwen2.5-7B-Instruct2026.05 | 58.1 | 28.8 | |
| ReflexionBackbone=Qwen2.5-7B-Instruct2026.05 | 58.1 | 28.8 | |
| ReflexionType=Prompting, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 55.8 | 21.9 | |
| VoyagerBackbone=Qwen2.5-7B-Instruct2026.05 | 50 | 26.4 | |
| ReActCategory=Prompt-based Agentic or Memory-based Methods2026.05 | 46.2 | 19.5 | |
| ReActType=Prompting2026.05 | 46.2 | 19.5 | |
| ReActType=Prompting, Base Model=Qwen2.5-7B-Instruct2026.05 | 46.2 | 19.5 | |
| ReActBackbone=Qwen2.5-7B-Instruct2026.05 | 46.2 | 19.5 | |
| Gemini-2.5-ProCategory=Closed-source LLMs2026.05 | 42.5 | 35.9 | |
| EvolveRCategory=Memory-Augmented RL-based Methods2026.05 | 42.5 | 17.6 | |
| Gemini-2.5-ProType=Prompting, Base Model=Closed-Source2026.05 | 42.5 | 35.9 | |
| EvoSkillBackbone=Qwen2.5-7B-Instruct2026.05 | 42.2 | 12.6 | |
| ReActType=Prompting, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 40.1 | 11.3 | |
| Skill-ProBackbone=Qwen2.5-7B-Instruct2026.05 | 38.7 | 23.2 | |
| SimpleMemCategory=Prompt-based Agentic or Memory-based Methods2026.05 | 33.2 | 8.59 | |
| SimpleMemBackbone=Qwen2.5-7B-Instruct2026.05 | 33.2 | 8.59 | |
| GPT-4oCategory=Closed-source LLMs2026.05 | 31.8 | 23.7 | |
| GPT-4oType=Prompting, Base Model=Closed-Source2026.05 | 31.8 | 23.7 | |
| ExpeLCategory=Prompt-based Agentic or Memory-based Methods2026.05 | 30.9 | 11.2 | |
| ExpeLBackbone=Qwen2.5-7B-Instruct2026.05 | 30.9 | 11.2 | |
| MemRLCategory=Memory-Augmented RL-based Methods2026.05 | 29.5 | 9.2 | |
| Qwen2.5-7B-InstructCategory=Prompt-based Agentic or Memory-based Methods2026.05 | 26.4 | 7.8 | |
| Qwen2.5-7B-InstructType=Prompting2026.05 | 26.4 | 7.8 | |
| Qwen2.5Type=Prompting, Base Model=Qwen2.5-7B-Instruct2026.05 | 26.4 | 7.8 | |
| Zero-ShotBackbone=Qwen2.5-7B-Instruct2026.05 | 26.4 | 7.8 | |
| MemPCategory=Prompt-based Agentic or Memory-based Methods2026.05 | 25.3 | 6.4 | |
| MemPBackbone=Qwen2.5-7B-Instruct2026.05 | 25.3 | 6.4 | |
| Mem0Category=Prompt-based Agentic or Memory-based Methods2026.05 | 23.9 | 2 | |
| Mem0Backbone=Qwen2.5-7B-Instruct2026.05 | 23.9 | 2 | |
| Qwen2.5Type=Prompting, Base Model=Qwen2.5-1.5B-Instruct2026.05 | 23.1 | 5.2 | |
| SDPOType=Self-Distillation2026.05 | 0 | 0 |