Web Shopping Agent on WebShop (Score, SR, Steps)
82.9Success Rate (SR)Skill1
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Skill1Training Protocol=Skill-Augmented RL Methods2026.05 | 82.9 | 89.7 | — | — | — | — | — | — | |
| SDARTraining Protocol=On-Policy Self-Distillation Methods2026.05 | 82.8 | 89.4 | — | — | — | — | — | — | |
| RetroAgentTraining Protocol=Skill-Augmented RL Methods2026.05 | 82.3 | 88.9 | — | — | — | — | — | — | |
| D2SkillTraining Protocol=Skill-Augmented RL Methods2026.05 | 80.5 | 91.1 | — | — | — | — | — | — | |
| RLSDTraining Protocol=On-Policy Self-Distillation Methods2026.05 | 77.3 | 87.4 | — | — | — | — | — | — | |
| GRPO+OPSDTraining Protocol=On-Policy Self-Distillation Methods2026.05 | 76.5 | 86.8 | — | — | — | — | — | — | |
| Skill-SDTraining Protocol=On-Policy Self-Distillation Methods2026.05 | 76.5 | 86.1 | — | — | — | — | — | — | |
| SKILLCTraining Protocol=Skill-Internalization RL Methods2026.05 | 74 | 85.6 | — | — | — | — | — | — | |
| GiGPOTraining Protocol=Direct RL Methods2026.05 | 72.8 | 84.4 | — | — | — | — | — | — | |
| SKILLRLTraining Protocol=Skill-Augmented RL Methods2026.05 | 72.7 | 85.2 | — | — | — | — | — | — | |
| SKILL0Training Protocol=Skill-Internalization RL Methods2026.05 | 70.9 | 83.3 | — | — | — | — | — | — | |
| AgentLM-13BCategory=Open LLMs, Parameters=13B2026.07 | 70.8 | — | — | — | — | — | — | — | |
| PPOTraining Protocol=Direct RL Methods2026.05 | 68.7 | 81.4 | — | — | — | — | — | — | |
| GRPOTraining Protocol=Direct RL Methods2026.05 | 66.1 | 79.3 | — | — | — | — | — | — | |
| RLOOTraining Protocol=Direct RL Methods2026.05 | 65.7 | 80.3 | — | — | — | — | — | — | |
| AgentLM-70BCategory=Open LLMs, Parameters=70B2026.07 | 64.9 | — | — | — | — | — | — | — | |
| Hephaestus-8B-IFTCategory=Open LLMs, Parameters=8B, Variant=IFT2026.07 | 63.9 | — | — | — | — | — | — | — | |
| AgentLM-7BCategory=Open LLMs, Parameters=7B2026.07 | 63.6 | — | — | — | — | — | — | — | |
| Hephaestus-8B-BaseCategory=Open LLMs, Parameters=8B, Variant=Base2026.07 | 60.5 | — | — | — | — | — | — | — | |
| EPPOCategory=RL Training2026.07 | 57.5 | — | — | — | — | — | — | — | |
| AgentRLCategory=RL Training2026.07 | 51.3 | — | — | — | — | — | — | — | |
| DCPOCategory=RL Training2026.07 | 50.5 | — | — | — | — | — | — | — | |
| BAPOCategory=RL Training2026.07 | 48.2 | — | — | — | — | — | — | — | |
| Reflexion2026.06 | 45 | 69 | — | — | — | — | — | — | |
| UCEExperience Level=peak2026.06 | 42 | 61.3 | — | — | — | — | — | — | |
| Plan-and-Act + ECUExperience Level=peak2026.06 | 42 | 62.7 | — | — | — | — | — | — | |
| SKillOSExecutor=Gemini-2.5-pro, Curator=Qwen3-8B2026.05 | 41.3 | 56 | 18.3 | — | — | — | — | — | |
| SKillOS-geminiExecutor=Gemini-2.5-pro, Curator=Gemini-2.5-pro2026.05 | 41 | 54.7 | 17.8 | — | — | — | — | — | |
| NoThinking + ECUExperience Level=peak2026.06 | 41 | 66.8 | — | — | — | — | — | — | |
| ReasoningBankExecutor=Gemini-2.5-pro, Curator=Gemini-2.5-pro2026.05 | 40.2 | 50.8 | 19.2 | — | — | — | — | — | |
| NoThinking + ECUExperience Level=initial2026.06 | 40 | 65.6 | — | — | — | — | — | — | |
| MemPExecutor=Gemini-2.5-pro, Curator=Gemini-2.5-pro2026.05 | 39.8 | 51.3 | 19.4 | — | — | — | — | — | |
| SKillOS-baseExecutor=Gemini-2.5-pro, Curator=Qwen3-8B2026.05 | 39.6 | 52.8 | 19 | — | — | — | — | — | |
| ReflAct + ECUExperience Level=peak2026.06 | 39 | 61.5 | — | — | — | — | — | — | |
| No MemoryExecutor=Gemini-2.5-pro, Curator=None2026.05 | 38.4 | 48.6 | 19.5 | — | — | — | — | — | |
| Claude-Sonnet-4 ThinkingCategory=API LLMs, Variant=Thinking2026.07 | 38.3 | — | — | — | — | — | — | — | |
| NoThinking2026.06 | 37 | 58.7 | — | — | — | — | — | — | |
| Plan-and-Act + ECUExperience Level=initial2026.06 | 37 | 56.8 | — | — | — | — | — | — | |
| Claude-Sonnet-4Category=API LLMs2026.07 | 34.6 | — | — | — | — | — | — | — | |
| ReflAct + ECUExperience Level=initial2026.06 | 34 | 57.4 | — | — | — | — | — | — | |
| GPT-5Category=API LLMs2026.07 | 33.7 | — | — | — | — | — | — | — | |
| o3-miniCategory=API LLMs2026.07 | 32.7 | — | — | — | — | — | — | — | |
| Plan-and-Act2026.06 | 31 | 52 | — | — | — | — | — | — | |
| DeepSeek-R1Category=Open LLMs2026.07 | 31 | — | — | — | — | — | — | — | |
| ReAct2026.06 | 30 | 45.1 | — | — | — | — | — | — | |
| UCEExperience Level=initial2026.06 | 29 | 51 | — | — | — | — | — | — | |
| ReflexionTraining Protocol=Training-free Methods2026.05 | 28.8 | 58.1 | — | — | — | — | — | — | |
| o4-miniCategory=API LLMs2026.07 | 28.5 | — | — | — | — | — | — | — | |
| Qwen2.5-32B-InstructCategory=Open LLMs, Parameters=32B, Variant=Instruct2026.07 | 27.5 | — | — | — | — | — | — | — | |
| ReflAct2026.06 | 24 | 39.6 | — | — | — | — | — | — | |
| DeepSeek-V3Category=Open LLMs2026.07 | 23.4 | — | — | — | — | — | — | — | |
| ReActTraining Protocol=Training-free Methods2026.05 | 19.5 | 46.2 | — | — | — | — | — | — | |
| ExpeL2026.06 | 18 | 31.6 | — | — | — | — | — | — | |
| EvolveRTraining Protocol=Skill-Augmented RL Methods2026.05 | 17.6 | 42.5 | — | — | — | — | — | — | |
| Qwen2.5-14B-InstructCategory=Open LLMs, Parameters=14B, Variant=Instruct2026.07 | 17.6 | — | — | — | — | — | — | — | |
| SKillOSExecutor=Qwen3-8B, Curator=Qwen3-8B2026.05 | 16.5 | 40.6 | 19.4 | — | — | — | — | — | |
| SKillOSExecutor=Qwen3-32B, Curator=Qwen3-8B2026.05 | 16.5 | 49.2 | 15.9 | — | — | — | — | — | |
| SKillOS-baseExecutor=Qwen3-8B, Curator=Qwen3-8B2026.05 | 13.6 | 38.6 | 20.1 | — | — | — | — | — | |
| SKillOS-geminiExecutor=Qwen3-8B, Curator=Gemini-2.5-pro2026.05 | 13.2 | 38.1 | 19.6 | — | — | — | — | — | |
| SKillOS-geminiExecutor=Qwen3-32B, Curator=Gemini-2.5-pro2026.05 | 13.2 | 45.2 | 16.6 | — | — | — | — | — | |
| SKillOS-baseExecutor=Qwen3-32B, Curator=Qwen3-8B2026.05 | 12.3 | 43.4 | 16.8 | — | — | — | — | — | |
| No MemoryExecutor=Qwen3-32B, Curator=None2026.05 | 12.2 | 41.5 | 17 | — | — | — | — | — | |
| MemPExecutor=Qwen3-8B, Curator=Qwen3-8B2026.05 | 12 | 35.7 | 21.3 | — | — | — | — | — | |
| ReasoningBankExecutor=Qwen3-8B, Curator=Qwen3-8B2026.05 | 11.4 | 35.4 | 20.5 | — | — | — | — | — | |
| ReasoningBankExecutor=Qwen3-32B, Curator=Qwen3-32B2026.05 | 11.2 | 40.4 | 17.9 | — | — | — | — | — | |
| ExpeLTraining Protocol=Training-free Methods2026.05 | 11.2 | 30.9 | — | — | — | — | — | — | |
| MemPExecutor=Qwen3-32B, Curator=Qwen3-32B2026.05 | 10.1 | 30.7 | 17.4 | — | — | — | — | — | |
| No MemoryExecutor=Qwen3-8B, Curator=None2026.05 | 9.8 | 33.3 | 20.3 | — | — | — | — | — | |
| Zero-ShotTraining Protocol=Training-free Methods2026.05 | 7.8 | 26.4 | — | — | — | — | — | — | |
| Qwen2.5-3B-InstructCategory=Open LLMs, Parameters=3B, Variant=Instruct2026.07 | 5.3 | — | — | — | — | — | — | — | |
| OPSDTraining Protocol=On-Policy Self-Distillation Methods2026.05 | 2.3 | 4.5 | — | — | — | — | — | — | |
| Mem0Training Protocol=Training-free Methods2026.05 | 2 | 23.9 | — | — | — | — | — | — | |
| EvolveRMethod Category=Memory-Augmented RL Methods2026.05 | — | — | — | 32.5 | 31.1 | 25 | 20.9 | 28 | |
| ExpeLMethod Category=Prompt-based Agentic or Memory-based Methods2026.05 | — | — | — | 6.2 | 12.5 | 9 | 21 | 12.1 | |
| Few-shotMethod Category=Prompt-based Methods2026.05 | — | — | — | 14.2 | 15.1 | 13.5 | 24 | 16.5 | |
| GRPOMethod Category=RL-based Methods2026.05 | — | — | — | 35.1 | 22.6 | 39 | 49.5 | 33.6 | |
| Mem0Method Category=Prompt-based Agentic or Memory-based Methods2026.05 | — | — | — | 8.9 | 9.2 | 2.3 | 11 | 8.2 | |
| Mem0+GRPOMethod Category=Memory-Augmented RL Methods2026.05 | — | — | — | 36.5 | 22 | 40.6 | 23.1 | 29.5 | |
| MemPMethod Category=Prompt-based Agentic or Memory-based Methods2026.05 | — | — | — | 15.9 | 13.2 | 9 | 19 | 14.3 | |
| MemRLMethod Category=Memory-Augmented RL Methods2026.05 | — | — | — | 22.2 | 15.2 | 25 | 48.3 | 26.2 | |
| ReActMethod Category=Prompt-based Agentic or Memory-based Methods2026.05 | — | — | — | 12.4 | 11.1 | 4.5 | 12 | 10.4 | |
| ReflexionMethod Category=Prompt-based Agentic or Memory-based Methods2026.05 | — | — | — | 3.5 | 8.6 | 1.1 | 7 | 5.5 | |
| RLOOMethod Category=RL-based Methods2026.05 | — | — | — | 34.9 | 23.9 | 41.5 | 32.1 | 31.1 | |
| SimpleMemMethod Category=Prompt-based Agentic or Memory-based Methods2026.05 | — | — | — | 11.5 | 13.8 | 6.7 | 13 | 11.7 | |
| SimpleMem+GRPOMethod Category=Memory-Augmented RL Methods2026.05 | — | — | — | 25.4 | 28 | 26.6 | 25.1 | 26.4 | |
| SKILL0Method Category=Skill-Augmented RL Methods2026.05 | — | — | — | 39.2 | 33 | 38.1 | 37.9 | 35.2 | |
| Skill0.5Method Category=Skill-Augmented RL Methods2026.05 | — | — | — | 39.1 | 37.3 | 41.1 | 50.9 | 40.4 | |
| SkillRLMethod Category=Skill-Augmented RL Methods2026.05 | — | — | — | 36 | 34.2 | 41.4 | 49.3 | 38.3 | |
| SLIMMethod Category=Skill-Augmented RL Methods2026.05 | — | — | — | 31.9 | 36.8 | 31.5 | 33 | 33.7 | |
| Zero-shotMethod Category=Prompt-based Methods2026.05 | — | — | — | 4.4 | 4.6 | 3.4 | 1 | 3.5 |