Web navigation on WebShop
90.5Average ScoreSAPO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| SAPOMethod Category=Memory-Augmented RL-based Methods2026.06 | 90.5 | — | 78.1 | |
| SKILLRLMethod Category=Memory-Augmented RL-based Methods2026.06 | 85.2 | — | 72.7 | |
| D2SkillMethod Category=Memory-Augmented RL-based Methods2026.06 | 83.4 | — | 73.4 | |
| Skill0Method Category=Memory-Augmented RL-based Methods2026.06 | 83.2 | — | 71.9 | |
| RLOO*Method Category=RL-based Methods2026.06 | 80.3 | — | 65.7 | |
| GRPO*Method Category=RL-based Methods2026.06 | 79.3 | — | 66.1 | |
| IPRRefinement Category=Process Refinement, Backbone=Llama-2-7B2025.12 | 71.3 | — | — | |
| Q-EvolveBackbone=Llama-2-7B-Chat2026.06 | 70.5 | — | — | |
| QLASSBackbone=Llama-2-7B-Chat2026.06 | 70.3 | — | — | |
| MACLARefinement Category=Process Refinement, Backbone=Llama-2-7B2025.12 | 70.2 | — | — | |
| DMPOBackbone=Llama-2-7B-Chat2026.06 | 70.1 | — | — | |
| Best-of-NBackbone=Llama-2-7B-Chat, N=62026.06 | 67.9 | — | — | |
| SimpleMem+GRPOMethod Category=Memory-Augmented RL-based Methods2026.06 | 67.8 | — | 46.9 | |
| ETORefinement Category=Outcome Refinement, Backbone=Llama-2-7B2025.12 | 67.4 | — | — | |
| ETOBackbone=Llama-2-7B-Chat2026.06 | 67.4 | — | — | |
| Co-Evolving AgentsBackbone=Qwen3-4B-Instruct-25072025.11 | 66.3 | 72.5 | — | |
| RFT-PPORefinement Category=Outcome Refinement, Backbone=Llama-2-7B2025.12 | 64.2 | — | — | |
| ReflexionPrompting=ReAct2026.06 | 64.2 | — | — | |
| PPOBackbone=Llama-2-7B-Chat2026.06 | 64.2 | — | — | |
| Step-PPORefinement Category=Process Refinement, Backbone=Llama-2-7B2025.12 | 64 | — | — | |
| RFT-CRRefinement Category=Outcome Refinement, Backbone=Llama-2-7B2025.12 | 63.6 | — | — | |
| RFTBackbone=Llama-2-7B-Chat2026.06 | 63.6 | — | — | |
| GPT-4Refinement Category=Prompt-based, Backbone=GPT-42025.12 | 63.2 | — | — | |
| GPT-4Prompting=ReAct2026.06 | 63.2 | — | — | |
| SFTBackbone=Llama-2-7B-Chat2026.06 | 63.1 | — | — | |
| GPT-3.5-TurboRefinement Category=Prompt-based, Backbone=GPT-3.5-Turbo2025.12 | 62.4 | — | — | |
| GPT-3.5-TurboPrompting=ReAct2026.06 | 62.4 | — | — | |
| SFTRefinement Category=Outcome Refinement, Backbone=Llama-2-7B2025.12 | 60.2 | — | — | |
| ETOBackbone=Qwen3-4B-Instruct-25072025.11 | 59.5 | 65.7 | — | |
| Reflexion*Method Category=Prompt-based Agentic or Memory-based Methods2026.06 | 58.1 | — | 28.8 | |
| Mem0+GRPOMethod Category=Memory-Augmented RL-based Methods2026.06 | 58.1 | — | 37.5 | |
| ReAct*Method Category=Prompt-based Agentic or Memory-based Methods2026.06 | 46.2 | — | 19.5 | |
| Gemini-2.5-ProMethod Category=Closed-source LLMs2026.06 | 42.5 | — | 35.9 | |
| EvolveRMethod Category=Memory-Augmented RL-based Methods2026.06 | 42.5 | — | 17.6 | |
| SFTBackbone=Qwen3-4B-Instruct-25072025.11 | 40.3 | 63.9 | — | |
| SimpleMemMethod Category=Prompt-based Agentic or Memory-based Methods2026.06 | 33.2 | — | 8.59 | |
| GPT-4oMethod Category=Closed-source LLMs2026.06 | 31.8 | — | 23.7 | |
| ExpeLMethod Category=Prompt-based Agentic or Memory-based Methods2026.06 | 30.9 | — | 11.2 | |
| MemRLMethod Category=Memory-Augmented RL-based Methods2026.06 | 29.5 | — | 9.2 | |
| Qwen2.5Method Category=Qwen2.5-7B-Instruct2026.06 | 26.4 | — | 7.8 | |
| MemPMethod Category=Prompt-based Agentic or Memory-based Methods2026.06 | 25.3 | — | 6.4 | |
| Mem0Method Category=Prompt-based Agentic or Memory-based Methods2026.06 | 23.9 | — | 2 | |
| InternVL2.5-8BModel=InternVL2.5-8B, Strategy=Zero-Shot2026.05 | 23.5 | — | 3.9 | |
| Qwen2.5-VL-7BModel=Qwen2.5-VL-7B, Strategy=Zero-Shot2026.05 | 19.4 | — | 4.1 | |
| Llama-2-7BRefinement Category=Prompt-based, Backbone=Llama-2-7B2025.12 | 17.9 | — | — | |
| Base AgentBackbone=Llama-2-7B-Chat2026.06 | 17.9 | — | — | |
| InternVL2.5-8BModel=InternVL2.5-8B, Strategy=GPT Trajectory Imitation2026.05 | 16.5 | — | 8.8 | |
| GPT-4oModel=GPT-4o, Strategy=Zero-Shot2026.05 | 15.5 | — | 17.2 | |
| ViGoRLModel=ViGoRL, Strategy=Zero-Shot2026.05 | 13.4 | — | 5.6 | |
| Qwen2.5-VL-7BModel=Qwen2.5-VL-7B, Strategy=OS-Genesis2026.05 | 12.7 | — | 11.2 | |
| SCALE-20kModel=LLaVA-NeXT-8B, Strategy=SCALE-20k2026.05 | 10.1 | — | 2.1 | |
| Qwen2.5-VL-7BModel=Qwen2.5-VL-7B, Strategy=GPT Trajectory Imitation2026.05 | 8.5 | — | 18.3 | |
| InternVL2.5-8BModel=InternVL2.5-8B, Strategy=OS-Genesis2026.05 | 6.6 | — | 11.6 | |
| SCALEModel=Qwen2.5-VL-7B, Strategy=SCALE(ours)2026.05 | 5.7 | — | 14.4 | |
| SCALEModel=InternVL2.5-8B, Strategy=SCALE(ours)2026.05 | 4.9 | — | 11 | |
| AugvisModel=Augvis, Strategy=Zero-Shot2026.05 | — | — | 4.3 | |
| Critique-GRPOBackbone=Qwen3-4B2026.05 | — | 87.1 | 72.5 | |
| Critique-GRPOBackbone=Qwen3-8B2026.05 | — | 86.8 | 73.5 | |
| Gemini-2.5-flashProtocol=Prompting2026.05 | — | 48.7 | 41.2 | |
| Gemini-3-FlashProtocol=Prompting2026.05 | — | 57.3 | 53 | |
| GRPOBackbone=Qwen3-4B2026.05 | — | 81.8 | 63.5 | |
| GRPOBackbone=Qwen3-8B2026.05 | — | 83.7 | 71.5 | |
| GSPOBackbone=Qwen3-4B2026.05 | — | 82.1 | 65 | |
| GSPOBackbone=Qwen3-8B2026.05 | — | 83.4 | 71 | |
| ICRLBackbone=Qwen3-4B2026.05 | — | 88.9 | 74.5 | |
| ICRLBackbone=Qwen3-8B2026.05 | — | 88.3 | 76 | |
| InternVL2.5-8BModel=InternVL2.5-8B, Strategy=Tree Search2026.05 | — | — | 3 | |
| LLaVA-NeXT-8BModel=LLaVA-NeXT-8B, Strategy=Zero-Shot2026.05 | — | — | 0 | |
| MATPOBackbone=Qwen3-4B2026.05 | — | 82.4 | 65 | |
| MATPOBackbone=Qwen3-8B2026.05 | — | 81.6 | 68 | |
| Qwen2.5-VL-7BModel=Qwen2.5-VL-7B, Strategy=Tree Search2026.05 | — | — | 4.1 | |
| Qwen3-30B-A3BProtocol=Prompting2026.05 | — | 25.6 | 3 | |
| Qwen3-4BProtocol=Prompting2026.05 | — | 9.6 | 1.5 | |
| Qwen3-8BProtocol=Prompting2026.05 | — | 10.3 | 2 | |
| ScalingInterBackbone=Qwen3-4B2026.05 | — | 84.7 | 66.5 | |
| ScalingInterBackbone=Qwen3-8B2026.05 | — | 86.2 | 73 |