Agentic Reasoning on WebShop (test)
82.3Success RateRETROAGENT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| RETROAGENTEvaluation Protocol=RL Training with Extrinsic and Dual Intrinsic Feedback, Reflection Strategy=RL-Trained Reflection, Base Algorithm=GRPO2026.03 | 82.3 | 88.9 | |
| RETROAGENTEvaluation Protocol=RL Training with Extrinsic and Dual Intrinsic Feedback, Reflection Strategy=In-Context Reflection, Base Algorithm=GRPO2026.03 | 78.9 | 87.6 | |
| GiGPOEvaluation Protocol=Fine-tuning with RL, Base Algorithm=GRPO2026.03 | 72.8 | 84.4 | |
| SKILLRLEvaluation Protocol=Fine-tuning with RL-based Frameworks, Teacher Model usage=true, Base Algorithm=GRPO2026.03 | 72.7 | 85.2 | |
| GRPO w/ EMPGEvaluation Protocol=Fine-tuning with RL-based Frameworks2026.03 | 69.3 | 81 | |
| GRPOEvaluation Protocol=Fine-tuning with RL2026.03 | 66.9 | 75.5 | |
| RLOOEvaluation Protocol=Fine-tuning with RL2026.03 | 65.7 | 80.3 | |
| LAMEREvaluation Protocol=Fine-tuning with Meta-RL Frameworks, Base Algorithm=GRPO2026.03 | 61.7 | — | |
| SimpleMemEvaluation Protocol=Fine-tuning with RL-based Frameworks, Base Algorithm=GRPO2026.03 | 46.9 | 67.8 | |
| Mem0Evaluation Protocol=Fine-tuning with RL-based Frameworks, Base Algorithm=GRPO2026.03 | 37.5 | 58.1 | |
| ReflexionEvaluation Protocol=Prompting-based2026.03 | 28.8 | 58.1 | |
| ReActEvaluation Protocol=Prompting-based2026.03 | 19.5 | 46.2 | |
| EvolveREvaluation Protocol=Fine-tuning with RL-based Frameworks, Base Algorithm=GRPO2026.03 | 17.6 | 42.5 | |
| MemRLEvaluation Protocol=Fine-tuning with RL-based Frameworks, Base Algorithm=GRPO2026.03 | 9.2 | 29.5 | |
| Qwen-2.5-7B-InstructEvaluation Protocol=Zero-Shot2026.03 | 0.8 | 4.5 |