Agentic Reasoning on ALFWorld (test)
97.7Success Rateteacher top-K local support matching
Evaluation Results
| Method | Links | |
|---|---|---|
| teacher top-K local support matchingmasking=true2026.03 | 97.7 | |
| RETROAGENTEvaluation Protocol=RL Training with Extrinsic and Dual Intrinsic Feedback, Reflection Strategy=RL-Trained Reflection, Base Algorithm=GRPO2026.03 | 95.6 | |
| GiGPO-Qwen2.5-7B-It-Alfworld2026.03 | 95.3 | |
| teacher top-K local support matchingmasking=false2026.03 | 95.3 | |
| Sampled-token OPDmasking=true2026.03 | 93.8 | |
| RETROAGENTEvaluation Protocol=RL Training with Extrinsic and Dual Intrinsic Feedback, Reflection Strategy=In-Context Reflection, Base Algorithm=GRPO2026.03 | 91.7 | |
| GiGPOEvaluation Protocol=Fine-tuning with RL, Base Algorithm=GRPO2026.03 | 90.8 | |
| Sampled-token OPDmasking=false2026.03 | 90.6 | |
| SKILLRLEvaluation Protocol=Fine-tuning with RL-based Frameworks, Teacher Model usage=true, Base Algorithm=GRPO2026.03 | 89.9 | |
| LAMEREvaluation Protocol=Fine-tuning with Meta-RL Frameworks, Base Algorithm=GRPO2026.03 | 82.3 | |
| GRPO w/ EMPGEvaluation Protocol=Fine-tuning with RL-based Frameworks2026.03 | 78.5 | |
| GRPOEvaluation Protocol=Fine-tuning with RL2026.03 | 77.3 | |
| RLOOEvaluation Protocol=Fine-tuning with RL2026.03 | 75.5 | |
| SimpleMemEvaluation Protocol=Fine-tuning with RL-based Frameworks, Base Algorithm=GRPO2026.03 | 62.5 | |
| Mem0Evaluation Protocol=Fine-tuning with RL-based Frameworks, Base Algorithm=GRPO2026.03 | 54.7 | |
| EvolveREvaluation Protocol=Fine-tuning with RL-based Frameworks, Base Algorithm=GRPO2026.03 | 43.8 | |
| ReflexionEvaluation Protocol=Prompting-based2026.03 | 42.7 | |
| ReActEvaluation Protocol=Prompting-based2026.03 | 31.2 | |
| Qwen2.5-7B-It2026.03 | 21.9 | |
| MemRLEvaluation Protocol=Fine-tuning with RL-based Frameworks, Base Algorithm=GRPO2026.03 | 21.4 | |
| Qwen-2.5-7B-InstructEvaluation Protocol=Zero-Shot2026.03 | 16.9 |