Interactive Decision-making on ALFWorld
99.6Overall Success RateDeepseek-V4-Pro
Evaluation Results
| Method | Links | ||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Deepseek-V4-ProParadigm=MAP, Model Category=Reasoning Models2026.05 | 99.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Plan-and-Act + ECUExperience Level=peak2026.06 | 97.8 | 95.8 | 100 | 100 | 100 | 94.4 | 94.1 | — | — | — | — | — | — | — | — | — | — | — | |
| SAMPODimension=Ours, Base Model=Qwen3-8B (SFT version)2026.02 | 97.71 | — | — | — | — | — | — | 8.98 | — | — | — | — | — | — | — | — | — | — | |
| HiPERBackbone=Qwen2.5-7B-Instruct2026.02 | 97.4 | 100 | 100 | 96.3 | 100 | 84.8 | 95.5 | — | — | — | — | — | — | — | — | — | — | — | |
| Deepseek-V4-ProParadigm=CoMAP, Model Category=Reasoning Models2026.05 | 97.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| G-MemoryBackbone=QWEN3-14B, Framework=AutoGen2026.05 | 96.69 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Deepseek-V4-ProParadigm=ReAct, Model Category=Reasoning Models2026.05 | 96.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Doubao-Seed-2.0-ProParadigm=MAP, Model Category=Reasoning Models2026.05 | 96.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| UCEExperience Level=peak2026.06 | 96.3 | 100 | 96.8 | 95.7 | 100 | 100 | 82.4 | — | — | — | — | — | — | — | — | — | — | — | |
| HIPIFType=RL2026.06 | 96.1 | 100 | 96.6 | 96.3 | 96.2 | 93.2 | 95.2 | — | — | — | — | — | — | — | — | — | — | — | |
| GAGPOType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.05 | 95.6 | 97.8 | 95.8 | 97.6 | 92.1 | 97.8 | 92.6 | — | — | — | — | — | — | — | — | — | — | — | |
| HiPERBackbone=Qwen2.5-1.5B-Instruct2026.02 | 95.3 | 98.9 | 97.5 | 90.9 | 96.7 | 91.7 | 91.3 | — | — | — | — | — | — | — | — | — | — | — | |
| Kimi-K2-ThinkingParadigm=MAP, Model Category=Reasoning Models2026.05 | 95 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ChatDevBackbone=QWEN3-14B, Framework=AutoGen2026.05 | 94.89 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Doubao-Seed-2.0-ProParadigm=CoMAP, Model Category=Reasoning Models2026.05 | 94.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MAP-4BParadigm=MAP, Model Category=Fine-Tuning Models2026.05 | 94.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GIGPOType=RL2026.06 | 93.8 | 95 | 96.4 | 100 | 88.5 | 100 | 85.7 | — | — | — | — | — | — | — | — | — | — | — | |
| Role-AgentType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.06 | 93.8 | 98.3 | 98.5 | 88.9 | 90 | 93.7 | 92.8 | — | — | — | — | — | — | — | — | — | — | — | |
| GAGPOType=RL Training, Backbone=Qwen2.5-1.5B-Instruct2026.05 | 93.5 | 99.2 | 97.3 | 95.1 | 84.9 | 83.8 | 89.8 | — | — | — | — | — | — | — | — | — | — | — | |
| ReflAct + ECUExperience Level=peak2026.06 | 93.3 | 95.8 | 96.8 | 87 | 100 | 100 | 76.5 | — | — | — | — | — | — | — | — | — | — | — | |
| DECENTMEMBackbone=QWEN3-14B, Framework=AutoGen2026.05 | 93.13 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Doubao-Seed-2.0-ProParadigm=ReAct, Model Category=Reasoning Models2026.05 | 93.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| STEP-HRLType=RL2026.06 | 92.9 | 100 | 100 | 90 | 89.7 | 100 | 77.8 | — | — | — | — | — | — | — | — | — | — | — | |
| NoThinking + ECUExperience Level=peak2026.06 | 92.5 | 100 | 96.8 | 73.9 | 100 | 88.9 | 94.1 | — | — | — | — | — | — | — | — | — | — | — | |
| SAPOMethod Category=Memory-Augmented RL-based Methods2026.06 | 92.2 | 98.7 | 98.1 | 92.6 | 85 | 73.9 | 89.2 | — | — | — | — | — | — | — | — | — | — | — | |
| G-MemoryBackbone=QWEN3-8B, Framework=AutoGen2026.05 | 92.11 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GIGPO w/ PaWType=RL Training, Base Model Series=Qwen2.5-7B-Instruct2026.06 | 91.8 | 98.2 | 98.6 | 84.5 | 91.5 | 85.6 | 84.3 | — | — | — | — | — | — | — | — | — | — | — | |
| DECENTMEMBackbone=QWEN3-8B, Framework=AutoGen2026.05 | 91.54 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Kimi-K2-ThinkingParadigm=CoMAP, Model Category=Reasoning Models2026.05 | 91.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Role-AgentType=RL Training, Backbone=Qwen2.5-1.5B-Instruct2026.06 | 90.9 | 95.8 | 95 | 97 | 87.5 | 78.3 | 91.7 | — | — | — | — | — | — | — | — | — | — | — | |
| GiGPOBackbone=Qwen2.5-7B-Instruct2026.02 | 90.8 | 97.7 | 98.8 | 83.7 | 89.3 | 82.7 | 79.2 | — | — | — | — | — | — | — | — | — | — | — | |
| GIGPOType=RL Training, Base Model Series=Qwen2.5-7B-Instruct2026.06 | 90.8 | 97.7 | 98.8 | 83.7 | 89.3 | 82.7 | 79.2 | — | — | — | — | — | — | — | — | — | — | — | |
| GiGPOType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.06 | 90.8 | 97.7 | 98.8 | 83.7 | 89.3 | 82.7 | 79.2 | — | — | — | — | — | — | — | — | — | — | — | |
| MemRLSkill Synthesis Model=None2026.05 | 90.7 | — | — | — | — | — | — | — | — | — | — | 11.85 | — | — | — | — | — | — | |
| SKILLGRAPHCategory=Ours2026.05 | 90.6 | 100 | 100 | 100 | 80 | 80 | 83.3 | — | — | — | — | — | — | — | — | — | — | — | |
| GIGPO w/ PaWType=RL Training, Base Model Series=Qwen2.5-1.5B-Instruct2026.06 | 90.4 | 95.3 | 91.8 | 89.5 | 89.1 | 83.3 | 84.5 | — | — | — | — | — | — | — | — | — | — | — | |
| AAWMBackbone=Qwen2.5-7B-Instruct, Training Stage=Imitation Learning then Reinforcement Learning initialized by World Modeling2026.06 | 90.1 | 95.2 | 93.1 | 94.1 | 81.5 | 85.3 | 83.4 | — | — | — | — | — | — | — | — | — | — | — | |
| SKILLRLCategory=Memory-Augmented RL-based Methods, Backbone=Qwen2.5-7B-Instruct2026.02 | 89.9 | 97.9 | 90 | 90 | 95.5 | 71.4 | 87.5 | — | — | — | — | — | — | — | — | — | — | — | |
| SkillRLBackbone=Qwen2.5-(VL)-7B-Instruct, Skill Augmentation=true2026.04 | 89.9 | 97.9 | 90 | 90 | 95.5 | 71.4 | 87.5 | — | — | — | — | — | — | — | — | — | — | — | |
| SkillRLCategory=Memory-Augmented RL-based Methods2026.05 | 89.9 | 97.9 | 90 | 90 | 95.5 | 71.4 | 87.5 | — | — | — | — | — | — | — | — | — | — | — | |
| SKILLRLMethod Category=Memory-Augmented RL-based Methods2026.06 | 89.9 | 97.9 | 90 | 90 | 95.5 | 71.4 | 87.5 | — | — | — | — | — | — | — | — | — | — | — | |
| SKILL0Backbone=Qwen2.5-(VL)-7B-Instruct, Reduced Token Overhead=true2026.04 | 89.8 | 100 | 94.6 | 81.9 | 85.7 | 85.8 | 80.1 | — | 0.41 | — | — | — | — | — | — | — | — | — | |
| MetaGPTBackbone=QWEN3-14B, Framework=AutoGen2026.05 | 89.43 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ChatDevBackbone=QWEN3-8B, Framework=AutoGen2026.05 | 89.14 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Kimi-K2-ThinkingParadigm=ReAct, Model Category=Reasoning Models2026.05 | 89.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GiGPOType=RL Training, Backbone=Qwen2.5-7B-Instruct2026.05 | 88.8 | 96.2 | 95.5 | 80.9 | 72.1 | 90.9 | 90.4 | — | — | — | — | — | — | — | — | — | — | — | |
| ITPRBackbone=Qwen3-8B, Type=Training2026.01 | 88.57 | 97.14 | 88.88 | 93.75 | 88 | 76.92 | 79.17 | — | — | — | — | — | — | — | — | — | — | — | |
| HiperType=RL2026.06 | 88.3 | 97.4 | 96.2 | 88.9 | 83.3 | 81.8 | 70 | — | — | — | — | — | — | — | — | — | — | — | |
| Minimax-M2.7Paradigm=MAP, Model Category=Reasoning Models2026.05 | 88.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GiGPOType=RL Training, Backbone=Qwen2.5-1.5B-Instruct2026.05 | 88.1 | 98.4 | 91.1 | 96.8 | 82.6 | 72.2 | 79.7 | — | — | — | — | — | — | — | — | — | — | — | |
| SKILL0Backbone=Qwen2.5-(VL)-3B-Instruct, Reduced Token Overhead=true2026.04 | 87.9 | 95.6 | 100 | 86.7 | 78.7 | 80.4 | 75.2 | — | 0.38 | — | — | — | — | — | — | — | — | — | |
| Minimax-M2.7Paradigm=CoMAP, Model Category=Reasoning Models2026.05 | 87.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SkillTTASkill Synthesis Model=GPT-4o-mini2026.05 | 87.9 | — | — | — | — | — | — | — | — | — | — | 9.03 | — | — | — | — | — | — | |
| Skill0Method Category=Memory-Augmented RL-based Methods2026.06 | 87.9 | 95.6 | 100 | 86.7 | 78.7 | 80.4 | 75.2 | — | — | — | — | — | — | — | — | — | — | — | |
| D2SkillMethod Category=Memory-Augmented RL-based Methods2026.06 | 87.8 | 93.8 | 95.5 | 77.8 | 95 | 94.7 | 72 | — | — | — | — | — | — | — | — | — | — | — | |
| GIGPOType=RL Training, Base Model Series=Qwen2.5-1.5B-Instruct2026.06 | 87.6 | 95.3 | 87.7 | 92.6 | 79.8 | 84.3 | 82.3 | — | — | — | — | — | — | — | — | — | — | — | |
| SkillTTASkill Synthesis Model=GPT-4o2026.05 | 87.2 | — | — | — | — | — | — | — | — | — | — | 8.88 | — | — | — | — | — | — | |
| ITPRBackbone=Llama3.1-8B, Type=Training2026.01 | 87.14 | 88.57 | 92.59 | 93.75 | 92 | 46.15 | 91.67 | — | — | — | — | — | — | — | — | — | — | — | |
| MAP-4BParadigm=ReAct, Model Category=Fine-Tuning Models2026.05 | 87.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GiGPOBackbone=Qwen2.5-1.5B-Instruct2026.02 | 86.7 | 94.4 | 94.8 | 94.4 | 79.8 | 67.5 | 76.4 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-8B + SelSkill Round3Backbone=Qwen3-8B, Training Stage=SelSkill Round32026.05 | 86.7 | — | — | — | — | — | — | — | — | — | — | 16.9 | 100 | 97 | 83.2 | 0.44 | — | — | |
| GiGPOType=RL Training, Backbone=Qwen2.5-1.5B-Instruct2026.06 | 86.7 | 94.4 | 94.8 | 94.4 | 79.8 | 67.5 | 76.4 | — | — | — | — | — | — | — | — | — | — | — | |
| IWMBackbone=Qwen2.5-7B-Instruct, Training Stage=Imitation Learning then Reinforcement Learning initialized by World Modeling2026.06 | 86.7 | 92.6 | 94.4 | 93.7 | 74.1 | 76.3 | 77.4 | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4o-0806Paradigm=MAP, Model Category=Non-reasoning Models2026.05 | 86.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| UCEExperience Level=initial2026.06 | 86.6 | 100 | 100 | 91.3 | 95.2 | 50 | 64.7 | — | — | — | — | — | — | — | — | — | — | — | |
| IWMBackbone=Llama3.1-8B, Type=Training2026.01 | 85.9 | 87.5 | 88.9 | 82.4 | 94.7 | 85.9 | 84.6 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-8B + SelSkill Round2Backbone=Qwen3-8B, Training Stage=SelSkill Round22026.05 | 85.9 | — | — | — | — | — | — | — | — | — | — | 16.3 | 96.6 | 91.4 | 83.9 | 0.46 | — | — | |
| Hiagent+GRPOType=RL2026.06 | 85.9 | 85.3 | 100 | 100 | 77.3 | 77.8 | 76.9 | — | — | — | — | — | — | — | — | — | — | — | |
| MAP-4BParadigm=CoMAP, Model Category=Fine-Tuning Models2026.05 | 85.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GRPOBackbone=Qwen2.5-7B-Instruct2026.02 | 85.4 | 97.2 | 86.4 | 81.1 | 84.1 | 68.4 | 75.9 | — | — | — | — | — | — | — | — | — | — | — | |
| ExpeL2026.06 | 85.1 | 91.7 | 93.5 | 73.9 | 95.2 | 77.8 | 70.6 | — | — | — | — | — | — | — | — | — | — | — | |
| ITPRBackbone=Qwen2.5-7B, Type=Training2026.01 | 85.07 | 94.29 | 88.89 | 87.5 | 53.84 | 76 | 91.67 | — | — | — | — | — | — | — | — | — | — | — | |
| BaseBackbone=Qwen2.5-7B-Instruct, Training Stage=Imitation Learning then Reinforcement Learning initialized by World Modeling2026.06 | 84.6 | 93 | 92.8 | 91.2 | 79.3 | 66.7 | 68.6 | — | — | — | — | — | — | — | — | — | — | — | |
| Minimax-M2.7Paradigm=ReAct, Model Category=Reasoning Models2026.05 | 84.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ACT-4BParadigm=MAP, Model Category=Fine-Tuning Models2026.05 | 84.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-32B-ThinkingParadigm=MAP, Model Category=Reasoning Models2026.05 | 84 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| HISRType=RL2026.03 | 83.6 | 83.3 | 87.1 | 65.2 | 85.7 | 100 | 82.4 | — | — | — | — | — | — | — | — | — | — | — | |
| DECENTMEMBackbone=QWEN3-14B, Framework=AgentNet2026.05 | 83.59 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MetaGPTBackbone=QWEN3-8B, Framework=AutoGen2026.05 | 83.25 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AAWMBackbone=Qwen2.5-1.5B-Instruct, Training Stage=Imitation Learning then Reinforcement Learning initialized by World Modeling2026.06 | 83.1 | 93.3 | 86.9 | 88.7 | 71.7 | 72 | 71.9 | — | — | — | — | — | — | — | — | — | — | — | |
| IWMBackbone=Qwen2.5-7B, Type=Training2026.01 | 82.8 | 90.6 | 85.2 | 88.2 | 84.2 | 42.9 | 76.9 | — | — | — | — | — | — | — | — | — | — | — | |
| PPOBackbone=Qwen2.5-7B-Instruct2026.02 | 82.8 | 98 | 82.5 | 95 | 52.5 | 68.8 | 75 | — | — | — | — | — | — | — | — | — | — | — | |
| HISR -w/o AGSType=RL2026.03 | 82.8 | 91.7 | 87.1 | 78.3 | 66.7 | 100 | 64.7 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-8B + SelSkill Round1Backbone=Qwen3-8B, Training Stage=SelSkill Round12026.05 | 82.8 | — | — | — | — | — | — | — | — | — | — | 19.7 | 94.1 | 90.7 | 78.8 | 0.66 | — | — | |
| GPT-4o-1120Paradigm=MAP, Model Category=Non-reasoning Models2026.05 | 82.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SkillRLBackbone=Qwen2.5-(VL)-3B-Instruct, Skill Augmentation=true2026.04 | 82.4 | 91.9 | 82.9 | 87.4 | 78.7 | 100 | 70 | — | 2.21 | — | — | — | — | — | — | — | — | — | |
| GPT-4o-0806Paradigm=CoMAP, Model Category=Non-reasoning Models2026.05 | 82.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| IWMBackbone=Qwen3-8B, Type=Training2026.01 | 82.14 | 85.71 | 85.19 | 87.5 | 84 | 46.15 | 87.5 | — | — | — | — | — | — | — | — | — | — | — | |
| HISR -w/o SPRType=RL2026.03 | 82.1 | 87.5 | 83.9 | 78.3 | 76.2 | 94.4 | 70.6 | — | — | — | — | — | — | — | — | — | — | — | |
| ReflAct2026.06 | 82.1 | 91.7 | 93.5 | 82.6 | 90.5 | 50 | 70.6 | — | — | — | — | — | — | — | — | — | — | — | |
| GRPOBackbone=Qwen2.5-(VL)-7B-Instruct2026.04 | 81.8 | 92.6 | 85.2 | 80 | 82.7 | 93.8 | 56.5 | — | 0.95 | — | — | — | — | — | — | — | — | — | |
| No memoryBackbone=QWEN3-14B, Framework=AutoGen2026.05 | 81.51 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Plan-and-Act + ECUExperience Level=initial2026.06 | 81.3 | 95.8 | 100 | 82.6 | 81 | 33.3 | 76.5 | — | — | — | — | — | — | — | — | — | — | — | |
| AgentOCRBackbone=Qwen2.5-(VL)-7B-Instruct, Reduced Token Overhead=true2026.04 | 81.2 | 95.6 | 78.1 | 73.2 | 72.4 | 96.2 | 72 | — | 0.43 | — | — | — | — | — | — | — | — | — | |
| ReDActSmall Model=Qwen3-80B, Large Model=GPT-5.2, UQ Measure=PPL2026.04 | 80.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-32B-ThinkingParadigm=CoMAP, Model Category=Reasoning Models2026.05 | 80.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ACT-4BParadigm=CoMAP, Model Category=Fine-Tuning Models2026.05 | 80.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| HISR -w/o HIMType=RL2026.03 | 80.6 | 87.5 | 83.9 | 78.3 | 66.7 | 94.4 | 70.6 | — | — | — | — | — | — | — | — | — | — | — | |
| GRPO w/ PaWType=RL Training, Base Model Series=Qwen2.5-7B-Instruct2026.06 | 80.6 | 90.4 | 86.8 | 82.9 | 76.5 | 80.7 | 67.3 | — | — | — | — | — | — | — | — | — | — | — | |
| PPO (with critic)Type=RL Training, Backbone=Qwen2.5-7B-Instruct2026.05 | 80.4 | 92.3 | 92.5 | 89.5 | 80.3 | 64 | 68.8 | — | — | — | — | — | — | — | — | — | — | — |