Task Planning on EB-ALFRED (Long)
74Success Rate (SR)BrainMem
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| BrainMemEvaluation Model=Qwen2.5-VL-72B-Ins2026.03 | 74 | 82.6 | |
| Claude-3.7-SonnetType=Closed-source2025.08 | 70 | — | |
| RoboMemoryType=Ours, Backbone=Qwen2.5-VL-72B-Ins2025.08 | 66 | 81.3 | |
| RoboMemoryEvaluation Model=Qwen2.5-VL-72B-Ins2026.03 | 66 | 81.3 | |
| Gemini-1.5-ProType=Closed-source2025.08 | 58 | 65 | |
| Gemini-2.0-flashType=Closed-source2025.08 | 58 | 62 | |
| GPT-4oType=Closed-source2025.08 | 54 | 62.5 | |
| Claude-3.5-SonnetType=Closed-source2025.08 | 52 | 54.5 | |
| InternVL2.5-78BType=Open-source2025.08 | 42 | 49 | |
| InternVL3-78BType=Open-source2025.08 | 36 | — | |
| Qwen2.5-VL-72B-InsType=Open-source2025.08 | 34 | — | |
| VoyagerType=Baselines, Backbone=Qwen2.5-VL-72B-Ins2025.08 | 32 | 54.2 | |
| CradleType=Baselines, Backbone=Qwen2.5-VL-72B-Ins2025.08 | 32 | 41 | |
| VoyagerEvaluation Model=Qwen2.5-VL-72B-Ins2026.03 | 32 | 54.2 | |
| CradleEvaluation Model=Qwen2.5-VL-72B-Ins2026.03 | 32 | 41 | |
| InternVL2.5-38BType=Open-source2025.08 | 26 | 36.5 | |
| Llama-3.2-90B-Vision-InsType=Open-source2025.08 | 16 | 24 | |
| RoboOSType=Baselines, Backbone=Qwen2.5-VL-72B-Ins2025.08 | 12 | 17.6 | |
| RoboOSEvaluation Model=Qwen2.5-VL-72B-Ins2026.03 | 12 | 17.6 | |
| ReflexionType=Baselines, Backbone=Qwen2.5-VL-72B-Ins2025.08 | 10 | 33 | |
| ReflexionEvaluation Model=Qwen2.5-VL-72B-Ins2026.03 | 10 | 33 | |
| RoboOSType=Baselines, Backbone=RoboBrain2-32B2025.08 | 8 | 13.2 | |
| GPT-4o-miniType=Closed-source2025.08 | 0 | 17 |