Embodied AI Navigation and Manipulation on ALFWorld (test)
94.2Pick & Place (P&P)Mem-π
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Mem-πBase Agent=gpt-5.4-mini, Memory Backbone=Qwen-2.5-7B-Instruct2026.05 | 94.2 | — | 90.2 | 92.3 | 91.5 | 91.6 | 86.7 | 91.6 | |
| Mem-π (Stage 1)Base Agent=gpt-5.4-mini, Memory Backbone=Qwen-2.5-7B-Instruct, Training Stage=Stage 12026.05 | 92.1 | — | 88 | 91.2 | 89 | 90 | 84.1 | 90 | |
| MemRLBase Agent=gpt-5.4-mini2026.05 | 90.8 | — | 85.5 | 88.7 | 86.5 | 87.1 | 81.2 | 88 | |
| Memory-R1Base Agent=gpt-5.4-mini, Memory Backbone=Qwen-2.5-7B-Instruct2026.05 | 89.9 | — | 85 | 87.7 | 86.3 | 87.7 | 81.1 | 87.9 | |
| Mem0Base Agent=gpt-5.4-mini, k=12026.05 | 89.7 | — | 84.1 | 87.3 | 85.8 | 87.2 | 80.5 | 87.5 | |
| RAGBase Agent=gpt-5.4-mini, k=12026.05 | 89.1 | — | 83.4 | 86.8 | 85.5 | 85.8 | 79.6 | 87.1 | |
| Base AgentBase Agent=gpt-5.4-mini2026.05 | 88.3 | — | 82.7 | 86.1 | 85 | 85.5 | 78.8 | 85.3 | |
| ACEModel Size=7B, Seeds=134×5, Evaluation Setting=single-pass2026.06 | — | 66.9 | — | — | — | — | — | — | |
| AWMModel Size=7B, Seeds=134×5, Evaluation Setting=single-pass2026.06 | — | 65.4 | — | — | — | — | — | — | |
| Dynamic CheatsheetModel Size=7B, Seeds=134×5, Evaluation Setting=single-pass2026.06 | — | 70.7 | — | — | — | — | — | — | |
| GEPAModel Size=7B, Seeds=134×5, Evaluation Setting=single-pass2026.06 | — | 63.9 | — | — | — | — | — | — | |
| ReActModel Size=7B, Seeds=134×5, Evaluation Setting=single-pass2026.06 | — | 64.6 | — | — | — | — | — | — | |
| RSEAModel Size=7B, Seeds=134×5, Evaluation Setting=single-pass2026.06 | — | 69.3 | — | — | — | — | — | — |