Embodied Agent on ALFWorld
100Success RateKintsugi
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Kintsugi2026.05 | 100 | — | |
| AutoAgentBackbone=gemini-3-pro-thinking2026.03 | 99.3 | 95 | |
| AutoAgentBackbone=gemini-3-pro2026.03 | 99.3 | 95 | |
| DeepAgentBackbone=gemini-3-pro2026.03 | 97.8 | 95.8 | |
| LLMKnowledge Base=tools2026.05 | 97.8 | — | |
| TTExploreBackbone=Qwen2.5, Know%=0%, Actor=trained, Thinker=trained2026.05 | 97.76 | — | |
| DeepAgentBackbone=gemini-3-pro-thinking2026.03 | 96.3 | 95.3 | |
| Dual Memory2026.05 | 94.78 | — | |
| TTExploreBackbone=Qwen2.5, Know%=0%, Actor=prompt, Thinker=trained2026.05 | 90.29 | — | |
| AWM2026.05 | 88.81 | — | |
| DeepAgentBackbone=gpt-4o2026.03 | 85.1 | 90.5 | |
| AutoAgentBackbone=gpt-4o2026.03 | 85.1 | 82.8 | |
| ExpeL2026.05 | 85.07 | — | |
| KnowselfKnow%=15%, Actor=trained, Thinker=–2026.05 | 84.33 | — | |
| WALL-E 2.02026.05 | 82.84 | — | |
| Reflexion2026.05 | 82.66 | — | |
| WKMKnow%=100%, Actor=trained, Thinker=–2026.05 | 77.61 | — | |
| TTExploreBackbone=LLaMA3, Know%=0%, Actor=trained, Thinker=trained2026.05 | 77.61 | — | |
| ReAct2026.05 | 76.87 | — | |
| KnowAgentKnow%=100%, Actor=trained, Thinker=–2026.05 | 75.37 | — | |
| ADaPT2026.05 | 72.39 | — | |
| StateAct2026.05 | 63.43 | — | |
| TTExploreBackbone=Qwen2.5, Know%=0%, Actor=prompt, Thinker=prompt2026.05 | 57.46 | — | |
| ReActBackbone=gemini-3-pro-thinking2026.03 | 53 | 64.9 | |
| LLMKnowledge Base=prompt2026.05 | 50 | — | |
| ReActBackbone=gemini-3-pro2026.03 | 46.3 | 63.9 | |
| ExpelKnow%=100%, Actor=prompt, Thinker=–2026.05 | 41.04 | — | |
| LLMKnowledge Base=None2026.05 | 39.5 | — | |
| TTExploreBackbone=LLaMA3, Know%=0%, Actor=prompt, Thinker=trained2026.05 | 32.08 | — | |
| ReActBackbone=gpt-4o2026.03 | 24.6 | 23.6 | |
| TTExploreBackbone=LLaMA3, Know%=0%, Actor=prompt, Thinker=prompt2026.05 | 11.94 | — |