Multi-agent planning on MAP-THOR averaged across all tasks 2-agent
66Success RateLLaMAR
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| LLaMARLM=GPT-4V, exploration=true2024.07 | 66 | 91 | 97 | 0.82 | 21.87 | |
| LLaMARLM=GPT-4V, exploration=false2024.07 | 62 | 87 | 95 | 0.82 | 23.44 | |
| LLaMARLM=CogVLM2024.07 | 61 | 89 | 95 | 0.8 | 23.21 | |
| LLaMARLM=IDEFICS-22024.07 | 57 | 86 | 94 | 0.78 | 25.27 | |
| LLaMARLM=LLaVA2024.07 | 54 | 84 | 91 | 0.75 | 26.21 | |
| LLaMARLM=GPT-42024.07 | 51 | 85 | 95 | 0.83 | 25.8 | |
| ReActLM=GPT-4V2024.07 | 34 | 72 | 92 | 0.67 | 24.08 | |
| ActLM=GPT-4V2024.07 | 33 | 67 | 91 | 0.59 | 24.92 | |
| CoELALM=GPT-4V2024.07 | 25 | 46 | 76 | 0.73 | 28.93 | |
| CoTLM=GPT-4V2024.07 | 14 | 59 | 87 | 0.62 | 28.4 | |
| SmartLLMLM=GPT-4V2024.07 | 11 | 23 | 91 | 0.45 | 29.87 |