Robotic Task Planning in Dynamic Environments on VirtualHome
92Success RateLookPlanGraph
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LookPlanGraphBackbone=Llama3.2, Ground-truth Augmentation=true2025.12 | 92 | 96 | |
| LookPlanGraphBackbone=GPT-4o, Ground-truth Augmentation=true2025.12 | 86 | 92 | |
| LookPlanGraphBackbone=Llama3.2, Ground-truth Augmentation=false2025.12 | 60 | 68 | |
| LookPlanGraphBackbone=GPT-4o, Ground-truth Augmentation=false2025.12 | 52 | 63 | |
| ReActBackbone=Llama3.2, Ground-truth Augmentation=true2025.12 | 50 | 74 | |
| SayPlan LiteBackbone=GPT-4o, Ground-truth Augmentation=false2025.12 | 48 | 67 | |
| LLM-as-PBackbone=GPT-4o, Ground-truth Augmentation=false2025.12 | 44 | 65 | |
| SayPlan LiteBackbone=Llama3.2, Ground-truth Augmentation=false2025.12 | 39 | 65 | |
| SayPlanBackbone=GPT-4o, Ground-truth Augmentation=false2025.12 | 38 | 59 | |
| ReActBackbone=GPT-4o, Ground-truth Augmentation=true2025.12 | 34 | 58 | |
| LLM+PBackbone=GPT-4o, Ground-truth Augmentation=false2025.12 | 32 | 58 | |
| ReActBackbone=Llama3.2, Ground-truth Augmentation=false2025.12 | 30 | 64 | |
| ReActBackbone=GPT-4o, Ground-truth Augmentation=false2025.12 | 22 | 50 | |
| SayPlanBackbone=Llama3.2, Ground-truth Augmentation=false2025.12 | 21 | 53 | |
| LLM-as-PBackbone=Llama3.2, Ground-truth Augmentation=false2025.12 | 16 | 53 | |
| LLM+PBackbone=Llama3.2, Ground-truth Augmentation=false2025.12 | 0 | 38 |