Embodied Task Planning on VirtualHome (Seen)
31.2Success RateGRPO w/ EVU
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| GRPO w/ EVUMethod Category=Training, Backbone Model=Qwen2.5-3B-Instruct, Base Method=GRPO, EVU Mechanism=Included2026.04 | 31.2 | — | — | — | — | — | |
| PPO w/ EVUMethod Category=Training, Backbone Model=Qwen2.5-3B-Instruct, Base Method=PPO, EVU Mechanism=Included2026.04 | 28.8 | — | — | — | — | — | |
| SFT w/ EVUMethod Category=Training, Backbone Model=Qwen2.5-3B-Instruct, Base Method=SFT, EVU Mechanism=Included2026.04 | 27.2 | — | — | — | — | — | |
| GRPOMethod Category=Training, Backbone Model=Qwen2.5-3B-Instruct, Base Method=GRPO, EVU Mechanism=None2026.04 | 25.6 | — | — | — | — | — | |
| PPOMethod Category=Training, Backbone Model=Qwen2.5-3B-Instruct, Base Method=PPO, EVU Mechanism=None2026.04 | 24 | — | — | — | — | — | |
| GRPO w/ EVUMethod Category=Training, Backbone Model=Qwen3-1.7B-Instruct, Base Method=GRPO, EVU Mechanism=Included2026.04 | 20 | — | — | — | — | — | |
| SFTMethod Category=Training, Backbone Model=Qwen2.5-3B-Instruct, Base Method=SFT, EVU Mechanism=None2026.04 | 20 | — | — | — | — | — | |
| PPO w/ EVUMethod Category=Training, Backbone Model=Qwen3-1.7B-Instruct, Base Method=PPO, EVU Mechanism=Included2026.04 | 17.6 | — | — | — | — | — | |
| ReAct w/ EVUMethod Category=Prompting, Backbone Model=DeepSeek V3.2, Base Method=ReAct, EVU Mechanism=Included2026.04 | 16 | — | — | — | — | — | |
| SFT w/ EVUMethod Category=Training, Backbone Model=Qwen3-1.7B-Instruct, Base Method=SFT, EVU Mechanism=Included2026.04 | 16 | — | — | — | — | — | |
| GRPOMethod Category=Training, Backbone Model=Qwen3-1.7B-Instruct, Base Method=GRPO, EVU Mechanism=None2026.04 | 15.7 | — | — | — | — | — | |
| Plan-and-Act w/ EVUMethod Category=Prompting, Backbone Model=DeepSeek V3.2, Base Method=Plan-and-Act, EVU Mechanism=Included2026.04 | 15.2 | — | — | — | — | — | |
| ReActMethod Category=Prompting, Backbone Model=DeepSeek V3.2, Base Method=ReAct, EVU Mechanism=None2026.04 | 13.6 | — | — | — | — | — | |
| NoThinking w/ EVUMethod Category=Prompting, Backbone Model=DeepSeek V3.2, Base Method=NoThinking, EVU Mechanism=Included2026.04 | 12.8 | — | — | — | — | — | |
| Plan-and-ActMethod Category=Prompting, Backbone Model=DeepSeek V3.2, Base Method=Plan-and-Act, EVU Mechanism=None2026.04 | 12.8 | — | — | — | — | — | |
| PPOMethod Category=Training, Backbone Model=Qwen3-1.7B-Instruct, Base Method=PPO, EVU Mechanism=None2026.04 | 10.4 | — | — | — | — | — | |
| NoThinkingMethod Category=Prompting, Backbone Model=DeepSeek V3.2, Base Method=NoThinking, EVU Mechanism=None2026.04 | 8 | — | — | — | — | — | |
| SFTMethod Category=Training, Backbone Model=Qwen3-1.7B-Instruct, Base Method=SFT, EVU Mechanism=None2026.04 | 7.2 | — | — | — | — | — | |
| finetuned GPT2 policybackbone=GPT-2, training_trajectories=10000, mode=fine-tuned2023.05 | — | 8,130 | 5,900 | 4,120 | 3,090 | 230 | |
| FLAREBackbone=Llama-3.2-3B2026.01 | — | 54.69 | 19.93 | — | — | — | |
| GPT3.5 Policybackbone=GPT-3.5, mode=few-shot, prompt_candidates=2002023.05 | — | 8,340 | 4,700 | 7,430 | 4,820 | 540 | |
| GPT3.5-MCTSbackbone=GPT-3.5, search=MCTS, mode=few-shot, prompt_candidates=2002023.05 | — | 9,140 | 7,120 | 8,810 | 7,260 | 3,360 | |
| LLM-PlannerBackbone=Llama-3.2-3B, Adaptation=Few-shot2026.01 | — | 50.98 | 17.85 | — | — | — | |
| LLM+FTBackbone=Llama-3.2-1B, Training=Fine-tuned2026.01 | — | 61.37 | 17.98 | — | — | — | |
| SayCanPaySay Backbone=Llama-3.2-3B, Pay Backbone=Llama-3.2-1B2026.01 | — | 64.14 | 15.7 | — | — | — | |
| TMoWBackbone=Llama-3.2-1B2026.01 | — | 83.61 | 11.07 | — | — | — | |
| UCTsearch=UCT, commonsense knowledge=false2023.05 | — | 0 | 0 | 0 | 0 | 0 | |
| ZSPBackbone=Llama-3.2-3B, Adaptation=Zero-shot2026.01 | — | 10.78 | 27.81 | — | — | — |