Long-horizon procedural planning on EgoPlan-Bench Out-of-Domain
54.37Success RateGPT-5.1
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5.1Base Model=-, Instruction Tuning=false, Preference Alignment=false2026.03 | 54.37 | |
| PlanAgent + Mem.Base Model=Qwen3-VL-8B, Instruction Tuning=true, Preference Alignment=true2026.03 | 54.3 | |
| PlanAgent + Mem.Base Model=Qwen3-VL-8B, Instruction Tuning=true, Preference Alignment=false2026.03 | 50.17 | |
| Video-LLaMABase Model=LLaMA2-Chat-7B, Instruction Tuning=true, Preference Alignment=true2026.03 | 44.42 | |
| PlanAgent + Mem.Base Model=Qwen3-VL-8B, Instruction Tuning=false, Preference Alignment=false2026.03 | 43.31 | |
| Video-LLaMABase Model=LLaMA2-Chat-7B, Instruction Tuning=true, Preference Alignment=false2026.03 | 40.52 | |
| GPT-4VBase Model=-, Instruction Tuning=false, Preference Alignment=false2026.03 | 36.9 | |
| PlanAgent (Ours)Base Model=Qwen3-VL-8B, Instruction Tuning=false, Preference Alignment=false2026.03 | 35.72 | |
| Video-LLaMABase Model=LLaMA2-Chat-7B, Instruction Tuning=false, Preference Alignment=false2026.03 | 30.44 |