Long-horizon procedural planning on EgoPlan-Bench In-Domain
62.46Success RatePlanAgent + Mem.
Evaluation Results
| Method | Links | |
|---|---|---|
| PlanAgent + Mem.Base Model=Qwen3-VL-8B, Instruction Tuning=true, Preference Alignment=true2026.03 | 62.46 | |
| PlanAgent + Mem.Base Model=Qwen3-VL-8B, Instruction Tuning=true, Preference Alignment=false2026.03 | 56.49 | |
| GPT-5.1Base Model=-, Instruction Tuning=false, Preference Alignment=false2026.03 | 55.08 | |
| Video-LLaMABase Model=LLaMA2-Chat-7B, Instruction Tuning=true, Preference Alignment=true2026.03 | 54.65 | |
| Video-LLaMABase Model=LLaMA2-Chat-7B, Instruction Tuning=true, Preference Alignment=false2026.03 | 52.14 | |
| PlanAgent + Mem.Base Model=Qwen3-VL-8B, Instruction Tuning=false, Preference Alignment=false2026.03 | 44.68 | |
| GPT-4VBase Model=-, Instruction Tuning=false, Preference Alignment=false2026.03 | 38.4 | |
| PlanAgent (Ours)Base Model=Qwen3-VL-8B, Instruction Tuning=false, Preference Alignment=false2026.03 | 36.6 | |
| Video-LLaMABase Model=LLaMA2-Chat-7B, Instruction Tuning=false, Preference Alignment=false2026.03 | 27.88 |