Long-horizon procedural planning on EgoPlan-Bench All
58.72Success RatePlanAgent + Mem.
Evaluation Results
| Method | Links | |
|---|---|---|
| PlanAgent + Mem.Base Model=Qwen3-VL-8B, Instruction Tuning=true, Preference Alignment=true2026.03 | 58.72 | |
| GPT-5.1Base Model=-, Instruction Tuning=false, Preference Alignment=false2026.03 | 54.78 | |
| PlanAgent + Mem.Base Model=Qwen3-VL-8B, Instruction Tuning=true, Preference Alignment=false2026.03 | 53.29 | |
| Video-LLaMABase Model=LLaMA2-Chat-7B, Instruction Tuning=true, Preference Alignment=true2026.03 | 51.83 | |
| Video-LLaMABase Model=LLaMA2-Chat-7B, Instruction Tuning=true, Preference Alignment=false2026.03 | 48.94 | |
| PlanAgent + Mem.Base Model=Qwen3-VL-8B, Instruction Tuning=false, Preference Alignment=false2026.03 | 43.63 | |
| GPT-4VBase Model=-, Instruction Tuning=false, Preference Alignment=false2026.03 | 37.98 | |
| PlanAgent (Ours)Base Model=Qwen3-VL-8B, Instruction Tuning=false, Preference Alignment=false2026.03 | 35.81 | |
| Gemini-Pro-VisionBase Model=-, Instruction Tuning=false, Preference Alignment=false2026.03 | 30.46 | |
| SEED-LLaMABase Model=LLaMA2-Chat-13B, Instruction Tuning=false, Preference Alignment=false2026.03 | 29.93 | |
| Video-LLaMABase Model=LLaMA2-Chat-7B, Instruction Tuning=false, Preference Alignment=false2026.03 | 28.58 | |
| Qwen-VL-ChatBase Model=Qwen-7B, Instruction Tuning=false, Preference Alignment=false2026.03 | 27.69 | |
| DeepSeek-VL-ChatBase Model=DeepSeek-LLM-7B, Instruction Tuning=false, Preference Alignment=false2026.03 | 27.57 |