Egocentric Action Planning on EgoPlan
38AccuracyGPT-4v
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-4v#Frames=322025.03 | 38 | |
| TGPOBackbone=Qwen2.5-VL 3B2026.03 | 36.8 | |
| EgoVLM GRPO (3B)*Model Scale=3B, Training Protocol=GRPO, Model Category=Open-source MLLMs2026.03 | 36.5 | |
| Qwen2-VL#Param=7B, #Frames=322025.03 | 34.3 | |
| LLaVA-Videos#Param=7B, #Frames=322025.03 | 33.6 | |
| LLaVA-VideoModel Scale=7B, Model Category=Open-source MLLMs2026.03 | 33.6 | |
| EgoGPT#Param=7B, #Frames=32, training_dataset=EgoIT+EgoLifeD12025.03 | 33.4 | |
| Oryx#Param=7B, #Frames=322025.03 | 33.2 | |
| Qwen2.5-VLModel Scale=7B, Model Category=Open-source MLLMs2026.03 | 33 | |
| EgoVLM Dr. GRPOModel Scale=3B, Training Protocol=Dr. GRPO, Model Category=Open-source MLLMs2026.03 | 33 | |
| Qwen2.5-VLModel Scale=3B, Model Category=Open-source MLLMs2026.03 | 32.9 | |
| GPT-4o#Frames=322025.03 | 32.8 | |
| Gemini 1.5-ProModel Category=Proprietary2026.03 | 32.8 | |
| GPT-4oModel Category=Proprietary2026.03 | 32.8 | |
| EgoGPT#Param=7B, #Frames=32, training_dataset=EgoIT2025.03 | 32.4 | |
| EgoVLM SFT (3B)*Model Scale=3B, Training Protocol=SFT, Model Category=Open-source MLLMs2026.03 | 32.1 | |
| Gemini-1.5-Pro#Frames=322025.03 | 31.3 | |
| LLaVA-OV#Param=7B, #Frames=322025.03 | 30.7 | |
| LongVA#Param=7B, #Frames=322025.03 | 29.9 | |
| IXC-2.5#Param=7B, #Frames=322025.03 | 29.4 | |
| LLaVA-Next-Video#Param=7B, #Frames=322025.03 | 29 | |
| InternVideo2#Param=8B, #Frames=322025.03 | 27.5 |