Question Answering on EgoSchema
58.4AccuracyQwen2.5-Omni
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-Omnievaluation_protocol=spoken questions2026.01 | 58.4 | |
| LLaVA-VideoBase Model=LLaVA-Video-7B, Input Frames=64, Mode=Upper Bound (Full Performance)2026.05 | 57.3 | |
| ST-SimDiffBase Model=LLaVA-Video-7B, Input Frames=64, Token Retain Ratio (r)=50%2026.05 | 57.3 | |
| FasterVLMBase Model=LLaVA-Video-7B, Input Frames=64, Token Retain Ratio (r)=50%2026.05 | 56.2 | |
| ST-SimDiffBase Model=LLaVA-Video-7B, Input Frames=64, Token Retain Ratio (r)=30%2026.05 | 56 | |
| FrameFusionBase Model=LLaVA-Video-7B, Input Frames=64, Token Retain Ratio (r)=50%2026.05 | 55.8 | |
| ROMAevaluation_protocol=spoken questions2026.01 | 55.4 | |
| MiniCPM-oevaluation_protocol=spoken questions2026.01 | 55.2 | |
| ROMAevaluation_protocol=spoken questions, wpos=42026.01 | 54.8 | |
| FastVBase Model=LLaVA-Video-7B, Input Frames=64, Token Retain Ratio (r)=50%2026.05 | 54.7 | |
| PruMergeBase Model=LLaVA-Video-7B, Input Frames=64, Token Retain Ratio (r)=50%2026.05 | 54.6 | |
| VisionZipBase Model=LLaVA-Video-7B, Input Frames=64, Token Retain Ratio (r)=50%2026.05 | 54.2 | |
| ROMAevaluation_protocol=spoken questions, K=12026.01 | 54 | |
| VisionZipBase Model=LLaVA-Video-7B, Input Frames=64, Token Retain Ratio (r)=30%2026.05 | 53 | |
| FrameFusionBase Model=LLaVA-Video-7B, Input Frames=64, Token Retain Ratio (r)=30%2026.05 | 53 | |
| ROMAevaluation_protocol=spoken questions, wpos=22026.01 | 52.6 | |
| FasterVLMBase Model=LLaVA-Video-7B, Input Frames=64, Token Retain Ratio (r)=30%2026.05 | 52.6 | |
| FastVBase Model=LLaVA-Video-7B, Input Frames=64, Token Retain Ratio (r)=30%2026.05 | 51.3 | |
| PruMergeBase Model=LLaVA-Video-7B, Input Frames=64, Token Retain Ratio (r)=30%2026.05 | 50.9 | |
| ROMAevaluation_protocol=spoken questions, ablation=Mixed Training2026.01 | 50.2 | |
| VITA-1.5evaluation_protocol=spoken questions2026.01 | 45.4 | |
| ROMAevaluation_protocol=spoken questions, ablation=without speak head2026.01 | 12.8 |