Long Video Understanding on LongVideoBench (test)
66.7AccuracyGPT-4o
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-4oSize=—2026.01 | 66.7 | |
| LLaVA-Video-7B + ST-GridPoolInput frames=64, Pooling Strategy=ST-GridPool2026.05 | 60.1 | |
| LLaVA-Video-7BBackbone=LLaVA-Video-7B, Retention Ratio=100%, Input Frames=642026.03 | 58.9 | |
| LinMU-NVSize=8B2026.01 | 58.8 | |
| NVILASize=8B2026.01 | 58.7 | |
| Apollo-7BBackbone=7B2026.05 | 58.5 | |
| LLaVA-Video-7BInput frames=642026.05 | 58.2 | |
| NVILA-8BBackbone=8B2026.05 | 57.7 | |
| V-CASTBackbone=LLaVA-Video-7B, Retention Ratio=25%, Input Frames=642026.03 | 57.3 | |
| HoliTomBackbone=LLaVA-Video-7B, Retention Ratio=25%, Input Frames=642026.03 | 57.1 | |
| VidCom2Backbone=LLaVA-Video-7B, Retention Ratio=25%, Input Frames=642026.03 | 57.1 | |
| Qwen2-VLSize=8B2026.01 | 56.8 | |
| LLaVA-OneVision-7B + ST-GridPoolInput frames=32, Pooling Strategy=ST-GridPool2026.05 | 56.7 | |
| LLaVA-OneVision-7BInput frames=322026.05 | 56.5 | |
| VisionZipBackbone=LLaVA-Video-7B, Retention Ratio=25%, Input Frames=642026.03 | 56.3 | |
| Oryx-1.5-7BBackbone=1.5B2026.05 | 56.3 | |
| FastVBackbone=LLaVA-Video-7B, Retention Ratio=25%, Input Frames=642026.03 | 54.8 | |
| SparseVLMBackbone=LLaVA-Video-7B, Retention Ratio=25%, Input Frames=642026.03 | 54.2 | |
| mPLUG-Owl3-8BBackbone=8B2026.05 | 52.1 |