General Video Understanding on NexT-QA (test)
83.8AccuracyLLaVA-Video-7B + ST-GridPool
Evaluation Results
| Method | Links | |
|---|---|---|
| LLaVA-Video-7B + ST-GridPoolInput frames=64, Pooling Strategy=ST-GridPool2026.05 | 83.8 | |
| LLaVA-Video-7BInput frames=642026.05 | 83.2 | |
| NVILA-8BBackbone=8B2026.05 | 82.2 | |
| Oryx-1.5-7BBackbone=1.5B2026.05 | 81.8 | |
| LLaVA-OneVision-7B + ST-GridPoolInput frames=32, Pooling Strategy=ST-GridPool2026.05 | 79.6 | |
| LLaVA-OneVision-7BInput frames=322026.05 | 79.4 | |
| mPLUG-Owl3-8BBackbone=8B2026.05 | 78.6 | |
| LongVA-7BBackbone=7B2026.05 | 68.3 |