Multi-modal Video Understanding on MVBench (test)
60.4MVBench ScoreLLaVA-Video-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| LLaVA-Video-7BBackbone=LLaVA-Video-7B, Retention Ratio=100%, Input Frames=642026.03 | 60.4 | |
| HoliTomBackbone=LLaVA-Video-7B, Retention Ratio=25%, Input Frames=642026.03 | 58.4 | |
| V-CASTBackbone=LLaVA-Video-7B, Retention Ratio=25%, Input Frames=642026.03 | 58 | |
| VisionZipBackbone=LLaVA-Video-7B, Retention Ratio=25%, Input Frames=642026.03 | 57.9 | |
| VidCom2Backbone=LLaVA-Video-7B, Retention Ratio=25%, Input Frames=642026.03 | 57 | |
| SparseVLMBackbone=LLaVA-Video-7B, Retention Ratio=25%, Input Frames=642026.03 | 55.4 | |
| FastVBackbone=LLaVA-Video-7B, Retention Ratio=25%, Input Frames=642026.03 | 52.1 |