Long Video Understanding on Video MME w/o sub (long)
71.4AccuracyLensWalk
Evaluation Results
| Method | Links | |
|---|---|---|
| LensWalkReasoner=o3, Observer=o32026.03 | 71.4 | |
| LensWalkReasoner=o3, Observer=GPT-4.12026.03 | 70 | |
| LensWalkReasoner=GPT-5, Observer=GPT-52026.03 | 69.2 | |
| GPT-52026.03 | 68.4 | |
| Gemini-1.5-Pro2026.03 | 67.4 | |
| Gemini 1.5 ProSize=-, Tokens per frame=-2026.04 | 67.4 | |
| Deep Video Discovery2026.03 | 67.3 | |
| LensWalkReasoner=o3, Observer=Qwen2.5-VL-72B2026.03 | 66.7 | |
| GPT-4o2026.03 | 65.3 | |
| GPT-4oSize=-, Tokens per frame=-2026.04 | 65.3 | |
| AdaReTaKe2026.03 | 65 | |
| Ego-R12026.03 | 64.9 | |
| o32026.03 | 64.7 | |
| GPT-4.12026.03 | 63.1 | |
| Qwen2.5-VL-72B2026.03 | 63.1 | |
| Gemini-2.0-Flash2026.03 | 63 | |
| InternVL2.5-78B2026.03 | 62.6 | |
| MR. Video2026.03 | 61.8 | |
| LLaVA-Video-72B2026.03 | 59.6 | |
| LongVUSize=7B, Tokens per frame=642026.04 | 59.5 | |
| Tempo*Size=6B, Tokens per frame=0.5–16, Visual Budget=4K, Actual tokens/frame=3.42026.04 | 57.8 | |
| Tempo*Size=6B, Tokens per frame=0.5–16, Visual Budget=8K, Actual tokens/frame=4.12026.04 | 57 | |
| VideoChat-FlashSize=7B, Tokens per frame=162026.04 | 55.4 | |
| VideoLLaMA3*Size=7B, Tokens per frame=≤ 912026.04 | 54.9 | |
| StormSize=7B, Tokens per frame=642026.04 | 53.4 | |
| LongVILASize=7B, Tokens per frame=1962026.04 | 47 | |
| KangarooSize=8B, Tokens per frame=2562026.04 | 46.7 | |
| LongLLaVASize=A13B, Tokens per frame=1442026.04 | 46.4 | |
| LongVASize=7B, Tokens per frame=1442026.04 | 46.2 | |
| VideoChat2-HDSize=7B, Tokens per frame=722026.04 | 39.8 |