Video Question Answering on Video-MME without subtitles
75.4Accuracy (Overall)GOPAgen
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| GOPAgenModel Category=Agentic video understanding framework2026.06 | 75.4 | 84.5 | 75.1 | 66.7 | |
| Gemini-1.5-ProModel Category=Closed-Source Models2026.06 | 75 | 81.7 | 74.3 | 67.4 | |
| AdaReTaKeModel Category=Open-Source Models2026.06 | 73.5 | 80.6 | 74.9 | 65 | |
| Qwen2.5-VL-72BZero-shot=true2026.01 | 73.3 | — | — | — | |
| Qwen2.5-VL-72BModel Category=Open-Source Models2026.06 | 73.3 | — | — | — | |
| LFS + Qwen3-VL-8BZero-shot=true2026.01 | 72.6 | — | — | — | |
| VideoLucyModel Category=Agentic video understanding framework2026.06 | 72.5 | 78.6 | 72.1 | 66.8 | |
| InternVL2.5-72BModel Category=Open-Source Models2026.06 | 72.1 | 82.8 | 70.9 | 62.6 | |
| GPT-4oModel Category=Closed-Source Models2026.06 | 71.9 | 80 | 70.3 | 65.3 | |
| Qwen3-VL-8BZero-shot=true2026.01 | 71.4 | — | — | — | |
| Gemini 2.5 Flash + ReFoCUSLLM Size=Closed, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 69.5 | 79.3 | 68.3 | 60.8 | |
| Qwen3-VL-4BZero-shot=true2026.01 | 69.3 | — | — | — | |
| Qwen3-VL + ReFoCUSLLM Size=8B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 68.5 | 79.6 | 67 | 58.9 | |
| InternVL3 + ReFoCUSLLM Size=8B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 67 | 75.8 | 66.8 | 58.3 | |
| InternVL3.5 + ReFoCUSLLM Size=8B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 66.7 | 76.2 | 64.9 | 58.9 | |
| Qwen3-VL + ReFoCUSLLM Size=4B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 66.4 | 76.7 | 65.7 | 57 | |
| Gemini 2.5 FlashLLM Size=Closed, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 66 | 77.6 | 63.7 | 56.8 | |
| InternVL3.5 + ReFoCUSLLM Size=4B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 65.9 | 78 | 62.3 | 57.4 | |
| Gemini2.5-Flash-LiteZero-shot=true2026.01 | 65 | — | — | — | |
| Qwen3-VLLLM Size=8B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 65 | 75.1 | 64.6 | 55.3 | |
| InternVL3.5LLM Size=8B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 64.4 | 77.4 | 62.4 | 53.2 | |
| InternVL3LLM Size=8B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 64.3 | 75.1 | 64.4 | 53.4 | |
| Logic-in-FramesModel Category=Agentic video understanding framework2026.06 | 63 | 71.9 | 61.9 | 55.2 | |
| InternVL3.5LLM Size=4B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 62.7 | 76.4 | 60.3 | 51.3 | |
| LLaVA-OneVision + ReFoCUSLLM Size=7B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 62.6 | 72.8 | 61.7 | 53.4 | |
| EvoGround2026.05 | 62.3 | — | — | — | |
| VideoLLaMA 3 + ReFoCUSLLM Size=7B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 62.2 | 72.2 | 60.1 | 54.3 | |
| Qwen3-VLLLM Size=4B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 62.1 | 74.1 | 61 | 51.3 | |
| Qwen2.5-VL-7Breproduced=true2026.05 | 61.4 | — | — | — | |
| Time-R1reproduced=true2026.05 | 61.2 | — | — | — | |
| LongVULLM Size=7B, Training Free=false2025.03 | 60.9 | 64.7 | 58.2 | 59.5 | |
| GPT-4o + ReFoCUSLLM Size=Closed, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 60.8 | 68.2 | 60.1 | 54 | |
| InternVL3 + ReFoCUSLLM Size=2B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 60.7 | 72.2 | 60.2 | 49.7 | |
| LLaVA-OneVision 32 frames + BOLTLLM Size=7B, Training Free=true, Frames=322025.03 | 59.9 | 70.1 | 60 | 49.6 | |
| LiveVLMFrames=0.5/0.2 fps2025.05 | 59.6 | — | 57 | 51.3 | |
| LLaVA-OneVision-7B + StreamMemFrames=0.5/0.2 fps2025.05 | 59.4 | — | 56.6 | 50.1 | |
| mPLUG-Owl3Model Category=Open-Source Models2026.06 | 59.3 | 70 | 57.7 | 50.1 | |
| VideoLLaMA 3LLM Size=7B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 59 | 70.4 | 57.7 | 48.9 | |
| GPT-4oLLM Size=Closed, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 58.8 | 68 | 55 | 53.3 | |
| LLaVA-OneVision 32 framesLLM Size=7B, Training Free=false, Frames=322025.03 | 58.5 | 70.3 | 56.6 | 48.8 | |
| InternVL3LLM Size=2B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 58.4 | 71 | 56.4 | 47.8 | |
| LLaVA-OneVisionLLM Size=7B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 58.4 | 70.9 | 55.7 | 48.8 | |
| LLaVA-OneVision-7B + ReKVFrames=0.5/0.2 fps2025.05 | 58.3 | — | 55.6 | 48.6 | |
| LLaVA-OneVision 16 frames + BOLTLLM Size=7B, Training Free=true, Frames=162025.03 | 57.8 | 69.2 | 56.8 | 47.3 | |
| Frame-VoyagerLLM Size=8B, Training Free=false2025.03 | 57.5 | 67.3 | 56.3 | 48.9 | |
| LLaVA-OneVision 16 framesLLM Size=7B, Training Free=false, Frames=162025.03 | 56.9 | 68.3 | 54 | 48.2 | |
| LLaVA-OneVision-7BFrames=322025.05 | 56.9 | — | 54.7 | 46.2 | |
| Dispider-7BFrames=1 fps†2025.05 | 56.5 | — | 53.7 | 49.7 | |
| LLaVA-OneVision 8 frames + BOLTLLM Size=7B, Training Free=true, Frames=82025.03 | 56.1 | 66.8 | 54.2 | 47.3 | |
| Kangaroo-8BFrames=642025.05 | 56 | — | 55.3 | 46.6 | |
| VideoXL-7BFrames=1282025.05 | 55.5 | — | 53.2 | 49.2 | |
| LLaVA-OneVision 8 framesLLM Size=7B, Training Free=false, Frames=82025.03 | 53.8 | 63.6 | 52 | 45.7 | |
| InternVL3 + ReFoCUSLLM Size=1B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 53.6 | 66.4 | 51.8 | 42.6 | |
| LongVA-7BFrames=1282025.05 | 52.6 | — | 50.4 | 46.2 | |
| InternVL3LLM Size=1B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 50 | 63.1 | 46.9 | 39.9 | |
| GPT5-NanoZero-shot=true2026.01 | 49.4 | — | — | — | |
| VideoLLaMA2LLM Size=7B, Training Free=false2025.03 | 47.9 | 56 | 45.4 | 42.1 | |
| LLaVA-OneVision + ReFoCUSLLM Size=0.5B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 47.1 | 58.3 | 44.6 | 38.3 | |
| VideoLLaMA 3 + ReFoCUSLLM Size=2B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 47.1 | 58.9 | 44.1 | 38.3 | |
| LLaVA-OneVisionLLM Size=0.5B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 43.5 | 53.7 | 39.9 | 37 | |
| VideoLLaMA 3LLM Size=2B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 43.1 | 55.2 | 38.8 | 35.2 | |
| Chat-UniVi-V1.5LLM Size=7B, Training Free=false2025.03 | 40.6 | 45.7 | 40.3 | 35.8 | |
| LLaVA-OneVisionSize=7B, Protocol=cloze format2026.04 | 40.1 | — | — | — | |
| Video-LLaVALLM Size=7B, Training Free=false2025.03 | 39.9 | 45.3 | 38 | 36.2 | |
| ShareGPT4VideoLLM Size=8B, Training Free=false2025.03 | 39.9 | 48.3 | 36.3 | 35 | |
| VideoChat2LLM Size=7B, Training Free=false2025.03 | 39.5 | 48.3 | 37 | 33.2 | |
| MovieChat-7BFrames=20482025.05 | 38.2 | — | — | 33.4 | |
| ROMAevaluation_protocol=spoken questions, K=12026.01 | 34.56 | — | — | — | |
| ROMAevaluation_protocol=spoken questions2026.01 | 33.3 | — | — | — | |
| ROMAevaluation_protocol=spoken questions, wpos=22026.01 | 33.2 | — | — | — | |
| ROMAevaluation_protocol=spoken questions, wpos=42026.01 | 33.1 | — | — | — | |
| ROMAevaluation_protocol=spoken questions, ablation=Mixed Training2026.01 | 33 | — | — | — | |
| Video-LLaVASize=7B, Protocol=cloze format2026.04 | 30.6 | — | — | — | |
| ABMambaSize=3.6B, Protocol=cloze format2026.04 | 29.4 | — | — | — | |
| VITA-1.5evaluation_protocol=spoken questions2026.01 | 28.56 | — | — | — | |
| Video-ChatGPTSize=7B, Protocol=cloze format2026.04 | 28 | — | — | — | |
| InternVL2.5Size=2.2B, Protocol=cloze format2026.04 | 27.6 | — | — | — | |
| VideoLLaMA3Size=2B, Protocol=cloze format2026.04 | 27.2 | — | — | — | |
| Qwen2.5-Omnievaluation_protocol=spoken questions2026.01 | 20.5 | — | — | — | |
| MiniCPM-oevaluation_protocol=spoken questions2026.01 | 19.37 | — | — | — | |
| ROMAevaluation_protocol=spoken questions, ablation=without speak head2026.01 | 9.11 | — | — | — | |
| Deep Video DiscoveryModel Category=Agentic video understanding framework2026.06 | — | — | — | 67.3 | |
| Gemini-2.0-FlashModel Category=Closed-Source Models2026.06 | — | — | — | 63 | |
| MR. VideoModel Category=Agentic video understanding framework2026.06 | — | — | — | 61.8 | |
| OpenAI o3Model Category=Closed-Source Models2026.06 | — | — | — | 63.2 | |
| VideoTreeModel Category=Agentic video understanding framework2026.06 | — | — | — | 54.2 |