Video Question Answering on DoraVQA
76.1AccuracyGemini-3.0-Flash
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini-3.0-FlashZero-shot evaluation=true2026.01 | 76.1 | |
| Qwen3-VL-8B + GRPOParams=8B, Training Scale=5.3K QA, Zero-shot evaluation=true2026.01 | 67.98 | |
| GPT-4VParams=~1.8T, Training Scale=~10T tokens, Zero-shot evaluation=true2026.01 | 67.79 | |
| Gemini-2.5-ProZero-shot evaluation=true2026.01 | 64.41 | |
| Qwen2-VL-7B + GRPOParams=7B, Training Scale=5.3K QA, Zero-shot evaluation=true2026.01 | 62.38 | |
| Qwen3-VL-8B (baseline)Params=8B, Training Scale=1.0T tokens2026.01 | 58.08 | |
| InternVideo2.5-8BParams=8B, Training Scale=16M clips2026.01 | 57.68 | |
| Qwen2-VL-7B (baseline)Params=7B, Training Scale=1.2T tokens2026.01 | 56.74 | |
| LLaVA-Video-7BParams=7B, Training Scale=1.3M videos2026.01 | 55.41 | |
| Qwen2-VL-2B + GRPOParams=2B, Training Scale=5.3K QA, Zero-shot evaluation=true2026.01 | 55.11 | |
| Qwen2-VL-2B (baseline)Params=2B, Training Scale=1.2T tokens2026.01 | 41.36 | |
| Video-LLaVA-7BParams=7B, Training Scale=760K videos2026.01 | 37.82 |