Grounded Video Question Answering on CGBench (test)
5.62mIoUGPT-4o
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-4oTraining Strategy=Proprietary Models, Zero-Shot=true2026.07 | 5.62 | 8.3 | 4.38 | |
| TimeThinkSize=7B, Training Strategy=Reinforcement Fine-Tuning Models, Zero-Shot=true2026.07 | 5.55 | 8.27 | 4.01 | |
| Qwen2.5-VL-GRPOSize=7B, Training Strategy=Reinforcement Fine-Tuning Models, Zero-Shot=true2026.07 | 4.88 | 7.39 | 3.85 | |
| Qwen2.5-VL-SFTSize=7B, Training Strategy=Supervised Fine-Tuning Models, Zero-Shot=true2026.07 | 4.01 | 5.54 | 3.31 | |
| Gemini-1.5-ProTraining Strategy=Proprietary Models, Zero-Shot=true2026.07 | 3.95 | 5.81 | 2.53 | |
| InternVL2Size=7B, Training Strategy=Supervised Fine-Tuning Models, Zero-Shot=true2026.07 | 3.91 | 5.05 | 2.64 | |
| Qwen2-VLSize=72B, Training Strategy=Supervised Fine-Tuning Models, Zero-Shot=true2026.07 | 3.58 | 5.32 | 2.54 | |
| LongVASize=7B, Training Strategy=Supervised Fine-Tuning Models, Zero-Shot=true2026.07 | 2.94 | 3.86 | 1.78 | |
| VideoCCAMSize=14B, Training Strategy=Supervised Fine-Tuning Models, Zero-Shot=true2026.07 | 2.63 | 3.53 | 1.76 | |
| MiniCPM-v2.6Size=8B, Training Strategy=Supervised Fine-Tuning Models, Zero-Shot=true2026.07 | 2.35 | 2.96 | 1.35 | |
| ShareGPT4VideoSize=16B, Training Strategy=Supervised Fine-Tuning Models, Zero-Shot=true2026.07 | 1.85 | 2.65 | 1.01 | |
| LLaVA-OVSize=13B, Training Strategy=Supervised Fine-Tuning Models, Zero-Shot=true2026.07 | 1.63 | 1.78 | 1.01 | |
| Videochat2Size=7B, Training Strategy=Supervised Fine-Tuning Models, Zero-Shot=true2026.07 | 1.28 | 1.98 | 0.94 |