Vision-Language Reasoning on CVBench
86.16AccuracyQwen3-VL-8B + GRPO
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-VL-8B + GRPOParams=8B, Training Scale=5.3K QA, Zero-shot evaluation=true2026.01 | 86.16 | |
| Qwen2-VL-7B + GRPOParams=7B, Training Scale=5.3K QA, Zero-shot evaluation=true2026.01 | 75.21 | |
| Gemini-3.0-FlashZero-shot evaluation=true2026.01 | 67.2 | |
| Gemini-2.5-ProZero-shot evaluation=true2026.01 | 62.4 | |
| Qwen2-VL-2B + GRPOParams=2B, Training Scale=5.3K QA, Zero-shot evaluation=true2026.01 | 60.31 | |
| InternVideo2.5-8BParams=8B, Training Scale=16M clips, Zero-shot evaluation=true2026.01 | 57.3 | |
| LLaVA-Video-7BParams=7B, Training Scale=1.3M videos, Zero-shot evaluation=true2026.01 | 52.6 | |
| GPT-4VParams=~1.8T, Training Scale=~10T tokens, Zero-shot evaluation=true2026.01 | 52.4 | |
| Qwen2-VL-7B (baseline)Params=7B, Training Scale=1.2T tokens2026.01 | 50.7 | |
| Qwen3-VL-8B (baseline)Params=8B, Training Scale=1.0T tokens2026.01 | 45.8 | |
| Qwen2-VL-2B (baseline)Params=2B, Training Scale=1.2T tokens2026.01 | 31.38 | |
| Video-LLaVA-7BParams=7B, Training Scale=760K videos, Zero-shot evaluation=true2026.01 | 28.1 |