Physical Reasoning on VideoPhy 2
0.386AccuracyVideoScoreV2
Evaluation Results
| Method | Links | |
|---|---|---|
| VideoScoreV2Model Type=Specialized VQA Model2026.02 | 0.386 | |
| ASOTraining Strategy=Analytic Score Optimization2026.02 | 0.344 | |
| SFTTraining Strategy=Supervised Fine-Tuning2026.02 | 0.306 | |
| Gemini-2.5ProModel Type=VLM-API2026.02 | 0.288 | |
| GPT-4.1Model Type=VLM-API2026.02 | 0.261 | |
| Q-AlignModel Type=Specialized VQA Model2026.02 | 0.232 | |
| Qwen2.5-VLModel Type=Open-source VLM2026.02 | 0.202 | |
| InternVL-3.5Model Type=Open-source VLM2026.02 | 0.181 |