Visual Social Reasoning on MoMentS
70.68DirectGPT-4o
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| GPT-4oModel Category=Standard VLMs2025.07 | 70.68 | 2.56 | 0.11 | 1.58 | |
| OpenAI o1-fullModel Category=Reasoning VLMs2025.07 | 67.6 | 3.41 | 1.1 | 1.55 | |
| Gemini-3.0-FlashModel Category=Reasoning VLMs2025.07 | 64.8 | 2 | 3.7 | 6.51 | |
| Claude-3.5-SonnetModel Category=Standard VLMs2025.07 | 63.75 | 2.4 | 1.75 | 0.8 | |
| GPT-5.2Model Category=Reasoning VLMs, Reasoning Effort=High2025.07 | 62.95 | 0.52 | 0.63 | 0.4 | |
| OpenAI o3-miniModel Category=Reasoning VLMs2025.07 | 56 | 3.06 | 0.5 | 2 | |
| Gemini-2.5-ProModel Category=Standard VLMs2025.07 | 55 | 2 | 0.5 | 5 | |
| Qwen2-VL-7BModel Category=Open-Source VLMs2025.07 | 49 | 11.91 | 8.13 | 7.8 | |
| LLaVA-OneVision-7BModel Category=Open-Source VLMs2025.07 | 45.82 | 2.39 | 0.32 | 0.4 |