Video Social Intelligence on VSIBench
36.1AccuracySDRL
Evaluation Results
| Method | Links | |
|---|---|---|
| SDRLTraining=RL, Training dataset=EventFlow2026.03 | 36.1 | |
| VideoRFTTraining=SFT+ RL, Input frames=16-frame2026.03 | 35.7 | |
| Video-R1Training=SFT+ RL2026.03 | 34.6 | |
| SDRLTraining=RL2026.03 | 32.9 | |
| LLaVA-OneVision-7BTraining=None2026.03 | 32.4 | |
| VideoRFTTraining=RL, Input frames=16-frame2026.03 | 32.1 | |
| Video-R1Training=SFT2026.03 | 31.8 | |
| Video-R1Training=RL2026.03 | 31.8 | |
| VideoRFTTraining=SFT, Input frames=16-frame2026.03 | 31.7 | |
| LongVA-7BTraining=None2026.03 | 29.2 | |
| Qwen2.5-VL-7BTraining=None2026.03 | 29.1 | |
| VILA-1.5-8BTraining=None2026.03 | 28.9 | |
| Qwen2.5-VL-7BTraining=None, Chain-of-Thought (CoT)=video-r1 CoT2026.03 | 27.7 | |
| Qwen2.5-VL-7BTraining=None, Chain-of-Thought (CoT)=ours CoT2026.03 | 26.6 |