Temporal Reasoning on TempCompass
74.4AccuracySDRL
Evaluation Results
| Method | Links | |
|---|---|---|
| SDRLTraining=RL, Training dataset=EventFlow2026.03 | 74.4 | |
| SDRLTraining=RL2026.03 | 73.4 | |
| TW-GRPOTraining=RL2026.03 | 73.3 | |
| VideoRFTTraining=SFT+ RL, Input frames=16-frame2026.03 | 73.1 | |
| VideoChat-R1Training=RL2026.03 | 72.9 | |
| Video-R1Training=SFT+ RL2026.03 | 72.6 | |
| Qwen2.5-VL-7BTraining=None2026.03 | 72.5 | |
| OmniJigsaw (CMM)Inference Mode=w/o Audio2026.04 | 72.34 | |
| VideoJigsawInference Mode=w/ Audio2026.04 | 72.28 | |
| Qwen2.5-VL-7BTraining=None, Chain-of-Thought (CoT)=video-r1 CoT2026.03 | 72.2 | |
| OmniJigsaw (SMS)Inference Mode=w/ Audio2026.04 | 72.03 | |
| OmniJigsaw (CMM)Inference Mode=w/ Audio2026.04 | 72.03 | |
| VideoJigsawInference Mode=w/o Audio2026.04 | 71.52 | |
| OmniJigsaw (SMS)Inference Mode=w/o Audio2026.04 | 71.52 | |
| OmniJigsaw (JMI)Inference Mode=w/o Audio2026.04 | 71.39 | |
| OmniJigsaw (JMI)Inference Mode=w/ Audio2026.04 | 71.08 | |
| Video-R1Training=RL2026.03 | 70.9 | |
| VideoRFTTraining=RL, Input frames=16-frame2026.03 | 70.8 | |
| Qwen3-Omni-30BInference Mode=w/o Audio2026.04 | 70.7 | |
| Qwen3-Omni-30BInference Mode=w/ Audio2026.04 | 70.63 | |
| Video-R1Inference Mode=w/o Audio2026.04 | 70 | |
| Qwen2.5-VL-7BTraining=None, Chain-of-Thought (CoT)=ours CoT2026.03 | 69.5 | |
| Video-R1Training=SFT2026.03 | 69.2 | |
| VideoRFTTraining=SFT, Input frames=16-frame2026.03 | 68.4 | |
| HumanOmniV2Inference Mode=w/ Audio2026.04 | 63.86 | |
| Omni-R1Inference Mode=w/ Audio2026.04 | 63.1 | |
| HumanOmniV2Inference Mode=w/o Audio2026.04 | 63.1 | |
| Omni-R1Inference Mode=w/o Audio2026.04 | 62.78 | |
| Kangaroo-8BTraining=None2026.03 | 62.5 | |
| Video-UTR-7BTraining=None2026.03 | 59.7 | |
| VILA-1.5-8BTraining=None2026.03 | 58.8 | |
| LongVA-7BTraining=None2026.03 | 56.9 | |
| LLaMA-VIDTraining=None2026.03 | 45.6 |