Video Understanding Reasoning on MLVU
73.46AccuracyOmniJigsaw (CMM)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| OmniJigsaw (CMM)Inference Mode=w/o Audio2026.04 | 73.46 | — | |
| VideoJigsawInference Mode=w/o Audio2026.04 | 73.14 | — | |
| OmniJigsaw (SMS)Inference Mode=w/ Audio2026.04 | 72.63 | — | |
| OmniJigsaw (JMI)Inference Mode=w/o Audio2026.04 | 72.45 | — | |
| OmniJigsaw (CMM)Inference Mode=w/ Audio2026.04 | 72.26 | — | |
| OmniJigsaw (SMS)Inference Mode=w/o Audio2026.04 | 72.17 | — | |
| VideoJigsawInference Mode=w/ Audio2026.04 | 71.9 | — | |
| OmniJigsaw (JMI)Inference Mode=w/ Audio2026.04 | 71.39 | — | |
| Qwen3-Omni-30BInference Mode=w/o Audio2026.04 | 70.98 | — | |
| Qwen3-Omni-30BInference Mode=w/ Audio2026.04 | 70.01 | — | |
| Video-R1Inference Mode=w/o Audio2026.04 | 67.07 | — | |
| HumanOmniV2Inference Mode=w/ Audio2026.04 | 66.7 | — | |
| HumanOmniV2Inference Mode=w/o Audio2026.04 | 66.65 | — | |
| Omni-R1Inference Mode=w/ Audio2026.04 | 65.69 | — | |
| Omni-R1Inference Mode=w/o Audio2026.04 | 64.86 | — | |
| STDBackbone=Qwen2-VL-7B-Instruct2025.05 | 0.661 | 1.94 | |
| StreamingBackbone=Qwen2-VL-7B-Instruct2025.05 | 0.539 | 1.61 | |
| STDBackbone=LLaVA-OneVision-7B2025.05 | 0.478 | 1.72 | |
| StreamingBackbone=LLaVA-OneVision-7B2025.05 | 0.347 | 1.34 | |
| LayerSkipBackbone=LLaVA-OneVision-7B2025.05 | 0.1 | 0.47 | |
| LayerSkipBackbone=Qwen2-VL-7B-Instruct2025.05 | 0.052 | 0.63 |