Audio-visual understanding on WorldSense
66.4AccuracyGemini-3-Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini-3-ProModel access type=Closed-source2026.02 | 66.4 | |
| OmniVideo-R1Model access type=Open-source2026.02 | 65.8 | |
| Gemini-2.5-ProModel access type=Closed-source2026.02 | 64.6 | |
| Qwen3.5-OmniOpen-Source=✗, Size=Flash2026.04 | 57.8 | |
| video-SALMONN 2+-72BModel access type=Open-source2026.02 | 56.5 | |
| Gemini-2.0-FlashModel access type=Closed-source2026.02 | 56.2 | |
| Nemotron 3 Nano OmniOpen-Source=✓, Size=30B-A3B, Mode=Reasoning on2026.04 | 55.4 | |
| Nemotron 3 Nano OmniOpen-Source=✓, Size=30B-A3B, Mode=Reasoning off2026.04 | 55.2 | |
| Qwen3-Omni-30B-A3B-InstructModel access type=Open-source2026.02 | 54 | |
| Qwen3-Omni-InstructModel Parameters=30B-A3B2026.02 | 54 | |
| Qwen3-OmniOpen-Source=✓, Size=30B-A3B, Mode=Instruct2026.04 | 54 | |
| D-ORCAModel Parameters=8B2026.02 | 53.7 | |
| Full TokensModel=Qwen2.5-Omni-7B, Retained Ratio=100%, FLOPs Ratio=100%2026.05 | 52.4 | |
| OmniRefineModel=Qwen2.5-Omni-3B, Retained Ratio=37%, FLOPs Ratio=22%2026.05 | 52.2 | |
| Full TokensModel=Qwen2.5-Omni-3B, Retained Ratio=100%, FLOPs Ratio=100%2026.05 | 51.5 | |
| video-SALMONN 2+-7BModel access type=Open-source2026.02 | 50.9 | |
| video-SALMONN 2+Model Parameters=7B2026.02 | 50.9 | |
| OmniRefineModel=Qwen2.5-Omni-7B, Retained Ratio=44%, FLOPs Ratio=31%2026.05 | 50.4 | |
| OmniRefineModel=Qwen2.5-Omni-7B, Retained Ratio=30%, FLOPs Ratio=20%2026.05 | 50.4 | |
| Qwen3-Omni-30BRecipe=Vanilla Baseline2026.05 | 50.3 | |
| Final 10K DPO Recipe (Ours)Recipe=Combined CTP, FV-D, and FV-A-L2026.05 | 50.3 | |
| OmniZipModel=Qwen2.5-Omni-7B, Retained Ratio=45%, FLOPs Ratio=39%2026.05 | 50.1 | |
| OmniZipModel=Qwen2.5-Omni-3B, Retained Ratio=45%, FLOPs Ratio=36%2026.05 | 50.1 | |
| FastVModel=Qwen2.5-Omni-3B, Retained Ratio=50%, FLOPs Ratio=49%2026.05 | 50 | |
| AVoCaDOModel Parameters=7B2026.02 | 49.9 | |
| DPO w/ OP + FV-D + LV-MCQARecipe=DPO with original-sync, video preference, and MCQA data2026.05 | 49.9 | |
| DPO w/ CTP + FV-D + FV-ARecipe=DPO with counterfactual, video preference, and audio preference data2026.05 | 49.9 | |
| DPO w/ SPRecipe=DPO with SFT-policy negatives2026.05 | 49.8 | |
| DPO w/ SP + FV-DRecipe=DPO with SFT-policy negatives and video preference data2026.05 | 49.8 | |
| DPO w/ CTP + FV-D + LV-MCQARecipe=DPO with counterfactual, video preference, and MCQA data2026.05 | 49.8 | |
| DPO w/ OP + SPRecipe=DPO with original-sync and SFT-policy negatives2026.05 | 49.7 | |
| OmniRefineModel=Qwen2.5-Omni-3B, Retained Ratio=23%, FLOPs Ratio=18%2026.05 | 49.6 | |
| DPO w/ CTP + FV-DRecipe=DPO with counterfactual and video preference data2026.05 | 49.5 | |
| OmniZipModel=Qwen2.5-Omni-3B, Retained Ratio=35%, FLOPs Ratio=26%2026.05 | 48.8 | |
| FastVModel=Qwen2.5-Omni-7B, Retained Ratio=50%, FLOPs Ratio=54%2026.05 | 48.8 | |
| DyCoke (V&A)Model=Qwen2.5-Omni-7B, Retained Ratio=50%, FLOPs Ratio=44%2026.05 | 48.4 | |
| OmniZipModel=Qwen2.5-Omni-7B, Retained Ratio=35%, FLOPs Ratio=29%2026.05 | 48.3 | |
| OmniVinciModel Parameters=7B2026.02 | 48.2 | |
| RandomModel=Qwen2.5-Omni-3B, Retained Ratio=55%, FLOPs Ratio=45%2026.05 | 48.2 | |
| SFT w/ CTP + FV-D + FV-ALRecipe=SFT with counterfactual, general video preference, and audio-visual data2026.05 | 48.2 | |
| DyCoke (V&A)Model=Qwen2.5-Omni-3B, Retained Ratio=50%, FLOPs Ratio=40%2026.05 | 48.1 | |
| Qwen3-Omni-30B-A3B-ThinkingModel access type=Open-source2026.02 | 48 | |
| Qwen2.5-OmniModel Parameters=7B2026.02 | 47.8 | |
| EchoingPixelsModel Size=7B, Token Budget=10%, Est. Tokens Per Min=1,0142025.12 | 47.4 | |
| HumanOmniV2-7BModel access type=Open-source2026.02 | 47.1 | |
| RandomModel=Qwen2.5-Omni-7B, Retained Ratio=55%, FLOPs Ratio=48%2026.05 | 47.1 | |
| Qwen2.5-Omni-7BModel Size=7B, Est. Tokens Per Min=10,1402025.12 | 46.1 | |
| Qwen2.5-Omni + AVATARReinforcement Learning Strategy=AVATAR2025.08 | 46 | |
| EcholnkReinforcement Learning Strategy=N/A2025.08 | 45.7 | |
| Qwen2.5-Omni-7BModel access type=Open-source2026.02 | 45.4 | |
| Qwen2.5-Omni-3BModel Size=3B, Est. Tokens Per Min=10,1402025.12 | 45.4 | |
| HumanOmniReinforcement Learning Strategy=N/A2025.08 | 45.4 | |
| ARC-Qwen-Video-NarratorModel Parameters=7B2026.02 | 45.1 | |
| Qwen2.5-Omni + GRPOReinforcement Learning Strategy=GRPO2025.08 | 45.1 | |
| EchoingPixelsModel Size=3B, Token Budget=20%, Est. Tokens Per Min=2,0282025.12 | 45 | |
| Ola-7B + AVATARReinforcement Learning Strategy=AVATAR2025.08 | 45 | |
| Ola-7B + GRPOReinforcement Learning Strategy=GRPO2025.08 | 44.7 | |
| AV-ReasonerReinforcement Learning Strategy=N/A2025.08 | 44.6 | |
| Ola-7BReinforcement Learning Strategy=Baseline2025.08 | 44.2 | |
| Qwen2.5-OmniReinforcement Learning Strategy=Baseline2025.08 | 44.2 | |
| Omni-R1Reinforcement Learning Strategy=N/A2025.08 | 44.1 | |
| EchoingPixelsModel Size=3B, Token Budget=10%, Est. Tokens Per Min=1,0142025.12 | 43.5 | |
| EchoingPixelsModel Size=3B, Token Budget=5%, Est. Tokens Per Min=5072025.12 | 40.9 | |
| IntraModalModel Size=7B, Token Budget=25%, Est. Tokens Per Min=2,5352025.12 | 40.6 | |
| FastVModel Size=3B, Token Budget=20%, Est. Tokens Per Min=2,0282025.12 | 40.1 | |
| IntraModalModel Size=3B, Token Budget=25%, Est. Tokens Per Min=2,5352025.12 | 37.4 | |
| VITA-1.5-7BModel access type=Open-source2026.02 | 36.9 | |
| Vita-1.5Model Size=7B, Est. Tokens Per Min=7,0962025.12 | 36.9 | |
| Unified-IO-2-XXLModel Size=7B, Est. Tokens Per Min=8962025.12 | 25.9 | |
| VideoLLaMA2-7BModel access type=Open-source2026.02 | 25.4 | |
| VideoLLaMA2Model Size=7B, Est. Tokens Per Min=∼4,1522025.12 | 25.4 | |
| Unified-IO-2-XLModel Size=3B, Est. Tokens Per Min=8962025.12 | 24.7 |