Omni-modal Understanding on WorldSense
48AccuracyGemini-1.5-Pro
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Gemini-1.5-Pro2025.08 | 48 | — | |
| Full TokensBackbone=Qwen2.5-Omni-7B, GPU Mem.=44G, Prefilling Time=2371ms (1.00×), Latency per Example=10.99s (1.00×)2026.05 | 46.8 | — | |
| OmniRefineBackbone=Qwen2.5-Omni-7B, Retention Ratio=30%, GPU Mem.=29G, Prefilling Time=451ms (5.26×), Latency per Example=9.59s (1.15×)2026.05 | 46.4 | — | |
| OmniZipBackbone=Qwen2.5-Omni-7B, Retention Ratio=45%, GPU Mem.=32G, Prefilling Time=894ms (2.65×), Latency per Example=7.99s (1.38×)2026.05 | 45.9 | — | |
| Qwen2.5-OmniParams (B)=7, Time (s)=6.02025.08 | 45.4 | — | |
| OmniZipBackbone=Qwen2.5-Omni-7B, Retention Ratio=35%, GPU Mem.=30G, Prefilling Time=649ms (3.65×), Latency per Example=7.46s (1.47×)2026.05 | 45.3 | — | |
| DyCoke (V&A)Backbone=Qwen2.5-Omni-7B, GPU Mem.=36G, Prefilling Time=1386ms (1.71×), Latency per Example=8.59s (1.28×)2026.05 | 44.6 | — | |
| Ours:base Qwen2.5Params (B)=14+72, Time (s)=3.22025.08 | 44.1 | — | |
| GPT-4oTime (s)=1.22025.08 | 42.6 | — | |
| InternVL-2Params (B)=82025.08 | 39.1 | — | |
| LLaVA-OVParams (B)=72025.08 | 37.7 | — | |
| VITA1.5Params (B)=7, Time (s)=3.72025.08 | 36.9 | — | |
| Gemini 2.5 FlashSize=-, Settings=simplex2026.04 | — | 52.6 | |
| MiniCPM-o 4.5Size=9B, Settings=simplex2026.04 | — | 55.7 | |
| Qwen3-OmniSize=30B-A3B, Settings=simplex2026.04 | — | 54 |