Multi-modal Understanding on MMBench (mean accuracy)
86.3Mean AccuracyOryx-1.5
Evaluation Results
| Method | Links | |
|---|---|---|
| Oryx-1.5Size=32B2024.09 | 86.3 | |
| LLaVA-OneVisionSize=72B2024.09 | 85.6 | |
| OryxSize=34B2024.09 | 84.5 | |
| Latent DenoisingArchitecture=Qwen-2.5-VL2026.04 | 84.5 | |
| Qwen3-VL-4BModel=Qwen3-VL, Parameter Count=4B2026.04 | 83.58 | |
| BaselineArchitecture=Qwen-2.5-VL2026.04 | 83.1 | |
| VILA-1.5Size=40B2024.09 | 82.4 | |
| Cambrian-1Size=34B2024.09 | 81.4 | |
| OryxSize=7B2024.09 | 81.4 | |
| Oryx-1.5Size=7B2024.09 | 81.3 | |
| LLaVA-OneVision-7BModel=LLaVA-OneVision, Parameter Count=7B2026.04 | 80.57 | |
| LLaVA-NeXTSize=34B2024.09 | 79.3 | |
| InfiniteVL-4BModel=InfiniteVL, Parameter Count=4B, Result Source=InfiniteVL (Tao et al., 2025) paper2026.04 | 79 | |
| SpB2.0-VL-5BModel=SpB2.0-VL, Parameter Count=5B2026.04 | 78.39 | |
| Bunny-LLama3Size=8B2024.09 | 77.2 | |
| Qwen2.5-VL-3BModel=Qwen2.5-VL, Parameter Count=3B2026.04 | 76.88 | |
| Idefics2Size=8B2024.09 | 76.7 | |
| Cambrian-1Size=8B2024.09 | 75.9 | |
| VILA-1.5Size=8B2024.09 | 75.3 | |
| InternVL2-4BModel=InternVL2, Parameter Count=4B, Result Source=OpenVLM Leaderboard (OpenCompass, 2025)2026.04 | 73.6 | |
| Deepseek-VLSize=7B2024.09 | 73.2 | |
| MonkeySize=7B2024.09 | 72.4 | |
| LLaVA-NeXTSize=8B2024.09 | 72.1 | |
| VITASize=8x7B2024.09 | 71.8 | |
| MaTVLM-3BModel=MaTVLM, Parameter Count=3B, Result Source=InfiniteVL (Tao et al., 2025) paper2026.04 | 69.4 | |
| Latent DenoisingArchitecture=LLaVA+SigLIP2026.04 | 68.7 | |
| VanillaBackbone=LLaVA-1.5-13B, Retained Tokens=5762026.06 | 68.5 | |
| PriorTRBackbone=LLaVA-1.5-13B, Retained Tokens=1922026.06 | 68.3 | |
| PDropBackbone=LLaVA-1.5-13B, Retained Tokens=1922026.06 | 67.5 | |
| SmolVLM2-2.2BModel=SmolVLM2, Parameter Count=2.2B2026.04 | 67.48 | |
| SparseVLMBackbone=LLaVA-1.5-13B, Retained Tokens=1922026.06 | 67.4 | |
| PriorTRBackbone=LLaVA-1.5-13B, Retained Tokens=1282026.06 | 67.3 | |
| VisPrunerBackbone=LLaVA-1.5-13B, Retained Tokens=1922026.06 | 66.8 | |
| CAMDBase Model=LLaVA-1.52026.03 | 66.6 | |
| VisPrunerBackbone=LLaVA-1.5-13B, Retained Tokens=1282026.06 | 66.5 | |
| FarSightBase Model=LLaVA-1.52026.03 | 66 | |
| SparseVLMBackbone=LLaVA-1.5-13B, Retained Tokens=1282026.06 | 65.8 | |
| PriorTRBackbone=LLaVA-1.5-13B, Retained Tokens=642026.06 | 65.8 | |
| Latent DenoisingArchitecture=LLaVA+CLIP2026.04 | 65.7 | |
| PDropBackbone=LLaVA-1.5-13B, Retained Tokens=1282026.06 | 65 | |
| LLaVA-1.5-7BRetained Tokens=576, Retention Ratio=100%, Base Model=LLaVA-1.5-7B2025.08 | 64.7 | |
| EVIANData Selection Strategy=10K subset2026.04 | 64.63 | |
| CGDBase Model=LLaVA-1.5, Decoding Strategy=CGD2026.03 | 64.5 | |
| BaselineArchitecture=LLaVA+SigLIP2026.04 | 64.5 | |
| OPERABase Model=LLaVA-1.5, Decoding Strategy=OPERA2026.03 | 64.4 | |
| LLaVA-1.5Decoding Strategy=Greedy2026.03 | 64.3 | |
| BaselineArchitecture=LLaVA+CLIP2026.04 | 64 | |
| VCDBase Model=LLaVA-1.5, Decoding Strategy=VCD2026.03 | 63.9 | |
| HiPrune++Retained Tokens=192, Retention Ratio=33.3%, Base Model=LLaVA-1.5-7B2025.08 | 63.5 | |
| VisPrunerBackbone=LLaVA-1.5-13B, Retained Tokens=642026.06 | 63.2 | |
| SCALEData Selection Strategy=10K subset2026.04 | 63.18 | |
| BLIP-2Data Selection Strategy=10K subset2026.04 | 63.17 | |
| ICDBase Model=LLaVA-1.5, Decoding Strategy=ICD2026.03 | 63.1 | |
| CAMDBase Model=Video-LLaVA2026.03 | 63.1 | |
| FarSightBase Model=Video-LLaVA2026.03 | 62.8 | |
| HiPruneRetained Tokens=192, Retention Ratio=33.3%, Base Model=LLaVA-1.5-7B2025.08 | 62.8 | |
| HiPrune++Retained Tokens=128, Retention Ratio=22.2%, Base Model=LLaVA-1.5-7B2025.08 | 62.3 | |
| HiPruneRetained Tokens=128, Retention Ratio=22.2%, Base Model=LLaVA-1.5-7B2025.08 | 62.2 | |
| BLIPData Selection Strategy=10K subset2026.04 | 61.83 | |
| SparseVLMBackbone=LLaVA-1.5-13B, Retained Tokens=642026.06 | 61.3 | |
| Video-LLaVADecoding Strategy=Greedy2026.03 | 60.9 | |
| CAMDBase Model=Chat-UniVi2026.03 | 60.6 | |
| HiPrune++Retained Tokens=64, Retention Ratio=11.1%, Base Model=LLaVA-1.5-7B2025.08 | 60.3 | |
| ALBEFData Selection Strategy=10K subset2026.04 | 60.03 | |
| FarSightBase Model=Chat-UniVi2026.03 | 59.8 | |
| Full DataData Selection Strategy=Full 300K pool2026.04 | 59.53 | |
| HiPruneRetained Tokens=64, Retention Ratio=11.1%, Base Model=LLaVA-1.5-7B2025.08 | 59.5 | |
| PDropBackbone=LLaVA-1.5-13B, Retained Tokens=642026.06 | 59.2 | |
| Qwen2.5-VLData Selection Strategy=10K subset2026.04 | 57.96 | |
| FastVBackbone=LLaVA-1.5-13B, Retained Tokens=1282026.06 | 57.9 | |
| CLIPScoreData Selection Strategy=10K subset2026.04 | 57.46 | |
| Chat-UniViDecoding Strategy=Greedy2026.03 | 56.3 | |
| Cobra-3BModel=Cobra, Parameter Count=3B, Result Source=InfiniteVL (Tao et al., 2025) paper2026.04 | 55.9 | |
| FastVBackbone=LLaVA-1.5-13B, Retained Tokens=1922026.06 | 54 | |
| RandomData Selection Strategy=10K subset2026.04 | 53.53 | |
| FastVBackbone=LLaVA-1.5-13B, Retained Tokens=642026.06 | 50.9 | |
| CAMDBase Model=InstructBLIP2026.03 | 46.8 | |
| FarSightBase Model=InstructBLIP2026.03 | 46.5 | |
| InstructBLIPDecoding Strategy=Greedy2026.03 | 43.4 |