multi-view Visual Question Answering on VSI-Bench (test)
52.9Average ScoreVaLR-M
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| VaLR-MMethod Group=Latent Reasoning Models, Encoder Alignment=Multiple encoder (DINOv3, SigLIPv2, pi3)2026.02 | 52.9 | 66.4 | 40.6 | 64.2 | 56.6 | 50 | 51.8 | 35.1 | 48.9 | |
| VaLR-SMethod Group=Latent Reasoning Models, Encoder Alignment=Single encoder (DINOv3)2026.02 | 41.5 | 49 | 24.5 | 53.9 | 38.2 | 43.9 | 41.9 | 34 | 39.2 | |
| LLaVA-NeXT-Video-7BMethod Group=Other Models, Model Architecture/Scale=7B2026.02 | 35.6 | 48.5 | 14 | 47.8 | 24.2 | 43.5 | 42.4 | 34 | 30.6 | |
| GPT-4oMethod Group=Other Models2026.02 | 34 | 46.2 | 5.3 | 43.8 | 38.2 | 37 | 41.3 | 31.5 | 28.5 | |
| Qwen2.5-VL-7B + vanilla SFTMethod Group=Base Model, Training Strategy=vanilla SFT2026.02 | 33.7 | 42.3 | 14.7 | 44.1 | 20.8 | 39.4 | 34.7 | 32.5 | 33.5 | |
| Qwen2.5-VL-7BMethod Group=Base Model, Model Architecture/Scale=7B2026.02 | 33 | 40.9 | 14.8 | 43.4 | 20.7 | 38.6 | 38.5 | 33 | 29.8 | |
| Ocean-R1-7BMethod Group=Reasoning Models, Model Architecture/Scale=7B2026.02 | 30.5 | 16.5 | 14.6 | 38.9 | 40.1 | 38 | 36.1 | 30.9 | 30.1 | |
| Qwen2.5-VL-7B + CoVTMethod Group=Latent Reasoning Models2026.02 | 18.6 | 16.5 | 2.3 | 1 | 7.3 | 35.9 | 33 | 25.8 | 30.4 | |
| Qwen2.5-VL-7B + LVRMethod Group=Latent Reasoning Models2026.02 | 18.4 | 21.4 | 3.6 | 1.4 | 9 | 35.1 | 30.9 | 32 | 23.1 | |
| R1-OneVision-7BMethod Group=Reasoning Models, Model Architecture/Scale=7B2026.02 | 16.1 | 15 | 1.7 | 0.5 | 2.8 | 26.5 | 40 | 24.2 | 14.7 | |
| Qwen2.5-VL-7B + MonetMethod Group=Latent Reasoning Models2026.02 | 14 | 1.9 | 0.1 | 0 | 0 | 38 | 20.5 | 24.2 | 31.2 |