Visual Question Answering on SEED-Bench 2-Plus
70.32AccuracyHyLaR-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| HyLaR-7BModel Type=Visual-Latent Model, Number of Parameters=7B2026.04 | 70.32 | |
| ZoomEyeModel Type=Thinking-with-Images Agent Model2026.04 | 70.27 | |
| LaserModel Type=Visual-Latent Model2026.04 | 70.05 | |
| InternVL3.5-8BModel Type=Open-Source Model, Number of Parameters=8B2026.04 | 69.78 | |
| HyLaR-SFTModel Type=Visual-Latent Model, Training Protocol=SFT2026.04 | 69.39 | |
| DeepEyesModel Type=Thinking-with-Images Agent Model2026.04 | 69.08 | |
| InternVL3.5-2BParameter Count=2B2025.12 | 68 | |
| Qwen3-VL-2BParameter Count=2B, Evaluation Toolkit=VLMEvalKit2025.12 | 67.3 | |
| jina-vlm2025.12 | 67.2 | |
| MonetModel Type=Visual-Latent Model2026.04 | 65.88 | |
| Qwen2.5-VL-7BModel Type=Open-Source Model, Number of Parameters=7B2026.04 | 65.31 | |
| InternVL3-2BParameter Count=2B2025.12 | 64.6 | |
| Qwen2-VL-2BParameter Count=2B2025.12 | 62.4 | |
| LLaVA-OneVisionModel Type=Open-Source Model2026.04 | 61.22 | |
| LVRModel Type=Visual-Latent Model2026.04 | 47.39 | |
| Omni-DiffusionModel Type=Any-to-Any, #Params=7B2026.03 | 34.5 | |
| Emu‡Model Type=Visual LLM, #Params=14B2026.03 | 33.5 | |
| mPLUG-OwlModel Type=Visual LLM†, #Params=7B2026.03 | 31.8 | |
| LLaVAModel Type=Visual LLM†, #Params=7B2026.03 | 30.1 | |
| InstructBLIPModel Type=Visual LLM†, #Params=14B2026.03 | 29.2 | |
| NExT-GPT‡Model Type=Any-to-Any, #Params=7B2026.03 | 26.2 |