Zoom-dependent visual question answering on ZoomBench
65.8AccuracyVision-OPD (Ours)
Evaluation Results
| Method | Links | |
|---|---|---|
| Vision-OPD (Ours)Param Size=9B, Inference Mode=Single Forward Pass2026.05 | 65.8 | |
| Gemini-3.1-ProParam Size=-, Inference Mode=Closed-Source2026.05 | 61.18 | |
| Vision-OPD (Ours)Param Size=4B, Inference Mode=Single Forward Pass2026.05 | 59.76 | |
| Qwen3.5Param Size=397B, Inference Mode=Single Forward Pass2026.05 | 57.16 | |
| ZwZParam Size=8B, Inference Mode=Single Forward Pass2026.05 | 56.69 | |
| Qwen3-VL-InstructParam Size=235B, Inference Mode=Single Forward Pass2026.05 | 56.09 | |
| GPT-5.4Param Size=-, Inference Mode=Closed-Source2026.05 | 52.66 | |
| Kimi-K2.5Param Size=1T, Inference Mode=Single Forward Pass2026.05 | 52.43 | |
| Qwen3.5Param Size=9B, Inference Mode=Single Forward Pass2026.05 | 52.07 | |
| GPT-5.2Param Size=-, Inference Mode=Closed-Source2026.05 | 50.89 | |
| GLM-4.6VParam Size=106B, Inference Mode=Single Forward Pass2026.05 | 50.06 | |
| SenseNova-MARSParam Size=8B, Inference Mode=Agentic2026.05 | 47.81 | |
| Qwen3.5Param Size=4B, Inference Mode=Single Forward Pass2026.05 | 47.69 | |
| DeepEyesParam Size=7B, Inference Mode=Agentic2026.05 | 46.51 | |
| MiMo-VL-RLParam Size=7B, Inference Mode=Single Forward Pass2026.05 | 45.68 | |
| ThymeParam Size=7B, Inference Mode=Agentic2026.05 | 45.09 | |
| DeepEyesV2Param Size=7B, Inference Mode=Agentic2026.05 | 44.97 | |
| Qwen3-VL-InstructParam Size=8B, Inference Mode=Single Forward Pass2026.05 | 42.96 | |
| MiniCPM-V-4.5Param Size=9B, Inference Mode=Single Forward Pass2026.05 | 42.6 |