Hallucination Evaluation on HallusionBench
82.02AccuracyMIRROR(ours)
Evaluation Results
| Method | Links | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MIRROR(ours)Param Size=7B2026.02 | 82.02 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VL-RethinkerReasoning Paradigm=Tool-use & RL Enhanced Reasoning2026.01 | 71.08 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ZoomEyeModel Type=Thinking-with-Images Agent Model2026.04 | 71.08 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Visual Para-ThinkerSize=7B2026.02 | 71 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VLSize=72B2026.02 | 70 | — | — | — | — | — | — | — | — | — | — | — | — | |
| EvoLMMBackbone=Qwen2.5-VL-7B-Instruct2026.02 | 69.72 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VisionZero-CLEVRBackbone=Qwen2.5-VL-7B-Instruct2026.02 | 69.51 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-30B+VRGABackbone=Qwen3-VL-30B, Enhancement=VRGA2026.03 | 69.4 | — | — | — | — | — | — | — | — | 51 | 61.4 | — | — | |
| Qwen3.5Parameters=9B, Internal Reasoning (Think mode)=false2026.05 | 69.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VisionZero-ChartBackbone=Qwen2.5-VL-7B-Instruct2026.02 | 69.09 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SenseNova-U1Parameters=30BA3B, Internal Reasoning (Think mode)=true2026.05 | 68.95 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Base ModelBackbone=Qwen2.5-VL-7B-Instruct2026.02 | 68.87 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7BParam Size=7B2026.02 | 68.66 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VisionZero-RealWorldBackbone=Qwen2.5-VL-7B-Instruct2026.02 | 68.56 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VisPlayBackbone=Qwen2.5-VL-7B-Instruct2026.02 | 68.56 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-30BBackbone=Qwen3-VL-30B2026.03 | 68.55 | — | — | — | — | — | — | — | — | 50.7 | 61.1 | — | — | |
| MIRROR(w/o tool)Param Size=7B2026.02 | 68.24 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Active-ZeroBackbone=Qwen2.5-VL-7B-Instruct2026.02 | 68.03 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3.5Parameters=35BA3B, Internal Reasoning (Think mode)=false2026.05 | 67.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SenseNova-U1Parameters=8B, Internal Reasoning (Think mode)=true2026.05 | 67.75 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LaserReasoning Paradigm=Latent Reasoning2026.01 | 67.72 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LaserModel Type=Visual-Latent Model2026.04 | 67.72 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Majority voting@4Size=7B2026.02 | 66.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SequentialSize=7B2026.02 | 66.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VLSize=7B2026.02 | 66 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3VLParameters=30BA3B, Internal Reasoning (Think mode)=true2026.05 | 66 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3VLParameters=8B, Internal Reasoning (Think mode)=true2026.05 | 65.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| TroL-7BParameters=7B2024.06 | 65.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LVRReasoning Paradigm=Latent Reasoning2026.01 | 65.19 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LVRModel Type=Visual-Latent Model2026.04 | 65.19 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VisPlayBackbone=Qwen2.5-VL-3B-Instruct2026.02 | 64.88 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Active-ZeroBackbone=Qwen2.5-VL-3B-Instruct2026.02 | 64.14 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ThinkLite-VLData Size=11k2026.04 | 63.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Vision-R1Reasoning Paradigm=Tool-use & RL Enhanced Reasoning2026.01 | 63.83 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Gemini2.5-ProSize=-2026.02 | 63.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| OpenVLThinkerData Size=59.2k2026.04 | 63.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| HyLaR-7BModel Type=Visual-Latent Model, Number of Parameters=7B2026.04 | 63.68 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT5-miniSize=-2026.02 | 63.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-3BParam Size=3B2026.02 | 63.09 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DeepEyesReasoning Paradigm=Tool-use & RL Enhanced Reasoning2026.01 | 62.57 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DeepEyesModel Type=Thinking-with-Images Agent Model2026.04 | 62.57 | — | — | — | — | — | — | — | — | — | — | — | — | |
| V-STARData Size=40k2026.04 | 62.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Visual Para-ThinkerSize=3B2026.02 | 62.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4oSize=-2026.02 | 61.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VLArchitecture Type=Modular, Parameter Scale=8B, Instruction Tuning=Instruct2026.05 | 61.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| R1-OnevisionData Size=155k2026.04 | 60.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5VL2026.04 | 60.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| HyLaR-SFTModel Type=Visual-Latent Model, Training Protocol=SFT2026.04 | 60.23 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Base ModelBackbone=Qwen2.5-VL-3B-Instruct2026.02 | 60.04 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Vision-R1Data Size=210k2026.04 | 59.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| NEO-ovArchitecture Type=Native, Parameter Scale=8B, Instruction Tuning=Instruct2026.05 | 59.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Claude-4-SonnetSize=-2026.02 | 59.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DenseMLLM-4BModel Scale=4B2026.02 | 59.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SequentialSize=3B2026.02 | 58.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Majority voting@4Size=3B2026.02 | 58.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-4BModel Scale=4B2026.02 | 57.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PAPOReasoning Paradigm=Tool-use & RL Enhanced Reasoning2026.01 | 57.52 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VAPOModel Scale=7B2025.09 | 57.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2-VL-7B+VRGABackbone=Qwen2-VL-7B, Enhancement=VRGA2026.03 | 57.1 | — | — | — | — | — | — | — | — | 43.8 | 57.7 | — | — | |
| Qwen2.5-VL-7BReasoning Paradigm=Zero-Shot VLMs2026.01 | 56.57 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7BModel Type=Open-Source Model, Number of Parameters=7B2026.04 | 56.57 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MonetReasoning Paradigm=Latent Reasoning2026.01 | 56.36 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MonetModel Type=Visual-Latent Model2026.04 | 56.36 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VL-CogitoData Size=80k2026.04 | 56.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL3.5-8BReasoning Paradigm=Zero-Shot VLMs2026.01 | 56.15 | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL3.5-8BModel Type=Open-Source Model, Number of Parameters=8B2026.04 | 56.15 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VLSize=3B2026.02 | 56.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VL-RethinkerData Size=39k2026.04 | 56.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7B+VRGABackbone=Qwen2.5-VL-7B, Enhancement=VRGA2026.03 | 55.9 | — | — | — | — | — | — | — | — | 44.4 | 54.9 | — | — | |
| PAPOModel Scale=7B, Decoding Strategy=greedy decoding2025.09 | 55.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ThymeModel Type=Thinking-with-Images Agent Model2026.04 | 55.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-3B+VRGABackbone=Qwen2.5-VL-3B, Enhancement=VRGA2026.03 | 55.4 | — | — | — | — | — | — | — | — | 45.9 | 43.3 | — | — | |
| VL-RethinkerModel Scale=7B2025.09 | 55.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2-VL-7BBackbone=Qwen2-VL-7B2026.03 | 55.1 | — | — | — | — | — | — | — | — | 42.5 | 58.3 | — | — | |
| GPT-4oModel Modality Type=Proprietary2024.10 | 55 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MM-EurekaModel Scale=7B2025.09 | 54.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| NEO-ovArchitecture Type=Native, Parameter Scale=2B, Instruction Tuning=Instruct2026.05 | 54.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL3.5Architecture Type=Modular, Parameter Scale=8B, Instruction Tuning=Instruct2026.05 | 54.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7BBackbone=Qwen2.5-VL-7B2026.03 | 54.2 | — | — | — | — | — | — | — | — | 42.7 | 54.7 | — | — | |
| SAILArchitecture Type=Native, Parameter Scale=8B, Instruction Tuning=Instruct2026.05 | 54.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-3BBackbone=Qwen2.5-VL-3B2026.03 | 53.4 | — | — | — | — | — | — | — | — | 44.2 | 44.5 | — | — | |
| Qwen2.5-VL-3B+CCOTBackbone=Qwen2.5-VL-3B, Enhancement=CCOT2026.03 | 53.21 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VLArchitecture Type=Modular, Parameter Scale=8B, Instruction Tuning=Instruct2026.05 | 52.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| F3ARatio=60%2026.05 | 52.83 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FastVRatio=60%2026.05 | 52.52 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-2BRatio=100%2026.05 | 52.41 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VisionZipRatio=60%2026.05 | 52.41 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SAIL-VLModel Size=8B2025.01 | 52.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SPHINX-Plus-13BParameters=13B2024.06 | 52.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FastVRatio=40%2026.05 | 52.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DivPruneRatio=60%2026.05 | 51.78 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VisionZipRatio=40%2026.05 | 51.68 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DivPruneRatio=40%2026.05 | 51.57 | — | — | — | — | — | — | — | — | — | — | — | — | |
| F3ARatio=40%2026.05 | 51.47 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VLArchitecture Type=Modular, Parameter Scale=2B, Instruction Tuning=Instruct2026.05 | 51.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| CDPrunerRatio=60%2026.05 | 51.36 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2-VL-7B+CCOTBackbone=Qwen2-VL-7B, Enhancement=CCOT2026.03 | 51.21 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-OneVisionReasoning Paradigm=Zero-Shot VLMs2026.01 | 51.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-OneVisionModel Type=Open-Source Model2026.04 | 51.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2 VLModel Scale=7B, Model Modality Type=Vision-language2024.10 | 50.6 | — | — | — | — | — | — | — | — | — | — | — | — |