Multimodal Understanding on MMBench (ECE, AUROC, AUPRC)
4.44ECEQwen2.5-VL-7B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Qwen2.5-VL-7BConf.=Token Confidence, Source=Open-source MLLMs2026.04 | 4.44 | 90.4 | 98.53 | 87.93 | |
| InternVL2.5-4BConf.=Token Confidence, Source=Open-source MLLMs2026.04 | 5.12 | 89.22 | 97.98 | 86.24 | |
| MiniCPM-V-2.6Conf.=Token Confidence, Source=Open-source MLLMs2026.04 | 5.14 | 89.97 | 98.4 | 88.01 | |
| InternVL3.5-8BConf.=Token Confidence, Source=Open-source MLLMs2026.04 | 5.63 | 88.86 | 98.66 | 86.92 | |
| GPT-4oConf.=Verbal Confidence, Source=Closed-source MLLMs2026.04 | 6.04 | 69.67 | 92.68 | 35.63 | |
| MiniCPM-V-2.6Conf.=Verbal Confidence, Source=Open-source MLLMs2026.04 | 7.45 | 74.71 | 93.62 | 52.27 | |
| GPT-4oConf.=Token Confidence, Source=Closed-source MLLMs2026.04 | 7.61 | 91.33 | 98.81 | 89.51 | |
| GPT-4o-miniConf.=Verbal Confidence, Source=Closed-source MLLMs2026.04 | 8.08 | 77.61 | 92.66 | 56.78 | |
| Phi-3.5-VisionConf.=Token Confidence, Source=Open-source MLLMs2026.04 | 9.16 | 86.22 | 96.5 | 82.2 | |
| InternVL3.5-8BConf.=Verbal Confidence, Source=Open-source MLLMs2026.04 | 9.35 | 72.93 | 95.18 | 53.12 | |
| InternVL2.5-4BConf.=Verbal Confidence, Source=Open-source MLLMs2026.04 | 10.99 | 56.67 | 86.5 | 7.94 | |
| Phi-3.5-VisionConf.=Verbal Confidence, Source=Open-source MLLMs2026.04 | 12.21 | 63.54 | 86.35 | 30.48 | |
| GPT-4o-miniConf.=Token Confidence, Source=Closed-source MLLMs2026.04 | 13.39 | 88.71 | 97.56 | 85.64 | |
| Qwen2.5-VL-7BConf.=Verbal Confidence, Source=Open-source MLLMs2026.04 | 37.64 | 45.92 | 87.54 | -2.19 |