Hallucination Evaluation on MME cognition-related 57 (test)
671.6AccuracyPTI
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| PTIBase Model=DeepSeek-VL-Chat, Decoding Strategy=greedy decoding2026.04 | 671.6 | 20 | |
| VTIBase Model=DeepSeek-VL-Chat, Decoding Strategy=greedy decoding2026.04 | 661.6 | 10 | |
| PAIBase Model=DeepSeek-VL-Chat, Decoding Strategy=greedy decoding2026.04 | 656.6 | 5 | |
| PTIBase Model=LLAVA-1.5, Decoding Strategy=greedy decoding2026.04 | 651.6 | 40 | |
| VanillaBase Model=DeepSeek-VL-Chat, Decoding Strategy=greedy decoding2026.04 | 651.6 | — | |
| VISTABase Model=DeepSeek-VL-Chat, Decoding Strategy=greedy decoding2026.04 | 646.6 | 5 | |
| PTIBase Model=Qwen-VL-Chat, Decoding Strategy=greedy decoding2026.04 | 638.3 | 40 | |
| VTIBase Model=LLAVA-1.5, Decoding Strategy=greedy decoding2026.04 | 633.3 | 21.7 | |
| VTIBase Model=Qwen-VL-Chat, Decoding Strategy=greedy decoding2026.04 | 626.6 | 28.3 | |
| PAIBase Model=LLAVA-1.5, Decoding Strategy=greedy decoding2026.04 | 625 | 13.4 | |
| VISTABase Model=LLAVA-1.5, Decoding Strategy=greedy decoding2026.04 | 615 | 3.4 | |
| VanillaBase Model=LLAVA-1.5, Decoding Strategy=greedy decoding2026.04 | 611.6 | — | |
| VISTABase Model=Qwen-VL-Chat, Decoding Strategy=greedy decoding2026.04 | 611.6 | 13.3 | |
| PAIBase Model=Qwen-VL-Chat, Decoding Strategy=greedy decoding2026.04 | 605 | 6.7 | |
| VanillaBase Model=Qwen-VL-Chat, Decoding Strategy=greedy decoding2026.04 | 598.3 | — |