Alignment on MIA-Bench
93.3AccuracyEVE (Ours-8B-iter4)
Evaluation Results
| Method | Links | |
|---|---|---|
| EVE (Ours-8B-iter4)Model Category=Pseudo-label-based Self-evolution Methods, Scale=8B, Iteration=42026.04 | 93.3 | |
| VisPlay-8B-iter3Model Category=Pseudo-label-based Self-evolution Methods, Scale=8B, Iteration=32026.04 | 93 | |
| MM-Zero-8B-iter3Model Category=Pseudo-label-based Self-evolution Methods, Scale=8B, Iteration=32026.04 | 92.9 | |
| GPT-5 mini (minimal)Model Category=Closed-Source MLLMs, Scale=minimal2026.04 | 92.3 | |
| Qwen3-VL-8B-InstructModel Category=Open-Source MLLMs, Scale=8B2026.04 | 92 | |
| Jigsaw-R1-8BModel Category=Template-based Self-evolution Methods, Scale=8B2026.04 | 91.3 | |
| GPT-5 nano (high)Model Category=Closed-Source MLLMs, Scale=high2026.04 | 89.9 |