Multimodal Visual Understanding on 839 Diverse Multimodal Evaluation Scenarios
0.3159Wins (Base)LLaMA-3 MM (stage-wise preference optimization)
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| LLaMA-3 MM (stage-wise preference optimization)Judge Model=LLaMA-4-Marvrick2026.05 | 0.3159 | 0.3778 | 0.2622 | 0.0226 | 0.0215 | 6.2 | |
| LLaMA-3 MM (stage-wise preference optimization)Judge Model=Gemini-2.5-Flash2026.05 | 0.3051 | 0.3766 | 0.2634 | 0.0238 | 0.031 | 7.2 | |
| LLaMA-3 MM (stage-wise preference optimization)Judge Model=GPT-4o2026.05 | 0.2467 | 0.329 | 0.3504 | 0.0584 | 0.0155 | 8.2 |