Multimodal Understanding on MMD-Bench Hard degradation 1.0
79.33MMBench ScoreGemini-2.5-Flash
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Gemini-2.5-FlashModel Category=Closed-source models2026.04 | 79.33 | 66.55 | 72.33 | 76.01 | 62 | 69.15 | 72.72 | 71.16 | |
| GPT-4.1-miniModel Category=Closed-source models2026.04 | 76.08 | 51.88 | 71 | 74.96 | 60.73 | 72.41 | 72.52 | 68.51 | |
| CLEAR-RLModel Category=CLEAR variants, Backbone=Bagel, Training Strategy=RL (Interleaved GRPO)2026.04 | 72.52 | 51.97 | 71.33 | 72.25 | 60.67 | 61.05 | 67.07 | 65.26 | |
| CLEAR-SFTModel Category=CLEAR variants, Backbone=Bagel, Training Strategy=SFT2026.04 | 72.06 | 47.56 | 70.33 | 70.51 | 57.67 | 60.13 | 65.65 | 63.42 | |
| BagelModel Category=Open-source unified models2026.04 | 67.88 | 45.09 | 65.66 | 64.81 | 55.53 | 58.43 | 61.64 | 60.15 | |
| GPT-4o-miniModel Category=Closed-source models2026.04 | 67.02 | 50.91 | 64 | 59.87 | 45.93 | 58.95 | 61.21 | 58.27 | |
| Text-only CoTModel Category=CLEAR variants, Backbone=Bagel2026.04 | 63.62 | 48.3 | 70.33 | 64.18 | 56.93 | 53.98 | 62.82 | 60.02 | |
| Janus-ProModel Category=Open-source unified models2026.04 | 55.57 | 31.33 | 52.66 | 66.75 | 41.53 | 43.52 | 49.09 | 48.64 | |
| Emu3Model Category=Open-source unified models2026.04 | 53.71 | 21.51 | 65 | 58.34 | 42.06 | 52.55 | 55.15 | 49.76 |