Image-to-Text Generation on Image-Text-Audio
91.36CIDErUnifiedIO2-L
Evaluation Results
| Method | Links | |
|---|---|---|
| UnifiedIO2-LCategory=Generalists2026.06 | 91.36 | |
| MoPoECategory=Multimodal VAEs, Backbone=same as MUNI, Compute=same as MUNI2026.06 | 59.74 | |
| MMVAECategory=Multimodal VAEs, Backbone=same as MUNI, Compute=same as MUNI2026.06 | 59.34 | |
| MUNI†Category=Generalists2026.06 | 58.31 | |
| MUNICategory=Generalists2026.06 | 57.72 | |
| LLaVA-NeXTCategory=Specialists2026.06 | 44.6 | |
| FlowBindCategory=Generalists2026.06 | 42.57 | |
| OmniFlowCategory=Generalists2026.06 | 26.6 | |
| CoDiCategory=Generalists2026.06 | 11.97 |