Image-to-Text Generation on One-to-one evaluation benchmarks Image-to-Text (CLIP)
32.14CLIP ScoreLLaVA-NeXT
Evaluation Results
| Method | Links | |
|---|---|---|
| LLaVA-NeXTCategory=Specialists2025.12 | 32.14 | |
| UnifiedIO2-LCategory=Generalists2025.12 | 30.73 | |
| FlowBindCategory=Generalists2025.12 | 29.74 | |
| OmniFlowCategory=Generalists2025.12 | 27.71 | |
| CoDiCategory=Generalists2025.12 | 26.24 |