Text-to-Image Generation on MSCOCO (CLIP-T, CLIP-I)
0.697CLIP-I (Image-Text Alignment)DreamLLM
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DreamLLMModel Type=Visual LLM, #Params=7B2026.03 | 0.697 | 0.238 | |
| NExT-GPT‡Model Type=Any-to-Any, #Params=7B2026.03 | 0.691 | 0.225 | |
| Omni-DiffusionModel Type=Any-to-Any, #Params=7B2026.03 | 0.667 | 0.235 | |
| Emu‡Model Type=Visual LLM, #Params=14B2026.03 | 0.656 | 0.286 | |
| AnyGPTModel Type=Any-to-Any, #Params=8B2026.03 | 0.65 | — |