Text-to-image synthesis evaluation (Human Correlation - Overall) on DrawBench
0.223Kendall's TauLLMScore
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LLMScore2023.05 | 0.223 | 0.3023 | |
| LLMScoreHuman ranking criteria=Overall2023.05 | 0.223 | 0.3023 | |
| LLMScoreHuman ranking criteria=Error Counting2023.05 | 0.2134 | 0.2865 | |
| BLIP-ITC2023.05 | 0.1569 | 0.2171 | |
| BLIP-ITCHuman ranking criteria=Overall2023.05 | 0.1569 | 0.2171 | |
| CLIP2023.05 | 0.153 | 0.2143 | |
| CLIPHuman ranking criteria=Overall2023.05 | 0.153 | 0.2143 | |
| BLIP-ITCHuman ranking criteria=Error Counting2023.05 | 0.1506 | 0.2029 | |
| NegCLIP2023.05 | 0.1463 | 0.1999 | |
| NegCLIPHuman ranking criteria=Overall2023.05 | 0.1463 | 0.1999 | |
| CLIPHuman ranking criteria=Error Counting2023.05 | 0.136 | 0.191 | |
| NegCLIPHuman ranking criteria=Error Counting2023.05 | 0.1179 | 0.1596 | |
| DescCLIPtype=LLMScore variant2023.05 | 0.1136 | 0.1557 | |
| BLIP-ITM2023.05 | 0.1044 | 0.1455 | |
| BLIP-ITMHuman ranking criteria=Overall2023.05 | 0.1044 | 0.1455 | |
| CapMETEORtype=LLMScore variant2023.05 | 0.0951 | 0.1312 | |
| BLIP-ITMHuman ranking criteria=Error Counting2023.05 | 0.0871 | 0.1189 | |
| CapCLIPtype=LLMScore variant2023.05 | 0.0056 | 0.0072 | |
| DescMETEORtype=LLMScore variant2023.05 | 0.0028 | 0.0048 |