Contextual Image Generation on VIST 28 (test)
0.641CLIP SimilarityGILL
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GILLinput_context=5 captions, 4 images2023.05 | 0.641 | 0.693 | |
| GILLinput_context=5 captions2023.05 | 0.612 | 0.696 | |
| Stable Diffusioninput_context=5 captions2023.05 | 0.598 | 0.704 | |
| Stable Diffusioninput_context=1 caption2023.05 | 0.592 | 0.703 | |
| GLIDEinput_context=5 captions2023.05 | 0.591 | 0.745 | |
| GLIDEinput_context=1 caption2023.05 | 0.582 | 0.753 | |
| GILLinput_context=1 caption2023.05 | 0.581 | 0.702 |