Event-enriched Image Captioning on EVENTA CondaBench 1.0 (test)
0.748CLIPScoreBeyond Vision
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Beyond VisionContextual Information=semantic chunk selection (rank 64)2025.12 | 0.748 | 0.5182 | 0.994 | 0.99 | 1 | 0.195 | |
| GemmaContextual Information=Article2025.12 | 0.6634 | — | — | — | — | 0.0184 | |
| GemmaContextual Information=None2025.12 | 0.5945 | — | — | — | — | 0.0111 | |
| QwenContextual Information=Article2025.12 | 0.5855 | — | — | — | — | 0.0565 | |
| SmolVLMContextual Information=Article2025.12 | 0.5552 | — | — | — | — | 0.017 | |
| QwenContextual Information=None2025.12 | 0.5283 | — | — | — | — | 0.0282 | |
| SmolVLMContextual Information=None2025.12 | 0.4609 | — | — | — | — | 0.0044 |