Event-Enriched Image Captioning on EVENTA Grand Challenge CondaBench (test)
0.748CLIPScoreBeyond Vision
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Beyond Visioncontext_strategy=semantic chunk selection, retrieval_rank=642025.12 | 0.748 | 0.5182 | 0.994 | 0.99 | 1 | 0.195 | — | — | |
| Gemmainclude_article_context=true2025.12 | 0.6634 | — | — | — | — | 0.0184 | 0.0341 | 0.1453 | |
| Gemmainclude_article_context=false2025.12 | 0.5945 | — | — | — | — | 0.0111 | 0.0243 | 0.1322 | |
| Qweninclude_article_context=true2025.12 | 0.5855 | — | — | — | — | 0.0565 | 0.0419 | 0.1383 | |
| SmolVLMinclude_article_context=true2025.12 | 0.5552 | — | — | — | — | 0.017 | 0.0229 | 0.0738 | |
| Qweninclude_article_context=false2025.12 | 0.5283 | — | — | — | — | 0.0282 | 0.0256 | 0.132 | |
| SmolVLMinclude_article_context=false2025.12 | 0.4609 | — | — | — | — | 0.0044 | 0.0155 | 0.0789 |