Vision-Language Captioning on xBD (test)
77.86CLIPScore (%)QwenVL
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| QwenVLDecoder context=Transformer, KG=No2025.09 | 77.86 | 81.24 | |
| VLCEDecoder=Transformer, Baseline=QwenVL, KG=Yes2025.09 | 60.6 | 69.86 | |
| VLCEDecoder=CNN-LSTM, Baseline=LLaVA, KG=No2025.09 | 55.34 | 66.41 | |
| VLCEDecoder=CNN-LSTM, Baseline=LLaVA, KG=Yes2025.09 | 51.1 | 66.56 | |
| LLaVADecoder context=CNN-LSTM, KG=Yes2025.09 | 48.9 | 33.44 | |
| LLaVADecoder context=CNN-LSTM, KG=No2025.09 | 44.66 | 33.59 | |
| QwenVLDecoder context=Transformer, KG=Yes2025.09 | 39.4 | 30.14 | |
| VLCEDecoder=Transformer, Baseline=QwenVL, KG=No2025.09 | 22.14 | 18.76 |