Image Captioning on COCO (CIDEr, RefPAC, CLIP-Score)
96CIDErMERCap
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| MERCap2025.10 | 96 | — | — | |
| EntroCap2025.10 | 94.3 | — | — | |
| T2D + Diffusionattention-based weighting=true2025.10 | 92.8 | 89 | 74.2 | |
| CapDec2025.10 | 91.8 | — | — | |
| Llava1.5 One-VisionAdaptation strategy=Visual Prompting, Number of Parameters=4B2025.10 | 91.6 | 91.1 | 80.1 | |
| Llava1.5 One-VisionAdaptation strategy=Crop, Number of Parameters=4B2025.10 | 91.6 | 91.1 | 80.1 | |
| ViECapreproduced=true2025.10 | 89.7 | 88.5 | 75.6 | |
| T2D + Mem.attention-based weighting=true2025.10 | 88.5 | 90.2 | 76 | |
| CLOSE2025.10 | 81.2 | — | — | |
| Qwen2.5 VLAdaptation strategy=Visual Prompting, Number of Parameters=3B2025.10 | 77.9 | 89.9 | 81.2 | |
| Qwen2.5 VLAdaptation strategy=Crop, Number of Parameters=3B2025.10 | 77.9 | 89.9 | 81.2 | |
| Patch-ionerConfiguration=T2D + Mem., Number of Parameters=0.21B, Adaptation strategy=Patch-based2025.10 | 69.2 | 87.4 | 72.8 | |
| T2D + Noisereproduced=true2025.10 | 65.5 | 86.2 | 70.9 | |
| MAGIC2025.10 | 49.3 | — | — | |
| Qwen3 VLAdaptation strategy=Visual Prompting, Number of Parameters=4B2025.10 | 19.4 | 84.2 | 85.3 | |
| Qwen3 VLAdaptation strategy=Crop, Number of Parameters=4B2025.10 | 19.4 | 84.2 | 85.3 | |
| ZeroCap2025.10 | 14.6 | — | — |