Dense Video Object Captioning on VidSTG
55.4CapA ScoreCaptionFormer
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| CaptionFormerPretraining set=COCO + LVIScap + LV-VIScap, temporal aggregation=true2025.10 | 55.4 | 66.8 | 71 | 64 | |
| CaptionFormerPretraining set=COCO + LVIScap + LV-VIScap2025.10 | 51 | 66.8 | 71 | 62.3 | |
| CaptionFormerPretraining set=COCO + VG + SMIT + AugCOCO2025.10 | 50.1 | 65 | 69.2 | 60.9 | |
| CaptionFormerPretraining set=COCO2025.10 | 44.3 | 65.1 | 70.2 | 58.7 | |
| OW-VISCaptorPretraining set=COCO2025.10 | 43.9 | 60.1 | 54 | 53 | |
| DVOC-DSPretraining set=COCO + VG + SMIT + AugCOCO2025.10 | 39.7 | 65.8 | 70.4 | 56.9 | |
| OVFormerPretraining set=LVIS + LV-VIS2025.10 | 12.8 | 64.3 | 50.1 | 34.6 | |
| Gemini-1.5-flashPretraining set=-2025.10 | 6.4 | 0.9 | 47.2 | 6.6 |