Dense Video Captioning on ActivityNet Captions
10.03METEOROurs
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| OursFeatures=TSP, Training protocol=with reinforcement training2023.03 | 10.03 | 7.11 | 33.33 | — | 1.11 | |
| SGRFeatures=TSN, Training protocol=with reinforcement training2023.03 | 9.37 | 5.29 | — | — | — | |
| OursFeatures=C3D, Training protocol=with reinforcement training2023.03 | 9.13 | 5.91 | 26 | — | 0.81 | |
| SGRFeatures=C3D, Training protocol=with reinforcement training2023.03 | 9.07 | — | 22.12 | — | 1.67 | |
| SDVC2020.11 | 8.82 | — | — | 2.94 | 0.93 | |
| SDVCFeatures=C3D, Training protocol=with reinforcement training2023.03 | 8.82 | — | 30.68 | — | 0.93 | |
| TSPAlgorithm=BMT2020.11 | 8.75 | — | — | 4.16 | 2.02 | |
| OursFeatures=TSP, Training protocol=with cross-entropy training2023.03 | 8.5 | 6.22 | 32.76 | — | 2.18 | |
| BMT2020.11 | 8.44 | — | — | 3.84 | 1.88 | |
| PDVCFeatures=TSP, Training protocol=with cross-entropy training2023.03 | 8.37 | 6.05 | 31.14 | — | 2.17 | |
| HiVid-NarratorEvaluation Protocol=supervised, SPA-Compressor=true2026.01 | 7.8 | 6.9 | 26.6 | — | — | |
| PDVCFeatures=C3D, Training protocol=with cross-entropy training2023.03 | 7.5 | 5.26 | 25.87 | — | 1.65 | |
| OursFeatures=C3D, Training protocol=with cross-entropy training2023.03 | 7.46 | 5.44 | 26.23 | — | 1.69 | |
| BMTFeatures=I3D, Training protocol=with cross-entropy training2023.03 | 7.43 | — | 11.94 | — | 1.88 | |
| UEDVCFeatures=C3D, Training protocol=with cross-entropy training2023.03 | 7.33 | 5.29 | 26.92 | — | 1.45 | |
| MDVC2020.11 | 7.31 | — | — | 2.6 | 1.07 | |
| ECHRFeatures=C3D, Training protocol=with cross-entropy training2023.03 | 7.19 | 3.22 | 14.71 | — | 1.29 | |
| MFT2020.11 | 7.08 | — | — | 2.82 | 1.24 | |
| MFTFeatures=TSN, Training protocol=with reinforcement training2023.03 | 7.08 | — | 21 | — | 1.24 | |
| TimeRefineparameters=7B2026.01 | 7 | 6.1 | 28.6 | — | — | |
| TA-Promptingparameters=7B2026.01 | 7 | 6.1 | 29.2 | — | — | |
| EvoGroundFT=false2026.05 | 7 | 4.1 | — | — | — | |
| DVC2020.11 | 6.93 | — | — | 2.27 | 0.73 | |
| DVCFeatures=C3D, Training protocol=with cross-entropy training2023.03 | 6.93 | — | 12.61 | — | 0.73 | |
| SDVCFeatures=C3D, Training protocol=with cross-entropy training2023.03 | 6.92 | — | — | — | — | |
| VTimeLLMEvaluation Protocol=supervised2026.01 | 6.8 | 5.8 | 27.6 | — | — | |
| VTimeLLMparameters=7B2026.01 | 6.8 | 5.8 | 27.6 | — | — | |
| VTimeLLMFT=true2026.05 | 6.8 | 5.8 | — | — | — | |
| HiVid-NarratorEvaluation Protocol=supervised, SPA-Compressor=false2026.01 | 6.7 | 6 | 24.4 | — | — | |
| TimeChatparameters=7B2026.01 | 6.7 | 4.7 | 19 | — | — | |
| TRACEEvaluation Protocol=supervised2026.01 | 6.4 | 6 | 25.9 | — | — | |
| Grounded-VideoLLMFT=true2026.05 | 6.4 | 6.2 | — | — | — | |
| EfficientFeatures=C3D, Training protocol=with cross-entropy training2023.03 | 6.21 | — | 13.82 | — | 1.35 | |
| Bi-SST2020.11 | 6.1 | — | — | 2.27 | 1.13 | |
| TRACEFT=true2026.05 | 6 | 6.4 | — | — | — | |
| VTG-LLMEvaluation Protocol=supervised2026.01 | 5.9 | 5.1 | 20.7 | — | — | |
| VTG-LLMparameters=7B2026.01 | 5.9 | 5.1 | 20.7 | — | — | |
| TDA-CGFeatures=C3D, Training protocol=with cross-entropy training2023.03 | 5.86 | — | 7.99 | — | 1.31 | |
| TimeChatEvaluation Protocol=supervised2026.01 | 5.7 | 4.7 | 19 | — | — | |
| DCEFeatures=C3D, Training protocol=with cross-entropy training2023.03 | 5.69 | — | 12.43 | — | 0.17 | |
| LITAparameters=13B2026.01 | 5.2 | 4.7 | 17.2 | — | — | |
| MTFeatures=TSN, Training protocol=with cross-entropy training2023.03 | 4.98 | — | 9.25 | — | 1.15 | |
| MonmentorEvaluation Protocol=supervised2026.01 | 4.7 | 2.3 | 14.9 | — | — | |
| Momentorparameters=7B2026.01 | 4.7 | 0.3 | 14.9 | — | — | |
| MomentorFT=false2026.05 | 4.7 | 2.3 | — | — | — | |
| VideoChatGPTparameters=7B2026.01 | 2.1 | 1.9 | 5.8 | — | — | |
| VideoChatparameters=7B2026.01 | 0.9 | 0.9 | 2.2 | — | — | |
| ValleyEvaluation Protocol=supervised2026.01 | 0.8 | 0.3 | 1.8 | — | — |