Video Summarization on TVSum (Rank, Kendall's τ, Spearman's ρ)
0.203Kendall's τDMASum
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| DMASumModel category=spatiotemporal, Feature extraction backbone=CNN2024.05 | 0.203 | 1 | 0.267 | |
| CSTAModel category=spatiotemporal, Feature extraction backbone=CNN2024.05 | 0.194 | 2 | 0.255 | |
| VSS-NetModel category=spatiotemporal, Feature extraction backbone=CNN2024.05 | 0.19 | 3 | 0.249 | |
| MSVAModel category=multi-modal, Feature extraction backbone=CNN2024.05 | 0.19 | 5.5 | 0.21 | |
| SSPVSModel category=multi-modal, Feature extraction backbone=CNN2024.05 | 0.181 | 4.5 | 0.238 | |
| MAAMModel category=temporal, Feature extraction backbone=CNN2024.05 | 0.179 | 5.5 | 0.236 | |
| HumanFeature extraction backbone=CNN2024.05 | 0.177 | — | 0.204 | |
| AAAMModel category=temporal, Feature extraction backbone=CNN2024.05 | 0.169 | 6.5 | 0.223 | |
| RR-STGModel category=spatiotemporal, Feature extraction backbone=CNN2024.05 | 0.162 | 7.5 | 0.212 | |
| VASNetModel category=temporal, Feature extraction backbone=CNN2024.05 | 0.16 | 9 | 0.17 | |
| A2SummModel category=multi-modal, Feature extraction backbone=CNN2024.05 | 0.137 | 10 | 0.165 | |
| iPTNetModel category=external dataset-based, Feature extraction backbone=CNN2024.05 | 0.134 | 11 | 0.163 | |
| GoogleNetModel category=spatiotemporal, Feature extraction backbone=CNN2024.05 | 0.129 | 11.5 | 0.163 | |
| DSNet-AFModel category=temporal, Feature extraction backbone=CNN2024.05 | 0.113 | 13.5 | 0.138 | |
| DSNet-ABModel category=temporal, Feature extraction backbone=CNN2024.05 | 0.108 | 15 | 0.129 | |
| CLIP-ItModel category=multi-modal, Feature extraction backbone=CNN2024.05 | 0.108 | 13.5 | 0.147 | |
| STVTModel category=spatiotemporal, Feature extraction backbone=CNN2024.05 | 0.1 | 15.5 | 0.131 | |
| VJMHTModel category=temporal, Feature extraction backbone=CNN2024.05 | 0.097 | 17.5 | 0.105 | |
| HMTModel category=multi-modal, Feature extraction backbone=CNN2024.05 | 0.096 | 17.5 | 0.107 | |
| HSA-RNNFeature extraction backbone=CNN2024.05 | 0.082 | 19.5 | 0.088 | |
| DANModel category=spatiotemporal, Feature extraction backbone=CNN2024.05 | 0.071 | 19.5 | 0.099 | |
| DACModel category=temporal, Feature extraction backbone=CNN2024.05 | 0.058 | 21 | 0.065 | |
| dppLSTMFeature extraction backbone=CNN2024.05 | 0.042 | 22 | 0.055 | |
| RandomFeature extraction backbone=CNN2024.05 | 0 | — | 0 |