Video Summarization on SumMe (Rank, Kendall's τ, Spearman's ρ)
0.253Kendall's τLLMVS
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| LLMVSModality=Visual + Text2025.04 | 0.253 | — | 0.282 | |
| CSTAModel category=spatiotemporal, Feature extraction backbone=CNN2024.05 | 0.246 | 1 | 0.274 | |
| CSTAModality=Visual2025.04 | 0.246 | — | 0.274 | |
| RR-STGModel category=spatiotemporal, Feature extraction backbone=CNN2024.05 | 0.211 | 2.5 | 0.234 | |
| HumanFeature extraction backbone=CNN2024.05 | 0.205 | — | 0.213 | |
| Human2025.04 | 0.205 | — | 0.213 | |
| MSVAModel category=multi-modal, Feature extraction backbone=CNN2024.05 | 0.2 | 3.5 | 0.23 | |
| MSVAModality=Visual2025.04 | 0.2 | — | 0.23 | |
| SSPVSModel category=multi-modal, Feature extraction backbone=CNN2024.05 | 0.192 | 3 | 0.257 | |
| SSPVSModality=Visual + Text2025.04 | 0.192 | — | 0.257 | |
| GoogleNetModel category=spatiotemporal, Feature extraction backbone=CNN2024.05 | 0.176 | 5 | 0.197 | |
| LLMModality=Visual + Text2025.04 | 0.17 | — | 0.189 | |
| VASNetModel category=temporal, Feature extraction backbone=CNN2024.05 | 0.16 | 6 | 0.17 | |
| VASNetModality=Visual2025.04 | 0.16 | — | 0.17 | |
| Argaw et al.Modality=Visual + Text2025.04 | 0.13 | — | 0.152 | |
| A2SummModel category=multi-modal, Feature extraction backbone=CNN2024.05 | 0.108 | 7 | 0.129 | |
| A2SummModality=Visual + Text2025.04 | 0.108 | — | 0.129 | |
| VJMHTModel category=temporal, Feature extraction backbone=CNN2024.05 | 0.106 | 8.5 | 0.108 | |
| iPTNetModel category=external dataset-based, Feature extraction backbone=CNN2024.05 | 0.101 | 8.5 | 0.119 | |
| iPTNetModality=Visual2025.04 | 0.101 | — | 0.119 | |
| HMTModel category=multi-modal, Feature extraction backbone=CNN2024.05 | 0.079 | 10.5 | 0.08 | |
| HSA-RNNFeature extraction backbone=CNN2024.05 | 0.064 | 11.5 | 0.066 | |
| DACModel category=temporal, Feature extraction backbone=CNN2024.05 | 0.063 | 12.5 | 0.059 | |
| DMASumModel category=spatiotemporal, Feature extraction backbone=CNN2024.05 | 0.063 | 11 | 0.089 | |
| DMASumModality=Visual2025.04 | 0.063 | — | 0.089 | |
| DSNet-ABModel category=temporal, Feature extraction backbone=CNN2024.05 | 0.051 | 13.5 | 0.059 | |
| DSNet-ABModality=Visual2025.04 | 0.051 | — | 0.059 | |
| dppLSTMFeature extraction backbone=CNN2024.05 | 0.04 | 15 | 0.049 | |
| DSNet-AFModel category=temporal, Feature extraction backbone=CNN2024.05 | 0.037 | 16 | 0.046 | |
| DSNet-AFModality=Visual2025.04 | 0.037 | — | 0.046 | |
| RandomFeature extraction backbone=CNN2024.05 | 0 | — | 0 | |
| Random2025.04 | 0 | — | 0 |