Video Segment Description on Ego4D HCap v1 (test)
46.88CIDErVideo ReCap
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Video ReCapVideo Encoder=TSF-B, Text Decoder=GPT2, Train Params=339M, Pseudo Ann.=true, Protocol=Finetuned2024.02 | 46.88 | 39.73 | 18.55 | |
| Video ReCap-UVideo Encoder=TSF-B, Text Decoder=GPT2, Train Params=113M, Pseudo Ann.=true, Protocol=Finetuned2024.02 | 45.6 | 39.33 | 18.17 | |
| Video ReCapVideo Encoder=TSF-B, Text Decoder=GPT2, Train Params=339M, Pseudo Ann.=false, Protocol=Finetuned2024.02 | 41.74 | 39.04 | 18.21 | |
| LaViLa + FLANT5Video Encoder=TSF-B, Text Decoder=FT5-XL, Train Params=586M, Pseudo Ann.=false, Protocol=Finetuned2024.02 | 39.13 | 38.77 | 16.88 | |
| LaViLa + GPT2Video Encoder=TSF-B, Text Decoder=GPT2, Train Params=336M, Pseudo Ann.=false, Protocol=Finetuned2024.02 | 38.22 | 38.1 | 16.58 | |
| LaViLaVideo Encoder=TSF-B, Text Decoder=GPT2, Train Params=258M, Pseudo Ann.=false, Protocol=Finetuned2024.02 | 24.63 | 33.31 | 15.3 | |
| LaViLa + GPT3.5Video Encoder=TSF-B, Text Decoder=GPT2, Train Params=0, Pseudo Ann.=false, Protocol=Zero-Shot2024.02 | 5.79 | 19.77 | 13.45 | |
| BLIP2 + GPT3.5Video Encoder=VIT-G, Text Decoder=FT5-XL, Train Params=0, Pseudo Ann.=false, Protocol=Zero-Shot2024.02 | 5.68 | 16.87 | 13.47 |