Model Efficiency Analysis on General 16 frames, 512 text tokens (inference)
20.74FPSMuLTI-S
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| MuLTI-SVideo Encoder=ViT-B/32, Text Encoder=6-layer BERT, Number of frames=16, Text length=512, Hardware=NVIDIA V100 16GB GPU2023.03 | 20.74 | — | — | |
| MuLTI-BVideo Encoder=ViT-B/16, Text Encoder=BERT-base (12-layer), Number of frames=16, Text length=512, Hardware=NVIDIA V100 16GB GPU2023.03 | 10.13 | — | — | |
| ALPRONumber of frames=16, Text length=512, Hardware=NVIDIA V100 16GB GPU2023.03 | 9.97 | — | — | |
| VIOLETNumber of frames=16, Text length=512, Hardware=NVIDIA V100 16GB GPU2023.03 | 9.05 | — | — | |
| MuLTI-LVideo Encoder=ViT-L/14, Text Encoder=BERT-large, Number of frames=16, Text length=512, Hardware=NVIDIA V100 16GB GPU2023.03 | 3.12 | — | — | |
| FrozenBiLMNumber of frames=16, Text length=512, Hardware=NVIDIA V100 16GB GPU2023.03 | 2.54 | — | — |