Text-to-Video Generation on UCF-101 (zero-shot)
200.2FVDSnap Video
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Snap VideoResolution=512 x 288 px, zero-shot=true2024.02 | 200.2 | 38.89 | 28.1 | |
| W.A.L.TNumber of parameters=3B, Training dataset=video + image2023.12 | 258.1 | 35.1 | — | |
| Snap VideoResolution=288 x 288 px, zero-shot=true2024.02 | 260.1 | 38.89 | 39 | |
| MicroCinematraining_data_source=WebVid-10M only, zero-shot=true2023.11 | 342.86 | 37.46 | — | |
| W.A.L.TNumber of parameters=419M, Training dataset=video + image2023.12 | 344.5 | 31.7 | — | |
| PYoCotraining_data_source=WebVid-10M and additional data, zero-shot=true2023.11 | 355.19 | 47.76 | — | |
| PYoCo2023.12 | 355.2 | 47.8 | — | |
| PYoCozero-shot=true2024.02 | 355.2 | 47.46 | — | |
| PYoCoResolution=16 x 256 x 256, Protocol=Zero-shot2024.06 | 355.2 | 47.46 | — | |
| VidRDtraining_data_source=WebVid-10M and additional data, zero-shot=true2023.11 | 363.19 | 39.37 | — | |
| Make-A-Video2023.12 | 367.2 | 33 | — | |
| Make-A-Videozero-shot=true2024.02 | 367.2 | 33 | — | |
| Make-A-VideoResolution=16 x 256 x 256, Protocol=Zero-shot2024.06 | 367.2 | 33 | — | |
| Make-A-Video#Params=∼9.6B, #Videos (train)=∼20M2025.05 | 367.2 | — | — | |
| Make-A-Videotraining_data_source=WebVid-10M and additional data, zero-shot=true2023.11 | 367.23 | 33 | — | |
| HPDM-T2VResolution=16 x 144 x 256, Protocol=Zero-shot2024.06 | 383.26 | 21.15 | — | |
| Show-1training_data_source=WebVid-10M only, zero-shot=true2023.11 | 394.46 | 35.42 | — | |
| VideoFactorytraining_data_source=WebVid-10M and additional data, zero-shot=true2023.11 | 410 | — | — | |
| ModelScopetraining_data_source=WebVid-10M and additional data, zero-shot=true2023.11 | 410 | — | — | |
| VideoFactoryzero-shot=true2024.02 | 410 | — | — | |
| VideoFactoryResolution=16 x 256 x 256, Protocol=Zero-shot2024.06 | 410 | — | — | |
| CAT-LVDM (SACN)#Params=∼2.3B, #Videos (train)=∼2M2025.05 | 440.3 | — | — | |
| Latte#Params=∼674M, #Videos (train)=∼25M2025.05 | 478 | — | — | |
| HPDM-T2VResolution=16 x 288 x 512, Protocol=Zero-shot2024.06 | 481.93 | 23.77 | — | |
| CMD#Params=∼1.6B, #Videos (train)=∼10.7M2025.05 | 504 | — | — | |
| CAT-LVDM (BCNI)#Params=∼2.3B, #Videos (train)=∼2M2025.05 | 505.5 | — | — | |
| Lavietraining_data_source=WebVid-10M and additional data, zero-shot=true2023.11 | 526.3 | — | — | |
| LaVie#Params=∼3B, #Videos (train)=∼35M2025.05 | 526.3 | — | — | |
| DEMO#Params=∼2.3B, #Videos (train)=∼10M2025.05 | 547.3 | — | — | |
| Video LDM2023.12 | 550.6 | 33.5 | — | |
| Video LDMzero-shot=true2024.02 | 550.6 | 33.45 | — | |
| Video LDMProtocol=Zero-shot2024.06 | 550.6 | 33.45 | — | |
| Video LDMtraining_data_source=WebVid-10M only, zero-shot=true2023.11 | 550.61 | 33.45 | — | |
| VideoGen#Params=∼1.7B, #Videos (train)=∼10M2025.05 | 554 | — | — | |
| W.A.L.TNumber of parameters=419M, Training dataset=video only2023.12 | 598.8 | 26.8 | — | |
| Uniform#Params=∼2.3B, #Videos (train)=∼2M2025.05 | 599.5 | — | — | |
| EMU Video#Params=∼8.6B, #Videos (train)=∼34M2025.05 | 606.2 | — | — | |
| ModelScopeT2V (Finetuned)#Params=∼1.7B, #Videos (train)=∼10M2025.05 | 612.5 | — | — | |
| Gaussian#Params=∼2.3B, #Videos (train)=∼2M2025.05 | 615.3 | — | — | |
| ModelScopeT2V#Params=∼1.7B, #Videos (train)=∼10M2025.05 | 628.2 | — | — | |
| VideoFusiontraining_data_source=WebVid-10M only, zero-shot=true2023.11 | 639.9 | 17.49 | — | |
| LVDMtraining_data_source=WebVid-10M only, zero-shot=true2023.11 | 641.8 | — | — | |
| LVDMzero-shot=true2024.02 | 641.8 | — | — | |
| LVDMResolution=16 x 256 x 256, Protocol=Zero-shot2024.06 | 641.8 | — | — | |
| Magic Videozero-shot=true2024.02 | 655 | — | — | |
| Magic VideoResolution=16 x 256 x 256, Protocol=Zero-shot2024.06 | 655 | — | — | |
| MagicVideo#Params=∼1.2B, #Videos (train)=∼17M2025.05 | 655 | — | — | |
| Video LDM#Params=∼1.3B, #Videos (train)=∼11M2025.05 | 656.5 | — | — | |
| Magic Video2023.12 | 699 | — | — | |
| MagicVideotraining_data_source=WebVid-10M only, zero-shot=true2023.11 | 699 | — | — | |
| CogVideotraining_data_source=WebVid-10M only, zero-shot=true2023.11 | 701.59 | 25.27 | — | |
| CogVideo (English)Language=English2023.12 | 701.6 | 25.3 | — | |
| CogVideoLanguage=English, zero-shot=true2024.02 | 701.6 | 25.27 | — | |
| CogVideoResolution=16 x 480 x 480, Protocol=Zero-shot2024.06 | 701.6 | 25.27 | — | |
| HPDM-T2VResolution=16 x 256 x 256, Protocol=Zero-shot2024.06 | 728.26 | 23.46 | — | |
| CogVideo (Chinese)Language=Chinese2023.12 | 751.3 | 23.6 | — | |
| CogVideoLanguage=Chinese, zero-shot=true2024.02 | 751.3 | 23.55 | — | |
| HPDM-T2VResolution=64 x 288 x 512, Protocol=Zero-shot2024.06 | 1,197.6 | — | — | |
| HPDM-T2VResolution=64 x 256 x 256, Protocol=Zero-shot2024.06 | 1,238.62 | — | — |