Video Generation on UCF-101
58FVDMAGVITv2
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| MAGVITv2Type=Mask., Params=307M2024.12 | 58 | — | — | — | — | |
| ELT (6N x 6L)Class=class-conditional, # params=76M, # steps=24, GFlops=∼13k, Pretrained on additional large video data=true2026.04 | 60.8 | — | — | — | 87.88 | |
| ELT (6N x 4L)Class=class-conditional, # params=76M, # steps=12, GFlops=∼4.3k, Pretrained on additional large video data=true2026.04 | 72.8 | — | — | — | 88.27 | |
| MAGVITType=Mask., Params=306M2024.12 | 76 | — | — | — | — | |
| MAGVIT-LClass=class-conditional, # params=306M, # steps=12, GFlops=∼4.3k, Pretrained on additional large video data=true2026.04 | 76 | — | — | — | 89.27 | |
| Make-A-VideoClass=class-conditional, # params=≫3.5B, # steps=≫250, Pretrained on additional large video data=true2026.04 | 81 | — | — | — | 82.55 | |
| ACDIT-H-LTType=AR+Diff., Params=954M, Training Strategy=Longer Training2024.12 | 90 | — | — | — | — | |
| MaGNeTSClass=class-conditional, # params=306M, # steps=12, GFlops=∼1.7k, Pretrained on additional large video data=true2026.04 | 96.4 | — | — | — | 88.53 | |
| PAR-4xClass=class-conditional, # params=792M, # steps=323, Pretrained on additional large video data=true2026.04 | 99.5 | — | — | — | — | |
| PAR-16xClass=class-conditional, # params=792M, # steps=95, Pretrained on additional large video data=true2026.04 | 103.4 | — | — | — | — | |
| ACDIT-HType=AR+Diff., Params=954M2024.12 | 104 | — | — | — | — | |
| MAGVITv2-ARType=AR, Params=307M2024.12 | 109 | — | — | — | — | |
| ACDIT-XLType=AR+Diff., Params=677M2024.12 | 111 | — | — | — | — | |
| SVDResolution=1024 × 576, NFE=50, TFLOPs=45.43, Latency (GPU)=376, Latency (Phone)=OOM2024.12 | 149 | — | — | — | — | |
| MobileVDResolution=512 × 256, NFE=1, TFLOPs=4.34, Latency (GPU)=45, Latency (Phone)=17802024.12 | 171 | — | — | — | — | |
| VideoFusionType=Diff., Params=510M2024.12 | 173 | — | — | — | — | |
| SF-VResolution=1024 × 576, NFE=1, TFLOPs=45.43, Latency (GPU)=376, Latency (Phone)=OOM2024.12 | 181 | — | — | — | — | |
| MobileVD-HDResolution=1024 × 576, NFE=1, TFLOPs=23.63, Latency (GPU)=227, Latency (Phone)=OOM2024.12 | 184 | — | — | — | — | |
| OmniTokenizerType=AR, Params=650M2024.12 | 191 | — | — | — | — | |
| MattenType=Diff., Params=853M2024.12 | 211 | — | — | — | — | |
| MAGVIT-ARType=AR, Params=306M2024.12 | 265 | — | — | — | — | |
| Video-LaVITType=Diff., Params=7B2024.12 | 281 | — | — | — | — | |
| AnimateLCMResolution=1024 × 576, NFE=8, TFLOPs=45.43, Latency (GPU)=376, Latency (Phone)=OOM2024.12 | 281 | — | — | — | — | |
| MMVGType=Mask., Params=230M2024.12 | 328 | — | — | — | — | |
| TATSType=AR, Params=331M2024.12 | 332 | — | — | — | — | |
| TATSClass=class-conditional, # params=321M, # steps=1024, Pretrained on additional large video data=false2026.04 | 332 | — | — | — | 79.28 | |
| MagDiffType=Diff., Params=2B2024.12 | 340 | — | — | — | — | |
| Make-A-VideoClass=class-conditional, Pretrained on additional large video data=false2026.04 | 367 | — | — | — | 33 | |
| LVDMType=Diff., Params=437M2024.12 | 372 | — | — | — | — | |
| CCVS+StyleGANClass=unconditional, Pretrained on additional large video data=false2026.04 | 386 | — | — | — | 24.47 | |
| TATSClass=unconditional, # params=321M, # steps=1024, Pretrained on additional large video data=false2026.04 | 420 | — | — | — | 57.63 | |
| SVDResolution=512 × 256, NFE=50, TFLOPs=8.60, Latency (GPU)=82, Latency (Phone)=OOM2024.12 | 476 | — | — | — | — | |
| LatteType=Diff., Params=674M2024.12 | 478 | — | — | — | — | |
| DIGANClass=unconditional, # steps=1, GFlops=∼148, Pretrained on additional large video data=false2026.04 | 577 | — | — | — | 32.7 | |
| CogVideoType=AR, Params=9.4B2024.12 | 626 | — | — | — | — | |
| CogVideoClass=class-conditional, # params=9.4B, Pretrained on additional large video data=true2026.04 | 626 | — | — | — | 50.46 | |
| LADDResolution=1024 × 576, NFE=1, TFLOPs=45.43, Latency (GPU)=376, Latency (Phone)=OOM2024.12 | 1,894 | — | — | — | — | |
| UFOGenResolution=1024 × 576, NFE=1, TFLOPs=45.43, Latency (GPU)=376, Latency (Phone)=OOM2024.12 | 1,917 | — | — | — | — | |
| AnimateDiff (V3)zero-shot=true2023.11 | — | 46.2 | 24.1 | 89.8 | — | |
| DVD-GANClass=class-conditional, # steps=1, Pretrained on additional large video data=false2026.04 | — | — | — | — | 32.97 | |
| I2VGen-XLzero-shot=true, Input Type=text&image-to-video, Training Videos=10M2023.11 | — | 44.1 | 21.3 | 89.4 | — | |
| MagDiffzero-shot=true, Input Type=text&image-to-video, Training Videos=5.3M + 76K2023.11 | — | 50.8 | 25.4 | 90.2 | — | |
| RaMViDClass=unconditional, # params=308M, # steps=500, Pretrained on additional large video data=false2026.04 | — | — | — | — | 21.71 | |
| StyleGAN-VClass=unconditional, # steps=1, Pretrained on additional large video data=false2026.04 | — | — | — | — | 23.94 | |
| Video DiffusionClass=unconditional, # params=1.1B, # steps=256, Pretrained on additional large video data=false2026.04 | — | — | — | — | 57 |