Video Generation on UCF-101 (test)
89.27Inception ScoreMAGVIT-L-CG
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| MAGVIT-L-CGExtra Video=false, Class (conditional)=true, Resolution=128x128, Used test set in training=false2022.12 | 89.27 | — | 76 | — | — | — | — | — | |
| MAGVIT-L-CGextra video data=false, class-conditional=true, resolution=128x128, parameters=464M2022.12 | 89.27 | — | 76 | — | — | — | — | — | |
| MAGVIT-B-CGExtra Video=false, Class (conditional)=true, Resolution=128x128, Used test set in training=false2022.12 | 83.55 | — | 159 | — | — | — | — | — | |
| MAGVIT-B-CGextra video data=false, class-conditional=true, resolution=128x128, parameters=128M2022.12 | 83.55 | — | 159 | — | — | — | — | — | |
| Make-A-VideoEvaluation Protocol=Finetuning Setting, Pretrain=Yes, Class=Yes, Resolution=256 x 2562022.09 | 82.55 | — | 81.25 | — | — | — | — | — | |
| Make-A-VideoExtra Video=true, Class (conditional)=true, Resolution=Custom, Used test set in training=false2022.12 | 82.55 | — | 81 | — | — | — | — | — | |
| Make-A-Video*extra video data=true, class-conditional=true, resolution=custom2022.12 | 82.55 | — | 81 | — | — | — | — | — | |
| TATS-baseEvaluation Protocol=Finetuning Setting, Pretrain=No, Class=Yes, Resolution=128 x 1282022.09 | 79.28 | — | 278 | — | — | — | — | — | |
| TATS-baseClass conditional=true, Trained on entire dataset=false2022.04 | 79.28 | — | 332 | — | — | — | — | — | |
| TATSExtra Video=false, Class (conditional)=true, Resolution=128x128, Used test set in training=false2022.12 | 79.28 | — | 332 | — | — | — | — | — | |
| TATSextra video data=false, class-conditional=true, resolution=128x1282022.12 | 79.28 | — | 332 | — | — | — | — | — | |
| VideoFusionResolution=16 x 128 x 1282023.03 | 72.22 | — | 220 | — | — | — | — | — | |
| WF-VAE-LChn=162024.11 | 71.86 | — | — | 947.18 | — | — | — | — | |
| VideoFusionResolution=16 x 64 x 642023.03 | 71.67 | — | 139 | — | — | — | — | — | |
| WF-VAE-LChn=42024.11 | 70.53 | — | — | 929.55 | — | — | — | — | |
| AllegroChn=42024.11 | 67.16 | — | — | 1,045.66 | — | — | — | — | |
| WF-VAE-SChn=42024.11 | 65.89 | — | — | 1,005.1 | — | — | — | — | |
| VIDMResolution=128 x 1282025.12 | 64.17 | 263 | — | — | — | — | — | — | |
| OD-VAEChn=42024.11 | 58.48 | — | — | 1,109.87 | — | — | — | — | |
| VDMEvaluation Protocol=Finetuning Setting, Pretrain=No, Class=No, Resolution=64 x 642022.09 | 57.8 | — | — | — | — | — | — | — | |
| TATS-baseClass conditional=false, Trained on entire dataset=false2022.04 | 57.63 | — | 420 | — | — | — | — | — | |
| TATSExtra Video=false, Class (conditional)=false, Resolution=128x128, Used test set in training=false2022.12 | 57.63 | — | 420 | — | — | — | — | — | |
| TATSextra video data=false, class-conditional=false, resolution=128x1282022.12 | 57.63 | — | 420 | — | — | — | — | — | |
| TATSResolution=16 x 128 x 1282023.03 | 57.63 | — | 420 | — | — | — | — | — | |
| CogVideoXChn=162024.11 | 57.47 | — | — | 1,117.57 | — | — | — | — | |
| Video DiffusionClass conditional=false, Trained on entire dataset=true2022.04 | 57 | — | — | — | — | — | — | — | |
| Video DiffusionExtra Video=false, Class (conditional)=false, Resolution=Custom, Used test set in training=true2022.12 | 57 | — | — | — | — | — | — | — | |
| Video Diffusion*extra video data=false, class-conditional=false, resolution=custom2022.12 | 57 | — | — | — | — | — | — | — | |
| VDMResolution=16 x 64 x 642023.03 | 57 | — | 295 | — | — | — | — | — | |
| CogVideoEvaluation Protocol=Finetuning Setting, Pretrain=Yes, Class=Yes, Resolution=160 x 1602022.09 | 50.46 | — | 626 | — | — | — | — | — | |
| CogVideoClass conditional=true, Trained on entire dataset=true2022.04 | 50.46 | — | 626 | — | — | — | — | — | |
| CogVideoExtra Video=true, Class (conditional)=true, Resolution=Custom, Used test set in training=false2022.12 | 50.46 | — | 626 | — | — | — | — | — | |
| CogVideo*extra video data=true, class-conditional=true, resolution=custom2022.12 | 50.46 | — | 626 | — | — | — | — | — | |
| EncGAN3Resolution=128 x 1282025.12 | 43.65 | 356 | — | — | — | — | — | — | |
| CCVSClass conditional=false, Trained on entire dataset=true, Initialization=Real frame2022.04 | 41.37 | — | 389 | — | — | — | — | — | |
| UCF-101 datasettype=Real data upper bound2016.11 | 34.49 | — | — | — | — | — | — | — | |
| MoCoGAN-HDEvaluation Protocol=Finetuning Setting, Pretrain=No, Class=No, Resolution=256 x 2562022.09 | 33.95 | — | 700 | — | — | — | — | — | |
| Make-A-VideoEvaluation Protocol=Zero-Shot Setting, Pretrain=No, Class=Yes, Resolution=256 x 2562022.09 | 33 | — | 367.23 | — | — | — | — | — | |
| Make-A-VideoExtra Video=false, Class (conditional)=true, Resolution=Custom, Used test set in training=false2022.12 | 33 | — | 367 | — | — | — | — | — | |
| Make-A-Video*extra video data=true, class-conditional=false, resolution=custom2022.12 | 33 | — | 367 | — | — | — | — | — | |
| DVD-GANExtra Video=false, Class (conditional)=true, Resolution=128x128, Used test set in training=true2022.12 | 32.97 | — | — | — | — | — | — | — | |
| DVD-GANextra video data=false, class-conditional=false, resolution=128x1282022.12 | 32.97 | — | — | — | — | — | — | — | |
| DIGANEvaluation Protocol=Finetuning Setting, Pretrain=No, Class=No2022.09 | 32.7 | — | 577 | — | — | — | — | — | |
| DIGANClass conditional=false, Trained on entire dataset=true2022.04 | 32.7 | — | 577 | — | — | — | — | — | |
| DIGANExtra Video=false, Class (conditional)=false, Resolution=128x128, Used test set in training=true2022.12 | 32.7 | — | 577 | — | — | — | — | — | |
| DIGANextra video data=false, class-conditional=false, resolution=128x1282022.12 | 32.7 | — | 577 | — | — | — | — | — | |
| DIGANResolution=16 x 128 x 1282023.03 | 32.7 | — | 577 | — | — | — | — | — | |
| DIGANResolution=128 x 1282025.12 | 32.7 | 577 | — | — | — | — | — | — | |
| StyleGAN-VResolution=128 x 1282025.12 | 32.7 | — | — | — | — | — | — | — | |
| MoCOGAN-HDClass conditional=false, Trained on entire dataset=true2022.04 | 32.36 | — | 838 | — | — | — | — | — | |
| MoCOGAN-HDExtra Video=false, Class (conditional)=false, Resolution=128x128, Used test set in training=true2022.12 | 32.36 | — | 838 | — | — | — | — | — | |
| MoCOGAN-HDResolution=16 x 128 x 1282023.03 | 32.36 | — | 838 | — | — | — | — | — | |
| MoCoGAN-HDResolution=128 x 1282025.12 | 32.36 | 838 | — | — | — | — | — | — | |
| DIGANClass conditional=false, Trained on entire dataset=false2022.04 | 29.71 | — | 655 | — | — | — | — | — | |
| DIGANExtra Video=false, Class (conditional)=false, Resolution=128x128, Used test set in training=false2022.12 | 29.71 | — | 655 | — | — | — | — | — | |
| TGANv2Class conditional=true, Trained on entire dataset=false2022.04 | 28.87 | — | 1,209 | — | — | — | — | — | |
| TGANv2Extra Video=false, Class (conditional)=true, Resolution=128x128, Used test set in training=false2022.12 | 28.87 | — | 1,209 | — | — | — | — | — | |
| DVD-GANClass conditional=true, Trained on entire dataset=true2022.04 | 27.38 | — | — | — | — | — | — | — | |
| DVD-GANResolution=128 x 1282025.12 | 27.38 | — | — | — | — | — | — | — | |
| TGAN2st=42018.11 | 26.6 | 3,431 | — | — | — | — | — | — | |
| TGANv2Evaluation Protocol=Finetuning Setting, Pretrain=No, Class=No, Resolution=128 x 1282022.09 | 26.6 | — | — | — | — | — | — | — | |
| CogVideo (English)Evaluation Protocol=Zero-Shot Setting, Pretrain=No, Class=Yes, Resolution=480 x 4802022.09 | 25.27 | — | 701.59 | — | — | — | — | — | |
| VideoGPTClass conditional=false, Trained on entire dataset=false2022.04 | 24.69 | — | — | — | — | — | — | — | |
| VideoGPTExtra Video=false, Class (conditional)=false, Resolution=128x128, Used test set in training=false2022.12 | 24.69 | — | — | — | — | — | — | — | |
| VideoGPTResolution=16 x 128 x 1282023.03 | 24.69 | — | — | — | — | — | — | — | |
| CCVSClass conditional=false, Trained on entire dataset=true, Initialization=StyleGAN2022.04 | 24.47 | — | 386 | — | — | — | — | — | |
| CCVS+StyleGANExtra Video=false, Class (conditional)=false, Resolution=128x128, Used test set in training=true2022.12 | 24.47 | — | 386 | — | — | — | — | — | |
| CCVS+StyleGANextra video data=false, class-conditional=false, resolution=128x1282022.12 | 24.47 | — | 386 | — | — | — | — | — | |
| StyleGAN-VClass conditional=false, Trained on entire dataset=true2022.04 | 23.94 | — | — | — | — | — | — | — | |
| StyleGAN-VExtra Video=false, Class (conditional)=false, Resolution=Custom, Used test set in training=true2022.12 | 23.94 | — | — | — | — | — | — | — | |
| StyleGAN-V*extra video data=false, class-conditional=false, resolution=custom2022.12 | 23.94 | — | — | — | — | — | — | — | |
| StyleGAN-VResolution=16 x 256 x 2562023.03 | 23.94 | — | — | — | — | — | — | — | |
| TGAN2st=22018.11 | 23.87 | 3,797 | — | — | — | — | — | — | |
| CogVideo (Chinese)Evaluation Protocol=Zero-Shot Setting, Pretrain=No, Class=Yes, Resolution=480 x 4802022.09 | 23.55 | — | 751.34 | — | — | — | — | — | |
| LDVD-GANClass conditional=false, Trained on entire dataset=false2022.04 | 22.91 | — | — | — | — | — | — | — | |
| LDVD-GANExtra Video=false, Class (conditional)=false, Resolution=128x128, Used test set in training=false2022.12 | 22.91 | — | — | — | — | — | — | — | |
| RaMViDExtra Video=false, Class (conditional)=false, Resolution=128x128, Used test set in training=false2022.12 | 21.71 | — | — | — | — | — | — | — | |
| RaMViDextra video data=false, class-conditional=false, resolution=128x1282022.12 | 21.71 | — | — | — | — | — | — | — | |
| Conditional TGANconstraint=SVC2016.11 | 15.83 | — | — | — | — | — | — | — | |
| TGANClass conditional=true, Trained on entire dataset=false2022.04 | 15.83 | — | — | — | — | — | — | — | |
| TGANExtra Video=false, Class (conditional)=true, Resolution=128x128, Used test set in training=false2022.12 | 15.83 | — | — | — | — | — | — | — | |
| ProgressiveVGAN w/ SWL2018.11 | 14.56 | — | — | — | — | — | — | — | |
| ProgressiveVGANClass conditional=true, Trained on entire dataset=false2022.04 | 14.56 | — | — | — | — | — | — | — | |
| ProgressiveVGANExtra Video=false, Class (conditional)=false, Resolution=128x128, Used test set in training=false2022.12 | 14.56 | — | — | — | — | — | — | — | |
| Progressive VGAN2018.11 | 13.59 | — | — | — | — | — | — | — | |
| Naive implementation2018.11 | 13.29 | 5,401 | — | — | — | — | — | — | |
| MoCoGAN2018.11 | 12.42 | — | — | — | — | — | — | — | |
| MoCOGANClass conditional=true, Trained on entire dataset=false2022.04 | 12.42 | — | — | — | — | — | — | — | |
| MoCOGANExtra Video=false, Class (conditional)=false, Resolution=Custom, Used test set in training=false2022.12 | 12.42 | — | — | — | — | — | — | — | |
| TGANconstraint=SVC2016.11 | 11.85 | — | — | — | — | — | — | — | |
| TGAN2018.11 | 11.85 | — | — | — | — | — | — | — | |
| TGANClass conditional=false, Trained on entire dataset=false2022.04 | 11.85 | — | — | — | — | — | — | — | |
| TGANExtra Video=false, Class (conditional)=false, Resolution=128x128, Used test set in training=false2022.12 | 11.85 | — | — | — | — | — | — | — | |
| TGANResolution=16 x 64 x 642023.03 | 11.85 | — | — | — | — | — | — | — | |
| TGANconstraint=Weight clipping2016.11 | 11.77 | — | — | — | — | — | — | — | |
| Single 3D discriminator only2018.11 | 11.1 | 8,358 | — | — | — | — | — | — | |
| 3D + 2D discriminators2018.11 | 10.47 | 8,304 | — | — | — | — | — | — | |
| TGANvariant=Normal GAN2016.11 | 9.18 | — | — | — | — | — | — | — | |
| Video GANconstraint=SVC2016.11 | 8.31 | — | — | — | — | — | — | — | |
| VGAN2018.11 | 8.31 | — | — | — | — | — | — | — |