Video Prediction on BAIR Push (test)
62FVDMAGVIT-L-FP
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| MAGVIT-L-FPModel size=Large (464M), Task=Frame Prediction2022.12 | 62 | 0.123 | 78.7 | — | — | 19.3 | |
| MCVD2022.12 | 90 | — | 78 | — | — | 16.9 | |
| Video Transformer2021.05 | 94 | — | — | — | — | — | |
| VideoFlowTemperature (T)=0.8, Training conditioning frames=3, Training total frames=13, Evaluation ground truth frames=3, Evaluation total frames=132019.03 | 95 | — | — | — | — | — | |
| CCVS2022.12 | 99 | — | 72.9 | — | — | — | |
| cINN (Ours)Context Frames=12021.05 | 99.3 | — | — | 0.98 | 1.93 | — | |
| DVD-GAN2021.05 | 109.8 | — | — | — | — | — | |
| SAVPTraining conditioning frames=2, Training total frames=16, Evaluation ground truth frames=2, Evaluation total frames=162019.03 | 116 | — | — | — | — | — | |
| SAVP2020.02 | 116.4 | — | — | — | — | — | |
| SAVPContext Frames=22021.05 | 116.4 | — | — | 0.98 | 1.7 | — | |
| IVRNNContext Frames=22021.05 | 121.3 | — | — | 0.69 | 1.13 | — | |
| LVT2021.05 | 125.8 | — | — | — | — | — | |
| VideoFlowTemperature (T)=0.8, Training conditioning frames=3, Training total frames=13, Evaluation ground truth frames=3, Evaluation total frames=162019.03 | 127 | — | — | — | — | — | |
| VideoFlowTemperature (T)=0.8, Training conditioning frames=3, Training total frames=13, Evaluation ground truth frames=2, Evaluation total frames=162019.03 | 131 | — | — | — | — | — | |
| Video Flow2021.05 | 131 | — | — | — | — | — | |
| cINN (Ours) w/o ADAINContext Frames=1, ablation=without ADAIN2021.05 | 131.2 | — | — | 0.78 | 1.73 | — | |
| cINN (Ours) w/o cINNContext Frames=1, ablation=without cINN2021.05 | 134.5 | — | — | 0.59 | 0.94 | — | |
| SRVPContext Frames=82021.05 | 141.7 | — | — | 0.93 | 1.65 | — | |
| Hierarchical VRNNHierarchy=with2019.04 | 143.4 | 0.055 | 82.2 | — | — | — | |
| SAVP2019.04 | 143.43 | 0.062 | 79.5 | — | — | — | |
| VideoFlowTemperature (T)=1.0, Training conditioning frames=3, Training total frames=13, Evaluation ground truth frames=3, Evaluation total frames=132019.03 | 149 | — | — | — | — | — | |
| Hierarchical VRNNHierarchy=without2019.04 | 149.22 | 0.058 | 82.9 | — | — | — | |
| STMFANet2020.02 | 159.6 | — | — | — | — | — | |
| VideoFlowTemperature (T)=1.0, Training conditioning frames=3, Training total frames=13, Evaluation ground truth frames=3, Evaluation total frames=162019.03 | 221 | — | — | — | — | — | |
| VideoFlowTemperature (T)=1.0, Training conditioning frames=3, Training total frames=13, Evaluation ground truth frames=2, Evaluation total frames=162019.03 | 251 | — | — | — | — | — | |
| SVG-LP2019.04 | 256.62 | 0.061 | 81.6 | — | — | — | |
| SV2P2020.02 | 262.5 | — | — | — | — | — | |
| SV2PTraining conditioning frames=2, Training total frames=16, Evaluation ground truth frames=2, Evaluation total frames=162019.03 | 263 | — | — | — | — | — | |
| cINN (Ours) w/o x0Context Frames=1, ablation=without initial frame context2021.05 | 272.6 | — | — | 2.4 | 2.48 | — | |
| SVG-FP2020.02 | 315.5 | — | — | — | — | — |