Video Prediction on RoboNet
51.5FVDForeDiff
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| ForeDiffnumber of samples=best of 1002025.05 | 51.5 | 28.2 | 90.4 | 4.5 | |
| FitVidConditioning=action-conditioned, Resolution=64x642024.05 | 62.5 | 28.2 | 89.3 | 2.4 | |
| FitVidnumber of samples=best of 1002025.05 | 62.5 | 28.2 | 89.3 | 2.4 | |
| iVideoGPTConditioning=action-conditioned, Resolution=64x642024.05 | 63.2 | 27.8 | 90.6 | 4.9 | |
| iVideoGPTnumber of samples=best of 1002025.05 | 63.2 | 27.8 | 90.6 | 4.9 | |
| GHVAEConditioning=action-conditioned, Resolution=64x642024.05 | 95.2 | 24.7 | 89.1 | 3.6 | |
| GHVAEnumber of samples=best of 1002025.05 | 95.2 | 24.7 | 89.1 | 3.6 | |
| SVGConditioning=action-conditioned, Resolution=64x642024.05 | 123.2 | 23.9 | 87.8 | 6 | |
| SVGnumber of samples=best of 1002025.05 | 123.2 | 23.9 | 87.8 | 6 | |
| MaskViTConditioning=action-conditioned, Resolution=64x642024.05 | 133.5 | 23.2 | 80.5 | 4.2 | |
| MaskViTnumber of samples=best of 1002025.05 | 133.5 | 23.2 | 80.5 | 4.2 | |
| iVideoGPTConditioning=action-conditioned, Resolution=256x2562024.05 | 197.9 | 23.8 | 80.8 | 14.7 | |
| MaskViTConditioning=action-conditioned, Resolution=256x2562024.05 | 211.7 | 20.4 | 67.1 | 17 |