Video Prediction on Cityscapes (Semantic & Generative Quality)
70.46mIoU (Avg)Re2Pix (Stage 1)
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Re2Pix (Stage 1)VFM forecasting only=true, Stage=12026.04 | 70.46 | 70.26 | 88.03 | 0.1208 | — | — | |
| Cosmos-Predict-2 (Finetuned)Pre-training=Large-scale internet-pretrained, Fine-tuning=true2026.04 | 65.69 | 64.85 | 86.19 | 0.1371 | 7.74 | 46.22 | |
| Re2Pix (ours)Training=from scratch2026.04 | 64.63 | 63.52 | 86.01 | 0.14 | 9.29 | 49.03 | |
| Vista (Finetuned)Pre-training=Large-scale internet-pretrained, Fine-tuning=true2026.04 | 63.88 | 62.35 | 85.91 | 0.1376 | 12.17 | 84.14 | |
| Baseline (Large)Training=from scratch, Model Scale=Large2026.04 | 63.32 | 61.73 | 85.63 | 0.1423 | 9.89 | 51.77 | |
| BaselineTraining=from scratch2026.04 | 62.77 | 60.86 | 85.62 | 0.1433 | 9.64 | 52.12 |