Video Semantic Segmentation on YouTube-VIS 2021
44.2mAPCLIP-VIS
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| CLIP-VISTraining=supervised2026.01 | 44.2 | 76.31 | |
| PyraTokTraining=zero-shot2026.01 | 24.54 | 66.56 | |
| UVISTraining=unsupervised2026.01 | 17.5 | 63.11 | |
| VideoCutLERTraining=unsupervised2026.01 | 17.1 | 62.23 | |
| OmniTokenizerTraining=zero-shot2026.01 | 14.54 | 51.12 | |
| VideoVae+Training=zero-shot2026.01 | 12.33 | 51.21 | |
| LARPTraining=zero-shot2026.01 | 10.52 | 49.37 |