Video Object Segmentation on YouTube-VOS 2018
92.4Score GMIVOS
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| MIVOSSegmentation Data=BL30K+VOS+DAVIS (4.88M), Type=Mask/Scribble, Refer-Type=Mask, Zero-Shot=false, Single Image=false2023.04 | 92.4 | — | — | — | — | |
| DeAOTBackbone=Swin-B2024.02 | 86.2 | — | — | — | — | |
| XMemvenue=ECCV'22, Training Data=with video data2023.04 | 86.1 | 85.1 | 89.8 | 80.3 | 89.2 | |
| XMemSegmentation Data=Image+VOS+DAVIS, Type=Video, Refer-Type=Mask, Zero-Shot=false2023.11 | 86.1 | 85.1 | 89.8 | 80.3 | 89.2 | |
| DeAOTBackbone=R502024.02 | 86 | — | — | — | — | |
| XMemBackbone=R502024.02 | 85.7 | — | — | — | — | |
| RDEvenue=CVPR'22, Training Data=with video data2023.04 | 83.3 | 81.9 | 86.3 | 78 | 86.9 | |
| SWEMvenue=CVPR'22, Training Data=with video data2023.04 | 82.8 | 82.4 | 86.9 | 77.1 | 85 | |
| SWEMSegmentation Data=Image+VOS+DAVIS (0.25M), Type=Mask, Refer-Type=Mask, Zero-Shot=false, Single Image=false2023.04 | 82.8 | 82.4 | 86.9 | 77.1 | 85 | |
| SWEMSegmentation Data=Image+VOS+DAVIS, Refer-Type=Mask, Zero-Shot=false2023.11 | 82.8 | 82.4 | 86.9 | 77.1 | 85 | |
| XMemSegmentation Data=Image+VOS+DAVIS (0.25M), Type=Video, Refer-Type=Mask, Zero-Shot=false, Single Image=false2023.04 | 82.6 | 81.1 | 85.6 | 77.7 | 86.2 | |
| MIVOSSegmentation Data=BL30K+VOS+DAVIS, Refer-Type=Mask, Zero-Shot=false2023.11 | 82.6 | 81.1 | 85.6 | 77.7 | 86.2 | |
| UniRefBackbone=SwinL, Joint Training=true, Universal=false2024.02 | 82.6 | — | — | — | — | |
| UniRefBackbone=ResNet50, Joint Training=true, Universal=false2024.02 | 81.4 | — | — | — | — | |
| AFB-URRvenue=NeurIPS'20, Training Data=with video data2023.04 | 79.6 | 78.8 | 83.1 | 74.1 | 82.6 | |
| STMvenue=ICCV'19, Training Data=with video data2023.04 | 79.4 | 79.7 | 84.2 | 72.8 | 80.9 | |
| UNINEXTBackbone=ViT-H, Joint Training=false, Universal=false2024.02 | 78.6 | — | — | — | — | |
| UNINEXT-LSegmentation Data=Image+Video, Type=Generalist, Refer-Type=Mask, Zero-Shot=false2023.11 | 78.1 | 79.1 | 83.5 | 71 | 78.9 | |
| UNINEXTBackbone=ConvNeXtL, Joint Training=false, Universal=false2024.02 | 78.1 | — | — | — | — | |
| UNINEXT-TSegmentation Data=Image+Video, Refer-Type=Mask, Zero-Shot=false2023.11 | 77 | 76.8 | 81 | 70.8 | 79.4 | |
| UNINEXTBackbone=ResNet50, Joint Training=false, Universal=false2024.02 | 77 | — | — | — | — | |
| SegGPTvenue=this work, Training Data=without video data2023.04 | 74.7 | 75.1 | 80.2 | 67.4 | 75.9 | |
| SegGPT-LSegmentation Data=COCO+ADE+VOC+... (0.25M), Type=Generalist, Refer-Type=Mask, Zero-Shot=false, Single Image=true2023.04 | 74.7 | 75.1 | 80.2 | 67.4 | 75.9 | |
| SegGPT-LSegmentation Data=COCO+ADE+VOC+..., Refer-Type=Mask, Zero-Shot=false2023.11 | 74.7 | 75.1 | 80.2 | 67.4 | 75.9 | |
| UniVSBackbone=SwinL, Joint Training=true, Universal=true2024.02 | 71.5 | — | — | — | — | |
| AGSSvenue=ICCV'19, Training Data=with video data2023.04 | 71.3 | 71.3 | 65.5 | 75.2 | 73.1 | |
| AGSSSegmentation Data=VOS+DAVIS (0.1M), Type=Mask, Refer-Type=Mask, Zero-Shot=false, Single Image=false2023.04 | 71.3 | 71.3 | 65.5 | 75.2 | 73.1 | |
| AGSSSegmentation Data=VOS+DAVIS, Refer-Type=Mask, Zero-Shot=false2023.11 | 71.3 | 71.3 | 65.5 | 75.2 | 73.1 | |
| UNINEXT-LSegmentation Data=Image+Video (3M), Type=Generalist, Refer-Type=Mask, Zero-Shot=false, Single Image=false2023.04 | 71 | 78.9 | — | — | — | |
| UniVSBackbone=SwinB, Joint Training=true, Universal=true2024.02 | 70.9 | — | — | — | — | |
| UNINEXT-TSegmentation Data=Image+Video (3M), Type=Generalist, Refer-Type=Mask, Zero-Shot=false, Single Image=false2023.04 | 70.8 | — | — | — | — | |
| UniVSBackbone=SwinT, Joint Training=true, Universal=true2024.02 | 70.3 | — | — | — | — | |
| UniVSBackbone=ResNet50, Joint Training=true, Universal=true2024.02 | 69.2 | — | — | — | — | |
| AGAMEvenue=CVPR'19, Training Data=with video data2023.04 | 66 | 66.9 | — | 61.2 | — | |
| AGAMESegmentation Data=(Synth)VOS+DAVIS (0.11M), Type=Mask, Refer-Type=Mask, Zero-Shot=false, Single Image=false2023.04 | 66 | 66.9 | 61.2 | — | — | |
| AGAMESegmentation Data=(Synth)VOS+DAVIS, Refer-Type=Mask, Zero-Shot=false2023.11 | 66 | 66.9 | — | 61.2 | — | |
| DINOv-TSegmentation Data=COCO+SAM, Refer-Type=Mask, Zero-Shot=true2023.11 | 60.9 | 65.3 | 70 | 52.3 | 57.9 | |
| SiamMaskSegmentation Data=COCO+VOS (0.21M), Type=Mask, Refer-Type=Box, Zero-Shot=false, Single Image=false2023.04 | 60.2 | 58.2 | 45.1 | 47.7 | — | |
| DINOv-LSegmentation Data=COCO+SAM, Refer-Type=Mask, Zero-Shot=true2023.11 | 59.6 | 61.7 | 65.7 | 52.3 | 58.8 | |
| SEEM-BSegmentation Data=COCO+LVIS (0.12M), Type=Generalist, Refer-Type=Mask/Single Point, Zero-Shot=true, Single Image=true2023.04 | 53.8 | 60 | 44.5 | 63.5 | 47.2 | |
| SEEM-TSegmentation Data=COCO+LVIS (0.12M), Type=Generalist, Refer-Type=Mask/Single Point, Zero-Shot=true, Single Image=true2023.04 | 51.4 | 55.6 | 44.1 | 59.2 | 46.9 | |
| SEEM-TSegmentation Data=COCO+LVIS, Type=Generalist, Refer-Type=Mask, Zero-Shot=true2023.11 | 51.4 | 55.6 | 44.1 | 59.2 | 46.9 | |
| SEEM-LSegmentation Data=COCO+LVIS (0.12M), Type=Generalist, Refer-Type=Mask/Single Point, Zero-Shot=true, Single Image=true2023.04 | 50 | 57.2 | 38.2 | 61.3 | 43.3 | |
| SEEM-LSegmentation Data=COCO+LVIS, Type=Generalist, Refer-Type=Mask, Zero-Shot=true2023.11 | 50 | 57.2 | 38.2 | 61.3 | 43.3 | |
| Paintervenue=CVPR'23, Training Data=without video data2023.04 | 24.1 | 27.6 | 35.8 | 14.3 | 18.7 | |
| Painter-LSegmentation Data=COCO+ADE+NYUv2 (0.16M), Type=Generalist, Refer-Type=Mask, Zero-Shot=false, Single Image=true2023.04 | 24.1 | 27.6 | 35.8 | 14.3 | 18.7 | |
| Painter-LSegmentation Data=COCO+ADE+NYUv2, Refer-Type=Mask, Zero-Shot=false2023.11 | 24.1 | 27.6 | 35.8 | 14.3 | 18.7 | |
| SiamMaskSegmentation Data=COCO+VOS, Refer-Type=Box, Zero-Shot=false2023.11 | — | 60.2 | 58.2 | 45.1 | 47.7 |