Pose Propagation on JHMDB
63.1PCK@0.1Ours - Sintel
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Ours - SintelArch.=PWC-Net2022.01 | 63.1 | 38 | 81.4 | |
| Ours - VOSArch.=PWC-Net2022.01 | 62.6 | 38.2 | 80.9 | |
| VFSArch.=ResNet502022.01 | 60.9 | — | 80.7 | |
| VFSBackbone=ResNet-50, Pre-training Dataset=K4002023.04 | 60.9 | — | 80.7 | |
| MoCoBackbone=ResNet-50, Pre-training Dataset=ImageNet2023.04 | 60.4 | — | 79.3 | |
| CRWArch.=ResNet182022.01 | 59.3 | 29.1 | 80.6 | |
| CRWBackbone=ResNet-18, Pre-training Dataset=K4002023.04 | 59.3 | — | 80.3 | |
| SupervisedBackbone=ResNet-50, Pre-training Dataset=ImageNet2023.04 | 59.2 | — | 78.3 | |
| UVCArch.=ResNet182022.01 | 58.6 | — | 79.6 | |
| UVCBackbone=ResNet-18, Pre-training Dataset=K4002023.04 | 58.6 | — | 79.6 | |
| SimSiamBackbone=ResNet-50, Pre-training Dataset=ImageNet2023.04 | 58.4 | — | 77.5 | |
| VINCEBackbone=ResNet-50, Pre-training Dataset=K4002023.04 | 58.2 | — | 76.3 | |
| DULBackbone=ResNet-18, Pre-training Dataset=K4002023.04 | 58.2 | — | 80.5 | |
| DropMAE + AdaptorBackbone=ViT-B/16, Pre-training Dataset=K400, frozen_backbone=true2023.04 | 57.8 | — | 80.8 | |
| TimeCycleBackbone=ResNet-50, Pre-training Dataset=VLOG2023.04 | 57.7 | — | 78.5 | |
| RegionTrackerBackbone=ResNet-50, Pre-training Dataset=TrackingNet2023.04 | 57.5 | — | 74.6 | |
| DropMAE + AdaptorBackbone=ViT-B/16, Pre-training Dataset=YT-VOS, frozen_backbone=true2023.04 | 57.3 | — | 80.2 | |
| DULBackbone=ResNet-18, Pre-training Dataset=YT-VOS2023.04 | 56.4 | — | 79.1 | |
| RAFTArch.=RAFT, Note=chained flow baseline2022.01 | 55.6 | 30.2 | 76 | |
| UFlowArch.=PWC-Net, Note=chained flow baseline2022.01 | 51.3 | 24.1 | 72.1 | |
| MAE +OursType=Mask Modeling, Backbone=ViT-B/16, Pre-training Dataset=K400, Epoch=+5, Evaluation Protocol=Zero-shot semi-supervised2026.03 | 48.4 | — | 76.7 | |
| DINO v2 +OursType=Self Distillation, Backbone=ViT-B/16, Pre-training Dataset=K400, Epoch=+5, Evaluation Protocol=Zero-shot semi-supervised2026.03 | 47.3 | — | 76.8 | |
| SiamMAEType=Video Pretrained, Backbone=ViT-B/16, Pre-training Dataset=K400, Epoch=400, Evaluation Protocol=Zero-shot semi-supervised2026.03 | 47 | — | 74.9 | |
| DINO v2Type=Self Distillation, Backbone=ViT-B/16, Pre-training Dataset=INet, Epoch=100, Evaluation Protocol=Zero-shot semi-supervised2026.03 | 46.6 | — | 76.3 | |
| DINO +OursType=Self Distillation, Backbone=ViT-B/16, Pre-training Dataset=K400, Epoch=+5, Evaluation Protocol=Zero-shot semi-supervised2026.03 | 46.2 | — | 75.2 | |
| iBOT +OursType=Self Distillation, Backbone=ViT-B/16, Pre-training Dataset=K400, Epoch=+5, Evaluation Protocol=Zero-shot semi-supervised2026.03 | 46.1 | — | 75.8 | |
| RSPType=Video Pretrained, Backbone=ViT-B/16, Pre-training Dataset=K400, Epoch=400, Evaluation Protocol=Zero-shot semi-supervised2026.03 | 46 | — | 74.6 | |
| iBOTType=Self Distillation, Backbone=ViT-B/16, Pre-training Dataset=INet, Epoch=400, Evaluation Protocol=Zero-shot semi-supervised2026.03 | 45.7 | — | 75.3 | |
| CropMAEType=Video Pretrained, Backbone=ViT-B/16, Pre-training Dataset=K400, Epoch=400, Evaluation Protocol=Zero-shot semi-supervised2026.03 | 45.3 | — | 73.3 | |
| MoCo v3 +OursType=Contrastive Learning, Backbone=ViT-B/16, Pre-training Dataset=K400, Epoch=+5, Evaluation Protocol=Zero-shot semi-supervised2026.03 | 45.3 | — | 75 | |
| MAE-STType=Video Pretrained, Backbone=ViT-L/16, Pre-training Dataset=K400, Epoch=1600, Evaluation Protocol=Zero-shot semi-supervised2026.03 | 44.4 | — | 72.5 | |
| I-JEPA +OursType=Mask Modeling, Backbone=ViT-B/16, Pre-training Dataset=K400, Epoch=+5, Evaluation Protocol=Zero-shot semi-supervised2026.03 | 44.4 | — | 73.2 | |
| DINOType=Self Distillation, Backbone=ViT-B/16, Pre-training Dataset=INet, Epoch=300, Evaluation Protocol=Zero-shot semi-supervised2026.03 | 44.4 | — | 74.2 | |
| MoCo v3Type=Contrastive Learning, Backbone=ViT-B/16, Pre-training Dataset=INet, Epoch=300, Evaluation Protocol=Zero-shot semi-supervised2026.03 | 43.6 | — | 73.5 | |
| VideoMAEType=Video Pretrained, Backbone=ViT-L/16, Pre-training Dataset=K400, Epoch=1600, Evaluation Protocol=Zero-shot semi-supervised2026.03 | 43.3 | — | 70.5 | |
| I-JEPAType=Mask Modeling, Backbone=ViT-B/16, Pre-training Dataset=INet, Epoch=800, Evaluation Protocol=Zero-shot semi-supervised2026.03 | 42.6 | — | 71.4 | |
| DropMAEType=Video Pretrained, Backbone=ViT-B/16, Pre-training Dataset=K400, Epoch=1600, Evaluation Protocol=Zero-shot semi-supervised2026.03 | 42.3 | — | 69.2 | |
| MAEType=Mask Modeling, Backbone=ViT-B/16, Pre-training Dataset=INet, Epoch=800, Evaluation Protocol=Zero-shot semi-supervised2026.03 | 41.6 | — | 69.3 | |
| CLIP +OursType=Contrastive Learning, Backbone=ViT-B/16, Pre-training Dataset=K400, Epoch=+5, Evaluation Protocol=Zero-shot semi-supervised2026.03 | 40.6 | — | 71.4 | |
| BLIP +OursType=Contrastive Learning, Backbone=ViT-B/16, Pre-training Dataset=K400, Epoch=+5, Evaluation Protocol=Zero-shot semi-supervised2026.03 | 38.9 | — | 70.2 | |
| CLIPType=Contrastive Learning, Backbone=ViT-B/16, Pre-training Dataset=WIT, Epoch=32, Evaluation Protocol=Zero-shot semi-supervised2026.03 | 36.9 | — | 67.7 | |
| BLIPType=Contrastive Learning, Backbone=ViT-B/16, Pre-training Dataset=LAION, Epoch=20, Evaluation Protocol=Zero-shot semi-supervised2026.03 | 35.1 | — | 65.9 |