Robotic Manipulation on Franka-Kitchen
93.75Avg Success RateBYOL
Evaluation Results
| Method | Links | |
|---|---|---|
| BYOLVisual Representation=Frozen2024.09 | 93.75 | |
| DynaMoVisual Representation=Frozen2024.09 | 91 | |
| RPTVisual Representation=Frozen2024.09 | 88.5 | |
| SlotMIMPre-training Dataset=Ego4D, Pre-training Scale=1.28M, Backbone=ViT-B/162025.03 | 86 | |
| BYOL-TVisual Representation=Frozen2024.09 | 83.25 | |
| RandomVisual Representation=Frozen2024.09 | 83 | |
| MoCo-v3Visual Representation=Frozen2024.09 | 82 | |
| MPI#Param=21.7M, Pre-training Dataset=Ego4D, #Seen frames=0.1B, Supervision Type=Supervision with Auxiliary Language Guidance, Uses multi-head attention pooling layers=true2025.07 | 76.5 | |
| ImageNetVisual Representation=Frozen2024.09 | 75.25 | |
| R3MVisual Representation=Frozen2024.09 | 71 | |
| Voltron#Param=21.7M, Pre-training Dataset=SS-v2, #Seen frames=0.3B, Supervision Type=Supervision with Auxiliary Language Guidance, Uses multi-head attention pooling layers=true2025.07 | 70.5 | |
| ToBo#Param=21.7M, Pre-training Dataset=Kinetics-400, #Seen frames=0.2B, Supervision Type=Self-supervised Learning2025.07 | 68 | |
| MAEVisual Representation=Frozen2024.09 | 67.5 | |
| VC-1Visual Representation=Frozen2024.09 | 65.75 | |
| DINOv2Pre-training Dataset=LVD, Pre-training Scale=142M, Backbone=ViT-B/142025.03 | 64 | |
| TCN-SVVisual Representation=Frozen2024.09 | 60.25 | |
| VC-1Pre-training Dataset=Ego4D+MNI, Pre-training Scale=5.6M, Backbone=ViT-B/162025.03 | 58.1 | |
| MVPVisual Representation=Frozen2024.09 | 57.75 | |
| R3M2024.05 | 57.6 | |
| data4robotics#Param=86.0M, Pre-training Dataset=Kinetics-700, #Seen frames=0.5B, Supervision Type=Self-supervised Learning2025.07 | 55 | |
| R3M#Param=25.6M, Pre-training Dataset=Ego4D, #Seen frames=0.8B, Supervision Type=Supervision with Auxiliary Language Guidance2025.07 | 53.1 | |
| SCR-FTfine-tuned=true2024.05 | 49.9 | |
| MVPPre-training Dataset=EgoSoup, Pre-training Scale=4.6M, Backbone=ViT-B/162025.03 | 49.7 | |
| VC-12024.05 | 47.5 | |
| R3M#Param=25.6M, Pre-training Dataset=Ego4D, #Seen frames=0.8B, Supervision Type=Self-supervised Learning, Excludes language guidance=true2025.07 | 47.2 | |
| SlotMIMPre-training Dataset=DetSoup, Pre-training Scale=4M, Backbone=ViT-B/162025.03 | 46.7 | |
| SCRfine-tuned=false2024.05 | 45 | |
| UniSplatMethod Category=Embodied-Specific2026.04 | 44.5 | |
| SD-VAE2024.05 | 43.7 | |
| MAEMethod Category=Vision-Centric2026.04 | 42.7 | |
| DINOv2Method Category=Vision-Centric2026.04 | 40.9 | |
| SPAMethod Category=Embodied-Specific2026.04 | 40.6 | |
| VC-1Method Category=Embodied-Specific2026.04 | 37.5 | |
| EVAMethod Category=Multi-Modal2026.04 | 37.3 | |
| CLIP2024.05 | 36.3 | |
| MVPMethod Category=Embodied-Specific2026.04 | 34.3 | |
| Voltron2024.05 | 33.5 | |
| CLIPMethod Category=Multi-Modal2026.04 | 30.8 | |
| InternViTMethod Category=Multi-Modal2026.04 | 28.5 |