Robotic Manipulation on Calvin ABC-D
100Task-1 ScoreCoVAR
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| CoVAR2025.12 | 100 | 80 | 86.7 | 90.9 | 92.9 | — | — | |
| HiF-VLAView=Multi-View2026.05 | 98.5 | 94.1 | 88.1 | 81.4 | 73.1 | — | 4.35 | |
| MoLA2026.05 | 98.5 | 95 | 91.1 | 88.1 | 82.6 | — | 4.55 | |
| DreamVLA2026.05 | 98.2 | 94.6 | 89.5 | 83.4 | 78.1 | — | 4.44 | |
| RoboVLMsView=Multi-View2026.05 | 98 | 93.6 | 85.4 | 77.8 | 70.4 | — | 4.25 | |
| ElasticFlowView=Multi-View2026.05 | 97.8 | 95.3 | 87.9 | 83.6 | 72.7 | — | 4.37 | |
| VPPView=Multi-View2026.05 | 96.5 | 90.9 | 86.6 | 82 | 76.9 | — | 4.33 | |
| VPP2026.05 | 96.5 | 90.9 | 86.6 | 82 | 76.9 | — | 4.33 | |
| OpenVLA-OFTView=Multi-View2026.05 | 96.3 | 89.1 | 82.4 | 75.8 | 66.5 | — | 4.1 | |
| SeerView=Multi-View2026.05 | 96.3 | 91.6 | 86.1 | 80.3 | 74 | — | 4.28 | |
| Seer2026.05 | 96.3 | 91.6 | 86.1 | 80.3 | 74 | — | 4.28 | |
| CLOVERView=Third View2026.05 | 96 | 83.5 | 70.8 | 57.5 | 45.4 | — | 3.53 | |
| CLOVER2026.05 | 96 | 83.5 | 70.8 | 57.5 | 45.4 | — | 3.53 | |
| ElasticFlowView=Third View2026.05 | 95.6 | 88.7 | 79.8 | 77.3 | 73.4 | — | 4.15 | |
| UniVLAView=Third View2026.05 | 95.5 | 85.8 | 74.8 | 66.9 | 56.5 | — | 3.8 | |
| UniVLA2026.05 | 95.5 | 85.8 | 75.4 | 66.9 | 56.5 | — | 3.8 | |
| π0.5reproduced_or_previously_reported=true2026.05 | 94.8 | 87.4 | 78.2 | 71.7 | 64.3 | — | 3.97 | |
| Qwen3VL-2BSize=2.1B, Framework=VLM4VLA2026.01 | 94.3 | 88.2 | 83.1 | 77.6 | 71 | 4.142 | — | |
| Qwen3VL-2BSize=2.1B, # Samples Seen=7.7M / 25.6M / 25.6M2026.04 | 94.3 | 88.2 | 83.1 | 77.6 | 71 | — | 4.142 | |
| GR00T N1reproduced_or_previously_reported=true2026.05 | 94.2 | 86.1 | 79.6 | 73.9 | 66.8 | — | 4.01 | |
| Qwen3VL-8BSize=8.8B, Framework=VLM4VLA2026.01 | 94 | 86.8 | 79.7 | 74.6 | 68.4 | 4.035 | — | |
| Qwen3VL-8BEvaluation Protocol=VLM with VLM4VLA Models (Full Fine-tuning)2026.03 | 94 | 86.8 | 79.7 | 74.6 | 68.4 | 4.035 | — | |
| Qwen3VL-8BSize=8.8B, # Samples Seen=7.7M / 25.6M / 25.6M2026.04 | 94 | 86.8 | 79.7 | 74.6 | 68.4 | — | 4.035 | |
| Qwen3VL-30B-A3BSize=31.1B, Framework=VLM4VLA2026.01 | 93.9 | 87.7 | 82 | 75.7 | 68.2 | 4.075 | — | |
| Qwen3VL-30B-A3BSize=30B-A3B, # Samples Seen=7.7M / 25.6M / 25.6M2026.04 | 93.9 | 87.7 | 82 | 75.7 | 68.2 | — | 4.075 | |
| π0View=Multi-View2026.05 | 93.8 | 85 | 76.7 | 68.1 | 59.9 | — | 3.92 | |
| π0reproduced_or_previously_reported=true2026.05 | 93.8 | 85 | 76.7 | 68.1 | 59.9 | — | 3.84 | |
| π0View=Third View2026.05 | 93.7 | 83.2 | 74 | 62.9 | 51 | — | 3.65 | |
| Qwen2.5VL-7BSize=8.3B, Framework=VLM4VLA2026.01 | 93.5 | 86.4 | 80.7 | 75.8 | 69.3 | 4.057 | — | |
| Qwen2.5VL-7BEvaluation Protocol=VLM with VLM4VLA Models (Full Fine-tuning)2026.03 | 93.5 | 86.4 | 80.7 | 75.8 | 69.3 | 4.057 | — | |
| Qwen2.5VL-7BSize=8.3B, # Samples Seen=7.7M / 25.6M / 25.6M2026.04 | 93.5 | 86.4 | 80.7 | 75.8 | 69.3 | — | 4.057 | |
| InternVL3.5-1B + EmbodiedMidtrainSize=1.1B, # Samples Seen=1.0M / 4.1M / 4.1M2026.04 | 93.5 | 83.8 | 73.7 | 65.3 | 55.1 | — | 3.714 | |
| HiF-VLAView=Third View2026.05 | 93.5 | 87.4 | 81.4 | 75.9 | 69.4 | — | 4.08 | |
| Qwen3VL-4BSize=4.4B, Framework=VLM4VLA2026.01 | 93.3 | 85.7 | 79 | 71.9 | 64.4 | 3.943 | — | |
| Qwen3VL-4BSize=4.4B, # Samples Seen=7.7M / 25.6M / 25.6M2026.04 | 93.3 | 85.7 | 79 | 71.9 | 64.4 | — | 3.943 | |
| OmniStream-7BEvaluation Protocol=VLM with VLM4VLA Models (Frozen Vision)2026.03 | 93.1 | 84.8 | 76.8 | 70.3 | 63.4 | 3.885 | — | |
| UP-VLAView=Multi-View2026.05 | 92.8 | 86.5 | 81.5 | 76.9 | 69.9 | — | 4.08 | |
| UP-VLA2026.05 | 92.8 | 86.5 | 81.5 | 76.9 | 69.9 | — | 4.08 | |
| Qwen2.5VL-3BSize=3.8B, Framework=VLM4VLA2026.01 | 92.2 | 84.2 | 76.6 | 70 | 62.6 | 3.856 | — | |
| Qwen2.5VL-3BSize=3.8B, # Samples Seen=7.7M / 25.6M / 25.6M2026.04 | 92.2 | 84.2 | 76.6 | 70 | 62.6 | — | 3.856 | |
| Qwen3VL-2B + EmbodiedMidtrainSize=2.1B, # Samples Seen=1.0M / 4.1M / 4.1M2026.04 | 92.2 | 80.8 | 70 | 62.3 | 53.3 | — | 3.584 | |
| 3D Diffusor Actor2026.05 | 92.2 | 78.7 | 63.9 | 51.2 | 41.2 | — | 3.27 | |
| VidmanView=Multi-View2026.05 | 91.5 | 76.4 | 68.2 | 59.2 | 46.7 | — | 3.42 | |
| Paligemma-1Size=2.9B, Framework=VLM4VLA2026.01 | 91.4 | 81.3 | 69.2 | 59.9 | 48.8 | 3.506 | — | |
| Paligemma-1-3BEvaluation Protocol=VLM with VLM4VLA Models (Full Fine-tuning)2026.03 | 91.4 | 81.3 | 69.2 | 59.9 | 48.8 | 3.506 | — | |
| Paligemma-1Size=2.9B, # Samples Seen=7.7M / 25.6M / 25.6M2026.04 | 91.4 | 81.3 | 69.2 | 59.9 | 48.8 | — | 3.506 | |
| OpenVLAView=Third View2026.05 | 91.3 | 77.8 | 62 | 52.1 | 43.5 | — | 3.27 | |
| OpenVLA2026.05 | 91.3 | 77.8 | 62 | 52.1 | 43.5 | — | 3.27 | |
| InternVL3.5-1BSize=1.1B, # Samples Seen=1.0M / 4.1M / 4.1M2026.04 | 90.9 | 75.4 | 60.6 | 49.8 | 40.6 | — | 3.173 | |
| VPPView=Third View2026.05 | 90.9 | 81.5 | 71.3 | 62 | 51.8 | — | 3.58 | |
| Paligemma-2Size=3.0B, Framework=VLM4VLA2026.01 | 90.1 | 77.5 | 66.9 | 57.5 | 48.6 | 3.406 | — | |
| Paligemma-2-3BEvaluation Protocol=VLM with VLM4VLA Models (Full Fine-tuning)2026.03 | 90.1 | 77.5 | 66.9 | 57.5 | 48.6 | 3.406 | — | |
| Qwen2.5VL-7BEvaluation Protocol=VLM with VLM4VLA Models (Frozen Vision)2026.03 | 90.1 | 70 | 53.6 | 43.3 | 33.4 | 2.905 | — | |
| Paligemma-2Size=3.0B, # Samples Seen=7.7M / 25.6M / 25.6M2026.04 | 90.1 | 77.5 | 66.9 | 57.5 | 48.6 | — | 3.406 | |
| pi0*VLM Backbone=Paligemma-1, Size=3.1B2026.01 | 89.6 | 78.5 | 78.6 | 61 | 53.2 | 3.509 | — | |
| pi0*Model (VLM Backbone)=Paligemma-1, Evaluation Protocol=Expert Vision-Language-Action Models2026.03 | 89.6 | 78.5 | 78.6 | 61 | 53.2 | 3.509 | — | |
| π0 (Paligemma-1)Size=3.1B, # Samples Seen=7.7M / 25.6M / 25.6M2026.04 | 89.6 | 78.5 | 78.6 | 61 | 53.2 | — | 3.509 | |
| Qwen3VL-2B (low budget)Size=2.1B, # Samples Seen=1.0M / 4.1M / 4.1M2026.04 | 88.7 | 74.7 | 61.2 | 52.7 | 43.2 | — | 3.205 | |
| KosMos-2Size=1.7B, Framework=VLM4VLA2026.01 | 87.8 | 72.1 | 59.1 | 49.8 | 40.8 | 3.096 | — | |
| KosMos-2-1.7BEvaluation Protocol=VLM with VLM4VLA Models (Full Fine-tuning)2026.03 | 87.8 | 72.1 | 59.1 | 49.8 | 40.8 | 3.096 | — | |
| KosMos-2Size=1.7B, # Samples Seen=7.7M / 25.6M / 25.6M2026.04 | 87.8 | 72.1 | 59.1 | 49.8 | 40.8 | — | 3.096 | |
| LLaVA-Video-7BEvaluation Protocol=VLM with VLM4VLA Models (Frozen Vision)2026.03 | 87.6 | 70.1 | 54.2 | 43.8 | 34 | 2.898 | — | |
| UVA2025.12 | 87.5 | 66.7 | 71.1 | 75.8 | 78.5 | — | — | |
| SuSIEView=Third View2026.05 | 87 | 69 | 49 | 38 | 26 | — | 2.69 | |
| GR-1View=Multi-View2026.05 | 85.4 | 71.2 | 59.6 | 49.7 | 40.1 | — | 3.06 | |
| UWM2025.12 | 81.3 | 73.3 | 64.4 | 57.6 | 71.4 | — | — | |
| OpenVLA*VLM Backbone=Llama-2, Size=7.7B2026.01 | 79.2 | 64.4 | 49.9 | 36.8 | 24.5 | 2.548 | — | |
| OpenVLA*Model (VLM Backbone)=Llama-2, Evaluation Protocol=Expert Vision-Language-Action Models2026.03 | 79.2 | 64.4 | 49.9 | 36.8 | 24.5 | 2.548 | — | |
| OpenVLA (Llama-2)Size=7.7B, # Samples Seen=7.7M / 25.6M / 25.6M2026.04 | 79.2 | 64.4 | 49.9 | 36.8 | 24.5 | — | 2.548 | |
| PAD2025.12 | 78.1 | 46.7 | 48.9 | 48.5 | 64.2 | — | — | |
| Unipi2025.12 | 46.9 | 26.7 | 28.9 | 18.2 | 45.2 | — | — |