6DoF object manipulation trajectory generation on HOT3D
0.2653D Positional ADEEgoFlow
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| EgoFlowZero-shot=true2026.04 | 0.265 | 0.027 | — | — | 1.49 | |
| PointLLM (7B)Input=Point cloud2025.06 | 0.271 | 0.458 | 0.208 | 0.327 | 0.541 | |
| BLIP-2 (2.7B)Input=Image + Depth2025.06 | 0.275 | 0.458 | 0.211 | 0.332 | 0.543 | |
| BLIP-2 (2.7B)Input=Image2025.06 | 0.28 | 0.465 | 0.218 | 0.344 | 0.545 | |
| MiniGPT-3D (2.7B)Input=Point cloud2025.06 | 0.281 | 0.467 | 0.218 | 0.342 | 0.544 | |
| BLIP-2 (6.7B)Input=Image + Depth2025.06 | 0.282 | 0.469 | 0.219 | 0.345 | 0.542 | |
| BLIP-2 (6.7B)Input=Image2025.06 | 0.286 | 0.475 | 0.219 | 0.349 | 0.543 | |
| VILA (8B)Input=Image2025.06 | 0.293 | 0.478 | 0.225 | 0.347 | 0.545 | |
| VILA (3B)Input=Image + Depth2025.06 | 0.294 | 0.48 | 0.223 | 0.344 | 0.541 | |
| GIMOZero-shot=true2026.04 | 0.299 | 0.436 | — | — | 2.06 | |
| VILA (8B)Input=Image + Depth2025.06 | 0.318 | 0.513 | 0.253 | 0.391 | 0.563 | |
| Seq2SeqInput=Image + Depth2025.06 | 0.341 | 0.558 | 0.28 | 0.452 | 0.59 | |
| EgoscalerZero-shot=true2026.04 | 0.351 | 0.54 | — | — | 0.856 | |
| Seq2SeqInput=Image2025.06 | 0.374 | 0.625 | 0.325 | 0.528 | 0.559 | |
| VILA (3B)Input=Image2025.06 | 0.477 | 0.619 | 0.39 | 0.489 | 0.826 | |
| CHOISZero-shot=true2026.04 | 0.513 | 0.571 | — | — | 2.46 | |
| SPOTZero-shot=true2026.04 | 1.018 | 1.082 | — | — | 2.535 | |
| DP3Zero-shot=true2026.04 | 1.019 | 1.096 | — | — | 2.541 | |
| M2DiffuserZero-shot=true2026.04 | 1.079 | 1.157 | — | — | 2.525 |