3DoF object manipulation trajectory generation on HOT3D
0.2693D ADESeq2Seq
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Seq2SeqInput=Image2025.06 | 0.269 | 0.475 | 0.215 | 0.372 | |
| PointLLM (7B)Input=Point cloud2025.06 | 0.274 | 0.459 | 0.21 | 0.328 | |
| BLIP-2 (2.7B)Input=Image2025.06 | 0.278 | 0.464 | 0.218 | 0.346 | |
| BLIP-2 (6.7B)Input=Image + Depth2025.06 | 0.283 | 0.477 | 0.218 | 0.351 | |
| BLIP-2 (6.7B)Input=Image2025.06 | 0.286 | 0.487 | 0.224 | 0.365 | |
| BLIP-2 (2.7B)Input=Image + Depth2025.06 | 0.288 | 0.481 | 0.227 | 0.363 | |
| VILA (8B)Input=Image + Depth2025.06 | 0.298 | 0.494 | 0.234 | 0.368 | |
| MiniGPT-3D (2.7B)Input=Point cloud2025.06 | 0.299 | 0.487 | 0.236 | 0.368 | |
| VILA (3B)Input=Image + Depth2025.06 | 0.301 | 0.492 | 0.229 | 0.355 | |
| Seq2SeqInput=Image + Depth2025.06 | 0.302 | 0.541 | 0.25 | 0.438 | |
| VILA (8B)Input=Image2025.06 | 0.316 | 0.512 | 0.249 | 0.388 | |
| VILA (3B)Input=Image2025.06 | 0.5 | 0.662 | 0.428 | 0.546 |