Visual Question Answering on Open3D-VQA (test)
63.5Judge AccuracyImage-only Adapter
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Image-only Adaptermodality=image2026.06 | 63.5 | 1.64 | 1 | |
| Depth-only Adaptermodality=depth2026.06 | 54 | 1.58 | 1 | |
| MASER routerselection=top-12026.06 | 47 | 1.56 | 1 | |
| MASERmode=cascade2026.06 | 47 | 1.57 | 1 | |
| Text-only Adaptermodality=text2026.06 | 44 | 1.87 | 1 | |
| Pointcloud-only Adaptermodality=pointcloud2026.06 | 44 | 1.88 | 1 | |
| Pose-only Adaptermodality=pose2026.06 | 40 | 1.68 | 1 | |
| Baseline VLMadapter=none2026.06 | 39 | 1.88 | 0 |