3D Object Captioning on 3D Object Captioning Dataset
0.1851BLEU-1ShapeLLM-Omni-7B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| ShapeLLM-Omni-7BInput modality=3D latent2026.01 | 0.1851 | 0.2137 | 0.1989 | |
| PointLLM-13BInput modality=3D latent2026.01 | 0.1709 | 0.2099 | 0.1645 | |
| 3D-LLMInput modality=3D latent2026.01 | 0.1691 | 0.1948 | 0.1973 | |
| Ours-2B MOT (CG-MLLM)Input modality=Image2026.01 | 0.1351 | 0.1913 | 0.1428 | |
| InstructBLIP-13BInput modality=Image2026.01 | 0.0465 | 0.0885 | 0.1323 | |
| LLaVA-13BInput modality=Image2026.01 | 0.0402 | 0.0815 | 0.1258 | |
| Qwen3-VL-2BInput modality=Image2026.01 | 0.0313 | 0.0721 | 0.1192 |