3D Object Captioning on Objaverse LVIS (test)
100S-BERT ScoreHuman
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| HumanInput=N/A2025.05 | 100 | 100 | 100 | 100 | 100 | |
| SVL-13B*Vision Encoder=Spike-driven PointFormer-L, LLM=SpikeLLM (Xing et al., 2025), Input=Point, Prompting=shorter captions (≤ 20 words)2025.05 | 51.21 | 50.18 | 18.45 | 21.32 | 18.4 | |
| PointLLM-13B*Vision Encoder=PointBert (Yu et al., 2022), LLM=Vicuna (Chiang et al., 2023), Input=Point, Prompting=shorter captions (≤ 20 words)2025.05 | 50.15 | 50.83 | 17.09 | 20.99 | 16.45 | |
| PointLLM-13BVision Encoder=PointBert (Yu et al., 2022), LLM=Vicuna (Chiang et al., 2023), Input=Point2025.05 | 47.91 | 49.12 | 3.83 | 7.23 | 12.26 | |
| SVL-13B*Vision Encoder=Spike-driven PointFormer-L, LLM=Vicuna (Chiang et al., 2023), Input=Point, Prompting=shorter captions (≤ 20 words)2025.05 | 47.8 | 47.08 | 11.45 | 14.69 | 16.4 | |
| LLaVA-13BVision Encoder=ViT (Dosovitskiy, 2020), LLM=Vicuna (Chiang et al., 2023), Input=Image2025.05 | 46.37 | 45.9 | 4.02 | 8.15 | 12.58 | |
| InstructBLIP-13BVision Encoder=ViT (Dosovitskiy, 2020), LLM=Vicuna (Chiang et al., 2023), Input=Image2025.05 | 45.9 | 48.86 | 4.65 | 8.85 | 13.23 | |
| SVL-13BVision Encoder=Spike-driven PointFormer-L, LLM=Vicuna (Chiang et al., 2023), Input=Point2025.05 | 44.87 | 45.91 | 3.77 | 6.85 | 12.25 |