3D Classification on Objaverse LVIS
59.5Top-1 AccTAMM
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| TAMMPre-Training Dataset=Ensembled, Evaluation Protocol=Linear Probing2024.02 | 59.5 | — | — | |
| HOLA-PointBERT#Params=72.1M, FLOPs=84G, FPS=152, Training data=Ensembled (875,665 triplets), Evaluation protocol=Zero-shot2026.05 | 57.3 | 84.9 | 79 | |
| HOLA-PointBERT#Params=32.3M, FLOPs=29G, FPS=202, Training data=Ensembled (875,665 triplets), Evaluation protocol=Zero-shot2026.05 | 56.7 | 84.4 | 78.4 | |
| 3DAlign-DAERZero-shot=true2025.11 | 55.8 | 83.1 | 77 | |
| HOLA-PointBERT#Params=26.0M, FLOPs=7G, FPS=264, Training data=Ensembled (875,665 triplets), Evaluation protocol=Zero-shot2026.05 | 55.7 | 83.6 | 77.3 | |
| UNI3D-G#Params=1020.0M, FLOPs=1130G, FPS=34, Training data=Ensembled (875,665 triplets), Evaluation protocol=Zero-shot2026.05 | 55.3 | 82.9 | 76.7 | |
| Uni3D-gVenue=ICLR 2024, Zero-shot=true2025.11 | 54.2 | 81.9 | 76.1 | |
| RECON++-L#Params=658.9M, FLOPs=423G, FPS=34, Training data=Ensembled (875,665 triplets), Evaluation protocol=Zero-shot2026.05 | 53.7 | 82 | 75.8 | |
| RECON++-B#Params=201.5M, FLOPs=137G, FPS=90, Training data=Ensembled (875,665 triplets), Evaluation protocol=Zero-shot2026.05 | 53.2 | 81.5 | 75.3 | |
| ReCon++-L (shapeLLM)Venue=ECCV 2024, Zero-shot=true2025.11 | 53.1 | 81.6 | 75.2 | |
| Uni3D-LVenue=ICLR 2024, Zero-shot=true2025.11 | 53 | 81.4 | 75.2 | |
| VIT-LENS-G#Params=1543.0M, FLOPs=806G, FPS=32, Training data=Ensembled (875,665 triplets), Evaluation protocol=Zero-shot2026.05 | 52 | 79.9 | 73.3 | |
| Uni3D-BVenue=ICLR 2024, Zero-shot=true2025.11 | 51.6 | 80.9 | 74 | |
| HOLA-PointBERT#Params=72.1M, FLOPs=84G, FPS=152, Training data=Ensembled no LVIS (829,460 triplets), Evaluation protocol=Zero-shot2026.05 | 50.7 | 78.9 | 72.1 | |
| TAMM-PointBERT#Params=35.4M, FLOPs=29G, FPS=85, Training data=Ensembled (875,665 triplets), Evaluation protocol=Zero-shot2026.05 | 50.7 | 80.6 | 73.2 | |
| Point-BERT (ULIP-2)Pre-train dataset=Objaverse + ShapeNet, Pre-train method=ULIP-2, Manual captions?=false, Zero-shot=true2023.05 | 50.6 | 79.1 | — | |
| MixCon3D-PointBERT#Params=30.9M, FLOPs=7G, FPS=153, Training data=Ensembled (875,665 triplets), Evaluation protocol=Zero-shot2026.05 | 50.4 | 79.1 | 72.2 | |
| HOLA-PointBERT#Params=32.3M, FLOPs=29G, FPS=202, Training data=Ensembled no LVIS (829,460 triplets), Evaluation protocol=Zero-shot2026.05 | 50.3 | 78 | 71.4 | |
| VIT-LENS-G#Params=2000.0M, FLOPs=1050G, FPS=27, Training data=Ensembled no LVIS (829,460 triplets), Evaluation protocol=Zero-shot2026.05 | 50.1 | 78.1 | 71.3 | |
| OpenShapePre-Training Dataset=Ensembled, Evaluation Protocol=Linear Probing2024.02 | 48.3 | — | — | |
| UNI3D-G#Params=1020.0M, FLOPs=1130G, FPS=34, Training data=Ensembled no LVIS (829,460 triplets), Evaluation protocol=Zero-shot2026.05 | 47.2 | 76.1 | 68.8 | |
| HOLA-SparseConvPre-training regime=Ensembled (875,665 triplets), Backbone=SparseConv (MinkowskiNet), Zero-shot=true2026.05 | 47.2 | 76.3 | 69.3 | |
| Point-BERT (OpenShape)Pre-train dataset=Objaverse + ShapeNet + (2 extra), Pre-train method=OpenShape, Manual captions?=true, Zero-shot=true2023.05 | 46.8 | 77 | — | |
| OpenShape-PointBERT#Params=32.3M, FLOPs=29G, FPS=86, Training data=Ensembled (875,665 triplets), Evaluation protocol=Zero-shot2026.05 | 46.8 | 77 | 69.1 | |
| OpenShape-PointBERTVenue=NeurIPS 2023, Zero-shot=true2025.11 | 46.7 | 77.1 | 69 | |
| Point-BERT (OpenShape)Pre-train dataset=Objaverse + ShapeNet, Pre-train method=OpenShape, Manual captions?=true, Zero-shot=true2023.05 | 46.5 | 76.3 | — | |
| Point-BERT (ULIP-2)Pre-train dataset=Objaverse(no LVIS) + ShapeNet, Pre-train method=ULIP-2, Manual captions?=false, Zero-shot=true2023.05 | 46.3 | 75 | — | |
| TAMM-SparseConvPre-training regime=Ensembled (875,665 triplets), Backbone=SparseConv (MinkowskiNet), Zero-shot=true2026.05 | 43.8 | 74.1 | 66.2 | |
| OpenShape-SparseConvPre-training regime=Ensembled (875,665 triplets), Backbone=SparseConv (MinkowskiNet), Zero-shot=true2026.05 | 43.4 | 72.4 | 64.8 | |
| HOLA-SparseConvPre-training regime=Ensembled no LVIS (829,460 triplets), Backbone=SparseConv (MinkowskiNet), Zero-shot=true2026.05 | 42.9 | 72.9 | 65.2 | |
| TAMM-PointBERT#Params=35.4M, FLOPs=29G, FPS=85, Training data=Ensembled no LVIS (829,460 triplets), Evaluation protocol=Zero-shot2026.05 | 42 | 71.7 | 63.6 | |
| PPTCross-dataset generalization=true2026.04 | 41 | — | — | |
| TAMM-SparseConvPre-training regime=Ensembled no LVIS (829,460 triplets), Backbone=SparseConv (MinkowskiNet), Zero-shot=true2026.05 | 39.8 | 70.4 | 62 | |
| P3TCross-dataset generalization=true2026.04 | 39.6 | — | — | |
| Point-PRCCross-dataset generalization=true2026.04 | 39.3 | — | — | |
| TAMMPre-Training Dataset=ShapeNet, Evaluation Protocol=Linear Probing2024.02 | 39.1 | — | — | |
| OpenShape-PointBERT#Params=32.3M, FLOPs=29G, FPS=86, Training data=Ensembled no LVIS (829,460 triplets), Evaluation protocol=Zero-shot2026.05 | 39.1 | 68.9 | 60.8 | |
| Point-BERT (OpenShape)Pre-train dataset=Objaverse(no LVIS) + ShapeNet, Pre-train method=OpenShape, Manual captions?=true, Zero-shot=true2023.05 | 38.8 | 68.8 | — | |
| ULIP-2Cross-dataset generalization=true2026.04 | 38.4 | — | — | |
| OpenShape-SparseConvPre-training regime=Ensembled no LVIS (829,460 triplets), Backbone=SparseConv (MinkowskiNet), Zero-shot=true2026.05 | 37 | 66.9 | 58.4 | |
| Point-BERT (ULIP)Pre-train dataset=Objaverse + ShapeNet, Pre-train method=ULIP, Manual captions?=true, Zero-shot=true2023.05 | 34.9 | 61 | — | |
| ULIPPre-Training Dataset=ShapeNet, Evaluation Protocol=Linear Probing2024.02 | 34.6 | — | — | |
| OpenShapePre-Training Dataset=ShapeNet, Evaluation Protocol=Linear Probing2024.02 | 29.3 | — | — | |
| ULIP-2Venue=CVPR 2024, Zero-shot=true2025.11 | 26.7 | 52.5 | 44.9 | |
| Point-BERT (ULIP)Pre-train dataset=Objaverse(no LVIS) + ShapeNet, Pre-train method=ULIP, Manual captions?=true, Zero-shot=true2023.05 | 21.4 | 41.9 | — | |
| ULIP-2 (ZS)Zero-shot setting=true, Cross-dataset generalization=true2026.04 | 18.1 | — | — | |
| HOLA-PointBERT#Params=32.3M, FLOPs=29G, FPS=202, Training data=ShapeNet (52,470 triplets), Evaluation protocol=Zero-shot2026.05 | 17.7 | 36.1 | 30 | |
| HOLA-SparseConvPre-training regime=ShapeNet (52,470 triplets), Backbone=SparseConv (MinkowskiNet), Zero-shot=true2026.05 | 17.5 | 36.1 | 30 | |
| ULIP-2Pre-train dataset=ShapeNet, Pre-train method=ULIP-2, Manual captions?=false, Zero-shot=true2023.05 | 16.4 | 34.3 | — | |
| TAMM-PointBERT#Params=35.4M, FLOPs=29G, FPS=85, Training data=ShapeNet (52,470 triplets), Evaluation protocol=Zero-shot2026.05 | 13.7 | 29.2 | 24.2 | |
| TAMM-SparseConvPre-training regime=ShapeNet (52,470 triplets), Backbone=SparseConv (MinkowskiNet), Zero-shot=true2026.05 | 13.6 | 29.3 | 24.2 | |
| OpenShape-SparseConvPre-training regime=ShapeNet (52,470 triplets), Backbone=SparseConv (MinkowskiNet), Zero-shot=true2026.05 | 11.6 | 27.1 | 21.8 | |
| Point-BERT (OpenShape)Pre-train dataset=ShapeNet, Pre-train method=OpenShape, Manual captions?=true, Zero-shot=true2023.05 | 10.8 | 25 | — | |
| OpenShape-PointBERT#Params=5.1M, FLOPs=1G, FPS=240, Training data=ShapeNet (52,470 triplets), Evaluation protocol=Zero-shot2026.05 | 10.8 | 25 | 20.2 | |
| ULIPVenue=CVPR 2023, Zero-shot=true2025.11 | 6.1 | 17.8 | 13.5 | |
| PointCLIPv2Zero-shot=true2023.05 | 4.7 | 12.9 | — | |
| CLIP2PointPre-train dataset=ShapeNet, Pre-train method=CLIP2Point, Manual captions?=false, Zero-shot=true2023.05 | 2.7 | 7.9 | — | |
| ULIPPre-train dataset=ShapeNet, Pre-train method=ULIP, Manual captions?=true, Zero-shot=true2023.05 | 2.6 | 8.1 | — | |
| PointCLIPZero-shot=true2023.05 | 1.9 | 5.8 | — | |
| PointCLIPVenue=CVPR 2022, Zero-shot=true2025.11 | 1.8 | 5.9 | 4 | |
| ReConPre-train dataset=ShapeNet, Pre-train method=ReCon, Zero-shot=true2023.05 | 1.1 | 3.7 | — |