Object Classification on ScanObjectNN OBJ_ONLY
97.59Overall AccuracyPointGST
Evaluation Results
| Method | Links | |
|---|---|---|
| PointGSTPre-trained model=PointGPT-L, Params. (M)=2.4, FLOPS (G)=67.952024.10 | 97.59 | |
| Fully fine-tunePre-trained model=PointGPT-L, Params. (M)=360.5, FLOPS (G)=67.712024.10 | 96.6 | |
| ReCon++-B*Evaluation Protocol=Full Fine-tune, PP (Post-pretraining)=true, CM (Cross-modal training)=true, #P(M)=177.42025.06 | 96.21 | |
| DAPTPre-trained model=PointGPT-L, Params. (M)=4.2, FLOPS (G)=71.642024.10 | 96.21 | |
| IDPTPre-trained model=PointGPT-L, Params. (M)=10.0, FLOPS (G)=75.192024.10 | 96.04 | |
| PointGPT-B*Evaluation Protocol=Full Fine-tune, PP (Post-pretraining)=true, CM (Cross-modal training)=false, #P(M)=92.02025.06 | 95.2 | |
| AsymDSD-S*Evaluation Protocol=Full Fine-tune, PP (Post-pretraining)=false, CM (Cross-modal training)=false, #P(M)=22.12025.06 | 94.83 | |
| Point-PEFTPre-trained model=PointGPT-L, Params. (M)=3.1, FLOPS (G)=73.622024.10 | 94.66 | |
| AsymDSD-S*Evaluation Protocol=Linear, PP (Post-pretraining)=false, CM (Cross-modal training)=false, #P(M)=21.82025.06 | 94.51 | |
| AsymDSD-B*Evaluation Protocol=Full Fine-tune, PP (Post-pretraining)=false, CM (Cross-modal training)=false, #P(M)=92.12025.06 | 94.32 | |
| PointGPT-LEvaluation Protocol=Full Fine-tune, PP (Post-pretraining)=false, CM (Cross-modal training)=false, #P(M)=310.02025.06 | 94.1 | |
| Point-MAE-B*Evaluation Protocol=Full Fine-tune, PP (Post-pretraining)=true, CM (Cross-modal training)=false, #P(M)=120.12025.06 | 93.9 | |
| Point-RAE#P(M)=29.2, ST (Standard Transformer)=false, SM (Single-modal training)=true, Evaluation Protocol=Full Fine-tune2025.06 | 93.63 | |
| AsymDSD-B*Evaluation Protocol=Linear, PP (Post-pretraining)=false, CM (Cross-modal training)=false, #P(M)=91.82025.06 | 93.44 | |
| Point-SRAParadigm=Cross-Modal Self-Supervised, Params (M)=40.12026.01 | 93.31 | |
| Point-FEMAE#P(M)=27.4, ST (Standard Transformer)=false, SM (Single-modal training)=true, Evaluation Protocol=Full Fine-tune2025.06 | 93.29 | |
| Point-FEMAEParadigm=Single-Modal Self-Supervised, Params (M)=41.52026.01 | 93.29 | |
| ReConParadigm=Cross-Modal Self-Supervised, Params (M)=44.32026.01 | 93.29 | |
| ReCon SM#P(M)=43.6, ST (Standard Transformer)=false, SM (Single-modal training)=true, Evaluation Protocol=Full Fine-tune2025.06 | 93.12 | |
| Point-MAE† w/ IDPT#TP (M)=1.7, Learning Paradigm=Self-Supervised Representation Learning (IDPT)2023.04 | 93.12 | |
| RECONPublication=ICML 23, Tunable Params.=43.6 M, Evaluation Protocol=Self-Supervised Representation Learning (Full Fine-Tuning)2025.04 | 93.12 | |
| PointGSTPre-trained model=RECON, Params. (M)=0.6, FLOPS (G)=4.812024.10 | 92.94 | |
| Fully fine-tunePre-trained model=RECON, Params. (M)=22.1, FLOPS (G)=4.762024.10 | 92.77 | |
| PointMamba#P(M)=12.3, ST (Standard Transformer)=false, SM (Single-modal training)=true, Evaluation Protocol=Full Fine-tune2025.06 | 92.6 | |
| PointGSTPre-trained model=ACT, Params. (M)=0.6, FLOPS (G)=4.812024.10 | 92.6 | |
| Mamba3DPT=Point-MAE, #P (M)=16.9, #F (G)=3.9, Voting Strategy=true2024.04 | 92.6 | |
| PointMambaParadigm=Single-Modal Self-Supervised, Params (M)=12.32026.01 | 92.6 | |
| PointNTPPretrain Framework=fully causal autoregression, Auxiliary Module=none2026.05 | 92.6 | |
| PointGPT-BEvaluation Protocol=Full Fine-tune, PP (Post-pretraining)=false, CM (Cross-modal training)=false, #P(M)=92.02025.06 | 92.5 | |
| DAPTPre-trained model=RECON, Params. (M)=1.1, FLOPS (G)=4.962024.10 | 92.43 | |
| Mamba3DPT=×, #P (M)=16.9, #F (G)=3.9, Voting Strategy=true2024.04 | 92.43 | |
| MVTNPT=×, #P (M)=11.2, #F (G)=43.72024.04 | 92.3 | |
| MVTNPretrain Framework=multi-view supervised learning, Auxiliary Module=N/A2026.05 | 92.3 | |
| IDPTPre-trained model=ACT, Params. (M)=1.7, FLOPS (G)=7.102024.10 | 92.26 | |
| ACT w/ IDPT#TP (M)=1.7, Learning Paradigm=Self-Supervised Representation Learning (IDPT)2023.04 | 92.26 | |
| Standard TransformerPre-train=UniPre3D2025.06 | 92.08 | |
| Point-MAE†#TP (M)=22.1, Learning Paradigm=Self-Supervised Representation Learning (Full Fine-tuning)2023.04 | 92.08 | |
| Mamba3DPT=×, #P (M)=16.9, #F (G)=3.9, Voting Strategy=false2024.04 | 92.08 | |
| Mamba3DPT=Point-MAE, #P (M)=16.9, #F (G)=3.9, Voting Strategy=false2024.04 | 92.08 | |
| UniPre3D (Std. Transformer)Pretrain Framework=cross-modal Gaussian splatting pretraining, Auxiliary Module=Gaussian predictor + fusion block2026.05 | 92.08 | |
| Standard TransformerPre-train=PointDif [82]2025.06 | 91.91 | |
| AsymDSD-S#P(M)=22.1, ST (Standard Transformer)=true, SM (Single-modal training)=true, Evaluation Protocol=Full Fine-tune2025.06 | 91.91 | |
| Fully fine-tunePre-trained model=ACT, Params. (M)=22.1, FLOPS (G)=4.762024.10 | 91.91 | |
| ACT#TP (M)=22.1, Learning Paradigm=Self-Supervised Representation Learning (Full Fine-tuning)2023.04 | 91.91 | |
| ACTPublication=ICLR'23, Tunable Params.=22.1 M, Evaluation Protocol=Self-Supervised Representation Learning (Full Fine-Tuning)2025.04 | 91.91 | |
| PointDifParadigm=Single-Modal Self-Supervised2026.01 | 91.91 | |
| ACTParadigm=Cross-Modal Self-Supervised, Params (M)=22.12026.01 | 91.91 | |
| PointDifPretrain Framework=diffusion-based pretraining, Auxiliary Module=conditional point generator2026.05 | 91.91 | |
| ACTPretrain Framework=cross-modal teacher-guided masked modeling, Auxiliary Module=teacher autoencoder + transformer decoder2026.05 | 91.91 | |
| Point-JEPAParadigm=Single-Modal Self-Supervised2026.01 | 91.9 | |
| IAE (M2AE)Pretrain Framework=implicit autoencoder, Auxiliary Module=implicit decoder2026.05 | 91.6 | |
| DAPTPre-trained model=ACT, Params. (M)=1.1, FLOPS (G)=4.962024.10 | 91.57 | |
| IDPTPre-trained model=RECON, Params. (M)=1.7, FLOPS (G)=7.102024.10 | 91.57 | |
| I2P-MAEParadigm=Cross-Modal Self-Supervised, Params (M)=15.32026.01 | 91.57 | |
| AsymSD-MPM-S#P(M)=22.1, ST (Standard Transformer)=true, SM (Single-modal training)=true, Evaluation Protocol=Full Fine-tune2025.06 | 91.39 | |
| Mamba3DPT=Point-BERT, #P (M)=16.9, #F (G)=3.9, Voting Strategy=false2024.04 | 91.05 | |
| AsymSD-CLS-S#P(M)=22.1, ST (Standard Transformer)=true, SM (Single-modal training)=true, Evaluation Protocol=Full Fine-tune2025.06 | 90.53 | |
| DAPTPre-trained model=Point-MAE, Params. (M)=1.1, FLOPS (G)=4.962024.10 | 90.19 | |
| PointGSTPre-trained model=Point-MAE, Params. (M)=0.6, FLOPS (G)=4.812024.10 | 90.19 | |
| Point-PEFTPre-trained model=RECON, Params. (M)=0.7, FLOPS (G)=7.612024.10 | 90.19 | |
| Point-MAE + DAPTPublication=CVPR 24, Tunable Params.=1.1 M (4.97%), Evaluation Protocol=Self-Supervised Representation Learning (Parameter-Efficient Fine-Tuning)2025.04 | 90.19 | |
| Point-MAE+DAPTPretrain Framework=masked autoencoder, Auxiliary Module=dynamic adapter + internal prompt2026.05 | 90.19 | |
| IDPTPre-trained model=Point-MAE, Params. (M)=1.7, FLOPS (G)=7.102024.10 | 90.02 | |
| Point-PEFTPre-trained model=ACT, Params. (M)=0.7, FLOPS (G)=7.612024.10 | 90.02 | |
| Point-MAE w/ IDPT#TP (M)=1.7, Learning Paradigm=Self-Supervised Representation Learning (IDPT)2023.04 | 90.02 | |
| Point-MAEPT=IDPT, #P (M)=22.1+1.7†, #F (G)=4.82024.04 | 90.02 | |
| Point-MAE + IDPTPublication=ICCV'23, Tunable Params.=1.7 M (7.69%), Evaluation Protocol=Self-Supervised Representation Learning (Parameter-Efficient Fine-Tuning)2025.04 | 90.02 | |
| Point-MAE+IDPTPretrain Framework=masked autoencoder, Auxiliary Module=dynamic prompt generator2026.05 | 90.02 | |
| PointGPT-S#P(M)=22.1, ST (Standard Transformer)=true, SM (Single-modal training)=true, Evaluation Protocol=Full Fine-tune2025.06 | 90 | |
| PointGPT-SPretrain Framework=autoregressive generation, Auxiliary Module=extractor-generator transformer decoder2026.05 | 90 | |
| GPMPretrain Framework=autoencoding + autoregressive pretraining, Auxiliary Module=dVAE tokenizer2026.05 | 90 | |
| ReCon#P(M)=43.3, ST (Standard Transformer)=false, SM (Single-modal training)=false, Evaluation Protocol=Linear2025.06 | 89.72 | |
| MaskPointPretrain Framework=masked modeling, Auxiliary Module=transformer decoder2026.05 | 89.7 | |
| Point-PEFTPre-trained model=Point-BERT, Params. (M)=0.7, FLOPS (G)=7.612024.10 | 89.67 | |
| DAPTPre-trained model=Point-BERT, Params. (M)=1.1, FLOPS (G)=4.962024.10 | 89.67 | |
| PointGSTPre-trained model=Point-BERT, Params. (M)=0.6, FLOPS (G)=4.812024.10 | 89.67 | |
| Standard TransformerPre-train=TAP [60]2025.06 | 89.5 | |
| TAP (Ours)Backbone=Standard Transformer, Pre-training=Generative2023.07 | 89.5 | |
| TAPPretrain Framework=cross-modal generative pre-training, Auxiliary Module=2D generator2026.05 | 89.5 | |
| Point-MAE + Point LoRAPublication=Ours, Tunable Params.=0.77 M (3.43%), Evaluation Protocol=Self-Supervised Representation Learning (Parameter-Efficient Fine-Tuning)2025.04 | 89.33 | |
| MaskPoint#TP (M)=22.1, Learning Paradigm=Self-Supervised Representation Learning (Full Fine-tuning)2023.04 | 89.3 | |
| MaskPointPublication=ECCV'22, Tunable Params.=22.1 M, Evaluation Protocol=Self-Supervised Representation Learning (Full Fine-Tuning)2025.04 | 89.3 | |
| Point-PEFTPre-trained model=Point-MAE, Params. (M)=0.7, FLOPS (G)=7.612024.10 | 88.98 | |
| Point-MAE + PPT*Publication=arXiv'24, Tunable Params.=1.04 M (4.57%), Evaluation Protocol=Self-Supervised Representation Learning (Parameter-Efficient Fine-Tuning)2025.04 | 88.98 | |
| Joint-MAEPretrain Framework=joint masked autoencoding, Auxiliary Module=transformer decoder2026.05 | 88.86 | |
| Point-M2AE#P(M)=15.3, ST (Standard Transformer)=false, SM (Single-modal training)=true, Evaluation Protocol=Full Fine-tune2025.06 | 88.81 | |
| Point-M2AELearning Paradigm=Self-Supervised Representation Learning (Full Fine-tuning)2023.04 | 88.81 | |
| Point-M2AEPT=Point-M2AE, #P (M)=15.3, #F (G)=3.62024.04 | 88.81 | |
| Point-M2AEPublication=NeurIPS 22, Tunable Params.=15.3 M, Evaluation Protocol=Self-Supervised Representation Learning (Full Fine-Tuning)2025.04 | 88.81 | |
| Point-M2AEParadigm=Single-Modal Self-Supervised, Params (M)=15.32026.01 | 88.81 | |
| Point-M2AEPretrain Framework=masked autoencoder, Auxiliary Module=transformer decoder2026.05 | 88.81 | |
| AsymDSD-S#P(M)=21.8, ST (Standard Transformer)=true, SM (Single-modal training)=true, Evaluation Protocol=Linear2025.06 | 88.73 | |
| SageMixBackbone=PointNet++2022.10 | 88.7 | |
| Standard TransformerPre-train=Point-CMAE [44]2025.06 | 88.64 | |
| PointMambaPT=Point-MAE, #P (M)=12.3, #F (G)=3.62024.04 | 88.47 | |
| AsymSD-CLS-S#P(M)=21.8, ST (Standard Transformer)=true, SM (Single-modal training)=true, Evaluation Protocol=Linear2025.06 | 88.31 | |
| IDPTPre-trained model=Point-BERT, Params. (M)=1.7, FLOPS (G)=7.102024.10 | 88.3 | |
| Point-BERT w/ IDPT#TP (M)=1.7, Learning Paradigm=Self-Supervised Representation Learning (IDPT)2023.04 | 88.3 | |
| Point-BERTPT=IDPT, #P (M)=22.1+1.7†, #F (G)=4.82024.04 | 88.3 | |
| Standard TransformerPre-train=Point-MAE [33]2025.06 | 88.29 |