Object Classification on ScanObjectNN OBJ_BG
98.97AccuracyPointGST
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| PointGSTPre-trained model=PointGPT-L, Params. (M)=2.4, FLOPS (G)=67.952024.10 | 98.97 | — | |
| ReCon++-B*Evaluation Protocol=Full Fine-tune, PP (Post-pretraining)=true, CM (Cross-modal training)=true, #P(M)=177.42025.06 | 98.62 | — | |
| IDPTPre-trained model=PointGPT-L, Params. (M)=10.0, FLOPS (G)=75.192024.10 | 98.11 | — | |
| DAPTPre-trained model=PointGPT-L, Params. (M)=4.2, FLOPS (G)=71.642024.10 | 98.11 | — | |
| Fully fine-tunePre-trained model=PointGPT-L, Params. (M)=360.5, FLOPS (G)=67.712024.10 | 97.2 | — | |
| AsymDSD-S*Evaluation Protocol=Full Fine-tune, PP (Post-pretraining)=false, CM (Cross-modal training)=false, #P(M)=22.12025.06 | 97.07 | — | |
| AsymDSD-B*Evaluation Protocol=Full Fine-tune, PP (Post-pretraining)=false, CM (Cross-modal training)=false, #P(M)=92.12025.06 | 96.73 | — | |
| Point-PEFTPre-trained model=PointGPT-L, Params. (M)=3.1, FLOPS (G)=73.622024.10 | 96.39 | — | |
| AsymDSD-B*Evaluation Protocol=Linear, PP (Post-pretraining)=false, CM (Cross-modal training)=false, #P(M)=91.82025.06 | 96.16 | — | |
| PointGPT-B*Evaluation Protocol=Full Fine-tune, PP (Post-pretraining)=true, CM (Cross-modal training)=false, #P(M)=92.02025.06 | 95.8 | — | |
| PointGPT-LEvaluation Protocol=Full Fine-tune, PP (Post-pretraining)=false, CM (Cross-modal training)=false, #P(M)=310.02025.06 | 95.7 | — | |
| Point-RAE#P(M)=29.2, ST (Standard Transformer)=false, SM (Single-modal training)=true, Evaluation Protocol=Full Fine-tune2025.06 | 95.53 | — | |
| Point-SRAParadigm=Cross-Modal Self-Supervised, Params (M)=40.12026.01 | 95.53 | — | |
| AsymDSD-S*Evaluation Protocol=Linear, PP (Post-pretraining)=false, CM (Cross-modal training)=false, #P(M)=21.82025.06 | 95.22 | — | |
| Point-FEMAE#P(M)=27.4, ST (Standard Transformer)=false, SM (Single-modal training)=true, Evaluation Protocol=Full Fine-tune2025.06 | 95.18 | — | |
| ReCon#P (M)=43.6, #F (G)=5.3, Training Paradigm=with Pretrained Cross-Modal Teacher Representation Learning2023.07 | 95.18 | — | |
| Mamba3DPT=Point-MAE, #P (M)=16.9, #F (G)=3.9, Voting Strategy=true2024.04 | 95.18 | — | |
| Point-FEMAEParadigm=Single-Modal Self-Supervised, Params (M)=41.52026.01 | 95.18 | — | |
| ReConParadigm=Cross-Modal Self-Supervised, Params (M)=44.32026.01 | 95.18 | — | |
| UtoniaParams Learn.=137.4M, Params Pct.=100%, Evaluation Protocol=fine-tuning2026.03 | 95 | — | |
| PointGSTPre-trained model=RECON, Params. (M)=0.6, FLOPS (G)=4.812024.10 | 94.49 | — | |
| Mamba3DPT=×, #P (M)=16.9, #F (G)=3.9, Voting Strategy=true2024.04 | 94.49 | — | |
| 3D-JEPAPre-training Epochs=300, Training Paradigm=SSRL2024.09 | 94.49 | — | |
| AsymDSD-S#P(M)=22.1, ST (Standard Transformer)=true, SM (Single-modal training)=true, Evaluation Protocol=Full Fine-tune2025.06 | 94.32 | — | |
| PointMamba#P(M)=12.3, ST (Standard Transformer)=false, SM (Single-modal training)=true, Evaluation Protocol=Full Fine-tune2025.06 | 94.32 | — | |
| Fully fine-tunePre-trained model=RECON, Params. (M)=22.1, FLOPS (G)=4.762024.10 | 94.32 | — | |
| DAPTPre-trained model=RECON, Params. (M)=1.1, FLOPS (G)=4.962024.10 | 94.32 | — | |
| PointMambaParadigm=Single-Modal Self-Supervised, Params (M)=12.32026.01 | 94.32 | — | |
| Point-MAE-B*Evaluation Protocol=Full Fine-tune, PP (Post-pretraining)=true, CM (Cross-modal training)=false, #P(M)=120.12025.06 | 94.2 | — | |
| ReCon SM#P(M)=43.6, ST (Standard Transformer)=false, SM (Single-modal training)=true, Evaluation Protocol=Full Fine-tune2025.06 | 94.15 | — | |
| I2P-MAE#P (M)=12.9, #F (G)=3.6, Training Paradigm=with Pretrained Cross-Modal Teacher Representation Learning2023.07 | 94.15 | — | |
| RECONPublication=ICML 23, Tunable Params.=43.6 M, Evaluation Protocol=Self-Supervised Representation Learning (Full Fine-Tuning)2025.04 | 94.15 | — | |
| I2P-MAEParadigm=Cross-Modal Self-Supervised, Params (M)=15.32026.01 | 94.15 | — | |
| 3D-JEPAPre-training Epochs=150, Training Paradigm=SSRL2024.09 | 93.8 | — | |
| PointNTPPretrain Framework=fully causal autoregression, Auxiliary Module=none2026.05 | 93.8 | — | |
| Point-MAE† w/ IDPT#TP (M)=1.7, Learning Paradigm=Self-Supervised Representation Learning (IDPT)2023.04 | 93.63 | — | |
| PointGPT-BEvaluation Protocol=Full Fine-tune, PP (Post-pretraining)=false, CM (Cross-modal training)=false, #P(M)=92.02025.06 | 93.6 | — | |
| PointGSTPre-trained model=ACT, Params. (M)=0.6, FLOPS (G)=4.812024.10 | 93.46 | — | |
| Standard TransformerPre-train=PointDif [82]2025.06 | 93.29 | — | |
| Fully fine-tunePre-trained model=ACT, Params. (M)=22.1, FLOPS (G)=4.762024.10 | 93.29 | — | |
| IDPTPre-trained model=RECON, Params. (M)=1.7, FLOPS (G)=7.102024.10 | 93.29 | — | |
| ACT#TP (M)=22.1, Learning Paradigm=Self-Supervised Representation Learning (Full Fine-tuning)2023.04 | 93.29 | — | |
| ACT#P (M)=22.1, #F (G)=4.8, Training Paradigm=with Pretrained Cross-Modal Teacher Representation Learning2023.07 | 93.29 | — | |
| ACTPre-training Epochs=300, Training Paradigm=SSRL2024.09 | 93.29 | — | |
| ACTPublication=ICLR'23, Tunable Params.=22.1 M, Evaluation Protocol=Self-Supervised Representation Learning (Full Fine-Tuning)2025.04 | 93.29 | — | |
| PointDifParadigm=Single-Modal Self-Supervised2026.01 | 93.29 | — | |
| ACTParadigm=Cross-Modal Self-Supervised, Params (M)=22.12026.01 | 93.29 | — | |
| PointDifPretrain Framework=diffusion-based pretraining, Auxiliary Module=conditional point generator2026.05 | 93.29 | — | |
| ACTPretrain Framework=cross-modal teacher-guided masked modeling, Auxiliary Module=teacher autoencoder + transformer decoder2026.05 | 93.29 | — | |
| Point-JEPAParadigm=Single-Modal Self-Supervised2026.01 | 93.2 | — | |
| IDPTPre-trained model=ACT, Params. (M)=1.7, FLOPS (G)=7.102024.10 | 93.12 | — | |
| ACT w/ IDPT#TP (M)=1.7, Learning Paradigm=Self-Supervised Representation Learning (IDPT)2023.04 | 93.12 | — | |
| Mamba3DPT=Point-MAE, #P (M)=16.9, #F (G)=3.9, Voting Strategy=false2024.04 | 93.12 | — | |
| VPP w/ vot.#P (M)=22.1, #F (G)=4.8, Training Paradigm=with Self-Supervised Representation Learning, Voting Strategy=Yes2023.07 | 93.11 | — | |
| Point-MAE†#TP (M)=22.1, Learning Paradigm=Self-Supervised Representation Learning (Full Fine-tuning)2023.04 | 92.94 | — | |
| Mamba3DPT=×, #P (M)=16.9, #F (G)=3.9, Voting Strategy=false2024.04 | 92.94 | — | |
| AsymSD-CLS-S#P(M)=22.1, ST (Standard Transformer)=true, SM (Single-modal training)=true, Evaluation Protocol=Full Fine-tune2025.06 | 92.77 | — | |
| AsymSD-MPM-S#P(M)=22.1, ST (Standard Transformer)=true, SM (Single-modal training)=true, Evaluation Protocol=Full Fine-tune2025.06 | 92.77 | — | |
| VPP w/o vot.#P (M)=22.1, #F (G)=4.8, Training Paradigm=with Self-Supervised Representation Learning, Voting Strategy=No2023.07 | 92.77 | — | |
| Standard TransformerPre-train=UniPre3D2025.06 | 92.6 | — | |
| DAPTPre-trained model=ACT, Params. (M)=1.1, FLOPS (G)=4.962024.10 | 92.6 | — | |
| MVTNPT=×, #P (M)=11.2, #F (G)=43.72024.04 | 92.6 | — | |
| MVTNPretrain Framework=multi-view supervised learning, Auxiliary Module=N/A2026.05 | 92.6 | — | |
| UniPre3D (Std. Transformer)Pretrain Framework=cross-modal Gaussian splatting pretraining, Auxiliary Module=Gaussian predictor + fusion block2026.05 | 92.6 | — | |
| IAE (M2AE)#Params(M)=15.3, GFLOPS=3.6, Evaluation Protocol=FULL SSL2025.12 | 92.5 | — | |
| IAE (M2AE)Pretrain Framework=implicit autoencoder, Auxiliary Module=implicit decoder2026.05 | 92.5 | — | |
| ConcertoParams Learn.=137.4M, Params Pct.=100%, Evaluation Protocol=fine-tuning2026.03 | 92.3 | — | |
| IAE (M2AE)#Params(M)=15.3, GFLOPS=3.6, Evaluation Protocol=FULL SSL, pre-training mesh=w/o mesh2025.12 | 92.3 | — | |
| Mamba3DPT=Point-BERT, #P (M)=16.9, #F (G)=3.9, Voting Strategy=false2024.04 | 92.25 | — | |
| SonataParams Learn.=124.8M, Params Pct.=100%, Evaluation Protocol=fine-tuning2026.03 | 92.2 | — | |
| Point-PEFTPre-trained model=RECON, Params. (M)=0.7, FLOPS (G)=7.612024.10 | 91.91 | — | |
| RI-MAEParadigm=Single-Modal Self-Supervised2026.01 | 91.9 | — | |
| PointGSTPre-trained model=Point-MAE, Params. (M)=0.6, FLOPS (G)=4.812024.10 | 91.74 | — | |
| PointGPT-S#P(M)=22.1, ST (Standard Transformer)=true, SM (Single-modal training)=true, Evaluation Protocol=Full Fine-tune2025.06 | 91.6 | — | |
| PointGPT-SPretrain Framework=autoregressive generation, Auxiliary Module=extractor-generator transformer decoder2026.05 | 91.6 | — | |
| PointGSTPre-trained model=Point-BERT, Params. (M)=0.6, FLOPS (G)=4.812024.10 | 91.39 | — | |
| Point-M2AE#P(M)=15.3, ST (Standard Transformer)=false, SM (Single-modal training)=true, Evaluation Protocol=Full Fine-tune2025.06 | 91.22 | — | |
| IDPTPre-trained model=Point-MAE, Params. (M)=1.7, FLOPS (G)=7.102024.10 | 91.22 | — | |
| Point-M2AELearning Paradigm=Self-Supervised Representation Learning (Full Fine-tuning)2023.04 | 91.22 | — | |
| Point-MAE w/ IDPT#TP (M)=1.7, Learning Paradigm=Self-Supervised Representation Learning (IDPT)2023.04 | 91.22 | — | |
| Point-M2AE#P (M)=14.8, #F (G)=3.6, Training Paradigm=with Self-Supervised Representation Learning2023.07 | 91.22 | — | |
| Point-MAEPT=IDPT, #P (M)=22.1+1.7†, #F (G)=4.82024.04 | 91.22 | — | |
| Point-M2AEPT=Point-M2AE, #P (M)=15.3, #F (G)=3.62024.04 | 91.22 | — | |
| Point-M2AEPre-training Epochs=300, Training Paradigm=SSRL2024.09 | 91.22 | — | |
| Point-M2AEPublication=NeurIPS 22, Tunable Params.=15.3 M, Evaluation Protocol=Self-Supervised Representation Learning (Full Fine-Tuning)2025.04 | 91.22 | — | |
| Point-MAE + IDPTPublication=ICCV'23, Tunable Params.=1.7 M (7.69%), Evaluation Protocol=Self-Supervised Representation Learning (Parameter-Efficient Fine-Tuning)2025.04 | 91.22 | — | |
| Point-M2AEParadigm=Single-Modal Self-Supervised, Params (M)=15.32026.01 | 91.22 | — | |
| Point-M2AEPretrain Framework=masked autoencoder, Auxiliary Module=transformer decoder2026.05 | 91.22 | — | |
| Point-MAE+IDPTPretrain Framework=masked autoencoder, Auxiliary Module=dynamic prompt generator2026.05 | 91.22 | — | |
| Point2VecPre-training Epochs=300, Training Paradigm=SSRL2024.09 | 91.2 | — | |
| Point-M2AE#Params(M)=15.3, GFLOPS=3.6, Evaluation Protocol=FULL SSL2025.12 | 91.2 | — | |
| Point-M2AEParams Learn.=12.9M, Params Pct.=100%, Evaluation Protocol=fine-tuning2026.03 | 91.2 | — | |
| DAPTPre-trained model=Point-BERT, Params. (M)=1.1, FLOPS (G)=4.962024.10 | 91.05 | — | |
| Joint-MAEPretrain Framework=joint masked autoencoding, Auxiliary Module=transformer decoder2026.05 | 90.94 | — | |
| DAPTPre-trained model=Point-MAE, Params. (M)=1.1, FLOPS (G)=4.962024.10 | 90.88 | — | |
| Point-MAE + DAPTPublication=CVPR 24, Tunable Params.=1.1 M (4.97%), Evaluation Protocol=Self-Supervised Representation Learning (Parameter-Efficient Fine-Tuning)2025.04 | 90.88 | — | |
| Point-MAE+DAPTPretrain Framework=masked autoencoder, Auxiliary Module=dynamic adapter + internal prompt2026.05 | 90.88 | — | |
| PointMambaPT=Point-MAE, #P (M)=12.3, #F (G)=3.62024.04 | 90.71 | — | |
| Point-MAE + Point LoRAPublication=Ours, Tunable Params.=0.77 M (3.43%), Evaluation Protocol=Self-Supervised Representation Learning (Parameter-Efficient Fine-Tuning)2025.04 | 90.71 | — | |
| ViPFormerPre-training Epochs=300, Training Paradigm=SSRL2024.09 | 90.7 | — |