Image Classification on ImageNet 1K (val) (Top-1 Accuracy Only)
87.1Top-1 AccuracyBamboo
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| BambooScale=Huge, Pre-train Data=IN1K2022.05 | 87.1 | — | |
| MAEScale=Huge, Pre-train Data=IN1K2022.05 | 86.9 | — | |
| BambooScale=Large, Pre-train Data=IN1K2022.05 | 86.3 | — | |
| MAEScale=Large, Pre-train Data=IN1K2022.05 | 85.9 | — | |
| MaskFeatScale=Large, Pre-train Data=IN1K2022.05 | 85.7 | — | |
| BEiTScale=Large, Pre-train Data=IN1K + DALL-E2022.05 | 85.2 | — | |
| IBOTScale=Large, Pre-train Data=IN1K2022.05 | 84.8 | — | |
| BambooScale=Base, Pre-train Data=IN1K2022.05 | 84.2 | — | |
| MoCo v3Scale=Large, Pre-train Data=IN1K2022.05 | 84.1 | — | |
| MaskFeatScale=Base, Pre-train Data=IN1K2022.05 | 84 | — | |
| IBOTScale=Base, Pre-train Data=IN1K2022.05 | 84 | — | |
| M2DTarget encoder=true, Target input=Masked patches only, Backbone=ViT-Base, Fine-tuning=true, Pre-training epochs=300, Batch size=2048, Masking ratio=0.752022.10 | 83.35 | — | |
| MAEScale=Base, Pre-train Data=IN1K2022.05 | 83.3 | — | |
| MAETarget encoder=N/A, Target input=N/A, Backbone=ViT-Base, Fine-tuning=true, Pre-training epochs=300, Batch size=2048, Masking ratio=0.752022.10 | 83.22 | — | |
| M2D variant (conventional)Target encoder=true, Target input=All patches, Backbone=ViT-Base, Fine-tuning=true, Pre-training epochs=300, Batch size=2048, Masking ratio=0.752022.10 | 83.22 | — | |
| MoCo v3Scale=Base, Pre-train Data=IN1K2022.05 | 83.2 | — | |
| BEiTScale=Base, Pre-train Data=IN1K + DALL-E2022.05 | 83.2 | — | |
| DINOScale=Base, Pre-train Data=IN1K2022.05 | 82.8 | — | |
| ViT (He et al., 2021)Scale=Base2022.05 | 82.1 | — | |
| UniFormer-XSInput Size=224, #Param (M)=16.5, FLOPs (G)=2.0, Throughput (images/s)=35062022.01 | 82 | — | |
| DeiTScale=Base2022.05 | 81.8 | — | |
| EfficientNet-B3Input Size=300, #Param (M)=12.2, FLOPs (G)=1.9, Throughput (images/s)=25682022.01 | 81.6 | — | |
| ViT (He et al., 2021)Scale=Large2022.05 | 81.5 | — | |
| UniFormer-XSInput Size=192, #Param (M)=16.5, FLOPs (G)=1.4, Throughput (images/s)=44922022.01 | 81.5 | — | |
| I-JEPAArch.=ViT-H/16_448, Epochs=300, Resolution=448x448, Evaluation Protocol=Linear-evaluation, Data Augmentation Setting=without view data augmentations2023.01 | 81.1 | — | |
| iBOTArch.=ViT-L/16, Epochs=250, Evaluation Protocol=Linear-evaluation, Data Augmentation Setting=using extra view data augmentations2023.01 | 81 | — | |
| DeepViTScale=Base2022.05 | 80.9 | — | |
| ViT (He et al., 2021)Scale=Huge2022.05 | 80.9 | — | |
| UniFormer-XXSInput Size=224, #Param (M)=10.2, FLOPs (G)=1.3, Throughput (images/s)=44462022.01 | 80.6 | — | |
| GLMCAugmentation=MixUp + CutMix, Backbone=ResNet-502023.05 | 80.2 | — | |
| DINOArch.=ViT-B/8, Epochs=300, Evaluation Protocol=Linear-evaluation, Data Augmentation Setting=using extra view data augmentations2023.01 | 80.1 | — | |
| EfficientNet-B2Input Size=260, #Param (M)=9.1, FLOPs (G)=1.1, Throughput (images/s)=42472022.01 | 80.1 | — | |
| UniFormer-XXSInput Size=192, #Param (M)=10.2, FLOPs (G)=0.96, Throughput (images/s)=57662022.01 | 79.9 | — | |
| I-JEPAArch.=ViT-H/14, Epochs=300, Evaluation Protocol=Linear-evaluation, Data Augmentation Setting=without view data augmentations2023.01 | 79.3 | — | |
| PaCoAugmentation=RandAugment, Backbone=ResNet-502023.05 | 79.3 | — | |
| MobileFormerInput Size=224, #Param (M)=14.0, FLOPs (G)=0.51, Throughput (images/s)=49532022.01 | 79.3 | — | |
| SimCLR v2Arch.=RN152 (2x), Epochs=800, Evaluation Protocol=Linear-evaluation, Data Augmentation Setting=using extra view data augmentations2023.01 | 79.1 | — | |
| UniFormer-XXSInput Size=160, #Param (M)=10.2, FLOPs (G)=0.67, Throughput (images/s)=93822022.01 | 79.1 | — | |
| EfficientNet-B1Input Size=240, #Param (M)=7.8, FLOPs (G)=0.74, Throughput (images/s)=58202022.01 | 79.1 | — | |
| DMST-S#PARAMETERS=22.25M, Configuration=Small2026.01 | 78.98 | — | |
| PaCoAugmentation=Simple Augment, Backbone=ResNet-502023.05 | 78.7 | — | |
| PVTv2-B1Input Size=224, #Param (M)=14.0, FLOPs (G)=2.1, Throughput (images/s)=48122022.01 | 78.7 | — | |
| vanillaAugmentation=CutMix, Backbone=ResNet-502023.05 | 78.6 | — | |
| SupconAugmentation=RandAugment, Backbone=ResNet-502023.05 | 78.4 | — | |
| MobileViT-SInput Size=256, #Param (M)=5.6, FLOPs (G)=2.0, Throughput (images/s)=33602022.01 | 78.3 | — | |
| CAEArch.=ViT-L/16, Epochs=1600, Evaluation Protocol=Linear-evaluation, Data Augmentation Setting=without view data augmentations2023.01 | 78.1 | — | |
| vanillaAugmentation=MixUp, Backbone=ResNet-502023.05 | 77.9 | — | |
| ViT (Dosovitskiy et al., 2020)Scale=Base2022.05 | 77.9 | — | |
| TSSA#PARAMETERS=22.20M, Configuration=Small2026.01 | 77.9 | — | |
| I-JEPAArch.=ViT-L/16, Epochs=600, Evaluation Protocol=Linear-evaluation, Data Augmentation Setting=without view data augmentations2023.01 | 77.5 | — | |
| data2vecArch.=ViT-L/16, Epochs=1600, Evaluation Protocol=Linear-evaluation, Data Augmentation Setting=without view data augmentations2023.01 | 77.3 | — | |
| MAEArch.=ViT-H/14, Epochs=1600, Evaluation Protocol=Linear-evaluation, Data Augmentation Setting=without view data augmentations2023.01 | 77.2 | — | |
| EfficientNet-B0Input Size=224, #Param (M)=5.3, FLOPs (G)=0.42, Throughput (images/s)=95012022.01 | 77.1 | — | |
| UniFormer-XXSInput Size=128, #Param (M)=10.2, FLOPs (G)=0.43, Throughput (images/s)=128862022.01 | 76.8 | — | |
| ViT (Dosovitskiy et al., 2020)Scale=Large2022.05 | 76.5 | — | |
| vanillaAugmentation=Simple Augment, Backbone=ResNet-502023.05 | 76.4 | — | |
| MAEArch.=ViT-L/16, Epochs=1600, Evaluation Protocol=Linear-evaluation, Data Augmentation Setting=without view data augmentations2023.01 | 76 | — | |
| MobileViT-XSInput Size=256, #Param (M)=2.3, FLOPs (G)=1.1, Throughput (images/s)=48222022.01 | 74.6 | — | |
| I-JEPAArch.=ViT-B/16, Epochs=600, Evaluation Protocol=Linear-evaluation, Data Augmentation Setting=without view data augmentations2023.01 | 72.9 | — | |
| PVTv2-B0Input Size=224, #Param (M)=3.7, FLOPs (G)=0.57, Throughput (images/s)=87372022.01 | 70.7 | — | |
| CAEArch.=ViT-B/16, Epochs=1600, Evaluation Protocol=Linear-evaluation, Data Augmentation Setting=without view data augmentations2023.01 | 70.4 | — | |
| MobileViT-XXSInput Size=256, #Param (M)=1.3, FLOPs (G)=0.43, Throughput (images/s)=97422022.01 | 68.9 | — | |
| MAEArch.=ViT-B/16, Epochs=1600, Evaluation Protocol=Linear-evaluation, Data Augmentation Setting=without view data augmentations2023.01 | 68 | — | |
| DMST-T#PARAMETERS=5.58M, Configuration=Tiny2026.01 | 66.87 | — | |
| TSSA#PARAMETERS=5.57M, Configuration=Tiny2026.01 | 65.42 | — | |
| BaselineModel=ResNet-502022.11 | — | 22.62 | |
| BaselineModel=ResNet-1012022.11 | — | 20.91 | |
| RandAugmentModel=ResNet-502022.11 | — | 22.02 | |
| RandAugmentModel=ResNet-1012022.11 | — | 20.39 | |
| Soft AugmentationModel=ResNet-50, softening curve k=22022.11 | — | 21.66 | |
| Soft AugmentationModel=ResNet-101, softening curve k=22022.11 | — | 20.63 | |
| Soft Augmentation + RandAugmentModel=ResNet-50, softening curve k=22022.11 | — | 21.27 | |
| Soft Augmentation + RandAugmentModel=ResNet-101, softening curve k=22022.11 | — | 19.86 |