Image Classification on ImageNet-1K (val) (Top-1 Score)
91Top-1 AccViT-g (CoCa)
Evaluation Results
| Method | Links | |
|---|---|---|
| ViT-g (CoCa)labeled_data_source=private, #param.=1.0B, extra labeled data=JFT-3B+ALIGN, image size=576^22022.11 | 91 | |
| ViT-Glabeled_data_source=private, #param.=1.8B, extra labeled data=JFT-3B, image size=518^22022.11 | 90.5 | |
| SimMIMPre-train Resolution=192x192, Fine-tune Resolution=640x640, Backbone=SwinV2-G, Parameters=3.0B, Pre-training Dataset=ImageNet-22K-ext2021.11 | 90.2 | |
| SwinV2-Glabeled_data_source=private, #param.=3.0B, extra labeled data=IN-21K-ext-70M, image size=640^22022.11 | 90.2 | |
| Florence#param=637M, #data=900M, tuning=fine-tuning (100%), Resolution=>=3842022.06 | 90 | |
| EVAlabeled_data_source=public, #param.=1.0B, extra labeled data=IN-21K (14M), image size=560^22022.11 | 89.7 | |
| BEIT-3labeled_data_source=public, #param.=2.0B, extra labeled data=35M img-txt pairs, image size=336^22022.11 | 89.6 | |
| EVAlabeled_data_source=public, #param.=1.0B, extra labeled data=IN-21K (14M), image size=336^22022.11 | 89.6 | |
| FD-CLIP-Llabeled_data_source=public, #param.=304M, extra labeled data=IN-21K (14M), image size=336^22022.11 | 89 | |
| MViTv2-Hlabeled_data_source=public, #param.=667M, extra labeled data=IN-21K (14M), image size=512^22022.11 | 88.8 | |
| MaxViT-XLlabeled_data_source=public, #param.=475M, extra labeled data=IN-21K (14M), image size=512^22022.11 | 88.7 | |
| ALIGN#param=480M, #data=1.8B, tuning=fine-tuning (100%), Resolution=2892022.06 | 88.6 | |
| CoAtNet-4labeled_data_source=public, #param.=275M, extra labeled data=IN-21K (14M), image size=512^22022.11 | 88.6 | |
| CoCa-B#param=86M, #data=4.8B, tuning=fine-tuning (100%), Resolution=5762022.06 | 88.3 | |
| SWAGBackbone=ViT-L/16, Evaluation Protocol=Fine-tuned, Pre-training Dataset=IG-3.6B2022.12 | 88.07 | |
| EfficientNetV2-L with Poly-1 lossBackbone=EfficientNetV2-L, Loss Function=L_Poly-1, epsilon_1=22022.04 | 87.2 | |
| SimMIMPre-train Resolution=192x192, Fine-tune Resolution=512x512, Backbone=SwinV2-H, Parameters=658M2021.11 | 87.1 | |
| Uni-Perceiver-L + Conditional MoEs#param=303M, #data=44.1M, tuning=fine-tuning (100%), Resolution=3842022.06 | 87 | |
| SoftmaxBackbone=SigLIP-L, Res.=512^2, Params (M)=316.74, FLOPS (G)=312.46, Peak Mem. (GB)=1.3636, Throughput (imgs/s)=40.862026.03 | 86.9 | |
| EfficientNetV2-L with Cross-entropy lossBackbone=EfficientNetV2-L, Loss Function=L_CE2022.04 | 86.8 | |
| SoftmaxBackbone=DINOv2-L, Res.=512^2, Params (M)=304.20, FLOPS (G)=310.60, Peak Mem. (GB)=1.3181, Throughput (imgs/s)=36.522026.03 | 86.8 | |
| Swin-B uparrow 384 (22k)Params (M)=88, MACs (B)=47.1, Input Resolution=384x3842022.04 | 86.4 | |
| SoftmaxBackbone=CLIP-L, Res.=512^2, Params (M)=304.15, FLOPS (G)=310.60, Peak Mem. (GB)=1.3179, Throughput (imgs/s)=36.822026.03 | 86.4 | |
| ViT-AdaLABackbone=SigLIP-L, Res.=512^2, Params (M)=316.74, FLOPS (G)=264.04↓15.5%, Peak Mem. (GB)=1.2620↓7.4%, Throughput (imgs/s)=46.91↑13.9%2026.03 | 86.4 | |
| MixedAEBackbone=ViT-Large, Pre-train Epochs=1600, Pre-train GPU-days=170.42023.03 | 86.2 | |
| MixedAEBackbone=ViT-Large, Pre-train Epochs=1500, Pre-train GPU-days=159.72023.03 | 86 | |
| ViT-AdaLABackbone=DINOv2-L, Res.=512^2, Params (M)=304.20, FLOPS (G)=262.19↓15.6%, Peak Mem. (GB)=1.2163↓7.7%, Throughput (imgs/s)=41.56↑16.1%2026.03 | 86 | |
| MAEBackbone=ViT-L/16, Evaluation Protocol=Fine-tuned2022.12 | 85.95 | |
| MAEBackbone=ViT-Large, Pre-train Epochs=1600, Pre-train GPU-days=151.12023.03 | 85.9 | |
| MAE+LLEBackbone=ViT-L/16, Evaluation Protocol=Fine-tuned2022.12 | 85.84 | |
| SimMIMPre-train Resolution=192x192, Fine-tune Resolution=224x224, Backbone=SwinV2-H, Parameters=658M2021.11 | 85.7 | |
| ViT-L#param=307M, #data=15.5M, tuning=fine-tuning (100%), Resolution=3842022.06 | 85.6 | |
| Mini-Swin-B uparrow 384Params (M)=47, MACs (B)=49.4, Input Resolution=384x3842022.04 | 85.5 | |
| ViT-B#param=86M, #data=15.5M, tuning=fine-tuning (100%), Resolution=3842022.06 | 85.5 | |
| ViT-AdaLABackbone=CLIP-L, Res.=512^2, Params (M)=304.15, FLOPS (G)=262.19↓15.6%, Peak Mem. (GB)=1.2161↓7.4%, Throughput (imgs/s)=42.14↑15.7%2026.03 | 85.5 | |
| SimMIMPre-train Resolution=192x192, Fine-tune Resolution=224x224, Backbone=Swin-L, Parameters=197M2021.11 | 85.4 | |
| SWAGBackbone=ViT-B/16, Evaluation Protocol=Fine-tuned, Pre-training Dataset=IG-3.6B2022.12 | 85.29 | |
| Swin-B (22k)Params (M)=88, MACs (B)=15.4, Input Resolution=224x2242022.04 | 85.2 | |
| TECBase=TEC_IBOT, Epoch=8002022.10 | 85.2 | |
| SWAGBackbone=ViT-L/16, Evaluation Protocol=Linear Probing, Pre-training Dataset=IG-3.6B2022.12 | 85.13 | |
| TECBase=iBOT, Epoch=8002022.10 | 85.1 | |
| iBOTBackbone=ViT-Large, Pre-train Epochs=1000*, Pre-train GPU-days=2852023.03 | 85 | |
| OFA#param=472M, #data=60.6M, tuning=fine-tuning (100%), Resolution=4802022.06 | 84.9 | |
| GG-SSM-BParams (M)=87, FLOPs (G)=14.1, Input size=224x2242024.12 | 84.9 | |
| ViT-AdaLA (Stage 2)Backbone=SigLIP-L, Res.=512^2, Params (M)=316.74, FLOPS (G)=264.04↓15.5%, Peak Mem. (GB)=1.2620↓7.4%, Throughput (imgs/s)=46.91↑13.9%2026.03 | 84.9 | |
| RDNet-LParam=186, FLOPS=34.72024.03 | 84.8 | |
| Mini-DeiT-B uparrow 384Params (M)=44, MACs (B)=56.9, Input Resolution=384x3842022.04 | 84.7 | |
| DeiT-B uparrow 384Params (M)=88, MACs (B)=55.7, Input Resolution=384x3842022.04 | 84.5 | |
| Swin-B uparrow 384Params (M)=88, MACs (B)=47.1, Input Resolution=384x3842022.04 | 84.5 | |
| ViT-AdaLA (Stage 2)Backbone=DINOv2-L, Res.=512^2, Params (M)=304.20, FLOPS (G)=262.19↓15.6%, Peak Mem. (GB)=1.2163↓7.7%, Throughput (imgs/s)=41.56↑16.1%2026.03 | 84.5 | |
| RDNet-BParam=87, FLOPS=15.42024.03 | 84.4 | |
| GG-SSM-SParams (M)=49, FLOPs (G)=6.6, Input size=224x2242024.12 | 84.4 | |
| Mini-Swin-BParams (M)=46, MACs (B)=15.7, Input Resolution=224x2242022.04 | 84.3 | |
| HorNet-BParam=87, FLOPS=15.62024.03 | 84.3 | |
| NAT-BParam=90, FLOPS=13.72024.03 | 84.3 | |
| ConvNeXt-LParam=198, FLOPS=34.42024.03 | 84.3 | |
| iBOTEpoch=16002022.10 | 84.1 | |
| SimMIMPre-train Resolution=192x192, Fine-tune Resolution=224x224, Backbone=Swin-B, Parameters=88M2021.11 | 84 | |
| ViT-B#param=86M, #data=15.5M, tuning=fine-tuning (100%)2022.06 | 84 | |
| ViT-L#param=307M, #data=15.5M, tuning=fine-tuning (100%)2022.06 | 84 | |
| HorNet-SParam=50, FLOPS=8.82024.03 | 84 | |
| SLaK-BParam=95, FLOPS=17.12024.03 | 84 | |
| VMamba-BParams (M)=89, FLOPs (G)=15.4, Input size=224x2242024.12 | 83.9 | |
| VMamba (baseline)Model Scale=base, GFlops=15.36, Training Setting=training-free2026.06 | 83.9 | |
| SLaK-SParam=55, FLOPS=9.82024.03 | 83.8 | |
| ConvNeXt-BParam=89, FLOPS=15.42024.03 | 83.8 | |
| HiViT-BParams (M)=66, FLOPs (G)=15.9, Input size=224x2242024.12 | 83.8 | |
| ConvNeXt-BParams (M)=89, FLOPs (G)=15.4, Input size=224x2242024.12 | 83.8 | |
| MAEBackbone=ViT-B/16, Evaluation Protocol=Fine-tuned2022.12 | 83.72 | |
| Florence#param=637M, #data=900M, tuning=without tuning, Resolution=3842022.06 | 83.7 | |
| NAT-SParam=51, FLOPS=7.82024.03 | 83.7 | |
| RDNet-SParam=50, FLOPS=8.72024.03 | 83.7 | |
| VMamba (baseline)Model Scale=small, GFlops=8.72, Training Setting=training-free2026.06 | 83.7 | |
| MAE+LLEBackbone=ViT-B/16, Evaluation Protocol=Fine-tuned2022.12 | 83.68 | |
| EfficientNet-B5Params (M)=30, MACs (B)=9.9, Input Resolution=456x4562022.04 | 83.6 | |
| Mini-Swin-SParams (M)=26, MACs (B)=8.9, Input Resolution=224x2242022.04 | 83.6 | |
| GG-SSM-TParams (M)=28, FLOPs (G)=4.4, Input size=224x2242024.12 | 83.6 | |
| VMamba-SParams (M)=50, FLOPs (G)=8.7, Input size=224x2242024.12 | 83.6 | |
| Swin-B#Param=88M, #FLOPs=15.4G, Input resolution=224x2242022.03 | 83.5 | |
| SNN-MLP-B#Param=88M, #FLOPs=15.2G, Input resolution=224x2242022.03 | 83.5 | |
| Swin-BParams (M)=88, MACs (B)=15.4, Input Resolution=224x2242022.04 | 83.5 | |
| SupervisedPre-train Resolution=192x192, Fine-tune Resolution=224x224, Backbone=Swin-L, Parameters=197M2021.11 | 83.5 | |
| Swin-BParam=88, FLOPS=15.42024.03 | 83.5 | |
| Swin-BParams (M)=88, FLOPs (G)=15.4, Input size=224x2242024.12 | 83.5 | |
| ViT-AdaLA (Stage 2)Backbone=CLIP-L, Res.=512^2, Params (M)=304.15, FLOPS (G)=262.19↓15.6%, Peak Mem. (GB)=1.2161↓7.4%, Throughput (imgs/s)=42.14↑15.7%2026.03 | 83.4 | |
| SEERBackbone=RG-32gf, Evaluation Protocol=Fine-tuned, Pre-training Dataset=IG-1B2022.12 | 83.35 | |
| SNN-MLP-S#Param=50M, #FLOPs=8.5G, Input resolution=224x2242022.03 | 83.3 | |
| AS-MLP-B#Param=88M, #FLOPs=15.2G, Input resolution=224x2242022.03 | 83.3 | |
| SupervisedPre-train Resolution=192x192, Fine-tune Resolution=224x224, Backbone=Swin-B, Parameters=88M2021.11 | 83.3 | |
| SupervisedPre-train Resolution=192x192, Fine-tune Resolution=224x224, Backbone=SwinV2-H, Parameters=658M2021.11 | 83.3 | |
| VMamba + STORM (ToMe)Model Scale=base, GFlops=11.00, Training Setting=training-free2026.06 | 83.3 | |
| ViP-Large/7#Param=88M, Input resolution=224x2242022.03 | 83.2 | |
| CycleMLP-B5#Param=76M, #FLOPs=12.3G, Input resolution=224x2242022.03 | 83.2 | |
| Mini-DeiT-BParams (M)=44, MACs (B)=17.7, Input Resolution=224x2242022.04 | 83.2 | |
| Swin-SParams (M)=50, MACs (B)=8.7, Input Resolution=224x2242022.04 | 83.2 | |
| NAT-TParam=28, FLOPS=4.32024.03 | 83.2 | |
| AS-MLP-S#Param=50M, #FLOPs=8.5G, Input resolution=224x2242022.03 | 83.1 | |
| DeiT-B#param=86M, #data=1.28M, tuning=fine-tuning (100%), Resolution=3842022.06 | 83.1 | |
| ConvNeXt-SParam=50, FLOPS=8.72024.03 | 83.1 | |
| Swin-S#Param=50M, #FLOPs=8.7G, Input resolution=224x2242022.03 | 83 |