Image Classification on ImageNet-1K 1.0 (test) (Top-1 Accuracy)
87.1Top-1 AccuracydBOT
Evaluation Results
| Method | Links | |
|---|---|---|
| dBOTBackbone=ViT-H/14, Evaluation Protocol=Full fine-tuning2024.02 | 87.1 | |
| dBOT-Ref.Backbone=ViT-H/14, Evaluation Protocol=Full fine-tuning2024.02 | 87.1 | |
| D2V2-Ref.Backbone=ViT-H/14, Evaluation Protocol=Full fine-tuning2024.02 | 86.8 | |
| D2V2-Ref.Backbone=ViT-L/16, Evaluation Protocol=Full fine-tuning2024.02 | 86.7 | |
| D2V2Backbone=ViT-L/16, Evaluation Protocol=Full fine-tuning2024.02 | 86.6 | |
| D2V2Backbone=ViT-H/14, Evaluation Protocol=Full fine-tuning2024.02 | 86.6 | |
| NFNet-F5Input Resolution=544x544, Parameters=377M, FLOPs=289.8B2021.06 | 86 | |
| CoAtNet-3Input Resolution=512x512, Parameters=168M, FLOPs=203.1B2021.06 | 86 | |
| dBOT-Ref.Backbone=ViT-L/16, Evaluation Protocol=Full fine-tuning2024.02 | 85.9 | |
| NFNet-F4Input Resolution=512x512, Parameters=316M, FLOPs=215.2B2021.06 | 85.9 | |
| CoAtNet-2Input Resolution=512x512, Parameters=75M, FLOPs=96.7B2021.06 | 85.9 | |
| dBOTBackbone=ViT-L/16, Evaluation Protocol=Full fine-tuning2024.02 | 85.8 | |
| CoAtNet-3Input Resolution=384x384, Parameters=168M, FLOPs=107.4B2021.06 | 85.8 | |
| NFNet-F3Input Resolution=416x416, Parameters=255M, FLOPs=114.8B2021.06 | 85.7 | |
| ENetV2-LInput Resolution=480x480, Parameters=121M, FLOPs=53B2021.06 | 85.7 | |
| CoAtNet-2Input Resolution=384x384, Parameters=75M, FLOPs=49.8B2021.06 | 85.7 | |
| MugsBackbone=ViT-L/16, Evaluation Protocol=Full fine-tuning2024.02 | 85.2 | |
| NFNet-F2Input Resolution=352x352, Parameters=194M, FLOPs=62.6B2021.06 | 85.1 | |
| ENetV2-MInput Resolution=480x480, Parameters=55M, FLOPs=24B2021.06 | 85.1 | |
| CoAtNet-1Input Resolution=384x384, Parameters=42M, FLOPs=27.4B2021.06 | 85.1 | |
| SuperClassPreTraining data=Datacomp-1B, #Seen Samples=13B, Backbone=ViT-Large, Evaluation Protocol=Linear probing2024.11 | 85 | |
| CaiT-S-36Input Resolution=384x384, Parameters=68M, FLOPs=48.0B2021.06 | 85 | |
| I-JEPABackbone=ViT-H/14, Evaluation Protocol=Full fine-tuning2024.02 | 84.9 | |
| iBOTBackbone=ViT-L/16, Evaluation Protocol=Full fine-tuning2024.02 | 84.8 | |
| LambdaResNet-420Input Resolution=320x3202021.06 | 84.8 | |
| NFNet-F1Input Resolution=320x320, Parameters=133M, FLOPs=35.5B2021.06 | 84.7 | |
| BotNet-T7Input Resolution=384x384, Parameters=75.1M, FLOPs=45.80B2021.06 | 84.7 | |
| DINOv2PreTraining data=LVD-142M, #Seen Samples=2B, Backbone=ViT-Large, Evaluation Protocol=Linear probing2024.11 | 84.5 | |
| CaiT-M-24Input Resolution=384x384, Parameters=185.9M, FLOPs=116.1B2021.06 | 84.5 | |
| CoAtNet-3Input Resolution=224x224, Parameters=168M, FLOPs=34.7B2021.06 | 84.5 | |
| ResNet-RS-420Input Resolution=320x320, Parameters=192M, FLOPs=128B2021.06 | 84.4 | |
| CaiT-S-24Input Resolution=384x384, Parameters=46.9M, FLOPs=32.2B2021.06 | 84.3 | |
| Swin-BInput Resolution=384x384, Parameters=88M, FLOPs=47.0B2021.06 | 84.2 | |
| CoAtNet-2Input Resolution=224x224, Parameters=75M, FLOPs=15.7B2021.06 | 84.1 | |
| CoAtNet-2#Param.=75M, Image Size=224^2, FLOPs (G)=15.72025.11 | 84.1 | |
| OpenCLIPPreTraining data=Datacomp-1B, #Seen Samples=13B, Backbone=ViT-Large, Evaluation Protocol=Linear probing2024.11 | 83.9 | |
| ENetV2-SInput Resolution=384x384, Parameters=24M, FLOPs=8.8B2021.06 | 83.9 | |
| CoAtNet-0Input Resolution=384x384, Parameters=25M, FLOPs=13.4B2021.06 | 83.9 | |
| Derfmodel=ViT-L2025.12 | 83.8 | |
| MobileMamba-B4FLOPs=4313, Params=17.1, Resolution=512x512, Training Strategy=true2024.11 | 83.6 | |
| DyTmodel=ViT-L2025.12 | 83.6 | |
| NFNet-F0Input Resolution=256x256, Parameters=72M, FLOPs=12.4B2021.06 | 83.6 | |
| Swin-B#Param.=88M, Image Size=224^2, FLOPs (G)=15.42025.11 | 83.5 | |
| CaiT-M-24Input Resolution=224x224, Parameters=185.9M, FLOPs=36.0B2021.06 | 83.4 | |
| MobileMamba-B2FLOPs=2427, Params=17.1, Resolution=384x384, Training Strategy=true2024.11 | 83.3 | |
| CaiT-S-36Input Resolution=224x224, Parameters=68.2M, FLOPs=13.9B2021.06 | 83.3 | |
| Swin-BInput Resolution=224x224, Parameters=88M, FLOPs=15.4B2021.06 | 83.3 | |
| CvT-21Input Resolution=384x384, Parameters=32M, FLOPs=24.9B2021.06 | 83.3 | |
| CoAtNet-1Input Resolution=224x224, Parameters=42M, FLOPs=8.4B2021.06 | 83.3 | |
| CoAtNet-1#Param.=42M, Image Size=224^2, FLOPs (G)=8.42025.11 | 83.3 | |
| LNmodel=ViT-L2025.12 | 83.1 | |
| GNmodel=ViT-L2025.12 | 83.1 | |
| DeiT-BInput Resolution=384x384, Parameters=86M, FLOPs=55.4B2021.06 | 83.1 | |
| DeepViT-LInput Resolution=224x224, Parameters=55M, FLOPs=12.5B2021.06 | 83.1 | |
| CappaPreTraining data=WebLI-1B, #Seen Samples=9B, Backbone=ViT-Large, Evaluation Protocol=Linear probing2024.11 | 83 | |
| RMSNormmodel=ViT-L2025.12 | 83 | |
| ResNet-RS-152Input Resolution=256x256, Parameters=87M, FLOPs=31B2021.06 | 83 | |
| Swin-SInput Resolution=224x224, Parameters=50M, FLOPs=8.7B2021.06 | 83 | |
| CvT-13Input Resolution=384x384, Parameters=20M, FLOPs=16.3B2021.06 | 83 | |
| Swin-S#Param.=50M, Image Size=224^2, FLOPs (G)=8.72025.11 | 83 | |
| Derfmodel=ViT-B2025.12 | 82.8 | |
| Openai CLIPPreTraining data=WIT-400M, #Seen Samples=13B, Backbone=ViT-Large, Evaluation Protocol=Linear probing2024.11 | 82.7 | |
| CaiT-S-24Input Resolution=224x224, Parameters=46.9M, FLOPs=9.4B2021.06 | 82.7 | |
| SuperClassPreTraining data=Datacomp-1B, #Seen Samples=1B, Backbone=ViT-Large, Evaluation Protocol=Linear probing2024.11 | 82.6 | |
| T2T-ViT-24Input Resolution=224x224, Parameters=64.1M, FLOPs=15.0B2021.06 | 82.6 | |
| MobileMamba-B4FLOPs=4313, Params=17.1, Resolution=512x512, Training Strategy=false2024.11 | 82.5 | |
| DyTmodel=ViT-B2025.12 | 82.5 | |
| GNmodel=ViT-B2025.12 | 82.5 | |
| CvT-21Input Resolution=224x224, Parameters=32M, FLOPs=7.1B2021.06 | 82.5 | |
| CvT-21#Param.=32M, Image Size=224^2, FLOPs (G)=7.12025.11 | 82.5 | |
| FTerViTW=2, A=8, Model=DeiT-III-S384, Size (MB)=6.09, Compression=14.6×, Regime=QAD, Epochs=2602026.05 | 82.43 | |
| ViL-BFLOPs=18600, Params=89.0, Resolution=224x224, Training Strategy=false2024.11 | 82.4 | |
| RMSNormmodel=ViT-B2025.12 | 82.4 | |
| DTTN†-LParams (M)=35.9, FLOPs(B)=12.3, Epoch=300, Activation=-, Attention=×, Reso.=22422025.02 | 82.4 | |
| ConViT-B#Param.=86M, Image Size=224^2, FLOPs (G)=17.02025.11 | 82.4 | |
| InceptionNeXt-TFLOPs=4200, Params=28.0, Resolution=224x224, Training Strategy=false2024.11 | 82.3 | |
| LNmodel=ViT-B2025.12 | 82.3 | |
| DeepViT-SInput Resolution=224x224, Parameters=27M, FLOPs=6.2B2021.06 | 82.3 | |
| MobileMamba-B1FLOPs=1080, Params=17.1, Resolution=256x256, Training Strategy=true2024.11 | 82.2 | |
| VMamba-TFLOPs=5600, Params=22.0, Resolution=224x224, Training Strategy=false2024.11 | 82.2 | |
| T2T-ViT-19Input Resolution=224x224, Parameters=39.2M, FLOPs=9.8B2021.06 | 82.2 | |
| SHViT-S4r512FLOPs=3973, Params=16.5, Resolution=512x512, Training Strategy=false2024.11 | 82 | |
| VRWKV-BFLOPs=18200, Params=93.7, Resolution=224x224, Training Strategy=false2024.11 | 82 | |
| CeiT-S#Param.=24M, Image Size=224^2, FLOPs (G)=4.52025.11 | 82 | |
| EfficientVMamba-BFLOPs=4000, Params=33.0, Resolution=224x224, Training Strategy=false2024.11 | 81.8 | |
| DeiT-BInput Resolution=224x224, Parameters=86M, FLOPs=17.5B2021.06 | 81.8 | |
| DeiT-B#Param.=87M, Image Size=224^2, FLOPs (G)=17.32025.11 | 81.8 | |
| PVT-LargeInput Resolution=224x224, Parameters=61.5M, FLOPs=9.8B2021.06 | 81.7 | |
| T2T-ViT-14Input Resolution=224x224, Parameters=21.5M, FLOPs=6.1B2021.06 | 81.7 | |
| PVT-L#Param.=61M, Image Size=224^2, FLOPs (G)=9.82025.11 | 81.7 | |
| MobileMamba-B2FLOPs=2427, Params=17.1, Resolution=384x384, Training Strategy=false2024.11 | 81.6 | |
| PlainMamba-L2FLOPs=8100, Params=25.0, Resolution=224x224, Training Strategy=false2024.11 | 81.6 | |
| CvT-13Input Resolution=224x224, Parameters=20M, FLOPs=4.5B2021.06 | 81.6 | |
| CoAtNet-0Input Resolution=224x224, Parameters=25M, FLOPs=4.2B2021.06 | 81.6 | |
| Fast-ConvNNBackbone=ViT-Base, k=25, #Param.=86M, Image Size=224^2, FLOPs (G)=34.42025.11 | 81.6 | |
| CvT-13#Param.=20M, Image Size=224^2, FLOPs (G)=4.52025.11 | 81.6 | |
| CoAtNet-0#Param.=25M, Image Size=224^2, FLOPs (G)=4.22025.11 | 81.6 | |
| ViL-SFLOPs=5100, Params=23.0, Resolution=224x224, Training Strategy=false2024.11 | 81.5 | |
| FasterNet-SFLOPs=4560, Params=31.1, Resolution=224x224, Training Strategy=false2024.11 | 81.3 | |
| Swin-TInput Resolution=224x224, Parameters=29M, FLOPs=4.5B2021.06 | 81.3 |