Image Classification on ImageNet 1K (val) (Standard Accuracy)
89.5Top-1 AccuracyViT-22B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ViT-22BParameters=21.7B, Evaluation Protocol=Linear Probing, Training Data=JFT-3B2023.12 | 89.5 | — | |
| InternViT-6BParameters=5.9B, Evaluation Protocol=Linear Probing2023.12 | 88.2 | — | |
| ERMoEBackbone=ViT-B2025.11 | 88.03 | 98.97 | |
| MAEBackbone=ViT-H/14, Input resolution=448x448, Pre-train data=IN1K, Evaluation protocol=fine-tuning, Pre-train epochs=16002021.11 | 87.8 | — | |
| MAWS-ViT-6.5BParameters=6.5B, Evaluation Protocol=Linear Probing2023.12 | 87.8 | — | |
| V-MoEBackbone=ViT-B2025.11 | 87.41 | 97.94 | |
| MAEBackbone=ViT-H/14, Input resolution=224x224, Pre-train data=IN1K, Evaluation protocol=fine-tuning, Pre-train epochs=16002021.11 | 86.9 | — | |
| DINOv2-gParameters=1.1B, Evaluation Protocol=Linear Probing2023.12 | 86.5 | — | |
| EVA-01-CLIP-gParameters=1.1B, Evaluation Protocol=Linear Probing2023.12 | 86.5 | — | |
| Swin-Base#Params (M)=88, GFLOPs=47.1, Pre-training=ImageNet-21K, Image size=384x3842021.03 | 86.4 | — | |
| EVT-XLParams(M)=205, FLOPs(G)=36.42026.04 | 86.3 | — | |
| ViL-Base-RPB#Params (M)=55.7, GFLOPs=43.7, Pre-training=ImageNet-21K, Image size=384x3842021.03 | 86.2 | — | |
| NesT-BPre-training=ImageNet-22K, Resolution=384x3842021.05 | 86.2 | — | |
| OpenCLIP-GParameters=1.8B, Evaluation Protocol=Linear Probing2023.12 | 86.2 | — | |
| Swin-BPre-training=ImageNet-22K, Resolution=384x3842021.05 | 86 | — | |
| MAEBackbone=ViT-L/16, Input resolution=224x224, Pre-train data=IN1K, Evaluation protocol=fine-tuning, Pre-train epochs=16002021.11 | 85.9 | — | |
| EVT-LParams(M)=101, FLOPs(G)=18.22026.04 | 85.8 | — | |
| LaplacianFormer-HugeFLOPs range=>14G, Params=78.5M, FLOPs=15.5G, Image Size=2242026.04 | 85.8 | — | |
| ITNet-LParams=307M, GFLOPs=61.6, Kernel type=Content + position2026.06 | 85.8 | — | |
| ViL-Medium-RPB#Params (M)=39.7, GFLOPs=28.4, Pre-training=ImageNet-21K, Image size=384x3842021.03 | 85.7 | — | |
| GC ViT-LParams(M)=201, FLOPs(G)=32.62026.04 | 85.7 | — | |
| DINOv2Learning type=ViT-B, Evaluation Protocol=Fine-tuning2026.05 | 85.7 | — | |
| UniNet-B6Family=Hybrid, Input Size=448, #FLOPs (G)=51, #Params (M)=1172022.07 | 85.6 | — | |
| LaplacianFormer-LargeFLOPs range=10∼14G, Params=63.1M, FLOPs=11.2G, Image Size=2242026.04 | 85.6 | — | |
| VLMo-BaseModel size=Base2021.11 | 85.5 | — | |
| BiT-152x4-M#Params (M)=928, GFLOPs=837, Pre-training=ImageNet-21K, Image size=480x4802021.03 | 85.4 | — | |
| BiFormer-BFLOPs (G)=9.8, Params (M)=58, Resolution=224x224, Token labeling=true2023.03 | 85.4 | — | |
| LV-VIT-LDepth=24, Embed dim.=768, MLP Ratio=3.0, #Heads=12, #Parameters=150M, Resolution=288x2882021.04 | 85.3 | — | |
| EVT-BParams(M)=57, FLOPs(G)=9.82026.04 | 85.3 | — | |
| LaplacianFormer-MediumFLOPs range=8∼10G, Params=46.3M, FLOPs=7.43G, Image Size=2242026.04 | 85.3 | — | |
| SViT-LFLOPs range=>14G, Params=95M, FLOPs=15.6G, Image Size=2242026.04 | 85.3 | — | |
| MLLA-BFLOPs range=>14G, Params=96M, FLOPs=16.2G, Image Size=2242026.04 | 85.3 | — | |
| ViT-Large/16#Params (M)=307, GFLOPs=191.1, Pre-training=ImageNet-21K, Image size=384x3842021.03 | 85.2 | — | |
| BEiTBackbone=ViT-L/16, Input resolution=224x224, Pre-train data=IN1K+DALLE, Evaluation protocol=fine-tuning2021.11 | 85.2 | — | |
| BEiT-BaseModel size=Base2021.11 | 85.2 | — | |
| EffNetV2-MFamily=Convolution, Input Size=480, #FLOPs (G)=24, #Params (M)=542022.07 | 85.1 | — | |
| NFNet-F2Family=Convolution, Input Size=352, #FLOPs (G)=62.6, #Params (M)=193.82022.07 | 85.1 | — | |
| CoAtNet-1Family=Hybrid, Input Size=384, #FLOPs (G)=27.4, #Params (M)=422022.07 | 85.1 | — | |
| Uniformer-BFLOPs (G)=8.3, Params (M)=50, Resolution=224x224, Token labeling=true2023.03 | 85.1 | — | |
| CSWin-SResolution=384x384, Protocol=finetuned, #Param.=35M, FLOPs=22.0G2021.07 | 85 | — | |
| Standard ViT-B (LN Fine-tuning)Fine-tuning Strategy=Fine-tuning, Normalization=LayerNorm (LN), Backbone=ViT-B, Epochs=20, Evaluation Runs=52026.05 | 84.94 | 97.43 | |
| UniNet-B5Family=Hybrid, Input Size=384, #FLOPs (G)=20.4, #Params (M)=72.92022.07 | 84.9 | — | |
| Wave-ViT-BFLOPs (G)=7.2, Params (M)=34, Resolution=224x224, Token labeling=true2023.03 | 84.8 | — | |
| TransNeXt-BaseParams(M)=90, FLOPs(G)=18.42026.04 | 84.8 | — | |
| SViT-BFLOPs range=8∼10G, Params=52M, FLOPs=9.9G, Image Size=2242026.04 | 84.8 | — | |
| BoTNet-T7Family=Transformer, Input Size=384, #FLOPs (G)=45.8, #Params (M)=75.12022.07 | 84.7 | — | |
| TransNeXt-SmallParams(M)=50, FLOPs(G)=10.32026.04 | 84.7 | — | |
| MogaNet-LFLOPs range=10∼14G, Params=82.5M, FLOPs=15.9G, Image Size=2242026.04 | 84.7 | — | |
| SE-CoTNetD-152Resolution=320, Params=55.8M, GFLOPs=26.5, Training Setup=Advanced2021.07 | 84.6 | 97.1 | |
| DINOv2Learning type=ViT-B, Evaluation Protocol=Linear Probing2026.05 | 84.5 | — | |
| DearKD-B-1000Params=86M, size=224x224, throughput=253.72022.04 | 84.4 | — | |
| UniNet-B4Family=Hybrid, Input Size=320, #FLOPs (G)=9.4, #Params (M)=43.82022.07 | 84.4 | — | |
| OpenCLIP-HParameters=0.6B, Evaluation Protocol=Linear Probing2023.12 | 84.4 | — | |
| EVT-SParams(M)=27, FLOPs(G)=4.62026.04 | 84.4 | — | |
| MaxViT-SmallParams(M)=69, FLOPs(G)=11.72026.04 | 84.4 | — | |
| BiFormer-BParams=87M, GFLOPs=16.5, Kernel type=Content (dynamic routing)2026.06 | 84.4 | — | |
| DeepVit-LParams. (M)=58, MAdds (G)=12.8, Training recipe=DeiT [37], Resolution=384x3842021.03 | 84.3 | — | |
| EfficientNet-B7Resolution=600, Params=66.0M, GFLOPs=37, Training Setup=Advanced2021.07 | 84.3 | 97 | |
| EffiNet-B7Params=66M, size=600x600, throughput=55.12022.04 | 84.3 | — | |
| HAT (+KD)Backbone=ViT-B, Training Set=ImageNet-1K, Knowledge Distillation=true2022.04 | 84.3 | — | |
| EffNet-B7Family=Convolution, Input Size=600, #FLOPs (G)=37, #Params (M)=662022.07 | 84.3 | — | |
| BiFormer-SFLOPs (G)=4.5, Params (M)=26, Resolution=224x224, Token labeling=true2023.03 | 84.3 | — | |
| BiFormer-BFLOPs (G)=9.8, Params (M)=57, Resolution=224x224, Token labeling=false2023.03 | 84.3 | — | |
| BiFormer-BParams(M)=57, FLOPs(G)=9.82026.04 | 84.3 | — | |
| BiFormer-BFLOPs range=8∼10G, Params=56.8M, FLOPs=9.8G, Image Size=2242026.04 | 84.3 | — | |
| StructViT-B-8-1FLOPs range=10∼14G, Params=52M, FLOPs=12G, Image Size=2242026.04 | 84.3 | — | |
| NAT-BFLOPs range=10∼14G, Params=90M, FLOPs=13.7G, Image Size=2242026.04 | 84.3 | — | |
| GP-DFine-tuning Strategy=Knowledge Distillation, Backbone=ViT-B, Epochs=20, Evaluation Runs=52026.05 | 84.25 | 97.18 | |
| BoTNet-S1-128Resolution=320, Params=75.1M, GFLOPs=30.9, Training Setup=Advanced2021.07 | 84.2 | 96.9 | |
| Swin-BResolution=384, Params=87.7M, GFLOPs=47, Training Setup=Advanced2021.07 | 84.2 | — | |
| ViP-BInput Size=384x384, Params (M)=87.8, FLOPS (G)=39.12021.07 | 84.2 | — | |
| DeiT-B-1000Params=87M, size=224x224, throughput=290.92022.04 | 84.2 | — | |
| Swin-BParams=88M, size=384x384, throughput=84.72022.04 | 84.2 | — | |
| CSWin-BFLOPs (G)=15.0, Params (M)=78, Resolution=224x224, Token labeling=false2023.03 | 84.2 | — | |
| Swin-V2-BParams=88M, GFLOPs=15.6, Kernel type=Content, local2026.06 | 84.2 | — | |
| ConvNeXt-V2-BParams=88M, GFLOPs=16.1, Kernel type=Position, local2026.06 | 84.2 | — | |
| LV-ViT-MDepth=20, Embed dim.=512, MLP Ratio=3.0, #Heads=8, #Parameters=56M, Resolution=224x2242021.04 | 84.1 | — | |
| MoCo v3Backbone=ViT-L/16, Input resolution=224x224, Pre-train data=IN1K, Evaluation protocol=fine-tuning2021.11 | 84.1 | — | |
| LV-ViT-MFLOPs (G)=162021.06 | 84.1 | — | |
| ScalableViT-BFLOPs (G)=8.6, Params (M)=81, Resolution=224x224, Token labeling=false2023.03 | 84.1 | — | |
| SOFT++-LargeFLOPs range=10∼14G, Params=64M, FLOPs=11G, Image Size=2242026.04 | 84.1 | — | |
| ViT-Base/16#Params (M)=86.6, GFLOPs=49.3, Pre-training=ImageNet-21K, Image size=384x3842021.03 | 84 | — | |
| SE-CoTNetD-152Resolution=224, Params=55.8M, GFLOPs=17, Training Setup=Advanced2021.07 | 84 | 97 | |
| EfficientNet-B6Resolution=528, Params=43.0M, GFLOPs=19, Training Setup=Advanced2021.07 | 84 | 96.8 | |
| ViT-B/16Pre-training=ImageNet-22K, Resolution=384x3842021.05 | 84 | — | |
| EffiNet-B6Params=43M, size=528x528, throughput=96.92022.04 | 84 | — | |
| CrossFormer-LFLOPs (G)=16.1, Params (M)=92, Resolution=224x224, Token labeling=false2023.03 | 84 | — | |
| TransNeXt-TinyParams(M)=28, FLOPs(G)=5.72026.04 | 84 | — | |
| EfficientVMamba-BParams=85M, GFLOPs=15.8, Kernel type=SSM-based2026.06 | 84 | — | |
| CViTc-18Resolution=384x384, Protocol=finetuned, #Param.=45M, FLOPs=32.4G2021.07 | 83.9 | — | |
| VAN-LargeParams (M)=44.82022.09 | 83.9 | — | |
| MSCAN-LParams (M)=45.22022.09 | 83.9 | — | |
| FocalNet-Bmode=LRF, #Params. (M)=88.7, FLOPs (G)=15.4, Throughput (imgs/s)=2692022.03 | 83.9 | — | |
| Wave-ViT-SFLOPs (G)=4.7, Params (M)=23, Resolution=224x224, Token labeling=true2023.03 | 83.9 | — | |
| ITNet-BParams=86M, GFLOPs=17.9, Kernel type=Content + position2026.06 | 83.9 | — | |
| SENet-350Resolution=384, Params=115.2M, GFLOPs=52.9, Training Setup=Advanced2021.07 | 83.8 | 96.6 | |
| ViP-BInput Size=224x224, Params (M)=87.8, FLOPS (G)=152021.07 | 83.8 | — | |
| SimMIMBackbone=ViT-B, Input Size=224x224, Protocol=Fine-tuning, Pre-training costs=1.0x2021.11 | 83.8 | — | |
| FocalAtt-Base#Params. (M)=89.8, FLOPs (G)=16.4, Throughput (imgs/s)=1382022.03 | 83.8 | — | |
| BiFormer-SFLOPs (G)=4.5, Params (M)=26, Resolution=224x224, Token labeling=false2023.03 | 83.8 | — |