Image Classification on ImageNet-1k
90.9Top-1 AccViT-e/14
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| ViT-e/14Evaluation Protocol=High-res fine-tuning2023.02 | 90.9 | — | — | — | — | — | — | — | — | — | |
| ViT-G/14Evaluation Protocol=High-res fine-tuning2023.02 | 90.45 | — | — | — | — | — | — | — | — | — | |
| CoCa-Ltuning_type=Fine-tuning (100%), #param=303M, #data=4.8B, resolution=576, high_resource_cost=true2022.06 | 90.2 | — | — | — | — | — | — | — | — | — | |
| Florencetuning_type=Fine-tuning (100%), #param=637M, #data=900M, resolution=>=384, high_resource_cost=true2022.06 | 90 | — | — | — | — | — | — | — | — | — | |
| MaxViT-XLEvaluation Protocol=High-res fine-tuning2023.02 | 89.53 | — | — | — | — | — | — | — | — | — | |
| ViT-22BEvaluation Protocol=Linear Probing (frozen), Resolution=224px2023.02 | 89.51 | — | — | — | — | — | — | — | — | — | |
| ViT-e/14Evaluation Protocol=Linear Probing (frozen), Resolution=224px2023.02 | 89.26 | — | — | — | — | — | — | — | — | — | |
| ViT-G/14Evaluation Protocol=Linear Probing (frozen), Resolution=224px2023.02 | 88.98 | — | — | — | — | — | — | — | — | — | |
| ALIGN-L2Evaluation Protocol=High-res fine-tuning2023.02 | 88.64 | — | — | — | — | — | — | — | — | — | |
| ALIGNtuning_type=Fine-tuning (100%), #param=480M, #data=1.8B, resolution=289, high_resource_cost=true2022.06 | 88.6 | — | — | — | — | — | — | — | — | — | |
| ViT-g/14Evaluation Protocol=Linear Probing (frozen), Resolution=224px2023.02 | 88.51 | — | — | — | — | — | — | — | — | — | |
| ViT-L/16Evaluation Protocol=High-res fine-tuning2023.02 | 88.5 | — | — | — | — | — | — | — | — | — | |
| FixNoisy-L2Evaluation Protocol=High-res fine-tuning2023.02 | 88.5 | — | — | — | — | — | — | — | — | — | |
| dBOTstudent backbone=ViT-H, teacher=CLIP-L [31]2022.09 | 88.5 | — | — | — | — | — | — | — | — | — | |
| CoCa-Btuning_type=Fine-tuning (100%), #param=86M, #data=4.8B, resolution=576, high_resource_cost=true2022.06 | 88.3 | — | — | — | — | — | — | — | — | — | |
| PeCoBackbone=ViT-H, pre-train dataset=IN-1K, pre-train epochs=800, Resolution=4482021.11 | 88.3 | — | — | — | — | — | — | — | — | — | |
| ConvNeXt-XLimage size=384^2, #param.=350M, FLOPs=179.0G, throughput (image/s)=30.2, Pre-training=ImageNet-22K2022.01 | 87.8 | — | — | — | — | — | — | — | — | — | |
| MAEBackbone=ViT-H, pre-train dataset=IN-1K, pre-train epochs=1600, Resolution=4482021.11 | 87.8 | — | — | — | — | — | — | — | — | — | |
| dBOTstudent backbone=ViT-L, teacher=CLIP-L [31]2022.09 | 87.8 | — | — | — | — | — | — | — | — | — | |
| ConvNeXt-Limage size=384^2, #param.=198M, FLOPs=101.0G, throughput (image/s)=50.4, Pre-training=ImageNet-22K2022.01 | 87.5 | — | — | — | — | — | — | — | — | — | |
| PeCoBackbone=ViT-H, pre-train dataset=IN-1K, pre-train epochs=8002021.11 | 87.5 | — | — | — | — | — | — | — | — | — | |
| ViT-H (random initialized)student backbone=ViT-H, teacher initialization=random2022.09 | 87.4 | — | — | — | — | — | — | — | — | — | |
| EffNetV2-XLimage size=480^2, #param.=208M, FLOPs=94.0G, throughput (image/s)=56.5, Pre-training=ImageNet-22K2022.01 | 87.3 | — | — | — | — | — | — | — | — | — | |
| Swin-Limage size=384^2, #param.=197M, FLOPs=103.9G, throughput (image/s)=46.0, Pre-training=ImageNet-22K2022.01 | 87.3 | — | — | — | — | — | — | — | — | — | |
| ELSA-VOLO-D5Params=298M, FLOPs=437G, Resolution=512x5122021.12 | 87.2 | — | — | — | — | — | — | — | — | — | |
| VOLO-D5Params=295M, FLOPs=407G, Resolution=512x5122021.12 | 87.1 | — | — | — | — | — | — | — | — | — | |
| ConvNeXt-XLimage size=224^2, #param.=350M, FLOPs=60.9G, throughput (image/s)=89.3, Pre-training=ImageNet-22K2022.01 | 87 | — | — | — | — | — | — | — | — | — | |
| Uni-Perceiver-Ltuning_type=Fine-tuning (100%), #param=303M, #data=44.1M, resolution=3842022.06 | 87 | — | — | — | — | — | — | — | — | — | |
| Uni-Perceiver-L + Conditional MoEstuning_type=Fine-tuning (100%), #param=303M, #data=44.1M, resolution=3842022.06 | 87 | — | — | — | — | — | — | — | — | — | |
| MAEBackbone=ViT-H, pre-train dataset=IN-1K, pre-train epochs=16002021.11 | 86.9 | — | — | — | — | — | — | — | — | — | |
| LLEarch=ViT-H/14, train data=IN-1k, base_framework=MAE2022.12 | 86.84 | — | — | — | — | — | — | — | — | — | |
| LLEarch=ViT-H/14, train data=IN-1k, base_framework=MAE, augmentation=Edge Aug2022.12 | 86.84 | — | — | — | — | — | — | — | — | — | |
| VOLO-D4Params=193M, FLOPs=194G, Resolution=448x4482021.12 | 86.8 | — | — | — | — | — | — | — | — | — | |
| EffNetV2-Limage size=480^2, #param.=120M, FLOPs=53.0G, throughput (image/s)=83.7, Pre-training=ImageNet-22K2022.01 | 86.8 | — | — | — | — | — | — | — | — | — | |
| ViT-L/16image size=384^2, #param.=305M, FLOPs=191.1G, throughput (image/s)=28.5, Pre-training=ImageNet-22K2022.01 | 86.8 | — | — | — | — | — | — | — | — | — | |
| ConvNeXt-Bimage size=384^2, #param.=89M, FLOPs=45.1G, throughput (image/s)=95.7, Pre-training=ImageNet-22K2022.01 | 86.8 | — | — | — | — | — | — | — | — | — | |
| ViT-L/16Evaluation Protocol=Linear Probing (frozen), Resolution=224px2023.02 | 86.66 | — | — | — | — | — | — | — | — | — | |
| ELSA-VOLO-D3Params=87M, FLOPs=98.6G, Resolution=448x4482021.12 | 86.6 | — | — | — | — | — | — | — | — | — | |
| ConvNeXt-Limage size=224^2, #param.=198M, FLOPs=34.4G, throughput (image/s)=146.8, Pre-training=ImageNet-22K2022.01 | 86.6 | — | — | — | — | — | — | — | — | — | |
| ViT-L (random initialized)student backbone=ViT-L, teacher initialization=random2022.09 | 86.6 | — | — | — | — | — | — | — | — | — | |
| CaiT-M48Params=356M, FLOPs=330G, Resolution=448x4482021.12 | 86.5 | — | — | — | — | — | — | — | — | — | |
| FAN-L-HybridParam.=77M, FLOPs=16.9G, Pre-trained=ImageNet-22K2022.04 | 86.5 | — | — | — | — | — | — | — | — | — | |
| PeCoBackbone=ViT-L, pre-train dataset=IN-1K, pre-train epochs=8002021.11 | 86.5 | — | — | — | — | — | — | — | — | — | |
| Swin-Bimage size=384^2, #param.=88M, FLOPs=47.0G, throughput (image/s)=85.1, Pre-training=ImageNet-22K2022.01 | 86.4 | — | — | — | — | — | — | — | — | — | |
| Uni-Perceiver-L + Conditional MoEstuning_type=Fine-tuning (100%), #param=303M, #data=44.1M, resolution=2242022.06 | 86.4 | — | — | — | — | — | — | — | — | — | |
| VOLO-D3Params=86M, FLOPs=92.9G, Resolution=448x4482021.12 | 86.3 | — | — | — | — | — | — | — | — | — | |
| CaiT-M36Params=271M, FLOPs=248G, Resolution=448x4482021.12 | 86.3 | — | — | — | — | — | — | — | — | — | |
| ELSA-VOLO-D5Params=298M, FLOPs=78.5G, Resolution=224x2242021.12 | 86.3 | — | — | — | — | — | — | — | — | — | |
| Swin-Limage size=224^2, #param.=197M, FLOPs=34.5G, throughput (image/s)=145.0, Pre-training=ImageNet-22K2022.01 | 86.3 | — | — | — | — | — | — | — | — | — | |
| ImageSwin-LBackbone=Swin-L2022.01 | 86.3 | 97.9 | — | — | — | — | — | — | — | — | |
| BEIT384-L/16Fine-tuning=true, Resolution=3842022.02 | 86.3 | — | — | — | — | — | — | — | — | — | |
| Uni-Perceiver-Ltuning_type=Fine-tuning (100%), #param=303M, #data=44.1M, resolution=2242022.06 | 86.2 | — | — | — | — | — | — | — | — | — | |
| VOLO-D5Params=295M, FLOPs=72.7G, Resolution=224x2242021.12 | 86.1 | — | — | — | — | — | — | — | — | — | |
| OMNIVORE (Swin-L)Backbone=Swin-L2022.01 | 86 | 97.7 | — | — | — | — | — | — | — | — | |
| MAE-L/16Fine-tuning=true2022.02 | 85.9 | — | — | — | — | — | — | — | — | — | |
| BootMAEBackbone=ViT-L, pre-train dataset=IN-1K, pre-train epochs=8002021.11 | 85.9 | — | — | — | — | — | — | — | — | — | |
| MAEBackbone=ViT-L, pre-train dataset=IN-1K, pre-train epochs=16002021.11 | 85.9 | — | — | — | — | — | — | — | — | — | |
| DINOv2 (dst)Arch=ViT-L/14, Data=LVD, Ep.=-, Probe=X-Blk2026.03 | 85.9 | — | — | — | — | — | — | — | — | — | |
| LLEarch=ViT-L/16, train data=IN-1k, base_framework=MAE2022.12 | 85.84 | — | — | — | — | — | — | — | — | — | |
| LLEarch=ViT-L/16, train data=IN-1k, base_framework=MAE, augmentation=Edge Aug2022.12 | 85.84 | — | — | — | — | — | — | — | — | — | |
| ConvNeXt-Simage size=384^2, #param.=50M, FLOPs=25.5G, throughput (image/s)=163.5, Pre-training=ImageNet-22K2022.01 | 85.8 | — | — | — | — | — | — | — | — | — | |
| ConvNeXt-Bimage size=224^2, #param.=89M, FLOPs=15.4G, throughput (image/s)=292.1, Pre-training=ImageNet-22K2022.01 | 85.8 | — | — | — | — | — | — | — | — | — | |
| Uni-Perceiver-B + Conditional MoEstuning_type=Fine-tuning (100%), #param=86M, #data=44.1M, resolution=3842022.06 | 85.8 | — | — | — | — | — | — | — | — | — | |
| ELSA-VOLO-D1Params=27M, FLOPs=23.3G, Resolution=384x3842021.12 | 85.7 | — | — | — | — | — | — | — | — | — | |
| ELSA-VOLO-D3Params=87M, FLOPs=22.3G, Resolution=224x2242021.12 | 85.7 | — | — | — | — | — | — | — | — | — | |
| VOLO-D4Params=193M, FLOPs=44.6G, Resolution=224x2242021.12 | 85.7 | — | — | — | — | — | — | — | — | — | |
| EffNetV2-Limage size=480^2, #param.=120M, FLOPs=53.0G, throughput (image/s)=83.7, Pre-training=ImageNet-1K2022.01 | 85.7 | — | — | — | — | — | — | — | — | — | |
| dBOTstudent backbone=ViT-B, teacher=CLIP-B [31]2022.09 | 85.7 | — | — | — | — | — | — | — | — | — | |
| OFALargeFine-tuning=true2022.02 | 85.6 | — | — | — | — | — | — | — | — | — | |
| ViT-Ltuning_type=Fine-tuning (100%), #param=307M, #data=15.5M, resolution=3842022.06 | 85.6 | — | — | — | — | — | — | — | — | — | |
| FAN-B-HybridParam.=50M, FLOPs=11.3G, Pre-trained=ImageNet-22K2022.04 | 85.6 | — | — | — | — | — | — | — | — | — | |
| ConvNeXt-Limage size=384^2, #param.=198M, FLOPs=101.0G, throughput (image/s)=50.4, Pre-training=ImageNet-1K2022.01 | 85.5 | — | — | — | — | — | — | — | — | — | |
| ViT-Btuning_type=Fine-tuning (100%), #param=86M, #data=15.5M, resolution=3842022.06 | 85.5 | — | — | — | — | — | — | — | — | — | |
| ALIGNEvaluation Protocol=Linear Probing (frozen), Resolution=360px2023.02 | 85.5 | — | — | — | — | — | — | — | — | — | |
| VOLO-D3Params=86M, FLOPs=20.9G, Resolution=224x2242021.12 | 85.4 | — | — | — | — | — | — | — | — | — | |
| R-152x4image size=480^2, #param.=937M, FLOPs=840.5G, Pre-training=ImageNet-22K2022.01 | 85.4 | — | — | — | — | — | — | — | — | — | |
| ViT-B/16image size=384^2, #param.=87M, FLOPs=55.5G, throughput (image/s)=93.1, Pre-training=ImageNet-22K2022.01 | 85.4 | — | — | — | — | — | — | — | — | — | |
| DINOv2 (dst)Arch=ViT-L/14, Data=LVD, Ep.=-, Probe=CLS2026.03 | 85.4 | — | — | — | — | — | — | — | — | — | |
| LLEarch=ViT-B/16, train data=IG-3.6B, base_framework=SWAG (FT)2022.12 | 85.37 | — | — | — | — | — | — | — | — | — | |
| ViT-B/16 + PyramidATBackbone=ViT-B/16, Resolution=512x512, Pre-training=ImageNet-21K, Fine-tuning=ImageNet-1K, Adversarial Training=PyramidAT2021.11 | 85.35 | — | — | — | — | — | — | — | — | — | |
| LLEarch=ViT-B/16, train data=IG-3.6B, base_framework=SWAG (FT), augmentation=Edge Aug2022.12 | 85.31 | — | — | — | — | — | — | — | — | — | |
| ViT-L/16Backbone=ViT-L, Patch size=162022.01 | 85.3 | — | — | — | — | — | — | — | — | — | |
| OMNIVORE (Swin-B)Backbone=Swin-B2022.01 | 85.3 | 97.5 | — | — | — | — | — | — | — | — | |
| VOLO-D1Params=27M, FLOPs=20.8G, Resolution=384x3842021.12 | 85.2 | — | — | — | — | — | — | — | — | — | |
| BEITArch.=ViT-L/16, Epochs=8002021.11 | 85.2 | — | — | — | — | — | — | — | — | — | |
| Swin-Bimage size=224^2, #param.=88M, FLOPs=15.4G, throughput (image/s)=286.6, Pre-training=ImageNet-22K2022.01 | 85.2 | — | — | — | — | — | — | — | — | — | |
| ImageSwin-BBackbone=Swin-B2022.01 | 85.2 | 97.5 | — | — | — | — | — | — | — | — | |
| Uni-Perceiver-Btuning_type=Fine-tuning (100%), #param=86M, #data=44.1M, resolution=3842022.06 | 85.2 | — | — | — | — | — | — | — | — | — | |
| BEiTBackbone=ViT-L, pre-train dataset=IN-1K, pre-train epochs=8002021.11 | 85.2 | — | — | — | — | — | — | — | — | — | |
| ConvNeXt-Bimage size=384^2, #param.=89M, FLOPs=45.0G, throughput (image/s)=95.7, Pre-training=ImageNet-1K2022.01 | 85.1 | — | — | — | — | — | — | — | — | — | |
| TinyMIMModel Size=ViT-B2023.01 | 85 | — | — | — | — | — | — | — | — | — | |
| DINO-Tok-AEType=AE, Latent=16 × 16, Linear probing=true2025.11 | 85 | — | — | — | — | — | — | — | — | — | |
| OFAtuning_type=Fine-tuning (100%), #param=472M, #data=60.6M, resolution=4802022.06 | 84.9 | — | — | — | — | — | — | — | — | — | |
| Uni-Perceiver-L + Conditional MoEstuning_type=Prompt Tuning (1%), #param=303M, #data=44.1M, resolution=2242022.06 | 84.9 | — | — | — | — | — | — | — | — | — | |
| DINOv3-bType=VFM, Latent=16 × 16, Linear probing=true2025.11 | 84.9 | — | — | — | — | — | — | — | — | — | |
| ViT-B/16 + PixelATBackbone=ViT-B/16, Resolution=512x512, Pre-training=ImageNet-21K, Fine-tuning=ImageNet-1K, Adversarial Training=PixelAT2021.11 | 84.82 | — | — | — | — | — | — | — | — | — | |
| iBOTArch.=ViT-L/16, Epochs=10002021.11 | 84.8 | — | — | — | — | — | — | — | — | — | |
| CoCa-Ltuning_type=Without Tuning (WT), #param=303M, #data=4.8B, resolution=576, high_resource_cost=true2022.06 | 84.8 | — | — | — | — | — | — | — | — | — | |
| ELSA-VOLO-D1Params=27M, FLOPs=8.0G, Resolution=224x2242021.12 | 84.7 | — | — | — | — | — | — | — | — | — | |
| Mini-DeiT-B↑384#Params=44M2022.04 | 84.7 | — | — | — | — | — | — | — | — | — |