Image Classification on ImageNet-1K (val) (Accuracy)
90.9Top-1 AccuracyCoAtNet-7
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| CoAtNet-7Model Type=Transformer, Resolution=512x512, Parameters=2.44B, FLOPs=2586G, Pre-training=Large-scale private dataset2022.11 | 90.9 | — | |
| ViT-G/14Model Type=Transformer, Resolution=518x518, Parameters=1.84B, FLOPs=5160G, Pre-training=Large-scale private dataset2022.11 | 90.5 | — | |
| CoAtNet-6Model Type=Transformer, Resolution=512x512, Parameters=1.47B, FLOPs=1521G, Pre-training=Large-scale private dataset2022.11 | 90.5 | — | |
| SwinV2-GParams=3.0 G, Pre-training Images=70 M, Pre-training Annotation=labeled2022.12 | 90.2 | — | |
| SwinV2-GModel Type=Transformer, Resolution=640x640, Parameters=3.00B, Pre-training=Large-scale private dataset2022.11 | 90.2 | — | |
| FlorenceParams=0.9 G, Pre-training Images=900 M, Pre-training Annotation=image-text2022.12 | 90.1 | — | |
| RevCol-HParams=2.1 G, Pre-training Images=168 M, Pre-training Annotation=semi-labeled2022.12 | 90 | — | |
| Florence-CoSwin-HModel Type=Transformer, Parameters=893M, Pre-training=Large-scale private dataset2022.11 | 90 | — | |
| BEiT3Params=1.0 G, Pre-training Images=35 M, Pre-training Annotation=labeled & image-text2022.12 | 89.6 | — | |
| InternImage-HModel Type=CNN, Resolution=640x640, Parameters=1.08B, FLOPs=1478G, Pre-training=Joint public dataset (Laion-400M, YFCC-15M, CC12M)2022.11 | 89.6 | — | |
| FD-SwinV2-GBackbone=SwinV2-G, Protocol=Fine-tuning2022.05 | 89.4 | — | |
| EVAEvaluation Protocol=fine-tuning2022.11 | 89.4 | — | |
| SwinV2-GBackbone=SwinV2-G, Protocol=Fine-tuning2022.05 | 89.2 | — | |
| dBOTEvaluation Protocol=fine-tuning2022.11 | 89.1 | — | |
| InternImage-HModel Type=CNN, Resolution=224x224, Parameters=1.08B, FLOPs=188G, Pre-training=Joint public dataset (Laion-400M, YFCC-15M, CC12M)2022.11 | 88.9 | — | |
| EVA-CLIP-18B#param=17.5B, Protocol=Linear Probing2024.02 | 88.9 | — | |
| EVA-CLIP-8B#param=7.5B, Protocol=Linear Probing2024.02 | 88.5 | — | |
| InternVL-C#param=5.9B, Protocol=Linear Probing2024.02 | 88.2 | — | |
| EVA-02-CLIP-E/14+#param=4.4B, Protocol=Linear Probing2024.02 | 88.1 | — | |
| ViT-H – cosubnb params=632.1M, throughput=112 im/s, FLOPs=167.4G, peak mem=6984 MB, Pre-training=ImageNet-21k, Resolution=224, Fine-tuning epochs=502022.12 | 88 | — | |
| InternImage-XLModel Type=CNN, Resolution=384x384, Parameters=335M, FLOPs=163G, Pre-training=ImageNet-22K2022.11 | 88 | — | |
| UniRepLKNet-XLType=ConvNet, Input size=384x384, Params (M)=386, FLOPs (G)=187, Throughput (img/s)=131, Pre-training=ImageNet-22K2023.11 | 88 | — | |
| InternImage-XLType=ConvNet, Input size=384x384, Params (M)=335, FLOPs (G)=163, Throughput (img/s)=114, Pre-training=ImageNet-22K2023.11 | 88 | — | |
| CoAtNet-4Model Type=Transformer, Resolution=384x384, Parameters=275M, FLOPs=190G, Pre-training=ImageNet-22K2022.11 | 87.9 | — | |
| UniRepLKNet-LType=ConvNet, Input size=384x384, Params (M)=218, FLOPs (G)=105.4, Throughput (img/s)=190, Pre-training=ImageNet-22K2023.11 | 87.9 | — | |
| CoAtNet-4Type=Transformer, Input size=384x384, Params (M)=275, FLOPs (G)=190, Throughput (img/s)=58, Pre-training=ImageNet-22K2023.11 | 87.9 | — | |
| ConvNeXt-XLModel Type=CNN, Resolution=384x384, Parameters=350M, FLOPs=179G, Pre-training=ImageNet-22K2022.11 | 87.8 | — | |
| RepLKNet-XLModel Type=CNN, Resolution=384x384, Parameters=335M, FLOPs=129G, Pre-training=Large-scale private dataset2022.11 | 87.8 | — | |
| ConvNeXt-XLType=ConvNet, Input size=384x384, Params (M)=350, FLOPs (G)=179, Throughput (img/s)=129, Pre-training=ImageNet-22K2023.11 | 87.8 | — | |
| DeiT III-LModel Type=Transformer, Resolution=384x384, Parameters=304M, FLOPs=191G, Pre-training=ImageNet-22K2022.11 | 87.7 | — | |
| HorNet-LModel Type=CNN, Resolution=384x384, Parameters=202M, FLOPs=102G, Pre-training=ImageNet-22K2022.11 | 87.7 | — | |
| InternImage-LModel Type=CNN, Resolution=384x384, Parameters=223M, FLOPs=108G, Pre-training=ImageNet-22K2022.11 | 87.7 | — | |
| InternImage-LType=ConvNet, Input size=384x384, Params (M)=223, FLOPs (G)=108, Throughput (img/s)=143, Pre-training=ImageNet-22K2023.11 | 87.7 | — | |
| HorNet-LType=ConvNet, Input size=384x384, Params (M)=202, FLOPs (G)=102, Throughput (img/s)=127, Pre-training=ImageNet-22K2023.11 | 87.7 | — | |
| DeiT III-LType=Transformer, Input size=384x384, Params (M)=305, FLOPs (G)=191, Throughput (img/s)=42, Pre-training=ImageNet-22K2023.11 | 87.7 | — | |
| InsertQuantBackbone=DINOv2 ViT-L, Bits (W-A-KV)=16-16-16, Evaluation Protocol=standard accuracy evaluation2026.06 | 87.7 | — | |
| CoAtNet-3Model Type=Transformer, Resolution=384x384, Parameters=168M, FLOPs=107G, Pre-training=ImageNet-22K2022.11 | 87.6 | — | |
| SwinV2-L/24Model Type=Transformer, Resolution=384x384, Parameters=197M, FLOPs=115G, Pre-training=ImageNet-22K2022.11 | 87.6 | — | |
| CoAtNet-3Type=Transformer, Input size=384x384, Params (M)=168, FLOPs (G)=107, Throughput (img/s)=103, Pre-training=ImageNet-22K2023.11 | 87.6 | — | |
| SwinV2-L/24Type=Transformer, Input size=384x384, Params (M)=197, FLOPs (G)=115, Throughput (img/s)=88, Pre-training=ImageNet-22K2023.11 | 87.6 | — | |
| ViT-L – cosubnb params=304.4M, throughput=277 im/s, FLOPs=61.6G, peak mem=3789 MB, Pre-training=ImageNet-21k, Resolution=224, Fine-tuning epochs=502022.12 | 87.5 | — | |
| ConvNeXt-LModel Type=CNN, Resolution=384x384, Parameters=198M, FLOPs=101G, Pre-training=ImageNet-22K2022.11 | 87.5 | — | |
| BiT-L-ResNet152x4Model Type=CNN, Resolution=480x480, Parameters=928M, Pre-training=Large-scale private dataset2022.11 | 87.5 | — | |
| ConvNeXt-LType=ConvNet, Input size=384x384, Params (M)=198, FLOPs (G)=101, Throughput (img/s)=185, Pre-training=ImageNet-22K2023.11 | 87.5 | — | |
| UniRepLKNet-BType=ConvNet, Input size=384x384, Params (M)=98, FLOPs (G)=47.2, Throughput (img/s)=314, Pre-training=ImageNet-22K2023.11 | 87.4 | — | |
| Swin-LModel Type=Transformer, Resolution=384x384, Parameters=197M, FLOPs=104G, Pre-training=ImageNet-22K2022.11 | 87.3 | — | |
| XCiT-L24 – cosubnb params=189.0M, throughput=334 im/s, FLOPs=36.1G, peak mem=3315 MB, Pre-training=ImageNet-21k, Resolution=224, Fine-tuning epochs=502022.12 | 87.2 | — | |
| ViT-H – DeiT-IIInb params=632.1M, throughput=112 im/s, FLOPs=167.4G, peak mem=6984 MB, Pre-training=ImageNet-21k, Resolution=2242022.12 | 87.2 | — | |
| Swin-L – cosubnb params=196.5M, throughput=337 im/s, FLOPs=34.5G, peak mem=7350 MB, Pre-training=ImageNet-21k, Resolution=224, Fine-tuning epochs=502022.12 | 87.1 | — | |
| CoAtNet-2Type=Transformer, Input size=384x384, Params (M)=75, FLOPs (G)=49.8, Throughput (img/s)=163, Pre-training=ImageNet-22K2023.11 | 87.1 | — | |
| ConvNeXt-XLnb params=350.2M, throughput=241 im/s, FLOPs=60.9G, peak mem=6951 MB, Pre-training=ImageNet-21k, Resolution=2242022.12 | 87 | — | |
| ViT-L – DeiT-IIInb params=304.4M, throughput=277 im/s, FLOPs=61.6G, peak mem=3789 MB, Pre-training=ImageNet-21k, Resolution=2242022.12 | 87 | — | |
| EfficientNetV2-Lnb params=118.5M, throughput=179 im/s, FLOPs=53.0G, peak mem=9540 MB, Pre-training=ImageNet-21k, Resolution=4802022.12 | 86.8 | — | |
| ConvNeXt-BType=ConvNet, Input size=384x384, Params (M)=89, FLOPs (G)=45.1, Throughput (img/s)=304, Pre-training=ImageNet-22K2023.11 | 86.8 | — | |
| MaxViT-LEval size=512, Params=212M, FLOPs=245.4G, throughput (img/s)=17.82022.04 | 86.7 | — | |
| DeiT III-BType=Transformer, Input size=384x384, Params (M)=87, FLOPs (G)=55.5, Throughput (img/s)=138, Pre-training=ImageNet-22K2023.11 | 86.7 | — | |
| MaxViT-BEval size=512, Params=120M, FLOPs=138.5G, throughput (img/s)=24.02022.04 | 86.66 | — | |
| data2vecBackbone=ViT-L, Model Setup=Single models, Training Epochs=16002022.02 | 86.6 | — | |
| ConvNeXtArchitecture Style=Multi-scale, Model Variant=Large2022.03 | 86.6 | — | |
| ConvNeXt-Lnb params=197.8M, throughput=344 im/s, FLOPs=34.4G, peak mem=4865 MB, Pre-training=ImageNet-21k, Resolution=2242022.12 | 86.6 | — | |
| ConvNeXt-L# of Params=198 M, FLOPs=34.4 G, Throughput (imgs/sec)=643, Memory (GB)=7.5, Pre-training=ImageNet-22K, Resolution=224x2242022.09 | 86.6 | — | |
| DINAT-L# of Params=200 M, FLOPs=30.6 G, Throughput (imgs/sec)=474, Memory (GB)=7.8, Pre-training=ImageNet-22K, Resolution=224x2242022.09 | 86.6 | — | |
| RepLKNet-31LModel Type=CNN, Resolution=384x384, Parameters=172M, FLOPs=96G, Pre-training=ImageNet-22K2022.11 | 86.6 | — | |
| RepLKNet-31LType=ConvNet, Input size=384x384, Params (M)=172, FLOPs (G)=96, Throughput (img/s)=158, Pre-training=ImageNet-22K2023.11 | 86.6 | — | |
| EVT-LModel Scale=large model (~18.0G), Resolution=384x384, Params (M)=101, FLOPs (G)=59.12026.04 | 86.6 | — | |
| PeCoBackbone=ViT-L, Model Setup=Multiple models2022.02 | 86.5 | — | |
| FocalNetArchitecture Style=Multi-scale, Model Variant=Large2022.03 | 86.5 | — | |
| EVAEvaluation Protocol=linear probing2022.11 | 86.5 | — | |
| XCiT-M24 – cosubnb params=84.4M, throughput=553 im/s, FLOPs=16.2G, peak mem=2010 MB, Pre-training=ImageNet-21k, Resolution=224, Fine-tuning epochs=502022.12 | 86.5 | — | |
| EVA-01-CLIP-g/14#param=1.0B, Protocol=Linear Probing2024.02 | 86.5 | — | |
| InsertQuantBackbone=DINOv2 ViT-L, Bits (W-A-KV)=8-8s-8, Evaluation Protocol=standard accuracy evaluation2026.06 | 86.5 | — | |
| MaxViT-LEval size=384, Params=212M, FLOPs=133.1G, throughput (img/s)=34.32022.04 | 86.4 | — | |
| UniRepLKNet-SType=ConvNet, Input size=384x384, Params (M)=56, FLOPs (G)=26.7, Throughput (img/s)=435, Pre-training=ImageNet-22K2023.11 | 86.4 | — | |
| MaxViT-BEval size=384, Params=120M, FLOPs=74.2G, throughput (img/s)=45.82022.04 | 86.34 | — | |
| BEIT384-LModel Size=307M, Resolution=384x384, Pre-training strategy=Self-Supervised Pre-Training on ImageNet-1K2021.06 | 86.3 | — | |
| Swin-Lnb params=196.5M, throughput=337 im/s, FLOPs=34.5G, peak mem=7350 MB, Pre-training=ImageNet-21k, Resolution=2242022.12 | 86.3 | — | |
| ViT-B – cosubnb params=86.6M, throughput=831 im/s, FLOPs=17.6G, peak mem=2078 MB, Pre-training=ImageNet-21k, Resolution=224, Fine-tuning epochs=502022.12 | 86.3 | — | |
| Swin-L# of Params=197 M, FLOPs=34.5 G, Throughput (imgs/sec)=478, Memory (GB)=10.4, Pre-training=ImageNet-22K, Resolution=224x2242022.09 | 86.3 | — | |
| EVT-XLModel Scale=XL model (~35.0G), Resolution=224x224, Params (M)=205, FLOPs (G)=36.42026.04 | 86.3 | — | |
| EfficientNetV2-Mnb params=54.1M, throughput=312 im/s, FLOPs=25.0G, peak mem=7127 MB, Pre-training=ImageNet-21k, Resolution=4802022.12 | 86.2 | — | |
| Swin-B – cosubnb params=87.8M, throughput=532 im/s, FLOPs=15.4G, peak mem=4695 MB, Pre-training=ImageNet-21k, Resolution=224, Fine-tuning epochs=502022.12 | 86.2 | — | |
| OpenCLIP-G/14#param=1.8B, Protocol=Linear Probing2024.02 | 86.2 | — | |
| EVT-BModel Scale=base model (~9.0G), Resolution=384x384, Params (M)=57, FLOPs (G)=32.82026.04 | 86.2 | — | |
| BaselineBackbone=DINOv2 ViT-L, Bits (W-A-KV)=16-16-16, Evaluation Protocol=standard accuracy evaluation2026.06 | 86.2 | — | |
| MaxViT-SEval size=512, Params=69M, FLOPs=67.6G, throughput (img/s)=43.32022.04 | 86.19 | — | |
| NFNet-F5Eval size=544, Params=377M, FLOPs=289.8G2022.04 | 86 | — | |
| CoAtNet-3Eval size=512, Params=168M, FLOPs=203.1G, throughput (img/s)=22.42022.04 | 86 | — | |
| NFNet-F4Eval size=512, Params=316M, FLOPs=215.2G, throughput (img/s)=51.72022.04 | 85.9 | — | |
| MAEBackbone=ViT-L, Model Setup=Single models2022.02 | 85.9 | — | |
| CoAtNet-3Eval size=384, Params=168M, FLOPs=107.4G, throughput (img/s)=48.52022.04 | 85.8 | — | |
| ConvNeXt-Bnb params=88.6M, throughput=563 im/s, FLOPs=15.4G, peak mem=3029 MB, Pre-training=ImageNet-21k, Resolution=2242022.12 | 85.8 | — | |
| ConvNeXt-B – cosubnb params=88.6M, throughput=563 im/s, FLOPs=15.4G, peak mem=3029 MB, Pre-training=ImageNet-21k, Resolution=224, Fine-tuning epochs=502022.12 | 85.8 | — | |
| PiT-B – cosubnb params=73.8M, throughput=615 im/s, FLOPs=12.5G, peak mem=7564 MB, Pre-training=ImageNet-21k, Resolution=224, Fine-tuning epochs=502022.12 | 85.8 | — | |
| ConvNeXt-SType=ConvNet, Input size=384x384, Params (M)=50, FLOPs (G)=25.5, Throughput (img/s)=415, Pre-training=ImageNet-22K2023.11 | 85.8 | — | |
| EVT-LModel Scale=large model (~18.0G), Resolution=224x224, Params (M)=101, FLOPs (G)=18.22026.04 | 85.8 | — | |
| iFormer-LModel Scale=large model (~18.0G), Resolution=384x384, Params (M)=87, FLOPs (G)=45.32026.04 | 85.8 | — | |
| MaxViT-SEval size=384, Params=69M, FLOPs=36.1G, throughput (img/s)=82.72022.04 | 85.74 | — | |
| MaxViT-TEval size=512, Params=31M, FLOPs=33.7G, throughput (img/s)=63.82022.04 | 85.72 | — | |
| NFNet-F3Eval size=416, Params=255M, FLOPs=114.7G, throughput (img/s)=78.82022.04 | 85.7 | — | |
| EffNetV2-LEval size=480, Params=121M, FLOPs=53.0G, throughput (img/s)=163.22022.04 | 85.7 | — |