Classification on ImageNet-1K 1.0 (val)
90.88Top-1 Accuracy (%)CoAtNet-7
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| CoAtNet-7Public=false2022.06 | 90.88 | — | — | — | — | — | |
| ViT-G/14Public=false2022.06 | 90.45 | — | — | — | — | — | |
| DaViT-Giant#Params (M)=1437, FLOPs (G)=1038, Resolution=512x512, Pre-training=1.5B image and text pairs2022.04 | 90.4 | — | — | — | — | — | |
| DaViT-Huge#Params (M)=362, FLOPs (G)=334, Resolution=512x512, Pre-training=1.5B image and text pairs2022.04 | 90.2 | — | — | — | — | — | |
| Meta Pseudo Labels (EffNet-L2)Public=false2022.06 | 90.2 | 98.8 | — | — | — | — | |
| MViTv2-HPre-training=ImageNet-21K, Input Resolution=512x512, Testing Protocol=resize crop, FLOPs (G)=763.5, Parameters (M)=6672021.12 | 88.8 | — | — | — | — | — | |
| ALIGN (EfficientNet-L2)Public=false2022.06 | 88.64 | 98.67 | — | — | — | — | |
| MViTv2-HPre-training=ImageNet-21K, Input Resolution=384x384, Testing Protocol=resize crop, FLOPs (G)=388.5, Parameters (M)=6672021.12 | 88.6 | — | — | — | — | — | |
| ViT-H/14Public=false2022.06 | 88.55 | — | — | — | — | — | |
| CoAtNet-4Pre-training=ImageNet-21K, Input Resolution=512x512, Testing Protocol=center crop, FLOPs (G)=360.9, Parameters (M)=2752021.12 | 88.4 | — | — | — | — | — | |
| MViTv2-LPre-training=ImageNet-21K, Input Resolution=384x384, Testing Protocol=resize crop, FLOPs (G)=140.7, Parameters (M)=2182021.12 | 88.4 | — | — | — | — | — | |
| CLIP (w/ Noisy Student EffNet-L2)Public=false2022.06 | 88.4 | — | — | — | — | — | |
| Top-k DiffSortNetsBackbone=Noisy Student EfficientNet-L22022.06 | 88.37 | 98.68 | — | — | — | — | |
| Noisy Student EfficientNet-L2Public=true, Source=Reported2022.06 | 88.35 | 98.65 | — | — | — | — | |
| Noisy Student EfficientNet-L2Source=Authors Re-run Baseline, Backbone=Noisy Student EfficientNet-L22022.06 | 88.33 | 98.65 | — | — | — | — | |
| Top-k SinkhornSortBackbone=Noisy Student EfficientNet-L22022.06 | 88.32 | 98.66 | — | — | — | — | |
| MViTv2-HPre-training=ImageNet-21K, Input Resolution=384x384, Testing Protocol=center crop, FLOPs (G)=388.5, Parameters (M)=6672021.12 | 88.3 | — | — | — | — | — | |
| MViTv2-HPre-training=ImageNet-21K, Input Resolution=512x512, Testing Protocol=center crop, FLOPs (G)=763.5, Parameters (M)=6672021.12 | 88.3 | — | — | — | — | — | |
| MViTv2-LPre-training=ImageNet-21K, Input Resolution=384x384, Testing Protocol=center crop, FLOPs (G)=140.7, Parameters (M)=2182021.12 | 88.2 | — | — | — | — | — | |
| MViTv2-HPre-training=ImageNet-21K, Input Resolution=224x224, Testing Protocol=center crop, FLOPs (G)=120.6, Parameters (M)=6672021.12 | 88 | — | — | — | — | — | |
| ConvNeXt-XLType=C, Image Size=384x384, Param. (M)=350, FLOPs (G)=179, Pre-trained=IN-21K2022.11 | 87.8 | — | — | — | — | — | |
| MogaNet-XLType=C, Image Size=384x384, Param. (M)=181, FLOPs (G)=102, Pre-trained=IN-21K2022.11 | 87.8 | — | — | — | — | — | |
| ConvNeXt-XLData/Size=22K/384^2, FLOPs (G)=179, Params (M)=350.22022.01 | 87.8 | — | — | — | — | — | |
| ConvNeXt-XLresolution=384, fine-tuned=true, nb params (x10^6)=350.2, throughput (im/s)=80, FLOPs (x10^9)=179.0, Peak Mem (MB)=162602022.04 | 87.8 | — | — | — | — | — | |
| ViT-L/16Public=true2022.06 | 87.76 | — | — | — | — | — | |
| DeiT III-LType=T, Image Size=384x384, Param. (M)=304, FLOPs (G)=191, Pre-trained=IN-21K2022.11 | 87.7 | — | — | — | — | — | |
| HorNet-LType=C, Image Size=384x384, Param. (M)=202, FLOPs (G)=102, Pre-trained=IN-21K2022.11 | 87.7 | — | — | — | — | — | |
| ViT-Lresolution=384, fine-tuned=true, nb params (x10^6)=304.8, throughput (im/s)=67, FLOPs (x10^9)=191.2, Peak Mem (MB)=128662022.04 | 87.7 | — | — | — | — | — | |
| CoAtNet-3Type=H, Image Size=384x384, Param. (M)=168, FLOPs (G)=107, Pre-trained=IN-21K2022.11 | 87.6 | — | — | — | — | — | |
| CvT-W24Pre-training=ImageNet-21K, Input Resolution=384x384, Testing Protocol=center crop, FLOPs (G)=193.2, Parameters (M)=2772021.12 | 87.6 | — | — | — | — | — | |
| CoAtNet-3#Params (M)=168.0, FLOPs (G)=107.4, Resolution=384x384, Pre-training=ImageNet-22k2022.04 | 87.6 | — | — | — | — | — | |
| BiT-LPublic=false2022.06 | 87.54 | 98.46 | — | — | — | — | |
| ConvNeXt-LType=C, Image Size=384x384, Param. (M)=198, FLOPs (G)=101, Pre-trained=IN-21K2022.11 | 87.5 | — | — | — | — | — | |
| ConvNeXt-LData/Size=22K/384^2, FLOPs (G)=101, Params (M)=197.82022.01 | 87.5 | — | — | — | — | — | |
| MViTv2-LPre-training=ImageNet-21K, Input Resolution=224x224, Testing Protocol=center crop, FLOPs (G)=42.1, Parameters (M)=2182021.12 | 87.5 | — | — | — | — | — | |
| CSwin-LPre-training=ImageNet-21K, Input Resolution=384x384, Testing Protocol=center crop, FLOPs (G)=96.8, Parameters (M)=1732021.12 | 87.5 | — | — | — | — | — | |
| CSwin-L#Params (M)=173.0, FLOPs (G)=96.8, Resolution=384x384, Pre-training=ImageNet-22k2022.04 | 87.5 | — | — | — | — | — | |
| DaViT-Large#Params (M)=196.8, FLOPs (G)=103.0, Resolution=384x384, Pre-training=ImageNet-22k2022.04 | 87.5 | — | — | — | — | — | |
| ConvNeXt-Lresolution=384, fine-tuned=true, nb params (x10^6)=197.8, throughput (im/s)=115, FLOPs (x10^9)=101, Peak Mem (MB)=119382022.04 | 87.5 | — | — | — | — | — | |
| Swin-LType=T, Image Size=384x384, Param. (M)=197, FLOPs (G)=104, Pre-trained=IN-21K2022.11 | 87.3 | — | — | — | — | — | |
| Swin-LPre-training=ImageNet-21K, Input Resolution=384x384, Testing Protocol=center crop, FLOPs (G)=103.9, Parameters (M)=1972021.12 | 87.3 | — | — | — | — | — | |
| EfficientNetV2-XLresolution=512, nb params (x10^6)=208.1, FLOPs (x10^9)=94.02022.04 | 87.3 | — | — | — | — | — | |
| Swin-Lresolution=384, nb params (x10^6)=196.7, throughput (im/s)=100, FLOPs (x10^9)=103.9, Peak Mem (MB)=334562022.04 | 87.3 | — | — | — | — | — | |
| ViT-Hresolution=224, nb params (x10^6)=632.1, throughput (im/s)=112, FLOPs (x10^9)=167.4, Peak Mem (MB)=69842022.04 | 87.2 | — | — | — | — | — | |
| CSwin-B#Params (M)=78.0, FLOPs (G)=47.0, Resolution=384x384, Pre-training=ImageNet-22k2022.04 | 87 | — | — | — | — | — | |
| ConvNeXt-XLresolution=224, nb params (x10^6)=350.2, throughput (im/s)=241, FLOPs (x10^9)=60.9, Peak Mem (MB)=69512022.04 | 87 | — | — | — | — | — | |
| ViT-Lresolution=224, nb params (x10^6)=304.4, throughput (im/s)=277, FLOPs (x10^9)=61.6, Peak Mem (MB)=37892022.04 | 87 | — | — | — | — | — | |
| DaViT-Base#Params (M)=87.9, FLOPs (G)=46.4, Resolution=384x384, Pre-training=ImageNet-22k2022.04 | 86.9 | — | — | — | — | — | |
| ConvNeXt-BData/Size=22K/384^2, FLOPs (G)=45.1, Params (M)=88.62022.01 | 86.8 | — | — | — | — | — | |
| EfficientNetV2-Lresolution=480, fine-tuned=true, nb params (x10^6)=118.5, throughput (im/s)=179, FLOPs (x10^9)=53.0, Peak Mem (MB)=95402022.04 | 86.8 | — | — | — | — | — | |
| ConvNeXt-Bresolution=384, fine-tuned=true, nb params (x10^6)=88.6, throughput (im/s)=190, FLOPs (x10^9)=45.1, Peak Mem (MB)=78512022.04 | 86.8 | — | — | — | — | — | |
| ViT-Bresolution=384, fine-tuned=true, nb params (x10^6)=86.9, throughput (im/s)=190, FLOPs (x10^9)=55.5, Peak Mem (MB)=89562022.04 | 86.7 | — | — | — | — | — | |
| RepLKNet-31LType=C, Image Size=384x384, Param. (M)=172, FLOPs (G)=96, Pre-trained=IN-21K2022.11 | 86.6 | — | — | — | — | — | |
| ConvNeXt-Lresolution=224, nb params (x10^6)=197.8, throughput (im/s)=344, FLOPs (x10^9)=34.4, Peak Mem (MB)=48652022.04 | 86.6 | — | — | — | — | — | |
| CaiT-M48↑448ΥNumber of parameters (M)=356, Throughput (im/s)=5.4, FLOPs (G)=329.6, Peak Memory (MB)=5477.8, Input Resolution=224x224, Batch size=32, Patch size=16x16, Number of patches=14x14, Hardware=V100-32GB GPU2021.05 | 86.5 | — | — | — | — | — | |
| NfNet-F6 SAMNumber of parameters (M)=438, Throughput (im/s)=16.0, FLOPs (G)=377.3, Peak Memory (MB)=5519.3, Input Resolution=224x224, Batch size=32, Hardware=V100-32GB GPU2021.05 | 86.5 | — | — | — | — | — | |
| MKDModel Variant=ViT-L, Teacher=BEIT-L2022.02 | 86.5 | — | — | — | — | — | |
| LV-ViT-L↑512Params=151M, FLOPs=214.8B, Train size=448, Test size=5122021.04 | 86.4 | — | — | — | — | — | |
| Swin-Large#Params (M)=197.0, FLOPs (G)=103.9, Resolution=384x384, Pre-training=ImageNet-22k2022.04 | 86.4 | — | — | — | — | — | |
| Swin-Bresolution=384, fine-tuned=true, nb params (x10^6)=87.9, throughput (im/s)=160, FLOPs (x10^9)=47.0, Peak Mem (MB)=193852022.04 | 86.4 | — | — | — | — | — | |
| CaiT-M36↑448Params=271M, FLOPs=247.8B, Train size=224, Test size=4482021.04 | 86.3 | — | — | — | — | — | |
| MViTv2-LTest Protocol=resize, Input Resolution=384x384, FLOPs (G)=140.2, Param (M)=2182021.12 | 86.3 | — | — | — | — | — | |
| Swin-LPre-training=ImageNet-21K, Input Resolution=224x224, Testing Protocol=center crop, FLOPs (G)=34.5, Parameters (M)=1972021.12 | 86.3 | — | — | — | — | — | |
| Swin-Lresolution=224, nb params (x10^6)=196.5, throughput (im/s)=337, FLOPs (x10^9)=34.5, Peak Mem (MB)=73502022.04 | 86.3 | — | — | — | — | — | |
| Top-k SinkhornSortBackbone=ResNeXt-101 32x48d WSL2022.06 | 86.22 | 97.99 | — | — | — | — | |
| Top-k DiffSortNetsBackbone=ResNeXt-101 32x48d WSL2022.06 | 86.21 | 98 | — | — | — | — | |
| LV-ViT-L↑448Params=150M, FLOPs=157.2B, Train size=448, Test size=4482021.04 | 86.2 | — | — | — | — | — | |
| ViL-B-RPBPre-training=ImageNet-21K, Input Resolution=384x384, Testing Protocol=center crop, FLOPs (G)=43.7, Parameters (M)=562021.12 | 86.2 | — | — | — | — | — | |
| EfficientNetV2-Mresolution=480, fine-tuned=true, nb params (x10^6)=54.1, throughput (im/s)=312, FLOPs (x10^9)=25.0, Peak Mem (MB)=71272022.04 | 86.2 | — | — | — | — | — | |
| ResNeXt-101 32x48d WSLSource=Authors Re-run Baseline, Backbone=ResNeXt-101 32x48d WSL2022.06 | 86.06 | 97.8 | — | — | — | — | |
| NFNet-F5Params=377M, FLOPs=289.8B, Train size=416, Test size=5442021.04 | 86 | — | — | — | — | — | |
| MViTv2-LTest Protocol=center, Input Resolution=384x384, FLOPs (G)=140.2, Param (M)=2182021.12 | 86 | — | — | — | — | — | |
| NFNet-F4Params=316M, FLOPs=215.3B, Train size=384, Test size=5122021.04 | 85.9 | — | — | — | — | — | |
| LV-ViT-L↑448Params=150M, FLOPs=157.2B, Train size=288, Test size=4482021.04 | 85.9 | — | — | — | — | — | |
| MAEModel Variant=ViT-L2022.02 | 85.9 | — | — | — | — | — | |
| NFNet-F4Test Protocol=center, Input Resolution=512x512, FLOPs (G)=215.3, Param (M)=3162021.12 | 85.9 | — | — | — | — | — | |
| CoAtNet-3Test Protocol=center, Input Resolution=384x384, FLOPs (G)=107.4, Param (M)=1682021.12 | 85.8 | — | — | — | — | — | |
| ConvNeXt-Bresolution=224, nb params (x10^6)=88.6, throughput (im/s)=563, FLOPs (x10^9)=15.4, Peak Mem (MB)=30292022.04 | 85.8 | — | — | — | — | — | |
| ViT-L↑384 (DeiT III)nb params=304.8M, throughput=67 im/s, FLOPs=191.2G, Peak Mem=12866 MB, Resolution=384, Fine-tuning=true2022.04 | 85.8 | — | — | — | — | — | |
| Fix-EfficientNet-B8Params=87M, FLOPs=89.5B, Train size=672, Test size=8002021.04 | 85.7 | — | — | — | — | — | |
| NFNet-F3Params=255M, FLOPs=114.8B, Train size=320, Test size=4162021.04 | 85.7 | — | — | — | — | — | |
| ViT-Bresolution=224, nb params (x10^6)=86.6, throughput (im/s)=831, FLOPs (x10^9)=17.6, Peak Mem (MB)=20782022.04 | 85.7 | — | — | — | — | — | |
| EfficientNetV2-Lnb params=118.5M, throughput=179 im/s, FLOPs=53.0G, Peak Mem=9540 MB2022.04 | 85.7 | — | — | — | — | — | |
| MViTv2-BTest Protocol=resize, Input Resolution=384x384, FLOPs (G)=36.7, Param (M)=522021.12 | 85.6 | — | — | — | — | — | |
| ViT-B/16resolution=384, fine-tuned=true, nb params (x10^6)=86.7, throughput (im/s)=190, FLOPs (x10^9)=55.5, Peak Mem (MB)=89562022.04 | 85.5 | — | — | — | — | — | |
| ViT-L/16resolution=384, fine-tuned=true, nb params (x10^6)=304.8, throughput (im/s)=67, FLOPs (x10^9)=191.1, Peak Mem (MB)=128662022.04 | 85.5 | — | — | — | — | — | |
| ConvNeXt-L-384nb params=197.8M, throughput=115 im/s, FLOPs=101.0G, Peak Mem=11938 MB, Resolution=3842022.04 | 85.5 | — | — | — | — | — | |
| ResNeXt-101 32x48d WSLPublic=true, Source=Reported2022.06 | 85.43 | 97.57 | — | — | — | — | |
| IG-940M-1.5k ResNeXt-101 32x48dImage size=224, Parameters=829M, Mult-adds=153B, Pre-training=Instagram hashtag data (940M images, 1.5k hashtags), Finetuning=train-IN-1k2018.05 | 85.4 | 97.6 | — | — | — | — | |
| CaiT-S36↑384Params=68M, FLOPs=48.0B, Train size=224, Test size=3842021.04 | 85.4 | — | — | — | — | — | |
| LV-ViT-M↑384Params=56M, FLOPs=42.2B, Train size=224, Test size=3842021.04 | 85.4 | — | — | — | — | — | |
| CSWin-B#Param.=78M, FLOPs=47.0G, Resolution=384x384, Training Protocol=finetuned, Model Complexity Group=Base Model2021.07 | 85.4 | — | — | — | — | — | |
| CSWin-BImage Size=384^2, #Param.=78M, FLOPs=47.0G, Throughput=85/s2021.07 | 85.4 | — | — | — | — | — | |
| CSWin-BResolution=384x384, Training Protocol=Finetuned, Parameters=78M, FLOPs=47.0G, Complexity Group=Base Model2021.07 | 85.4 | — | — | — | — | — | |
| CaiT-SResolution=384x384, Protocol=finetuned, #param. (M)=68, FLOPs (G)=48.02021.11 | 85.4 | — | — | — | — | — | |
| R-152x4resolution=480, fine-tuned=true, nb params (x10^6)=937, FLOPs (x10^9)=840.52022.04 | 85.4 | — | — | — | — | — | |
| LV-ViT-LParams=150M, FLOPs=59.0B, Train size=288, Test size=2882021.04 | 85.3 | — | — | — | — | — | |
| MViTv2-LTest Protocol=center, Input Resolution=224x224, FLOPs (G)=42.1, Param (M)=2182021.12 | 85.3 | — | — | — | — | — | |
| BEITModel Variant=ViT-L2022.02 | 85.2 | — | — | — | — | — | |
| MViTv2-BTest Protocol=center, Input Resolution=384x384, FLOPs (G)=36.7, Param (M)=522021.12 | 85.2 | — | — | — | — | — |