Image Classification on ImageNet-C (mCE)
28.2mCEDINOv2
Evaluation Results
| Method | Links | |
|---|---|---|
| DINOv2Architecture=ViT-g/14, Pretraining Data=LVD-142M, Resolution=224, Protocol=linear probe on frozen features2023.04 | 28.2 | |
| Noisy Student Trainingunlabeled data=300M images, backbone=EfficientNet2019.11 | 28.3 | |
| CAFormer-B36Params (M)=99, MACS (G)=72.2, Resolution=384, Pre-training=ImageNet-21K2022.10 | 30.8 | |
| MAE-H + DAT2022.09 | 31.4 | |
| DINOv2Architecture=ViT-L/14, Pretraining Data=LVD-142M, Resolution=224, Protocol=linear probe on frozen features2023.04 | 31.5 | |
| QUESTModel=ViT-L/16†, Epochs=400 + 20, Training Recipe=DeiT-32026.03 | 32.3 | |
| StandardModel=ViT-L/16†, Epochs=400 + 20, Training Recipe=DeiT-32026.03 | 32.5 | |
| QUESTModel=ViT-H/14†, Epochs=400 + 20, Training Recipe=DeiT-32026.03 | 33.7 | |
| MAE-H2022.09 | 33.92 | |
| StandardModel=ViT-H/14†, Epochs=400 + 20, Training Recipe=DeiT-32026.03 | 34.1 | |
| QUESTModel=ViT-B/16, Epochs=400 + 20, Training Recipe=DeiT-32026.03 | 35.7 | |
| StandardModel=ViT-B/16, Epochs=400 + 20, Training Recipe=DeiT-32026.03 | 36.5 | |
| VOLO-D5HAT=true, Resolution=224x2242022.04 | 38.4 | |
| Dr. ViTPre-trained=ImageNet-21K, Fine-tuned on=ImageNet, Resolution=384x384, Discrete representation=true2021.11 | 38.74 | |
| Dr. ViTPre-trained=ImageNet-21K, Fine-tuned on=ImageNet, Resolution=512x512, Discrete representation=true2021.11 | 38.97 | |
| EVT-XLParams(M)=205, FLOPs(G)=36.42026.04 | 41.1 | |
| VOLO-D5HAT=false, Resolution=224x2242022.04 | 41.8 | |
| ViT-BHAT=true, Resolution=224x2242022.04 | 42.2 | |
| DINOv2Architecture=ViT-B/14, Pretraining Data=LVD-142M, Resolution=224, Protocol=linear probe on frozen features2023.04 | 42.7 | |
| EVT-LParams(M)=101, FLOPs(G)=18.22026.04 | 42.7 | |
| QUESTModel=ViT-B/16, Epochs=100, Training Recipe=DeiT-12026.03 | 42.9 | |
| GC ViT-LParams(M)=201, FLOPs(G)=32.62026.04 | 42.9 | |
| EVT-BParams(M)=57, FLOPs(G)=9.82026.04 | 43.1 | |
| RVT-B-AdaNCAParams (M)=91, FLOPS (G)=19, Pre-training=ImageNet-1K2024.06 | 43.2 | |
| QUESTModel=ViT-S/16, Epochs=200, Training Recipe=DeiT-12026.03 | 43.2 | |
| TransNeXt-BaseParams(M)=90, FLOPs(G)=18.42026.04 | 43.5 | |
| DATBackbone=ViT-B/16, Pre-training=ImageNet-21K, Fine-tuning=ImageNet-1K, Input Resolution=224x2242023.11 | 43.6 | |
| VOLO-D1HAT=true, Resolution=224x2242022.04 | 43.7 | |
| TAPADL-FANParams (M)=50.7, FLOPS (G)=11.8, Pre-training=ImageNet-1K2024.06 | 43.7 | |
| iBOTArchitecture=ViT-L/16, Pretraining Data=INet-22k, Resolution=224, Protocol=linear probe on frozen features2023.04 | 43.9 | |
| TransNeXt-SmallParams(M)=50, FLOPs(G)=10.32026.04 | 43.9 | |
| EVT-SParams(M)=27, FLOPs(G)=4.62026.04 | 44.2 | |
| ConViT-B-AdaNCAParams (M)=89, FLOPS (G)=19, Pre-training=ImageNet-1K2024.06 | 44.3 | |
| QKNorm-DS‡Model=ViT-B/16, Epochs=100, Training Recipe=DeiT-12026.03 | 44.4 | |
| AugReg-ViT + DAT2022.09 | 44.65 | |
| TAPADL-RVTParams (M)=89.4, FLOPS (G)=17.9, Pre-training=ImageNet-1K2024.06 | 44.7 | |
| FAN-B-AdaNCAParams (M)=51.7, FLOPS (G)=12.4, Pre-training=ImageNet-1K2024.06 | 44.7 | |
| StandardModel=ViT-S/16, Epochs=200, Training Recipe=DeiT-12026.03 | 44.8 | |
| ConViT-BParams (M)=93.6, FLOPS (G)=19.2, Pre-training=ImageNet-22K2024.06 | 45.2 | |
| OpenCLIPArchitecture=ViT-G/14, Pretraining Data=LAION-2B, Resolution=224, Protocol=linear probe on frozen features2023.04 | 45.3 | |
| Prev. SOTAweakly labeled data=3.5B Instagram images2019.11 | 45.7 | |
| Dr. ViTPre-trained=ImageNet-21K, Fine-tuned on=ImageNet, Resolution=384x384, Discrete representation=discrete only2021.11 | 45.86 | |
| FAN-BParams (M)=50.4, FLOPS (G)=11.7, Pre-training=ImageNet-1K2024.06 | 46.1 | |
| CAFormer-S18Params (M)=26, MACS (G)=13.4, Resolution=384, Pre-training=ImageNet-1K2022.10 | 46.1 | |
| Dr. ViTBackbone=ViT-B, Discrete Only=false2021.11 | 46.22 | |
| DrViT2022.09 | 46.22 | |
| Dr. ViTBackbone=ViT-B, Resolution=384x3842021.11 | 46.32 | |
| ViT-BHAT=false, Resolution=224x2242022.04 | 46.4 | |
| MaxViT-SmallParams(M)=69, FLOPs(G)=11.72026.04 | 46.4 | |
| TransNeXt-TinyParams(M)=28, FLOPs(G)=5.72026.04 | 46.5 | |
| VOLO-D1HAT=false, Resolution=224x2242022.04 | 46.8 | |
| RVT-BParams (M)=88.5, FLOPS (G)=17.7, Pre-training=ImageNet-1K2024.06 | 46.8 | |
| ConvNeXt-BParams(M)=89, FLOPs(G)=15.42026.04 | 46.8 | |
| Swin-BHAT=true, Resolution=224x2242022.04 | 46.9 | |
| ConViT-BParams (M)=86.5, FLOPS (G)=17.7, Pre-training=ImageNet-1K2024.06 | 46.9 | |
| BiFormer-BParams(M)=57, FLOPs(G)=9.82026.04 | 47.2 | |
| BiFormer-SParams(M)=26, FLOPs(G)=4.52026.04 | 48.5 | |
| Swin-SHAT=true, Resolution=224x2242022.04 | 48.6 | |
| EVT-TParams(M)=15, FLOPs(G)=2.52026.04 | 49.2 | |
| ViT-BPre-trained=ImageNet-21K, Fine-tuned on=ImageNet, Resolution=384x384, Discrete representation=false2021.11 | 49.62 | |
| ViT-SHAT=true, Resolution=224x2242022.04 | 49.7 | |
| QUESTModel=ViT-L/16, Epochs=100, Training Recipe=DeiT-12026.03 | 50.3 | |
| TransNeXt-MicroParams(M)=13, FLOPs(G)=2.72026.04 | 50.8 | |
| DeepAugment + Augmix + DAT2022.09 | 50.82 | |
| Swin-B-AdaNCAParams (M)=90.7, FLOPS (G)=16.3, Pre-training=ImageNet-1K2024.06 | 51.5 | |
| Swin-BHAT=false, Resolution=224x2242022.04 | 51.7 | |
| ConvFormer-S18Params (M)=27, MACS (G)=3.9, Resolution=224, Pre-training=ImageNet-1K2022.10 | 51.7 | |
| Swin-SHAT=false, Resolution=224x2242022.04 | 51.8 | |
| DADBackbone=ViT-B/16, Pre-training=ImageNet-21K, Fine-tuning=ImageNet-1K, Input Resolution=224x2242023.11 | 52 | |
| Swin-BParams (M)=94.1, FLOPS (G)=16.7, Pre-training=ImageNet-22K2024.06 | 53.2 | |
| ViT-SHAT=false, Resolution=224x2242022.04 | 53.3 | |
| ViT-BBackbone=ViT-B2021.11 | 53.51 | |
| DeepAugment + Augmix2022.09 | 53.55 | |
| DeepAugment + AugMixBackbone=ResNet-502020.10 | 53.6 | |
| Swin-THAT=true, Resolution=224x2242022.04 | 53.9 | |
| Swin-BParams (M)=87.8, FLOPS (G)=15.4, Pre-training=ImageNet-1K2024.06 | 54.3 | |
| DINOv2Architecture=ViT-S/14, Pretraining Data=LVD-142M, Resolution=224, Protocol=linear probe on frozen features2023.04 | 54.4 | |
| QKNorm-DS‡Model=ViT-L/16, Epochs=100, Training Recipe=DeiT-12026.03 | 54.4 | |
| AugReg-ViT2022.09 | 54.5 | |
| Swin-BParams(M)=88, FLOPs(G)=15.42026.04 | 54.5 | |
| FocalNet-TParams(M)=29, FLOPs(G)=4.52026.04 | 55 | |
| Dr. ViTBackbone=ViT-B, Discrete Only=true2021.11 | 55.21 | |
| SWAD2021.02 | 55.7 | |
| BiFormer-TParams(M)=13, FLOPs(G)=2.22026.04 | 55.7 | |
| DINOArchitecture=ViT-B/8, Pretraining Data=INet-1k, Resolution=224, Protocol=linear probe on frozen features2023.04 | 56.6 | |
| SWA2021.02 | 56.8 | |
| Dr. ViTBackbone=ViT-S, Discrete Only=false2021.11 | 56.89 | |
| AENIBBackbone=ViT-B/162023.03 | 57.5 | |
| ERM2021.02 | 57.6 | |
| Swin-THAT=false, Resolution=224x2242022.04 | 58 | |
| BaselineBackbone=ViT-B/162023.03 | 58.6 | |
| ViT-BBackbone=ViT-B, Resolution=384x3842021.11 | 59.01 | |
| DeepAugment2022.09 | 60.37 | |
| MAEArchitecture=ViT-H/14, Pretraining Data=INet-1k, Resolution=224, Protocol=linear probe on frozen features2023.04 | 61.4 | |
| ViTBackbone=ViT-B/16, Pre-training=ImageNet-21K, Fine-tuning=ImageNet-1K, Input Resolution=224x2242023.11 | 61.9 | |
| ViT-SBackbone=ViT-S2021.11 | 61.99 | |
| Swin-TParams (M)=29, MACS (G)=4.5, Resolution=224, Pre-training=ImageNet-1K2022.10 | 62 | |
| Swin-TParams(M)=29, FLOPs(G)=4.52026.04 | 62 | |
| PVTv2-B1Params(M)=14, FLOPs(G)=2.12026.04 | 62.6 | |
| AENIBBackbone=ViT-S/162023.03 | 65.2 |