Image Classification on ImageNet matched frequency V2 (test)
76.9Top-1 AccuracyCaiT-M48↑448Y
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| CaiT-M48↑448YParameters=356M, FLOPs=329.6B, Training Resolution=224, Test Resolution=4482021.03 | 76.9 | — | |
| CaiT-M36↑448YParameters=271M, FLOPs=247.8B, Training Resolution=224, Test Resolution=4482021.03 | 76.7 | — | |
| CaiT-M36↑384YParameters=271M, FLOPs=173.3B, Training Resolution=224, Test Resolution=3842021.03 | 76.3 | — | |
| CaiT-S36↑384YParameters=68M, FLOPs=48.0B, Training Resolution=224, Test Resolution=3842021.03 | 76.2 | — | |
| RobustViTBackbone=AR-L, Annotated segmentation=false2022.06 | 76.1 | 93.6 | |
| EfficientNet-B7 AdvPropParameters=66M, FLOPs=37.0B, Training Resolution=600, Test Resolution=6002021.03 | 76 | — | |
| Fix-EfficientNet-B8Parameters=87M, FLOPs=89.5B, Training Resolution=672, Test Resolution=8002021.03 | 75.9 | — | |
| NFNet-F6+SAMParameters=438M, FLOPs=377.3B, Training Resolution=448, Test Resolution=5762021.03 | 75.8 | — | |
| OriginalBackbone=AR-L, Annotated segmentation=false2022.06 | 75.8 | 93.4 | |
| RobustViTBackbone=AR-L, Annotated segmentation=true2022.06 | 75.8 | 93.4 | |
| CaiT-S48↑384Parameters=89M, FLOPs=63.8B, Training Resolution=224, Test Resolution=3842021.03 | 75.5 | — | |
| NFNet-F3Parameters=255M, FLOPs=114.8B, Training Resolution=320, Test Resolution=4162021.03 | 75.2 | — | |
| NFNet-F4Parameters=316M, FLOPs=215.3B, Training Resolution=384, Test Resolution=5122021.03 | 75.2 | — | |
| DeiT-B↑384Y 1000 epochsParameters=87M, FLOPs=55.5B, Training Resolution=224, Test Resolution=3842021.03 | 75.2 | — | |
| CaiT-S36↑384Parameters=68M, FLOPs=48.0B, Training Resolution=224, Test Resolution=3842021.03 | 75 | — | |
| NFNet-F5Parameters=377M, FLOPs=289.8B, Training Resolution=416, Test Resolution=5442021.03 | 74.6 | — | |
| NFNet-F1Parameters=133M, FLOPs=35.5B, Training Resolution=224, Test Resolution=3202021.03 | 74.4 | — | |
| NFNet-F2Parameters=194M, FLOPs=62.6B, Training Resolution=256, Test Resolution=3522021.03 | 74.3 | — | |
| Right for the Right Reason (RRR)Backbone=AR-B, Annotated segmentation=true2022.06 | 74.3 | 92.6 | |
| CaiT-S36YParameters=68M, FLOPs=13.9B, Training Resolution=224, Test Resolution=2242021.03 | 74.1 | — | |
| GradMaskBackbone=AR-B, Annotated segmentation=true2022.06 | 74 | 92.6 | |
| OriginalBackbone=AR-B, Annotated segmentation=false2022.06 | 73.8 | 92.3 | |
| RobustViTBackbone=AR-B, Annotated segmentation=false2022.06 | 73.7 | 92.4 | |
| EfficientNet-B5Parameters=30M, FLOPs=9.9B, Training Resolution=456, Test Resolution=4562021.03 | 73.6 | — | |
| RobustViTBackbone=AR-B, Annotated segmentation=true2022.06 | 73.5 | 92 | |
| NFNet-F0Parameters=72M, FLOPs=12.4B, Training Resolution=192, Test Resolution=2562021.03 | 72.6 | — | |
| CaiT-S36Parameters=68M, FLOPs=13.9B, Training Resolution=224, Test Resolution=2242021.03 | 72.5 | — | |
| RegNetY-16GFParameters=84M, FLOPs=16.0B, Training Resolution=224, Test Resolution=2242021.03 | 72.4 | — | |
| DeiT-B↑384Parameters=86M, FLOPs=55.4B, Training Resolution=224, Test Resolution=3842021.03 | 72.4 | — | |
| RobustViTBackbone=ViT-L, Annotated segmentation=false2022.06 | 72.1 | 91.2 | |
| OriginalBackbone=ViT-L, Annotated segmentation=false2022.06 | 71.8 | 90.7 | |
| DeiT-BParameters=86M, FLOPs=17.5B, Training Resolution=224, Test Resolution=2242021.03 | 71.5 | — | |
| GradMaskBackbone=ViT-B, Annotated segmentation=true2022.06 | 71.4 | 90.5 | |
| Right for the Right Reason (RRR)Backbone=ViT-B, Annotated segmentation=true2022.06 | 71.4 | 90.5 | |
| RobustViTBackbone=ViT-L, Annotated segmentation=true2022.06 | 71.3 | 90.6 | |
| OriginalBackbone=ViT-B, Annotated segmentation=false2022.06 | 71.1 | 89.9 | |
| Right for the Right Reason (RRR)Backbone=AR-S, Annotated segmentation=true2022.06 | 70.3 | 90.1 | |
| GradMaskBackbone=AR-S, Annotated segmentation=true2022.06 | 70.1 | 90.3 | |
| RobustViTBackbone=ViT-B, Annotated segmentation=true2022.06 | 70 | 89.4 | |
| OriginalBackbone=AR-S, Annotated segmentation=false2022.06 | 69.9 | 90.1 | |
| RobustViTBackbone=ViT-B, Annotated segmentation=false2022.06 | 69.8 | 89.4 | |
| OriginalBackbone=DeiT-B, Annotated segmentation=false2022.06 | 69.7 | 86.8 | |
| GradMaskBackbone=DeiT-B, Annotated segmentation=true2022.06 | 69.7 | 88.7 | |
| RobustViTBackbone=AR-S, Annotated segmentation=true2022.06 | 69.6 | 90 | |
| RobustViTBackbone=AR-S, Annotated segmentation=false2022.06 | 69.6 | 90.1 | |
| Right for the Right Reason (RRR)Backbone=DeiT-B, Annotated segmentation=true2022.06 | 69.5 | 88.6 | |
| RobustViTBackbone=DeiT-B, Annotated segmentation=false2022.06 | 69.3 | 88.5 | |
| RobustViTBackbone=DeiT-B, Annotated segmentation=true2022.06 | 69.1 | 88.3 | |
| DeiT-SParameters=22M, FLOPs=4.6B, Training Resolution=224, Test Resolution=2242021.03 | 68.5 | — | |
| RobustViTBackbone=DeiT-S, Annotated segmentation=true2022.06 | 67.3 | 87.3 | |
| RobustViTBackbone=DeiT-S, Annotated segmentation=false2022.06 | 67.1 | 87.4 | |
| OriginalBackbone=DeiT-S, Annotated segmentation=false2022.06 | 66.5 | 86.6 | |
| Right for the Right Reason (RRR)Backbone=DeiT-S, Annotated segmentation=true2022.06 | 66 | 86.7 | |
| RELICv2Protocol=Linear evaluation, Backbone=ResNet-50, Pre-training=ImageNet, Evaluation=Zero-shot2022.01 | 65.3 | — | |
| SupervisedProtocol=Linear evaluation, Backbone=ResNet-50, Pre-training=ImageNet, Evaluation=Zero-shot2022.01 | 65.1 | — | |
| GradMaskBackbone=DeiT-S, Annotated segmentation=true2022.06 | 64.5 | 85.6 | |
| RELICProtocol=Linear evaluation, Backbone=ResNet-50, Pre-training=ImageNet, Evaluation=Zero-shot2022.01 | 63.1 | — | |
| BYOLProtocol=Linear evaluation, Backbone=ResNet-50, Pre-training=ImageNet, Evaluation=Zero-shot2022.01 | 62.2 | — | |
| HO-HEBackbone=ViT-B/162026.07 | 61.25 | — | |
| RealismBackbone=ViT-B/162026.07 | 60.61 | — | |
| Random SamplingBackbone=ViT-B/162026.07 | 60.57 | — | |
| SBSimBackbone=ViT-B/162026.07 | 60.16 | — | |
| RELICv2Epochs=5000, Pre-training Dataset=JFT-300M, Evaluation Protocol=linear evaluation2022.01 | 59.1 | — | |
| RELICv2Epochs=3000, Pre-training Dataset=JFT-300M, Evaluation Protocol=linear evaluation2022.01 | 58.6 | — | |
| RELICv2Epochs=1000, Pre-training Dataset=JFT-300M, Evaluation Protocol=linear evaluation2022.01 | 57.6 | — | |
| CLIP-AlignBackbone=ViT-B/162026.07 | 56.3 | — | |
| SimCLRProtocol=Linear evaluation, Backbone=ResNet-50, Pre-training=ImageNet, Evaluation=Zero-shot2022.01 | 53.2 | — | |
| OriginalBackbone=ViT-B/162026.07 | 49.86 | — |