Image Classification on Birdsnap
84.3Top-1 AccuracyEfficientNet-B7
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| EfficientNet-B7Number of Parameters=64M, Pre-training=ImageNet, Transfer Learning=true2019.05 | 84.3 | — | |
| FixResBackbone=SENet-1542019.06 | 84.3 | — | |
| EfficientNet-B72019.06 | 84.3 | — | |
| GPipeNumber of Parameters=556M, Pre-training=ImageNet, Transfer Learning=true2019.05 | 83.6 | — | |
| SENet-154Evaluation Protocol=Baseline2019.06 | 83.4 | — | |
| MLCDPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 83.4 | — | |
| UNICOMPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 82.4 | — | |
| EfficientNet-B5Number of Parameters=28M, Pre-training=ImageNet, Transfer Learning=true2019.05 | 82 | — | |
| Inception-v4Number of Parameters=41M, Pre-training=ImageNet, Transfer Learning=true2019.05 | 81.8 | — | |
| JFT - Adaptive TransferPre-training Source=JFT, Backbone=Inception v3, Evaluation Protocol=Fine-tuning2018.11 | 81.7 | — | |
| JFT - BirdPre-training Source=JFT, Subset=Bird, Backbone=Inception v3, Evaluation Protocol=Fine-tuning2018.11 | 80.7 | — | |
| DFN5B-CLIP-H/14+Zero-shot=true2024.02 | 80.5 | — | |
| EVA-CLIP-18BZero-shot=true2024.02 | 79.9 | — | |
| EVA-CLIP-8BZero-shot=true2024.02 | 79 | — | |
| SimCLRProtocol=Fine-tuned, Backbone=ResNet-50 (4x), Pre-training=ImageNet2020.02 | 78.2 | — | |
| OpenCLIPPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 78 | — | |
| JFT - AnimalPre-training Source=JFT, Subset=Animal, Backbone=Inception v3, Evaluation Protocol=Fine-tuning2018.11 | 77.8 | — | |
| SupervisedProtocol=Fine-tuned, Backbone=ResNet-50 (4x), Pre-training=ImageNet2020.02 | 77.8 | — | |
| CLIP+Pre-training Data=WIT-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 77.8 | — | |
| EVA-02-CLIP-E/14+Zero-shot=true2024.02 | 77.6 | — | |
| DFN5B-CLIP-H/14Zero-shot=true2024.02 | 77.4 | — | |
| ImageNet - Entire DatasetPre-training Source=ImageNet, Backbone=Inception v3, Evaluation Protocol=Fine-tuning2018.11 | 77.2 | — | |
| Random initProtocol=Fine-tuned, Backbone=ResNet-50 (4x), Pre-training=None2020.02 | 77 | — | |
| NNCLR + SupArch=ViT-B/82021.04 | 77 | — | |
| CLIPPre-training Data=WIT-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 77 | — | |
| ImageNet - Adaptive TransferPre-training Source=ImageNet, Backbone=Inception v3, Evaluation Protocol=Fine-tuning2018.11 | 76.6 | — | |
| BYOLArchitecture=ResNet-50, Pre-training Dataset=ImageNet, Evaluation Protocol=Fine-tuned2020.06 | 76.3 | — | |
| BYOLEvaluation Protocol=Fine-tuned, Backbone=ResNet, Pre-training Dataset=ImageNet2020.11 | 76.3 | — | |
| FNCEvaluation Protocol=Fine-tuned, Backbone=ResNet, Pre-training Dataset=ImageNet2020.11 | 76.3 | — | |
| Random initEvaluation protocol=Fine-tuned, Backbone=ResNet-502020.02 | 76.1 | — | |
| Random initArchitecture=ResNet-50, Pre-training Dataset=ImageNet, Evaluation Protocol=Fine-tuned2020.06 | 76.1 | — | |
| SimCLREvaluation protocol=Fine-tuned, Pre-training=ImageNet, Backbone=ResNet-502020.02 | 75.9 | — | |
| SimCLRArchitecture=ResNet-50, Pre-training Dataset=ImageNet, Evaluation Protocol=Fine-tuned2020.06 | 75.9 | — | |
| SimCLR v1Evaluation Protocol=Fine-tuned, Backbone=ResNet, Pre-training Dataset=ImageNet2020.11 | 75.9 | — | |
| SupervisedEvaluation protocol=Fine-tuned, Pre-training=ImageNet, Backbone=ResNet-502020.02 | 75.8 | — | |
| Supervised-INArchitecture=ResNet-50, Pre-training Dataset=ImageNet, Evaluation Protocol=Fine-tuned2020.06 | 75.8 | — | |
| Random InitializationPre-training Source=None, Backbone=Inception v3, Evaluation Protocol=Fine-tuning2018.11 | 75.2 | — | |
| SimCLR (repro)Architecture=ResNet-50, Pre-training Dataset=ImageNet, Evaluation Protocol=Fine-tuned2020.06 | 75 | — | |
| JFT - FoodPre-training Source=JFT, Subset=Food, Backbone=Inception v3, Evaluation Protocol=Fine-tuning2018.11 | 74.9 | — | |
| SimCLR v2Evaluation Protocol=Fine-tuned, Backbone=ResNet, Pre-training Dataset=ImageNet2020.11 | 74.9 | — | |
| JFT - TransportPre-training Source=JFT, Subset=Transport, Backbone=Inception v3, Evaluation Protocol=Fine-tuning2018.11 | 74.4 | — | |
| Entire JFT DatasetPre-training Source=JFT, Backbone=Inception v3, Evaluation Protocol=Fine-tuning2018.11 | 74.2 | — | |
| JFT - VehiclePre-training Source=JFT, Subset=Vehicle, Backbone=Inception v3, Evaluation Protocol=Fine-tuning2018.11 | 74.2 | — | |
| JFT - CarPre-training Source=JFT, Subset=Car, Backbone=Inception v3, Evaluation Protocol=Fine-tuning2018.11 | 73.4 | — | |
| JFT - AircraftPre-training Source=JFT, Subset=Aircraft, Backbone=Inception v3, Evaluation Protocol=Fine-tuning2018.11 | 73.4 | — | |
| OpenCLIP-G/14Zero-shot=true2024.02 | 73 | — | |
| NNCLRArch=ViT-B/82021.04 | 71.7 | — | |
| EVA-01-CLIP-g/14+Zero-shot=true2024.02 | 70.2 | — | |
| InternVL-CZero-shot=true2024.02 | 69.2 | — | |
| nCLIPLinear probing=true, Backbone=ViT-B/16, Pre-train=IT35M2022.10 | 66.8 | — | |
| EVA-01-CLIP-g/14Zero-shot=true2024.02 | 65.8 | — | |
| NNCLRArch=ViT-B/162021.04 | 63.5 | — | |
| FLIPPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot classification2024.07 | 63 | — | |
| MLCDPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot classification2024.07 | 62.9 | — | |
| Sup. INArch=ViT-B/82021.04 | 61.9 | — | |
| NNCLRBackbone=ResNet-50, Pre-training=ImageNet2021.04 | 61.4 | — | |
| NNCLRArch=R502021.04 | 61.4 | — | |
| OpenCLIPPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot classification2024.07 | 61.3 | — | |
| BASIC-LBackbone=BASIC-L, Evaluation Protocol=Zero-shot, Resolution=224x2242021.11 | 59.2 | — | |
| xCLIPLinear probing=true, Backbone=ViT-B/16, Pre-train=IT35M2022.10 | 58.4 | — | |
| BYOLArchitecture=ResNet-50, Pre-training Dataset=ImageNet, Evaluation Protocol=Linear evaluation2020.06 | 57.2 | — | |
| BYOLBackbone=ResNet-50, Pre-training=ImageNet2021.04 | 57.2 | — | |
| BYOLArch=R502021.04 | 57.2 | — | |
| BYOLEvaluation Protocol=Linear eval, Backbone=ResNet, Pre-training Dataset=ImageNet2020.11 | 57.2 | — | |
| SupervisedProtocol=Linear evaluation, Backbone=ResNet-50 (4x), Pre-training=ImageNet2020.02 | 56.4 | — | |
| CLIPLinear probing=true, Backbone=ViT-B/16, Pre-train=IT35M2022.10 | 54.4 | — | |
| FNCEvaluation Protocol=Linear eval, Backbone=ResNet, Pre-training Dataset=ImageNet2020.11 | 54 | — | |
| SupervisedEvaluation protocol=Linear evaluation, Pre-training=ImageNet, Backbone=ResNet-502020.02 | 53.7 | — | |
| Supervised-INArchitecture=ResNet-50, Pre-training Dataset=ImageNet, Evaluation Protocol=Linear evaluation2020.06 | 53.7 | — | |
| Sup.-INBackbone=ResNet-50, Pre-training=ImageNet2021.04 | 53.7 | — | |
| Sup. INArch=R502021.04 | 53.7 | — | |
| ViC-MAEBackbone=ViT/L-16, Pre-training=K710+MiT+IN1K, Evaluation Protocol=Linear Evaluation2023.03 | 53.5 | — | |
| Sup. INArch=ViT-B/162021.04 | 52.8 | — | |
| ViC-MAEBackbone=ViT/L-16, Pre-training=IN1K+K400, Evaluation Protocol=Linear Evaluation2023.03 | 52.8 | — | |
| CLIPPre-training Data=WIT-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot classification2024.07 | 51.2 | — | |
| OmniMAEBackbone=ViT/L-16, Pre-training=SSv2+IN1K, Evaluation Protocol=Linear Evaluation2023.03 | 50.1 | — | |
| MAEBackbone=ViT/L-16, Pre-training=IN1K, Evaluation Protocol=Linear Evaluation2023.03 | 49.8 | — | |
| CLIPBackbone=ViT-L/14-336, Evaluation Protocol=Zero-shot, Resolution=336x3362021.11 | 49.5 | — | |
| BASIC-MBackbone=BASIC-M, Evaluation Protocol=Zero-shot, Resolution=224x2242021.11 | 49.4 | — | |
| AWTTrain=false2024.07 | 48.75 | — | |
| SimCLRProtocol=Linear evaluation, Backbone=ResNet-50 (4x), Pre-training=ImageNet2020.02 | 48.4 | — | |
| CLIP (reported)Pre-training Data=WIT-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot classification2024.07 | 48.3 | — | |
| ViC-MAEBackbone=ViT/B-16, Pre-training=MiT, Evaluation Protocol=Linear Evaluation2023.03 | 48.21 | — | |
| MAEBackbone=ViT/B-16, Pre-training=MiT, Evaluation Protocol=Linear Evaluation2023.03 | 47.98 | — | |
| ViC-MAEBackbone=ViT/B-16, Pre-training=K400, Evaluation Protocol=Linear Evaluation2023.03 | 47.56 | — | |
| MAEBackbone=ViT/B-16, Pre-training=K400, Evaluation Protocol=Linear Evaluation2023.03 | 46.51 | — | |
| SuS-X-SDTrain=false2024.07 | 45.53 | — | |
| PromptAlignTrain=true2024.07 | 45.28 | — | |
| SimCLR v2Evaluation Protocol=Linear eval, Backbone=ResNet, Pre-training Dataset=ImageNet2020.11 | 44.7 | — | |
| TPTTrain=true2024.07 | 44.56 | — | |
| MaPLeTrain=true2024.07 | 44.06 | — | |
| POMPTrain=true2024.07 | 43.94 | — | |
| WaffleCLIPTrain=false2024.07 | 43.92 | — | |
| CoCoOpTrain=true2024.07 | 43.75 | — | |
| VisDescTrain=false2024.07 | 43.64 | — | |
| CLIPTrain=false2024.07 | 42.8 | — | |
| SimCLR (repro)Architecture=ResNet-50, Pre-training Dataset=ImageNet, Evaluation Protocol=Linear evaluation2020.06 | 42.4 | — | |
| SimCLRBackbone=ResNet-50, Pre-training=ImageNet2021.04 | 42.4 | — | |
| SimCLRArch=R502021.04 | 42.4 | — | |
| CoOpTrain=true2024.07 | 41.43 | — |