Classification on Food101
96.2Top-1 AccuracyFlorence-CoSwin-H
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Florence-CoSwin-HResolution=384px, Backbone=CoSwin-H, Evaluation Protocol=Linear Probing2021.11 | 96.2 | — | — | |
| CLIP-ViT-L/14Resolution=336px, Evaluation Protocol=Linear Probing2021.11 | 95.9 | — | — | |
| Food-R12026.06 | 95.21 | — | — | |
| Florence-CoSwin-HBackbone=CoSwin-H, Resolution=384px, Evaluation Protocol=Zero-shot2021.11 | 95.1 | — | — | |
| CLIP-ResNet-50x64Evaluation Protocol=Linear Probing2021.11 | 94.8 | — | — | |
| RoDE2026.06 | 94.02 | — | — | |
| FoodLMM2026.06 | 93.93 | — | — | |
| CLIP-ViT-L/14Backbone=ViT-L/14, Resolution=336px, Evaluation Protocol=Zero-shot2021.11 | 93.8 | — | — | |
| GrafitBackbone=RegNetY-8.0GF, Pre-training=ImageNet-1k, Resolution=384x384, Number of Parameters=39M, Classifier=MLP2020.11 | 93.7 | — | — | |
| GrafitBackbone=RegNetY-8.0GF, Parameters=39M, Resolution=384x3842020.11 | 93.7 | — | — | |
| CAE*Supervised=false, Self-supervised=true, Backbone=ViT-B2022.02 | 93.32 | — | — | |
| MAESupervised=false, Self-supervised=true, Backbone=ViT-B2022.02 | 93.19 | — | — | |
| EfficientNet-B7Pre-training=ImageNet-1k, Resolution=600x600, Number of Parameters=64M2020.11 | 93 | — | — | |
| EfficientNet-B7Parameters=64M, Resolution=600x6002020.11 | 93 | — | — | |
| SigLIP2Data size=10B, Zero-shot=true2026.03 | 92.8 | — | — | |
| PEcoreData size=5B, Zero-shot=true2026.03 | 92.5 | — | — | |
| FLIP-ViT-L/14Backbone=ViT-L/14, Evaluation Protocol=Zero-shot2021.11 | 92.2 | — | — | |
| EfficientNet-L2Resolution=800px, Evaluation Protocol=Linear Probing2021.11 | 92 | — | — | |
| DeiTSupervised=true, Self-supervised=false, Backbone=ViT-B2022.02 | 91.81 | — | — | |
| CLIP-ResNet-50x64Backbone=ResNet-50x64, Evaluation Protocol=Zero-shot2021.11 | 91.8 | — | — | |
| DINOSupervised=false, Self-supervised=true, Backbone=ViT-B2022.02 | 91.67 | — | — | |
| SigLIPData size=10B, Zero-shot=true2026.03 | 91.6 | — | — | |
| PRENet2026.06 | 91.13 | — | — | |
| ITOPre-training Dataset=DataComp-1B, Backbone=ViT-B/16, Training Epochs=10 epochs, Zero-shot protocol=EVA-CLIP2026.03 | 90.8 | — | — | |
| CLIPPre-training Dataset=DataComp-1B, Backbone=ViT-B/16, Training Epochs=10 epochs, Zero-shot protocol=EVA-CLIP2026.03 | 90.2 | — | — | |
| LSA+OnZetaBackbone=ViT-B/162026.06 | 89.85 | — | — | |
| ITO sub2Pre-training Dataset=DataComp-1B, Backbone=ViT-B/16, Training Epochs=10 epochs, Zero-shot protocol=EVA-CLIP2026.03 | 89.7 | — | — | |
| LSABackbone=ViT-B/162026.06 | 89.51 | — | — | |
| GrafitBackbone=ResNet-50, Pre-training=ImageNet, Resolution=224x224, Evaluation Protocol=Single center crop, Classifier=MLP2020.11 | 89.5 | — | — | |
| OnZetaBackbone=ViT-B/162026.06 | 89.07 | — | — | |
| LlipData size=2.5B, Zero-shot=true2026.03 | 89 | — | — | |
| BaselineBackbone=ResNet-50, Pre-training=ImageNet, Resolution=224x224, Evaluation Protocol=Single center crop, Classifier=Linear FC2020.11 | 88.9 | — | — | |
| ClusterFit+Backbone=ResNet-50, Pre-training=ImageNet, Resolution=224x224, Evaluation Protocol=Single center crop, Classifier=Linear FC2020.11 | 88.9 | — | — | |
| CLIP baselineBackbone=ViT-B/162026.06 | 88.86 | — | — | |
| SNCA+Backbone=ResNet-50, Pre-training=ImageNet, Resolution=224x224, Evaluation Protocol=Single center crop, Classifier=MLP2020.11 | 88.8 | — | — | |
| Grafit FCBackbone=ResNet-50, Pre-training=ImageNet, Resolution=224x224, Evaluation Protocol=Single center crop, Classifier=Linear FC2020.11 | 88.7 | — | — | |
| MetaCLIPData size=2.5B, Zero-shot=true2026.03 | 88.3 | — | — | |
| Noun SubmanifoldBackbone=CLIP ViT-L-14, zero-shot=true2023.05 | 87.7 | — | — | |
| POS PGABackbone=CLIP ViT-L-14, zero-shot=true2023.05 | 87.5 | — | — | |
| ViT-L/16Resolution=384px, Evaluation Protocol=Linear Probing2021.11 | 87.4 | — | — | |
| CLIPBackbone=CLIP ViT-L-14, zero-shot=true2023.05 | 87.4 | — | — | |
| POS PCABackbone=CLIP ViT-L-14, zero-shot=true2023.05 | 87.4 | — | — | |
| TPTBackbone=ViT-B/162026.06 | 86.97 | — | — | |
| OpenCLIPData size=2B, Zero-shot=true2026.03 | 86.2 | — | — | |
| SimCLRv2Backbone=ResNet-152x3, Evaluation Protocol=Linear Probing2021.11 | 83.6 | — | — | |
| Random Init.Supervised=false, Self-supervised=false, Backbone=ViT-B2022.02 | 82.77 | — | — | |
| LSA+OnZetaBackbone=ResNet-502026.06 | 82.01 | — | — | |
| LSABackbone=ResNet-502026.06 | 81.58 | — | — | |
| OnZetaBackbone=ResNet-502026.06 | 80.85 | — | — | |
| CLIP baselineBackbone=ResNet-502026.06 | 80.59 | — | — | |
| RényiCLProtocol=linear evaluation2022.08 | 78 | — | — | |
| NNCLRBackbone=ResNet-50, Pre-training=ImageNet2021.04 | 76.7 | — | — | |
| NNCLRBackbone=ResNet-50, Pre-training Dataset=ImageNet-1K, Pre-training Epochs=8002022.03 | 76.7 | — | — | |
| NNCLRProtocol=linear evaluation2022.08 | 76.7 | — | — | |
| CaCoBackbone=ResNet-50, Pre-training Dataset=ImageNet-1K, Pre-training Epochs=8002022.03 | 75.4 | — | — | |
| DreamLIPData size=30M, Zero-shot=true2026.03 | 75.4 | — | — | |
| BYOLBackbone=ResNet-50, Pre-training=ImageNet2021.04 | 75.3 | — | — | |
| BYOLBackbone=ResNet-50, Pre-training Dataset=ImageNet-1K, Pre-training Epochs=8002022.03 | 75.3 | — | — | |
| BYOLProtocol=linear evaluation2022.08 | 75.3 | — | — | |
| TPTBackbone=ResNet-502026.06 | 74.89 | — | — | |
| GoldiCLIPData size=30M, Zero-shot=true2026.03 | 74 | — | — | |
| COSMOSData size=30M, Zero-shot=true2026.03 | 73.9 | — | — | |
| AdCo +Backbone=ResNet-50, Pre-training Dataset=ImageNet-1K, Pre-training Epochs=8002022.03 | 73.8 | — | — | |
| SimCLRBackbone=ResNet-50, Pre-training=ImageNet2021.04 | 72.8 | — | — | |
| SimCLRBackbone=ResNet-50, Pre-training Dataset=ImageNet-1K, Pre-training Epochs=8002022.03 | 72.8 | — | — | |
| FLAIRData size=30M, Zero-shot=true2026.03 | 72.5 | — | — | |
| Sup.-INBackbone=ResNet-50, Pre-training=ImageNet2021.04 | 72.3 | — | — | |
| SupervisedProtocol=linear evaluation2022.08 | 72.3 | — | — | |
| MoCo v2+Backbone=ResNet-50, Pre-training Dataset=ImageNet-1K, Pre-training Epochs=8002022.03 | 71.5 | — | — | |
| SimCLRProtocol=linear evaluation2022.08 | 68.4 | — | — | |
| SigLIPData size=30M, Zero-shot=true2026.03 | 64.2 | — | — | |
| CLIPData size=30M, Zero-shot=true2026.03 | 61.3 | — | — | |
| CLIP-PGSImage Encoder=ViT-B/16, p=0.3, Protocol=Zero-shot2025.03 | 46.5 | — | — | |
| CLIP-PGSImage Encoder=ViT-B/16, p=0.5, Protocol=Zero-shot2025.03 | 42.8 | — | — | |
| CLIPImage Encoder=ViT-B/16, Protocol=Zero-shot2025.03 | 42.3 | — | — | |
| E-CLIPImage Encoder=ViT-B/16, Protocol=Zero-shot2025.03 | 42.1 | — | — | |
| A-CLIPImage Encoder=ViT-B/16, Protocol=Zero-shot2025.03 | 41.8 | — | — | |
| FLIPImage Encoder=ViT-B/16, Protocol=Zero-shot2025.03 | 39.9 | — | — | |
| CLIP-PGSImage Encoder=ViT-S/16, p=0.3, Protocol=Zero-shot2025.03 | 39.1 | — | — | |
| CLIP-PGSImage Encoder=ViT-S/16, p=0.5, Protocol=Zero-shot2025.03 | 38.7 | — | — | |
| ITO sub2Backbone=ViT-B/16, Pre-training Dataset=CC3M-recap, Evaluation Protocol=Zero-shot2026.03 | 23.1 | — | — | |
| ITO sub3Backbone=ViT-B/16, Pre-training Dataset=CC3M-recap, Evaluation Protocol=Zero-shot2026.03 | 23.1 | — | — | |
| FLAIRBackbone=ViT-B/16, Pre-training Dataset=CC3M-recap, Evaluation Protocol=Zero-shot2026.03 | 19.8 | — | — | |
| BiFTAProtocol=Zero-shot, Backbone=RN-1012026.01 | — | 84.02 | — | |
| BiFTAZero-shot=true, Backbone=Average over five CLIP variants2026.01 | — | 87.18 | — | |
| BiFTABackbone=ViT-B/32, Evaluation Protocol=Zero-shot2026.01 | — | 86.43 | — | |
| BiFTABackbone=RN-50, Zero-shot evaluation protocol=true2026.01 | — | 81.39 | — | |
| BYOLArchitecture=ResNet-50, Pre-training Dataset=ImageNet, Evaluation Protocol=Linear evaluation2020.06 | — | 75.3 | — | |
| BYOLArchitecture=ResNet-50, Pre-training Dataset=ImageNet, Evaluation Protocol=Fine-tuned2020.06 | — | 88.5 | — | |
| C-TPTBackbone=ResNet50, Optim.=true2026.07 | — | 74.83 | — | |
| CaBMBottleneck type=Strictly text bottleneck, Protocol=text-only classifier2026.07 | — | 93.68 | — | |
| CLIPBackbone=ViT-B/16, Pre-trained on=IT35M, Mode=Zero-shot2022.10 | — | 61.2 | — | |
| CLIPProtocol=Zero-shot, Backbone=RN-1012026.01 | — | 82.44 | — | |
| CLIPBackbone=ViT-B/32, Evaluation Protocol=Zero-shot2026.01 | — | 82.6 | — | |
| CLIPBackbone=RN-50, Zero-shot evaluation protocol=true2026.01 | — | 78.62 | — | |
| CLIPShots=0, Venue=ICML’222025.12 | — | 77.3 | — | |
| CLIPWay=5, Shot=52026.03 | — | 86 | — | |
| CLIPBackbone=ResNet50, Optim.=false2026.07 | — | 73.94 | — | |
| CLIP-DProtocol=Zero-shot, Backbone=RN-1012026.01 | — | 83.25 | — | |
| CLIP-DBackbone=ViT-B/32, Evaluation Protocol=Zero-shot2026.01 | — | 84.12 | — |