Image Classification on Flowers-102
99.7Top-1 AccFlorence-CoSwin-H
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Florence-CoSwin-HResolution=384px, Backbone=CoSwin-H, Evaluation Protocol=Linear Probing2021.11 | 99.7 | — | |
| DinoV2Architecture=ViT-S/14+reg (22M), Dataset (Size)=LVD (142M), Evaluation Protocol=Frozen features2025.02 | 99.6 | 99.9 | |
| ViT-12/16Architecture=ViT-12/16, Pre-training Dataset=ImageNet-21K (14.2M), Evaluation Protocol=Fine-tuned2025.02 | 99.6 | — | |
| LOUPEevaluation=linear probing2022.08 | 99.5 | — | |
| ConvMLP-S# Params (M)=9.0, Fine-tuning=true, Pre-trained on=ImageNet-1k2021.09 | 99.5 | — | |
| ConvMLP-M# Params (M)=17.4, Fine-tuning=true, Pre-trained on=ImageNet-1k2021.09 | 99.5 | — | |
| ConvMLP-L# Params (M)=42.7, Fine-tuning=true, Pre-trained on=ImageNet-1k2021.09 | 99.5 | — | |
| ViT-L/16Resolution=384px, Evaluation Protocol=Linear Probing2021.11 | 99.3 | — | |
| CLIP-ViT-L/14Resolution=336px, Evaluation Protocol=Linear Probing2021.11 | 99.2 | — | |
| CLIPevaluation=linear probing2022.08 | 99.2 | — | |
| GrafitBackbone=RegNetY-8.0GF, Pre-training=ImageNet-1k, Resolution=384x384, Number of Parameters=39M, Classifier=MLP2020.11 | 99.1 | — | |
| GrafitBackbone=RegNetY-8.0GF, Parameters=39M, Resolution=384x3842020.11 | 99.1 | — | |
| Grafit RegNetY-8GFim/sec=591.62020.12 | 99 | — | |
| DeiT-B⚗ ↑384im/sec=85.9, resolution=384, distillation=true2020.12 | 98.9 | — | |
| CLIP-ResNet-50x64Evaluation Protocol=Linear Probing2021.11 | 98.9 | — | |
| DeiT-III-LParam (M)=304, FLOPs (G)=61.6, Pre-training=ImageNet-1k2024.03 | 98.9 | — | |
| DeiT-B# Params (M)=86.6, Fine-tuning=true, Pre-trained on=ImageNet-1k2021.09 | 98.9 | — | |
| EfficientNet-B7Pre-training=ImageNet-1k, Resolution=600x600, Number of Parameters=64M2020.11 | 98.8 | — | |
| EfficientNet-B7Parameters=64M, Resolution=600x6002020.11 | 98.8 | — | |
| EfficientNet-B7im/sec=55.12020.12 | 98.8 | — | |
| DeiT-B⚗im/sec=290.9, resolution=224, distillation=true2020.12 | 98.8 | — | |
| EfficientNet-B7FLOPs=37.0B, Resolution=600, Pre-training Dataset=ImageNet-1k, Evaluation Protocol=Fine-tune2021.05 | 98.8 | — | |
| RDNet-TParam (M)=24, FLOPs (G)=5.0, Pre-training=ImageNet-1k2024.03 | 98.6 | — | |
| DeiT-III-BParam (M)=87, FLOPs (G)=17.5, Pre-training=ImageNet-1k2024.03 | 98.6 | — | |
| RDNet-BParam (M)=87, FLOPs (G)=15.4, Pre-training=ImageNet-1k2024.03 | 98.6 | — | |
| RDNet-LParam (M)=186, FLOPs (G)=34.7, Pre-training=ImageNet-1k2024.03 | 98.6 | — | |
| DeiT-B ↑384im/sec=85.9, resolution=3842020.12 | 98.5 | — | |
| EfficientNet-B5Params (M)=302022.02 | 98.5 | — | |
| DeiT-BArchitecture=DeiT-B, Pre-training Dataset=ImageNet-1K (1.3M), Evaluation Protocol=Fine-tuned2025.02 | 98.5 | — | |
| DeiT-Bim/sec=292.3, resolution=2242020.12 | 98.4 | — | |
| Deit-B/16FLOPs=17.5B, Resolution=224, Pre-training Dataset=ImageNet-1k, Evaluation Protocol=Fine-tune2021.05 | 98.4 | — | |
| DeiT-BParams (M)=86.62022.02 | 98.4 | — | |
| RDNet-SParam (M)=50, FLOPs (G)=8.7, Pre-training=ImageNet-1k2024.03 | 98.4 | — | |
| DeiT-BParam (M)=87, FLOPs (G)=17.5, Pre-training=ImageNet-1k2024.03 | 98.4 | — | |
| ViT-B# Params (M)=86.6, Fine-tuning=true, Pre-trained on=ImageNet-1k2021.09 | 98.4 | — | |
| SNCA+Backbone=ResNet-50, Pre-training=ImageNet, Resolution=224x224, Evaluation Protocol=Single center crop, Classifier=MLP2020.11 | 98.2 | — | |
| GrafitBackbone=ResNet-50, Pre-training=ImageNet, Resolution=224x224, Evaluation Protocol=Single center crop, Classifier=MLP2020.11 | 98.2 | — | |
| Grafit ResNet-50im/sec=1226.12020.12 | 98.2 | — | |
| Grafit ResNet-50Params (M)=25.62022.02 | 98.2 | — | |
| Grafit ResNet-50Param (M)=26, FLOPs (G)=4.1, Pre-training=ImageNet-1k2024.03 | 98.2 | — | |
| CLIP*Image Encoder=ViT-B/16, Pretrain Dataset Size=400M, Evaluation Protocol=Linear probe2022.04 | 98.1 | — | |
| ResMLP-S24FLOPs=6.0B, Resolution=224, Pre-training Dataset=ImageNet-1k, Evaluation Protocol=Fine-tune2021.05 | 97.9 | — | |
| EfficientNet-L2Resolution=800px, Evaluation Protocol=Linear Probing2021.11 | 97.9 | — | |
| LP-FTEvaluation Protocol=Linear Probing then Fine-tuning2022.12 | 97.9 | — | |
| ViTAE-SParams (M)=23.62022.02 | 97.8 | — | |
| FLYPEvaluation Protocol=Fine-tuning2022.12 | 97.7 | — | |
| SCOTT-12/16Features=MIM-JEPA, Dataset (Size)=†, Evaluation Protocol=Frozen features2025.02 | 97.7 | 99.2 | |
| SCOTT-12/16Architecture=SCOTT-12/16, Pre-training Dataset=Target dataset (unlabeled), Evaluation Protocol=Frozen, Training Epochs=12002025.02 | 97.7 | — | |
| Grafit FCBackbone=ResNet-50, Pre-training=ImageNet, Resolution=224x224, Evaluation Protocol=Single center crop, Classifier=Linear FC2020.11 | 97.6 | — | |
| Grafit/ResNet50FLOPs=4.1B, Resolution=224, Pre-training Dataset=ImageNet-1k, Evaluation Protocol=Fine-tune2021.05 | 97.6 | — | |
| ViTAE-TParams (M)=4.82022.02 | 97.5 | — | |
| ResMLP-S12FLOPs=3.0B, Resolution=224, Pre-training Dataset=ImageNet-1k, Evaluation Protocol=Fine-tune2021.05 | 97.4 | — | |
| ResMLP-S12# Params (M)=15.4, Fine-tuning=true, Pre-trained on=ImageNet-1k2021.09 | 97.4 | — | |
| ResMLP-S24# Params (M)=30.0, Fine-tuning=true, Pre-trained on=ImageNet-1k2021.09 | 97.4 | — | |
| SCOTT-12/16Architecture=SCOTT-12/16, Pre-training Dataset=Target dataset (unlabeled), Evaluation Protocol=Frozen2025.02 | 97.1 | — | |
| CLIP*Image Encoder=ViT-B/32, Pretrain Dataset Size=400M, Evaluation Protocol=Linear probe2022.04 | 96.9 | — | |
| SCOTT-7/16Features=MIM-JEPA, Dataset (Size)=†, Evaluation Protocol=Frozen features2025.02 | 96.9 | 99.3 | |
| SCOTT-7/16Architecture=SCOTT-7/16, Pre-training Dataset=Target dataset (unlabeled), Evaluation Protocol=Frozen, Training Epochs=12002025.02 | 96.9 | — | |
| Baseline#seeds=5, Backbone=ViT, Training Protocol=Pre-trained2026.04 | 96.83 | — | |
| DeiT-III-SParam (M)=22, FLOPs (G)=4.6, Pre-training=ImageNet-1k2024.03 | 96.4 | — | |
| SimCLRv2Backbone=ResNet-152x3, Evaluation Protocol=Linear Probing2021.11 | 96.3 | — | |
| BaselineBackbone=ResNet-50, Pre-training=ImageNet, Resolution=224x224, Evaluation Protocol=Single center crop, Classifier=Linear FC2020.11 | 96.2 | — | |
| ClusterFit+Backbone=ResNet-50, Pre-training=ImageNet, Resolution=224x224, Evaluation Protocol=Single center crop, Classifier=Linear FC2020.11 | 96.2 | — | |
| ResNet50FLOPs=4.1B, Resolution=224, Pre-training Dataset=ImageNet-1k, Evaluation Protocol=Fine-tune2021.05 | 96.2 | — | |
| ITOPre-training Dataset=DataComp-1B, Backbone=ViT-B/16, Training Epochs=10, Evaluation Protocol=Linear Probing2026.03 | 96.16 | — | |
| ViT-B/16 [71]Param (M)=86, FLOPs (G)=35.1, Pre-training=ImageNet-1k2024.03 | 96 | — | |
| LPEvaluation Protocol=Linear Probing2022.12 | 95.9 | — | |
| ITOPre-training Dataset=DataComp-1B, Backbone=ViT-L/16, Training Epochs=1, Evaluation Protocol=Linear Probing2026.03 | 95.8 | — | |
| ViT-12/16Architecture=ViT-12/16, Pre-training Dataset=ImageNet-1K (1.3M), Evaluation Protocol=Fine-tuned2025.02 | 95.7 | — | |
| SCOTT-7/16Architecture=SCOTT-7/16, Pre-training Dataset=Target dataset (unlabeled), Evaluation Protocol=Frozen2025.02 | 95.7 | — | |
| PyramidCLIPImage Encoder=ViT-B/16, Pretrain Dataset Size=143M, Evaluation Protocol=Linear probe2022.04 | 95.6 | — | |
| PEFTShots=102023.05 | 95.2 | — | |
| ITO sub2Pre-training Dataset=DataComp-1B, Backbone=ViT-B/16, Training Epochs=10, Evaluation Protocol=Linear Probing2026.03 | 94.96 | — | |
| ITO sub2Pre-training Dataset=DataComp-1B, Backbone=ViT-L/16, Training Epochs=1, Evaluation Protocol=Linear Probing2026.03 | 94.89 | — | |
| CLIPPre-training Dataset=DataComp-1B, Backbone=ViT-B/16, Training Epochs=10, Evaluation Protocol=Linear Probing2026.03 | 94.67 | — | |
| LottaLoRARank=1, #seeds=5, % trainable=1.19%, Backbone=ViT, Training Protocol=Pre-trained2026.04 | 94.53 | — | |
| PyramidCLIPImage Encoder=ViT-B/32, Pretrain Dataset Size=143M, Evaluation Protocol=Linear probe2022.04 | 93.9 | — | |
| I-JEPAArchitecture=ViT-H/14 (630M), Dataset (Size)=ImageNet-1k (1.3M), Evaluation Protocol=Frozen features2025.02 | 93.7 | 98.5 | |
| CLIPPre-training Dataset=DataComp-1B, Backbone=ViT-L/16, Training Epochs=1, Evaluation Protocol=Linear Probing2026.03 | 93.04 | — | |
| DFN5B-CLIP-H/14Zero-shot=true2024.02 | 92.5 | — | |
| ITOPre-training Dataset=DataComp-1B, Backbone=ViT-B/16, Training Epochs=1, Evaluation Protocol=Linear Probing2026.03 | 92.47 | — | |
| ITO sub2Pre-training Dataset=YFCC15M, Backbone=ViT-B/16, Training Epochs=30, Evaluation Protocol=Linear Probing2026.03 | 92.37 | — | |
| ITO sub2Pre-training Dataset=DataComp-1B, Backbone=ViT-B/16, Training Epochs=1, Evaluation Protocol=Linear Probing2026.03 | 92.24 | — | |
| ITOPre-training Dataset=YFCC15M, Backbone=ViT-B/16, Training Epochs=30, Evaluation Protocol=Linear Probing2026.03 | 92.18 | — | |
| CoOpShots=102023.05 | 92.1 | — | |
| DFN5B-CLIP-H/14+Zero-shot=true2024.02 | 91.6 | — | |
| ViT-L/16 [71]Param (M)=303, FLOPs (G)=122.9, Pre-training=ImageNet-1k2024.03 | 91.4 | — | |
| PEFTShots=52023.05 | 91.1 | — | |
| FTEvaluation Protocol=Fine-tuning2022.12 | 90.4 | — | |
| ITOPre-training Dataset=Laion100M, Backbone=ViT-B/16, Training Epochs=30, Evaluation Protocol=Linear Probing2026.03 | 90.21 | — | |
| FLIP-ViT-L/14Backbone=ViT-L/14, Evaluation Protocol=Zero-shot2021.11 | 90.1 | — | |
| CLIPPre-training Dataset=DataComp-1B, Backbone=ViT-B/16, Training Epochs=1, Evaluation Protocol=Linear Probing2026.03 | 90.08 | — | |
| ViT-L/16im/sec=27.32020.12 | 89.7 | — | |
| ViT-L/16FLOPs=190.7B, Resolution=384, Pre-training Dataset=ImageNet-1k, Evaluation Protocol=Fine-tune2021.05 | 89.7 | — | |
| ViT-L/16Params (M)=304.32022.02 | 89.7 | — | |
| ViT-L/16 [21]Param (M)=303, FLOPs (G)=122.9, Pre-training=ImageNet-1k2024.03 | 89.7 | — | |
| ViT-B/16im/sec=85.92020.12 | 89.5 | — | |
| ViT-B/16FLOPs=55.5B, Resolution=384, Pre-training Dataset=ImageNet-1k, Evaluation Protocol=Fine-tune2021.05 | 89.5 | — | |
| ViT-B/16Params (M)=86.52022.02 | 89.5 | — | |
| ViT-B/16 [21]Param (M)=86, FLOPs (G)=35.1, Pre-training=ImageNet-1k2024.03 | 89.5 | — |