Image Classification on ImageNet (Accuracy)
91AccuracyCoCa
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| CoCaEvaluation Protocol=finetuned2022.05 | 91 | — | |
| CoAtNet2022.05 | 90.9 | — | |
| ViT-G + Model Soups2022.05 | 90.9 | — | |
| CoCaEvaluation Protocol=frozen2022.05 | 90.6 | — | |
| ViT-G2022.05 | 90.5 | — | |
| MetaPseudoLabels2022.05 | 90.2 | — | |
| Florence2022.05 | 90.1 | — | |
| ALIGN2022.05 | 88.6 | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=End-to-End, Mixing Coefficient (alpha)=optimal2021.09 | 87.1 | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=End-to-End, Mixing Coefficient (alpha)=0.52021.09 | 86.8 | — | |
| Fine-tuned E2EModel=CLIP ViT-L/14@336px, Evaluation Protocol=End-to-End, Implementation=ours2021.09 | 86.2 | — | |
| Linear ProbeBackbone=DINOv2 ViT-g/142024.06 | 86.2 | — | |
| Fine-tuned LCModel=CLIP ViT-L/14@336px, Evaluation Protocol=Linear Classifier, Source=[82]2021.09 | 85.4 | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=Linear Classifier, Mixing Coefficient (alpha)=optimal2021.09 | 85.3 | — | |
| Fine-tuned LCModel=CLIP ViT-L/14@336px, Evaluation Protocol=Linear Classifier, Implementation=ours2021.09 | 85.2 | — | |
| LiTTraining Data Source=Private, Evaluation Protocol=Zero-shot2021.11 | 85.2 | — | |
| Dr. ViTPre-trained=ImageNet-21K, Fine-tuned on=ImageNet, Resolution=512x512, Discrete representation=true2021.11 | 85.07 | — | |
| Dr. ViTPre-trained=ImageNet-21K, Fine-tuned on=ImageNet, Resolution=384x384, Discrete representation=true2021.11 | 84.43 | — | |
| Knowledge-CLIPtraining_mode=fine-tuned2022.10 | 84.4 | — | |
| ViT-BPre-trained=ImageNet-21K, Fine-tuned on=ImageNet, Resolution=384x384, Discrete representation=false2021.11 | 84.2 | — | |
| CLIPtraining_mode=fine-tuned2022.10 | 84.2 | — | |
| SoViT-400m/14N-shot=10, Probe=linear regression2023.05 | 84.1 | — | |
| ViT-g/14N-shot=10, Probe=linear regression2023.05 | 84 | — | |
| Linear Probing2024.04 | 83.9 | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=Linear Classifier, Mixing Coefficient (alpha)=0.52021.09 | 83.7 | — | |
| Dr. ViTPre-trained=ImageNet-21K, Fine-tuned on=ImageNet, Resolution=384x384, Discrete representation=discrete only2021.11 | 83.4 | — | |
| DeiTtraining_mode=fine-tuned2022.10 | 83.4 | — | |
| ViT-g/14mode=Zero-shot transfer2023.05 | 82.4 | — | |
| SoViT-400Mmode=Zero-shot transfer2023.05 | 82.2 | — | |
| ViT-L/16N-shot=10, Probe=linear regression2023.05 | 81.5 | — | |
| SimVLMbasepublic data=false, protocol=linear eval2021.12 | 80.6 | — | |
| SparKPre-train task=Generative, Eff. epoch=1600, Backbone=ResNet-50, Fine-tuning resolution=2242023.01 | 80.6 | — | |
| ISyNet-N3Latency, ms.=1.55, # Params, x10^6=20.47, MACs, x10^9=7.32, MEM=0.3942021.09 | 80.43 | — | |
| CLIP-ViT-B/16public data=false, protocol=linear eval2021.12 | 80.2 | — | |
| ResNet-50+Latency, ms.=1.64, # Params, x10^6=25.56, MACs, x10^9=5.19, MEM=0.2862021.09 | 80.18 | — | |
| SwAVPre-train task=Contrastive, Eff. epoch=1200, Backbone=ResNet-50, Fine-tuning resolution=2242023.01 | 80.1 | — | |
| SimCLRPre-train task=Contrastive, Eff. epoch=4000, Backbone=ResNet-50, Fine-tuning resolution=2242023.01 | 80 | — | |
| BYOLPre-train task=Contrastive, Eff. epoch=1600, Backbone=ResNet-50, Fine-tuning resolution=2242023.01 | 80 | — | |
| ViT-L/16mode=Zero-shot transfer2023.05 | 79.9 | — | |
| SupervisedBackbone=ResNet-50, Fine-tuning resolution=2242023.01 | 79.8 | — | |
| MoCov2Pre-train task=Contrastive, Eff. epoch=1600, Backbone=ResNet-50, Fine-tuning resolution=2242023.01 | 79.8 | — | |
| SimSiamPre-train task=Contrastive, Eff. epoch=800, Backbone=ResNet-50, Fine-tuning resolution=2242023.01 | 79.1 | — | |
| ISyNet-N2Latency, ms.=1.1, # Params, x10^6=19.43, MACs, x10^9=4.93, MEM=0.3512021.09 | 79.07 | — | |
| ISyNet-N1-S3Latency, ms.=0.97, # Params, x10^6=10.81, MACs, x10^9=4.12, MEM=0.3992021.09 | 78.25 | — | |
| ResNet-34+Latency, ms.=1.05, # Params, x10^6=21.8, MACs, x10^9=4.63, MEM=0.4972021.09 | 77.95 | — | |
| Concept Matrix Search2024.04 | 77.82 | — | |
| ISyNet-N1-S2Latency, ms.=0.83, # Params, x10^6=8.86, MACs, x10^9=3.34, MEM=0.3952021.09 | 77.45 | — | |
| BCAVisual backbone=ViT-L/142025.03 | 77.09 | — | |
| ISyNet-N1-S1Latency, ms.=0.74, # Params, x10^6=7.82, MACs, x10^9=2.88, MEM=0.3912021.09 | 76.78 | — | |
| Zero-shotModel=CLIP ViT-L/14@336px, Evaluation Protocol=Zero-shot, Source=PyTorch2021.09 | 76.6 | — | |
| ISyNet-N1Latency, ms.=0.72, # Params, x10^6=7.42, MACs, x10^9=2.85, MEM=0.3992021.09 | 76.41 | — | |
| ALIGNTraining Data Source=Private, Evaluation Protocol=Zero-shot2021.11 | 76.4 | — | |
| TDAVisual backbone=ViT-L/142025.03 | 76.28 | — | |
| Zero-shotModel=CLIP ViT-L/14@336px, Evaluation Protocol=Zero-shot, Source=[82]2021.09 | 76.2 | — | |
| CLIPTraining Data Source=Private, Evaluation Protocol=Zero-shot2021.11 | 76.2 | — | |
| Zero-Shot2024.04 | 76.2 | — | |
| ResNet50Training Data Source=Multiple datasets, Evaluation Protocol=Supervised fine-tuned2021.11 | 75.8 | — | |
| LiTTraining Data Source=Public, Evaluation Protocol=Zero-shot2021.11 | 75.7 | — | |
| FLAVApublic data=true, protocol=linear eval2021.12 | 75.5 | — | |
| UnpatchedBackbone=ViT-L/142022.08 | 75.5 | — | |
| Multiple modelsBackbone=ViT-L/142022.08 | 75.5 | — | |
| Joint patchingBackbone=ViT-L/142022.08 | 75.2 | — | |
| Parallel patchingBackbone=ViT-L/142022.08 | 75.1 | — | |
| ISyNet-N0Latency, ms.=0.43, # Params, x10^6=9.59, MACs, x10^9=1.13, MEM=0.2142021.09 | 75.03 | — | |
| DescriptionCLS2024.04 | 75 | — | |
| ResNet-18+Latency, ms.=0.63, # Params, x10^6=11.69, MACs, x10^9=2.28, MEM=0.4392021.09 | 74.3 | — | |
| CLIPVisual backbone=ViT-L/142025.03 | 74.04 | — | |
| Seq. patchingBackbone=ViT-L/14, Seed=22022.08 | 73.5 | — | |
| Seq. patchingBackbone=ViT-L/14, Seed=02022.08 | 73.3 | — | |
| Seq. patchingBackbone=ViT-L/14, Seed=12022.08 | 73.1 | — | |
| Linear ProbeBackbone=CLIP RN502024.06 | 73.1 | — | |
| CLIP-ViT-B/16 (PMD)public data=true, pretraining dataset=PMD, protocol=linear eval2021.12 | 73 | — | |
| TURTLEBackbone=DINOv2 + CLIP ViT-L/14, Setting=2-spaces2024.06 | 72.9 | — | |
| Label-Free CBM2024.04 | 71.95 | — | |
| MobileNetV2 QATQuantization=INT82022.08 | 71.82 | — | |
| TransCLIP-FSShots=16, Backbone=ViT-B/16, Text Regularization=KL penalty2024.06 | 71.8 | — | |
| MobileNetV2 QAT + m learningQuantization=FP8 (Best flexible)2022.08 | 71.74 | — | |
| MobileNetV2Quantization=FP322022.08 | 71.7 | — | |
| MobileNetV2 QAT + m learningQuantization=FP8 (Best flex bias)2022.08 | 71.69 | — | |
| MobileNetV2 QAT + c learningQuantization=FP8 (Best flex bias)2022.08 | 71.67 | — | |
| MobileNetV2 QAT + c learningQuantization=FP8 (Best flexible)2022.08 | 71.66 | — | |
| Sparse-CBM2024.04 | 71.61 | — | |
| PDC-NL#p x 10^6=11.51, variant=[4]2021.04 | 71.6 | — | |
| MobileNetV2 QATQuantization=FP8 (Best flexible)2022.08 | 71.6 | — | |
| MobileNetV2 QATQuantization=FP8 (Best flex bias)2022.08 | 71.54 | — | |
| CoOpLearnable=true, Number of training images per class=162022.03 | 71.51 | — | |
| MobileNetV2 QAT + m learningQuantization=FP8 (Best fixed)2022.08 | 71.34 | — | |
| MobileNetV2 PTQQuantization=FP8 (Best flexible)2022.08 | 71.28 | — | |
| TETSpiking Network=VGG-16, Time-steps=52023.04 | 71.24 | — | |
| PDC-NL#p x 10^6=11.35, variant=[3]2021.04 | 71.2 | — | |
| MobileNetV2 PTQQuantization=FP8 (Best flex bias)2022.08 | 71.06 | — | |
| MobileNetV2 QATQuantization=FP8 (Best fixed)2022.08 | 71.03 | — | |
| CoCoOpLearnable=true, Number of training images per class=162022.03 | 71.02 | — | |
| l1-CBM2024.04 | 71.02 | — | |
| PDC#p x 10^6=10.692021.04 | 71 | — | |
| MobileNetV2 PTQQuantization=INT82022.08 | 70.94 | — | |
| MobileNetV2 QAT + c learningQuantization=FP8 (Best fixed)2022.08 | 70.93 | — | |
| PSNSpiking Network=SEW ResNet-34, Time-steps=42023.04 | 70.54 | — | |
| ResNet18 QATQuantization=INT82022.08 | 70.43 | — | |
| LaBosupervision=full-supervised2024.04 | 70.4 | — |