Image Classification on ImageNet (val) (ACC)
91.1AccuracyLion
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LionBackbone=BASIC-L, Evaluation Protocol=Fine-tune2023.02 | 91.1 | — | |
| Previous SOTA (Yu et al., 2022)Evaluation Protocol=Fine-tune2023.02 | 91 | — | |
| AdafactorBackbone=BASIC-L, Evaluation Protocol=Fine-tune2023.02 | 90.9 | — | |
| Full fine-tuningBackbone=ViT-g (1B)2024.02 | 89 | — | |
| LoSABackbone=ViT-g (1B)2024.02 | 89 | — | |
| Finetune only mlpBackbone=ViT-g (1B)2024.02 | 88.9 | — | |
| Finetune last layerBackbone=ViT-g (1B)2024.02 | 88.8 | — | |
| Finetune last 2 layersBackbone=ViT-g (1B)2024.02 | 88.8 | — | |
| Finetune only attentionBackbone=ViT-g (1B)2024.02 | 88.8 | — | |
| LoRABackbone=ViT-g (1B)2024.02 | 88.6 | — | |
| Linear probingBackbone=ViT-g (1B)2024.02 | 88.4 | — | |
| BitfitBackbone=ViT-g (1B)2024.02 | 88.4 | — | |
| LionBackbone=BASIC-L, Evaluation Protocol=Zero-shot2023.02 | 88.3 | — | |
| Prompt TuningBackbone=ViT-g (1B)2024.02 | 88.3 | — | |
| LSTBackbone=ViT-g (1B)2024.02 | 88.2 | — | |
| Previous SOTA (Yu et al., 2022)Evaluation Protocol=Zero-shot2023.02 | 86.3 | — | |
| AdafactorBackbone=BASIC-L, Evaluation Protocol=Zero-shot2023.02 | 85.7 | — | |
| ALIGNProtocol=Linear evaluation2021.08 | 85.5 | — | |
| CLIPProtocol=Linear evaluation2021.08 | 85.4 | — | |
| SimVLMProtocol=Linear evaluation, Size=huge2021.08 | 83.6 | — | |
| Ours (multi-fill)Backbone=CLIP ViT-B/162023.03 | 82.5 | — | |
| Ours (single-fill)Backbone=CLIP ViT-B/162023.03 | 82.4 | — | |
| SimVLMProtocol=Linear evaluation, Size=large2021.08 | 82.3 | — | |
| DIFFQBackbone=DeiT, Penalty Level (lambda)=1e-2, Model Size (MB)=45.72021.04 | 82 | — | |
| UNCOMPRESSEDBackbone=DeiT, Model Size (MB)=371.42021.04 | 81.8 | — | |
| LP-FTBackbone=CLIP ViT-B/162023.03 | 81.7 | — | |
| WiSE-FTBackbone=CLIP ViT-B/162023.03 | 81.7 | — | |
| QATBackbone=DeiT, Quantization Bits=8, Model Size (MB)=82.92021.04 | 81.6 | — | |
| DIFFQBackbone=DeiT, Penalty Level (lambda)=0.1, Model Size (MB)=33.022021.04 | 81.5 | — | |
| CCT-14/7×2 Distilled# Params=22.36 M, MACs=5.53 G, Training Epochs=300, knowledge distillation=true2021.04 | 81.34 | — | |
| DeiT-S# Params=22.44 M, MACs=4.63 G, Training Epochs=3002021.04 | 81.16 | — | |
| GreedyNASv2-LBackbone=GreedyNASv2-L2021.11 | 81.1 | — | |
| Vanilla fine-tuningBackbone=CLIP ViT-B/162023.03 | 80.7 | — | |
| CCT-14/7×2# Params=22.36 M, MACs=5.53 G, Training Epochs=300, knowledge distillation=false2021.04 | 80.67 | — | |
| SimVLMProtocol=Linear evaluation, Size=base2021.08 | 80.6 | — | |
| AC/DCSparsity=0.5, Backbone=DeiT-S2023.11 | 80.15 | — | |
| RPGSparsity=0.5, Backbone=DeiT-S2023.11 | 80.15 | — | |
| DINOProtocol=Linear evaluation2021.08 | 80.1 | — | |
| Uniform soupBackbone=CLIP ViT-B/322023.03 | 80 | — | |
| LaCLIPModel Architecture=ViT-B/16, Pre-training Data=LAION-400M, Evaluation Protocol=Linear Probing2023.05 | 79.9 | — | |
| RPGSparsity=0.6, Backbone=DeiT-S2023.11 | 79.89 | — | |
| ViT-S# Params=22.05 M, MACs=4.61 G, Training Epochs=3002021.04 | 79.85 | — | |
| DeiT-SSparsity=0, Backbone=DeiT-S2023.11 | 79.85 | — | |
| SimCLRv2Protocol=Linear evaluation2021.08 | 79.8 | — | |
| ResNet50 (2021)# Params=25.55 M, MACs=4.15 G, Training Epochs=3002021.04 | 79.8 | — | |
| SVIT-ESparsity=0.5, Backbone=DeiT-S2023.11 | 79.72 | — | |
| AC/DCSparsity=0.6, Backbone=DeiT-S2023.11 | 79.69 | — | |
| SVIT-ESparsity=0.6, Backbone=DeiT-S2023.11 | 79.41 | — | |
| QATBackbone=DeiT, Quantization Bits=4, Model Size (MB)=41.72021.04 | 79.2 | — | |
| CLIPModel Architecture=ViT-B/16, Pre-training Data=LAION-400M, Evaluation Protocol=Linear Probing2023.05 | 78.6 | — | |
| AlignMixBackbone=ResNet-50, Epochs=100, Cost per Epoch=1.05/12022.12 | 78 | — | |
| AutoMixBackbone=ResNet-50, Epochs=100, Cost per Epoch=2/12022.12 | 77.91 | — | |
| Ours (multi-fill)Backbone=CLIP ViT-B/322023.03 | 77.9 | — | |
| Concept Matrix Search2024.04 | 77.82 | — | |
| Co-MixupBackbone=ResNet-50, Epochs=100, Cost per Epoch=3/12022.12 | 77.61 | — | |
| PuzzleMixBackbone=ResNet-50, Epochs=100, Cost per Epoch=2.9/12022.12 | 77.51 | — | |
| GreedyNASv2-SBackbone=GreedyNASv2-S2021.11 | 77.5 | — | |
| Ours (single-fill)Backbone=CLIP ViT-B/322023.03 | 77.5 | — | |
| RPGSparsity=0.8, Backbone=DeiT-S2023.11 | 77.42 | — | |
| R-MixBackbone=ResNet-50, Epochs=100, Cost per Epoch=2/12022.12 | 77.41 | — | |
| RTFormer-Base#Params=20.5M, FLOPs=3.0G, Crop Size=224 x 224, Input Scale=224 x 2242022.10 | 77.4 | — | |
| ResNet50# Params=25.55 M, MACs=4.15 G, Training Epochs=1202021.04 | 77.15 | — | |
| SaliencyMixBackbone=ResNet-50, Epochs=100, Cost per Epoch=2/12022.12 | 77.14 | — | |
| UncompressedBackbone=ResNet-50, M.S. (MB)=97.52021.04 | 77.1 | — | |
| ResNet-50 (DyRep)Base Model=ResNet-50, Rep method=DyRep, Training Cost (GPU days)=8.5, Avg. FLOPs (G)=5.05, Avg. params (M)=31.52022.03 | 77.08 | — | |
| DyRepModel=ResNet-50, Cost (GPU days)=8.5, Avg. FLOPs (G)=5.05, Avg. params (M)=31.52022.03 | 77.08 | — | |
| CutMixBackbone=ResNet-50, Epochs=100, Cost per Epoch=1/12022.12 | 77.08 | — | |
| Input Mix-upBackbone=ResNet-50, Epochs=100, Cost per Epoch=1/12022.12 | 77.03 | — | |
| DIFFQBackbone=ResNet-50, M.S. (MB)=142021.04 | 76.9 | — | |
| LSQBackbone=ResNet-50, Bit-width=8 bits, M.S. (MB)=24.52021.04 | 76.8 | — | |
| LSQ*Backbone=ResNet-50, Bit-width=8 bits, M.S. (MB)=24.52021.04 | 76.8 | — | |
| MobileCLIP-BEvaluation Protocol=zero-shot2023.11 | 76.8 | — | |
| ResNet-50 (DBB)Base Model=ResNet-50, Rep method=DBB, Training Cost (GPU days)=13.7, Avg. FLOPs (G)=6.79, Avg. params (M)=40.72022.03 | 76.71 | — | |
| DBBModel=ResNet-50, Cost (GPU days)=13.7, Avg. FLOPs (G)=6.79, Avg. params (M)=40.72022.03 | 76.71 | — | |
| LSQ*Backbone=ResNet-50, Bit-width=4 bits, M.S. (MB)=12.32021.04 | 76.7 | — | |
| Manifold Mix-upBackbone=ResNet-50, Epochs=100, Cost per Epoch=1/12022.12 | 76.7 | — | |
| DIFFQBackbone=ResNet-50, M.S. (MB)=10.52021.04 | 76.6 | — | |
| WiSE-FTBackbone=CLIP ViT-B/322023.03 | 76.6 | — | |
| STDC2#Params=12.5M, FLOPs=1.4G, Crop Size=224 x 224, Input Scale=224 x 2242022.10 | 76.4 | — | |
| ALIGNTraining Dataset=ALIGN-1800M, Evaluation Protocol=Zero-shot, Resolution=640x640, Implementation Source=Official2022.10 | 76.4 | — | |
| Efficient-Net-B0#Params=5.3M, FLOPs=0.4G, Crop Size=224 x 224, Input Scale=224 x 2242022.10 | 76.3 | — | |
| DIFFQBackbone=ResNet-50, M.S. (MB)=8.82021.04 | 76.3 | — | |
| AC/DCSparsity=0.8, Backbone=DeiT-S2023.11 | 76.24 | — | |
| LSQBackbone=ResNet-50, Bit-width=4 bits, M.S. (MB)=12.32021.04 | 76.2 | — | |
| Zero-Shot CLIPMode=Zero-shot2024.04 | 76.2 | — | |
| DSBackbone=ResNet-50, % params rm=80.47, % FLOPS rm=72.132022.07 | 76.15 | — | |
| GMPBackbone=ResNet-50, % params rm=80.082022.07 | 76.15 | — | |
| ResNet-50 (Origin)Base Model=ResNet-50, Rep method=Origin, Training Cost (GPU days)=7.5, Avg. FLOPs (G)=4.09, Avg. params (M)=25.62022.03 | 76.14 | — | |
| OriginModel=ResNet-50, Cost (GPU days)=7.5, Avg. FLOPs (G)=4.09, Avg. params (M)=25.62022.03 | 76.14 | — | |
| DenseModel=ResNet50, N:M Pruning=None2022.08 | 76.13 | — | |
| ResNet-50Backbone=ResNet-502021.11 | 76.1 | — | |
| STRBackbone=ResNet-50, % params rm=79.69, % FLOPS rm=81.172022.07 | 76 | — | |
| VanillaBackbone=ResNet-50, Epochs=100, Cost per Epoch=1/12022.12 | 75.97 | — | |
| DDRNet-23#Params=28.2M, FLOPs=3.9G, Crop Size=224 x 224, Input Scale=224 x 2242022.10 | 75.9 | — | |
| Vanilla fine-tuningBackbone=CLIP ViT-B/322023.03 | 75.9 | — | |
| LSQ*Backbone=ResNet-50, Bit-width=3 bits, M.S. (MB)=9.32021.04 | 75.8 | — | |
| LSQBackbone=ResNet-50, Bit-width=3 bits, M.S. (MB)=9.32021.04 | 75.6 | — | |
| Sequential Attention++Backbone=ResNet-50, Block Size=64x64, Sparsity=58%2024.02 | 75.53 | — | |
| ResNet-50#Params=23.5M, FLOPs=3.7G, Crop Size=224 x 224, Input Scale=224 x 2242022.10 | 75.3 | — | |
| RényiCLEvaluation Protocol=linear evaluation, Multi-crop Augmentation=true2022.08 | 75.3 | — |