Image Classification on ImageNet 1K (val) (Single Accuracy Metric)
86.6Accuracy4M-L
Evaluation Results
| Method | Links | |
|---|---|---|
| 4M-LProtocol=Fine-tuning (FT), Extra Labels=Yes2026.03 | 86.6 | |
| MambaVision-L+MAPResolution=224x224, Parameters=241M, Throughput=229, Memory=78.6G, Pretraining=MAP2024.10 | 86.4 | |
| HybridNet-L + MAPResolution=384x384, Parameters=443M, Throughput=63, Pretraining=MAP2024.10 | 86.2 | |
| ViT-L/16 + MAPResolution=224x224, Parameters=307M, Throughput=149, Pretraining=MAP2024.10 | 86.1 | |
| ViT-L/16 + MAEResolution=224x224, Parameters=307M, Throughput=1492024.10 | 85.9 | |
| HybridNet-B + MAPResolution=384x384, Parameters=128M, Throughput=244, Memory=76.1G, Pretraining=MAP2024.10 | 85.5 | |
| MambaVision-LResolution=224x224, Parameters=241M, Throughput=229, Memory=78.6G2024.10 | 85.3 | |
| HybridNet-L + MAPResolution=224x224, Parameters=443M, Throughput=63, Memory=78.3G, Pretraining=MAP2024.10 | 85 | |
| MambaVision-B+MAPResolution=224x224, Parameters=97M, Throughput=826, Memory=50.8G, Pretraining=MAP2024.10 | 84.9 | |
| HybridNet-B + MAPResolution=224x224, Parameters=128M, Throughput=244, Memory=30.0G, Pretraining=MAP2024.10 | 84.9 | |
| SparKArch.=ConvX-B, Eff. epoch=1600, Fine-tuning resolution=2242023.01 | 84.8 | |
| MambaR-L+MAPResolution=224x224, Parameters=341M, Throughput=92, Memory=55.5G, Pretraining=MAP2024.10 | 84.8 | |
| STL (FAN-L-Hybrid)Params (M)=77.32024.01 | 84.7 | |
| HybridNet-LResolution=384x384, Parameters=443M, Throughput=632024.10 | 84.6 | |
| STL (FAN-B-Hybrid)Params (M)=50.92024.01 | 84.5 | |
| ARM-L (Mamba+AR)Resolution=224x224, Parameters=297M, Throughput=111, Memory=53.1G2024.10 | 84.5 | |
| HybridNet-BResolution=384x384, Parameters=128M, Throughput=244, Memory=76.1G2024.10 | 84.5 | |
| FAN-L-HybridParams (M)=76.82024.01 | 84.3 | |
| MambaVision-BResolution=224x224, Parameters=97M, Throughput=826, Memory=50.8G2024.10 | 84.2 | |
| SimMIMArch.=Swin-B, Eff. epoch=800, Fine-tuning resolution=2242023.01 | 84 | |
| LV-ViT-MParams (M)=562024.01 | 84 | |
| MambaR-B+MAPResolution=224x224, Parameters=99M, Throughput=315, Memory=20.3G, Pretraining=MAP2024.10 | 84 | |
| FAN-B-HybridParams (M)=50.42024.01 | 83.9 | |
| M2PT-PointSetting=Pretrained, Evaluation Protocol=tune acc, Backbone=ViT-B2024.01 | 83.9 | |
| VMamba-BResolution=224x224, Parameters=89M, Throughput=246, Memory=37.1G2024.10 | 83.9 | |
| HybridNet-B + MAEResolution=224x224, Parameters=128M, Throughput=244, Memory=30.0G, Pretraining=MAE2024.10 | 83.9 | |
| Supervised (Liu et al., 2022)Arch.=ConvX-B, Eff. epoch=300, Fine-tuning resolution=2242023.01 | 83.8 | |
| ConvNext-BParams (M)=88.62024.01 | 83.8 | |
| ConvNeXt-BResolution=224x224, Parameters=89M, Throughput=334, Memory=17.9G2024.10 | 83.8 | |
| HybridNet-B + ARResolution=224x224, Parameters=128M, Throughput=244, Memory=30.0G, Pretraining=AR2024.10 | 83.8 | |
| M2PT-AudioSetting=Pretrained, Evaluation Protocol=tune acc, Backbone=ViT-B2024.01 | 83.7 | |
| MambaR-B+ARResolution=224x224, Parameters=99M, Throughput=315, Memory=20.3G2024.10 | 83.7 | |
| MAEArch.=ViT-B, Eff. epoch=1600, Fine-tuning resolution=2242023.01 | 83.6 | |
| MAE-ViT-BParams (M)=862024.01 | 83.6 | |
| MFFSetting=Pretrained, Evaluation Protocol=tune acc, Backbone=ViT-B2024.01 | 83.6 | |
| M2PT-VideoSetting=Pretrained, Evaluation Protocol=tune acc, Backbone=ViT-B2024.01 | 83.6 | |
| ViT-B/16 + MAEResolution=224x224, Parameters=86M, Throughput=284, Memory=63.8G2024.10 | 83.6 | |
| ViT-B/16 + MAPResolution=224x224, Parameters=86M, Throughput=284, Memory=63.8G, Pretraining=MAP2024.10 | 83.6 | |
| VMamba-SResolution=224x224, Parameters=50M, Throughput=313, Memory=27.6G2024.10 | 83.6 | |
| Supervised (Liu et al., 2021)Arch.=Swin-B, Eff. epoch=300, Fine-tuning resolution=2242023.01 | 83.5 | |
| FAN-S-HybridParams (M)=26.32024.01 | 83.5 | |
| STL (FAN-S-Hybrid)Params (M)=26.52024.01 | 83.4 | |
| Swin-SParams (M)=502024.01 | 83.4 | |
| Swin-BParams (M)=87.82024.01 | 83.4 | |
| SemMAESetting=Pretrained, Evaluation Protocol=tune acc, Backbone=ViT-B2024.01 | 83.4 | |
| LV-ViT-SParams (M)=262024.01 | 83.3 | |
| MAESetting=Pretrained, Evaluation Protocol=tune acc, Backbone=ViT-B2024.01 | 83.3 | |
| MambaVision-SResolution=224x224, Parameters=51M, Throughput=1058, Memory=36.6G2024.10 | 83.3 | |
| MoCov3Arch.=ViT-B, Eff. epoch=1600, Fine-tuning resolution=2242023.01 | 83.2 | |
| BEITArch.=ViT-B, Eff. epoch=800, Fine-tuning resolution=2242023.01 | 83.2 | |
| MambaR-LResolution=224x224, Parameters=341M, Throughput=92, Memory=55.5G2024.10 | 83.2 | |
| ARM-B (Mamba+AR)Resolution=224x224, Parameters=85M, Throughput=325, Memory=19.7G2024.10 | 83.2 | |
| HybridNet-LResolution=224x224, Parameters=443M, Throughput=63, Memory=78.3G2024.10 | 83.2 | |
| ConvNeXt-SResolution=224x224, Parameters=50M, Throughput=444, Memory=13.1G2024.10 | 83.1 | |
| MambaR-B+MAEResolution=224x224, Parameters=99M, Throughput=315, Memory=20.3G2024.10 | 83.1 | |
| HybridNet-BResolution=224x224, Parameters=128M, Throughput=244, Memory=30.0G2024.10 | 83.1 | |
| HybridNet-B + CLResolution=224x224, Parameters=128M, Throughput=244, Memory=30.0G, Pretraining=CL2024.10 | 83.1 | |
| MambaR-BResolution=224x224, Parameters=99M, Throughput=315, Memory=20.3G2024.10 | 82.9 | |
| MambaVision-TResolution=224x224, Parameters=35M, Throughput=1349, Memory=10.7G2024.10 | 82.7 | |
| DREAMProtocol=Fine-tuning (FT)2026.03 | 82.7 | |
| XCIT-S24Params (M)=47.72024.01 | 82.6 | |
| RVT-BParams (M)=91.82024.01 | 82.6 | |
| E2E-FT + Weight-space ensembleBackbone=CLIP ViT-B/162024.11 | 82.5 | |
| E2E-FT + Weight-space ensembleBackbone=ViT-B/16, Ensemble=Weight-space2024.11 | 82.5 | |
| ViT-B/16 + ARResolution=224x224, Parameters=86M, Throughput=284, Memory=63.8G2024.10 | 82.5 | |
| VMamba-TResolution=224x224, Parameters=31M, Throughput=464, Memory=7.6G2024.10 | 82.5 | |
| HybridNet-S + MAPResolution=224x224, Parameters=37M, Throughput=512, Memory=14.6G, Pretraining=MAP2024.10 | 82.5 | |
| LP-FT + Weight-space ensembleBackbone=CLIP ViT-B/162024.11 | 82.4 | |
| LP-FT + Weight-space ensembleBackbone=ViT-B/16, Ensemble=Weight-space2024.11 | 82.4 | |
| Supervised (He et al., 2021)Arch.=ViT-B, Eff. epoch=300, Fine-tuning resolution=2242023.01 | 82.3 | |
| E2E-FT + VRFBackbone=CLIP ViT-B/162024.11 | 82.3 | |
| E2E-FT + VRFBackbone=ViT-B/16, Method=Variance Reduction Fine-tuning2024.11 | 82.3 | |
| E2E-FT + Output-space ensembleBackbone=CLIP ViT-B/162024.11 | 82.2 | |
| E2E-FT + Output-space ensembleBackbone=ViT-B/16, Ensemble=Output-space2024.11 | 82.2 | |
| ConvNext-TParams (M)=28.62024.01 | 82.1 | |
| ConvNext-SParams (M)=50.22024.01 | 82.1 | |
| LP-FT + Output-space ensembleBackbone=CLIP ViT-B/162024.11 | 82.1 | |
| LP-FT + VRFBackbone=CLIP ViT-B/162024.11 | 82.1 | |
| LP-FT + Output-space ensembleBackbone=ViT-B/16, Ensemble=Output-space2024.11 | 82.1 | |
| LP-FT + VRFBackbone=ViT-B/16, Method=Variance Reduction Fine-tuning2024.11 | 82.1 | |
| ConvNeXt-TResolution=224x224, Parameters=29M, Throughput=701, Memory=8.3G2024.10 | 82.1 | |
| RVT-SParams (M)=23.32024.01 | 81.9 | |
| XCIT-S12Params (M)=26.32024.01 | 81.9 | |
| M2PT-PointSetting=From-scratch, Evaluation Protocol=tune acc, Backbone=ViT-B2024.01 | 81.9 | |
| EfficientNet-B3Resolution=300x300, Parameters=12M, Throughput=496, Memory=19.7G2024.10 | 81.6 | |
| DAT-AugReg-ViTParams (M)=862024.01 | 81.5 | |
| LP-FTBackbone=CLIP ViT-B/162024.11 | 81.5 | |
| LP-FTBackbone=ViT-B/16, Protocol=Linear-probing then fine-tuning2024.11 | 81.5 | |
| E2E-FTBackbone=CLIP ViT-B/162024.11 | 81.3 | |
| E2E-FTBackbone=ViT-B/16, Protocol=End-to-end fine-tuning2024.11 | 81.3 | |
| HybridNet-SResolution=224x224, Parameters=37M, Throughput=512, Memory=14.6G2024.10 | 81.3 | |
| Swin-TParams (M)=28.32024.01 | 81.2 | |
| MambaR-SResolution=224x224, Parameters=28M, Throughput=608, Memory=9.9G2024.10 | 81.1 | |
| iBOTEvaluation Protocol=Fine-tuning2026.02 | 80.72 | |
| Vim-SResolution=224x224, Parameters=26M, Throughput=612, Memory=9.4G2024.10 | 80.5 | |
| STELLAREvaluation Protocol=Fine-tuning2026.02 | 80.05 | |
| DINOEvaluation Protocol=Fine-tuning2026.02 | 79.58 | |
| Linear classifierBackbone=CLIP ViT-B/162024.11 | 79.3 | |
| Linear classifierBackbone=ViT-B/16, Protocol=Linear probing2024.11 | 79.3 | |
| HybridNet-T + MAPResolution=224x224, Parameters=12M, Throughput=910, Memory=7.6G, Pretraining=MAP2024.10 | 78.6 |