Image Classification on ImageNet-1k 1.0 (val) (Accuracy, Latency, Throughput)
0.841Top-1 AccConvNext-T
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| ConvNext-TType=CNN, Neural search=false, Extra data=ImageNet-21k, Image size=384x384, Params=28.6M, FLOPs=13.1G2022.06 | 0.841 | 0.0086 | 645 | |
| MobileViTv2-2.0Type=Hybrid, Neural search=false, Extra data=ImageNet-21k-P, Image size=384x384, Params=18.5M, FLOPs=16.1G2022.06 | 0.834 | 0.017 | 488 | |
| ConvNext-TType=CNN, Neural search=false, Extra data=ImageNet-21k, Image size=224x224, Params=28.6M, FLOPs=4.5G2022.06 | 0.829 | 0.0037 | 1,800 | |
| PoolFormer-M48General Arch.=MetaFormer, Token Mixer=Pooling, Image Size=224, Params (M)=73, MACs (G)=11.62021.11 | 0.825 | — | — | |
| MobileViTv2-2.0Type=Hybrid, Neural search=false, Extra data=ImageNet-21k-P, Image size=256x256, Params=18.5M, FLOPs=7.5G2022.06 | 0.824 | 0.0075 | 1,105 | |
| ConvNext-TType=CNN, Neural search=false, Extra data=None, Image size=224x224, Params=28.6M, FLOPs=4.5G2022.06 | 0.821 | 0.0037 | 1,800 | |
| PoolFormer-M36General Arch.=MetaFormer, Token Mixer=Pooling, Image Size=224, Params (M)=56, MACs (G)=8.82021.11 | 0.821 | — | — | |
| DeiT-BaseType=Transformer, Neural search=false, Extra data=None, Image size=224x224, Params=86.6M, FLOPs=17.6G2022.06 | 0.818 | 0.0132 | 958 | |
| RSB-ResNet-152General Arch.=Convolutional Neural Networks, Token Mixer=-, Image Size=224, Params (M)=60, MACs (G)=11.62021.11 | 0.818 | — | — | |
| DeiT-BGeneral Arch.=Attention, Token Mixer=Attention, Image Size=224, Params (M)=86, MACs (G)=17.52021.11 | 0.818 | — | — | |
| PVT-LargeGeneral Arch.=Attention, Token Mixer=Attention, Image Size=224, Params (M)=61, MACs (G)=9.82021.11 | 0.817 | — | — | |
| gMLP-BGeneral Arch.=MetaFormer, Token Mixer=Spatial MLP, Image Size=224, Params (M)=73, MACs (G)=15.82021.11 | 0.816 | — | — | |
| PoolFormer-S36General Arch.=MetaFormer, Token Mixer=Pooling, Image Size=224, Params (M)=31, MACs (G)=52021.11 | 0.814 | — | — | |
| Swin-TType=Hybrid, Neural search=false, Extra data=None, Image size=224x224, Params=28.3M, FLOPs=4.5G2022.06 | 0.813 | — | 1,390 | |
| RSB-ResNet-101General Arch.=Convolutional Neural Networks, Token Mixer=-, Image Size=224, Params (M)=45, MACs (G)=7.92021.11 | 0.813 | — | — | |
| Swin-Mixer-B/D24General Arch.=MetaFormer, Token Mixer=Spatial MLP, Image Size=224, Params (M)=61, MACs (G)=10.42021.11 | 0.813 | — | — | |
| MobileViTv2-2.0Type=Hybrid, Neural search=false, Extra data=None, Image size=256x256, Params=18.5M, FLOPs=7.5G2022.06 | 0.812 | 0.0075 | 1,105 | |
| PVT-MediumGeneral Arch.=Attention, Token Mixer=Attention, Image Size=224, Params (M)=44, MACs (G)=6.72021.11 | 0.812 | — | — | |
| ResMLP-B24General Arch.=MetaFormer, Token Mixer=Spatial MLP, Image Size=224, Params (M)=116, MACs (G)=232021.11 | 0.81 | — | — | |
| MobileViTv2-1.5Type=Hybrid, Neural search=false, Extra data=None, Image size=256x256, Params=10.6M, FLOPs=4.0G2022.06 | 0.804 | 0.0051 | 1,418 | |
| PoolFormer-S24General Arch.=MetaFormer, Token Mixer=Pooling, Image Size=224, Params (M)=21, MACs (G)=3.42021.11 | 0.803 | — | — | |
| EfficientNet-b2Type=CNN, Neural search=true, Extra data=None, Image size=288x288, Params=9.1M, FLOPs=1.2G2022.06 | 0.801 | 0.0038 | 2,032 | |
| RSB-ResNet-50General Arch.=Convolutional Neural Networks, Token Mixer=-, Image Size=224, Params (M)=26, MACs (G)=4.12021.11 | 0.798 | — | — | |
| DeiT-SGeneral Arch.=Attention, Token Mixer=Attention, Image Size=224, Params (M)=22, MACs (G)=4.62021.11 | 0.798 | — | — | |
| PVT-SmallGeneral Arch.=Attention, Token Mixer=Attention, Image Size=224, Params (M)=25, MACs (G)=3.82021.11 | 0.798 | — | — | |
| ViT-B/16*General Arch.=Attention, Token Mixer=Attention, Image Size=224, Params (M)=86, MACs (G)=17.62021.11 | 0.797 | — | — | |
| Swin-Mixer-T/D6General Arch.=MetaFormer, Token Mixer=Spatial MLP, Image Size=256, Params (M)=23, MACs (G)=42021.11 | 0.797 | — | — | |
| gMLP-SGeneral Arch.=MetaFormer, Token Mixer=Spatial MLP, Image Size=224, Params (M)=20, MACs (G)=4.52021.11 | 0.796 | — | — | |
| ResMLP-S24General Arch.=MetaFormer, Token Mixer=Spatial MLP, Image Size=224, Params (M)=30, MACs (G)=62021.11 | 0.794 | — | — | |
| Swin-Mixer-T/D24General Arch.=MetaFormer, Token Mixer=Spatial MLP, Image Size=256, Params (M)=20, MACs (G)=42021.11 | 0.794 | — | — | |
| MobileViT-SType=Hybrid, Neural search=false, Extra data=None, Image size=256x256, Params=5.6M, FLOPs=2.0G2022.06 | 0.784 | 0.0034 | 1,986 | |
| MobileViTv2-1.0Type=Hybrid, Neural search=false, Extra data=None, Image size=256x256, Params=4.9M, FLOPs=1.8G2022.06 | 0.781 | 0.0034 | 2,351 | |
| MobileFormer-294Type=Hybrid, Neural search=false, Extra data=None, Image size=224x224, Params=11.8M, FLOPs=294M2022.06 | 0.779 | 0.0407 | 1,402 | |
| PoolFormer-S12General Arch.=MetaFormer, Token Mixer=Pooling, Image Size=224, Params (M)=12, MACs (G)=1.82021.11 | 0.772 | — | — | |
| EfficientNet-b0Type=CNN, Neural search=true, Extra data=None, Image size=224x224, Params=5.3M, FLOPs=422M2022.06 | 0.771 | 0.0016 | 4,619 | |
| ResMLP-S12General Arch.=MetaFormer, Token Mixer=Spatial MLP, Image Size=224, Params (M)=15, MACs (G)=32021.11 | 0.766 | — | — | |
| MLP-Mixer-B/16General Arch.=MetaFormer, Token Mixer=Spatial MLP, Image Size=224, Params (M)=59, MACs (G)=12.72021.11 | 0.764 | — | — | |
| ViT-L/16*General Arch.=Attention, Token Mixer=Attention, Image Size=224, Params (M)=307, MACs (G)=63.62021.11 | 0.761 | — | — | |
| RSB-ResNet-34General Arch.=Convolutional Neural Networks, Token Mixer=-, Image Size=224, Params (M)=22, MACs (G)=3.72021.11 | 0.755 | — | — | |
| PVT-TinyGeneral Arch.=Attention, Token Mixer=Attention, Image Size=224, Params (M)=13, MACs (G)=1.92021.11 | 0.751 | — | — | |
| DeiT-TinyType=Transformer, Neural search=false, Extra data=None, Image size=224x224, Params=5.5M, FLOPs=1.3G2022.06 | 0.722 | 0.0034 | 4,541 | |
| RSB-ResNet-18General Arch.=Convolutional Neural Networks, Token Mixer=-, Image Size=224, Params (M)=12, MACs (G)=1.82021.11 | 0.706 | — | — | |
| MobileViTv2-0.5Type=Hybrid, Neural search=false, Extra data=None, Image size=256x256, Params=1.4M, FLOPs=0.5G2022.06 | 0.702 | 0.0016 | 4,595 | |
| MobileViT-XXSType=Hybrid, Neural search=false, Extra data=None, Image size=256x256, Params=1.3M, FLOPs=0.4G2022.06 | 0.69 | 0.0048 | 4,225 | |
| MobileFormer-52Type=Hybrid, Neural search=false, Extra data=None, Image size=224x224, Params=3.6M, FLOPs=52M2022.06 | 0.687 | 0.0071 | 4,445 |