Image Classification on ImageNet 1k (train)
88.58Top-1 AccuracyEVA
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| EVAModel=BEIT-L/14, Pre-training Dataset=CLIP OpenAI, IN-22K, IN-1K, Evaluation Protocol=Linear Probing2024.11 | 88.58 | — | — | — | — | — | — | |
| ScaleKDModel=ViT-B/16, Pre-training Dataset=IN-1K, Evaluation Protocol=Linear Probing2024.11 | 86.43 | — | — | — | — | — | — | |
| CLIPModel=ViT-B/16, Pre-training Dataset=LAION-2B, IN-12K, IN-1K, Evaluation Protocol=Linear Probing2024.11 | 86.17 | — | — | — | — | — | — | |
| CLIPModel=ViT-B/16, Pre-training Dataset=CLIP OpenAI, IN-12K, IN-1K, Evaluation Protocol=Linear Probing2024.11 | 85.99 | — | — | — | — | — | — | |
| CLIPModel=ViT-B/16, Pre-training Dataset=LAION-2B, IN-1K, Evaluation Protocol=Linear Probing2024.11 | 85.49 | — | — | — | — | — | — | |
| From-scratchModel=ViT-B/16, Pre-training Dataset=IN-1K, Evaluation Protocol=Linear Probing2024.11 | 81.8 | — | — | — | — | — | — | |
| DARTS-AParams (M)=5.52020.09 | 77.8 | — | — | — | — | — | — | |
| FairDARTS-CParams (M)=5.32020.09 | 77.2 | — | — | — | — | — | — | |
| MixNet-MParams (M)=5.02020.09 | 77 | — | — | — | — | — | — | |
| SupervisedBackbone=ResNet-50, Pre-training epochs=2002022.02 | 76.1 | — | — | — | — | — | — | |
| MnasNet-A2Params (M)=4.82020.09 | 75.6 | — | — | — | — | — | — | |
| MobileNetV3Params (M)=5.42020.09 | 75.2 | — | — | — | — | — | — | |
| SingPath NASParams (M)=4.32020.09 | 75 | — | — | — | — | — | — | |
| MobileNetV2Params (M)=3.42020.09 | 72 | — | — | — | — | — | — | |
| InfoMinBackbone=ResNet-50, Pre-training epochs=2002022.02 | 70.1 | — | — | — | — | — | — | |
| FPModel=ViT-B2025.03 | 70.01 | — | — | — | — | — | — | |
| HOTModel=ViT-B2025.03 | 69.4 | — | — | — | — | — | — | |
| MoCoV2 + ContrastiveCropBackbone=ResNet-50, Pre-training epochs=2002022.02 | 67.8 | — | — | — | — | — | — | |
| LPLDIPC=200, Pruning rate=1x, Backbone=ResNet-502024.10 | 67.8 | — | — | — | — | — | — | |
| MoCoV2Backbone=ResNet-50, Pre-training epochs=2002022.02 | 67.5 | — | — | — | — | — | — | |
| LPLDIPC=200, Pruning rate=10x, Backbone=ResNet-502024.10 | 67.1 | — | — | — | — | — | — | |
| LUQModel=ViT-B2025.03 | 67.1 | — | — | — | — | — | — | |
| LPLDIPC=200, Pruning rate=20x, Backbone=ResNet-502024.10 | 66.7 | — | — | — | — | — | — | |
| LPLDIPC=100, Pruning rate=1x, Backbone=ResNet-502024.10 | 65.7 | — | — | — | — | — | — | |
| LPLDIPC=200, Pruning rate=30x, Backbone=ResNet-502024.10 | 65.4 | — | — | — | — | — | — | |
| LPLDIPC=100, Pruning rate=10x, Backbone=ResNet-502024.10 | 65.1 | — | — | — | — | — | — | |
| LPLDIPC=200, Pruning rate=50x, Backbone=ResNet-502024.10 | 64.1 | — | — | — | — | — | — | |
| LPLDIPC=100, Pruning rate=20x, Backbone=ResNet-502024.10 | 63.9 | — | — | — | — | — | — | |
| MoCoV1 + ContrastiveCropBackbone=ResNet-50, Pre-training epochs=2002022.02 | 63 | — | — | — | — | — | — | |
| LPLDIPC=50, Pruning rate=1x, Backbone=ResNet-502024.10 | 62.2 | — | — | — | — | — | — | |
| LPLDIPC=100, Pruning rate=30x, Backbone=ResNet-502024.10 | 62 | — | — | — | — | — | — | |
| LPLDIPC=50, Pruning rate=10x, Backbone=ResNet-502024.10 | 61.2 | — | — | — | — | — | — | |
| MoCoV1Backbone=ResNet-50, Pre-training epochs=2002022.02 | 60.6 | — | — | — | — | — | — | |
| LPLDIPC=200, Pruning rate=100x, Backbone=ResNet-502024.10 | 60.1 | — | — | — | — | — | — | |
| LPLDIPC=100, Pruning rate=50x, Backbone=ResNet-502024.10 | 59.8 | — | — | — | — | — | — | |
| CaffeHardware=1 NVIDIA K20, Net=AlexNet, Epochs=100, Batch size=256, Initial Learning Rate=0.012015.10 | 58.9 | — | — | 1 | — | — | — | |
| CaffeHardware=1 NVIDIA K20, Net=NiN, Epochs=47, Batch size=256, Initial Learning Rate=0.012015.10 | 58.9 | — | — | 1 | — | — | — | |
| FireCaffeHardware=32 NVIDIA K20s (Titan supercomputer), Net=NiN, Epochs=47, Batch size=256, Initial Learning Rate=0.012015.10 | 58.9 | — | 11 | 13 | — | — | — | |
| LPLDIPC=50, Pruning rate=20x, Backbone=ResNet-502024.10 | 58.8 | — | — | — | — | — | — | |
| FireCaffeHardware=32 NVIDIA K20s (Titan supercomputer), Net=NiN, Epochs=47, Batch size=1024, Initial Learning Rate=0.042015.10 | 58.6 | — | 6 | 23 | — | — | — | |
| FireCaffeHardware=128 NVIDIA K20s (Titan supercomputer), Net=NiN, Epochs=47, Batch size=1024, Initial Learning Rate=0.042015.10 | 58.6 | — | 3.6 | 39 | — | — | — | |
| Google cuda-convnet2Hardware=8 NVIDIA K20s (1 node), Net=AlexNet, Epochs=100, Batch size=varies, Initial Learning Rate=0.022015.10 | 57.1 | — | 16 | 7.7 | — | — | — | |
| LPLDIPC=50, Pruning rate=30x, Backbone=ResNet-502024.10 | 56.2 | — | — | — | — | — | — | |
| LPLDIPC=20, Pruning rate=1x, Backbone=ResNet-502024.10 | 54.4 | — | — | — | — | — | — | |
| LPLDIPC=100, Pruning rate=100x, Backbone=ResNet-502024.10 | 54.2 | — | — | — | — | — | — | |
| LPLDIPC=20, Pruning rate=10x, Backbone=ResNet-502024.10 | 52.3 | — | — | — | — | — | — | |
| LPLDIPC=50, Pruning rate=50x, Backbone=ResNet-502024.10 | 52.3 | — | — | — | — | — | — | |
| LPLDIPC=20, Pruning rate=20x, Backbone=ResNet-502024.10 | 48.9 | — | — | — | — | — | — | |
| LPLDIPC=20, Pruning rate=30x, Backbone=ResNet-502024.10 | 45.4 | — | — | — | — | — | — | |
| LPLDIPC=50, Pruning rate=100x, Backbone=ResNet-502024.10 | 44.7 | — | — | — | — | — | — | |
| LPLDIPC=10, Pruning rate=1x, Backbone=ResNet-502024.10 | 41.7 | — | — | — | — | — | — | |
| LPLDIPC=20, Pruning rate=50x, Backbone=ResNet-502024.10 | 39.5 | — | — | — | — | — | — | |
| LPLDIPC=10, Pruning rate=10x, Backbone=ResNet-502024.10 | 37.7 | — | — | — | — | — | — | |
| LPLDIPC=10, Pruning rate=20x, Backbone=ResNet-502024.10 | 35.4 | — | — | — | — | — | — | |
| LPLDIPC=10, Pruning rate=30x, Backbone=ResNet-502024.10 | 27.5 | — | — | — | — | — | — | |
| LPLDIPC=20, Pruning rate=100x, Backbone=ResNet-502024.10 | 24 | — | — | — | — | — | — | |
| LPLDIPC=10, Pruning rate=50x, Backbone=ResNet-502024.10 | 22.6 | — | — | — | — | — | — | |
| LPLDIPC=10, Pruning rate=100x, Backbone=ResNet-502024.10 | 11 | — | — | — | — | — | — | |
| AdamWModel=ResNet-50, Optimizer=AdamW2026.06 | — | — | — | — | 1.5629 | 1.2283 | 1.0755 | |
| AdamWf=N/A, τ=N/A, ϵ0=10−9, ϵmax=N/A2026.06 | — | 0.635 | 0.8736 | — | — | — | — | |
| DR-Shampoof=20, τ=0.75, ϵ0=10−9, ϵmax=N/A2026.06 | — | 0.531 | 0.9097 | — | — | — | — | |
| FOAMf=20, τ=0.75, ϵ0=10−9, ϵmax=3 × 10−72026.06 | — | 0.458 | 0.9028 | — | — | — | — | |
| FourierBackbone=ConvNeXt-T (28M), Act.=Fourier, Deg.=6, FLOPs=4.83G, FLOPs/Act.=7d+1 = 432025.02 | — | 2.756 | — | — | — | — | — | |
| GELUBackbone=ConvNeXt-T (28M), Act.=GELU, FLOPs=4.57G, FLOPs/Act.=122025.02 | — | 2.825 | — | — | — | — | — | |
| HermiteBackbone=ConvNeXt-T (28M), Act.=Hermite, Deg.=3, FLOPs=4.58G, FLOPs/Act.=4d+1 = 132025.02 | — | 2.8 | — | — | — | — | — | |
| LionModel=ResNet-50, Optimizer=Lion2026.06 | — | — | — | — | 1.7574 | 1.4263 | 1.1685 | |
| LPSGDMModel=ResNet-50, Optimizer=LPSGDM2026.06 | — | — | — | — | 1.6447 | 1.1491 | 1.1253 | |
| SGD w/ MomentumModel=ResNet-50, Optimizer=SGD w/ Momentum2026.06 | — | — | — | — | 2.5096 | 1.8504 | 1.5996 | |
| SOAPf=20, τ=N/A, ϵ0=10−9, ϵmax=N/A2026.06 | — | 0.485 | 0.9444 | — | — | — | — | |
| stale Shampoof=20, τ=N/A, ϵ0=10−9, ϵmax=N/A2026.06 | — | 0.498 | 0.9826 | — | — | — | — | |
| TropicalBackbone=ConvNeXt-T (28M), Act.=Tropical, Deg.=6, FLOPs=4.62G, FLOPs/Act.=3d + 1 = 192025.02 | — | 2.854 | — | — | — | — | — |