Image Classification on Fashion-MNIST (val)
95.56AccuracyCCT-7/3x1
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| CCT-7/3x1# Params=3.76 M, MACs=1.19 G2021.04 | 95.56 | — | — | — | — | — | — | — | |
| ResNet110# Params=1.73 M, MACs=0.26 G2021.04 | 95.32 | — | — | — | — | — | — | — | |
| CVT-7/4# Params=3.72 M, MACs=0.25 G2021.04 | 95.32 | — | — | — | — | — | — | — | |
| MobileNetV2/2.0# Params=8.72 M, MACs=0.02 G2021.04 | 95.26 | — | — | — | — | — | — | — | |
| ResNet56# Params=0.85 M, MACs=0.13 G2021.04 | 95.25 | — | — | — | — | — | — | — | |
| ViT-Lite-7/4# Params=3.72 M, MACs=0.26 G2021.04 | 95.16 | — | — | — | — | — | — | — | |
| CCT-7/3x2# Params=3.85 M, MACs=0.29 G2021.04 | 95.16 | — | — | — | — | — | — | — | |
| ResNet18# Params=11.18 M, MACs=0.04 G2021.04 | 94.78 | — | — | — | — | — | — | — | |
| ResNet34# Params=21.29 M, MACs=0.08 G2021.04 | 94.78 | — | — | — | — | — | — | — | |
| CVT-7/8# Params=3.74 M, MACs=0.06 G2021.04 | 94.5 | — | — | — | — | — | — | — | |
| ViT-Lite-7/8# Params=3.74 M, MACs=0.06 G2021.04 | 94.49 | — | — | — | — | — | — | — | |
| PLIFOptimizable parameters=13.2M2025.12 | 94.38 | — | — | — | — | — | — | — | |
| CCT-2/3x2# Params=0.28 M, MACs=0.04 G2021.04 | 94.08 | — | — | — | — | — | — | — | |
| MobileNetV2/0.5# Params=0.70 M, MACs=< 0.01 G2021.04 | 93.93 | — | — | — | — | — | — | — | |
| ViT-12/16# Params=85.63 M, MACs=0.43 G2021.04 | 93.61 | — | — | — | — | — | — | — | |
| ViT-Lite-7/16# Params=3.89 M, MACs=0.02 G2021.04 | 93.24 | — | — | — | — | — | — | — | |
| LISNNOptimizable parameters=272,4162025.12 | 92.07 | — | — | — | — | — | — | — | |
| ANN-ResNet34Optimizable parameters=21.8M2025.12 | 91.13 | — | — | — | — | — | — | — | |
| BPNumber of hidden layers=32023.10 | 90.3 | — | — | — | — | — | — | — | |
| ST-RSBP (400-R400)Optimizable parameters=478,8412025.12 | 90.13 | — | — | — | — | — | — | — | |
| ANN-ResNet18Optimizable parameters=11.8M2025.12 | 90.12 | — | — | — | — | — | — | — | |
| BPNumber of hidden layers=12023.10 | 90.1 | — | — | — | — | — | — | — | |
| Exact inverseNumber of hidden layers=12023.10 | 90 | — | — | — | — | — | — | — | |
| Exact inverseNumber of hidden layers=32023.10 | 90 | — | — | — | — | — | — | — | |
| Exact inverse (Avg J)Number of hidden layers=12023.10 | 89.8 | — | — | — | — | — | — | — | |
| Exact inverse (Avg J)Number of hidden layers=32023.10 | 89.2 | — | — | — | — | — | — | — | |
| SEW-ResNet18 (ADD)Optimizable parameters=11.8M2025.12 | 88.83 | — | — | — | — | — | — | — | |
| Linear thresholdNumber of hidden layers=12023.10 | 88.8 | — | — | — | — | — | — | — | |
| Linear thresholdNumber of hidden layers=32023.10 | 88.1 | — | — | — | — | — | — | — | |
| Linear threshold (Avg J)Number of hidden layers=12023.10 | 87.9 | — | — | — | — | — | — | — | |
| SPT-slicSegmentation=SLIC2026.05 | 87.9 | — | — | — | — | — | — | — | |
| ViTGraph=Fully-connected, Embedding=BERT2026.05 | 86.5 | — | — | — | — | — | — | — | |
| Linear threshold (Avg J)Number of hidden layers=32023.10 | 86.2 | — | — | — | — | — | — | — | |
| SQDR-CNN[4b-18q]Optimizable parameters=1,2982025.12 | 82.18 | — | — | — | — | — | — | — | |
| SQDR-CNN[4b-9q]Optimizable parameters=8572025.12 | 79.99 | — | — | — | — | — | — | — | |
| SQDR-CNN[2b-18q]Optimizable parameters=1,1902025.12 | 77.8 | — | — | — | — | — | — | — | |
| SQDR-CNN[2b-9q]Optimizable parameters=8032025.12 | 76.18 | — | — | — | — | — | — | — | |
| Spiking-ResNet34Optimizable parameters=21.8M2025.12 | 69.59 | — | — | — | — | — | — | — | |
| Spiking-ResNet50Optimizable parameters=25.6M2025.12 | 12.8 | — | — | — | — | — | — | — | |
| AdamModel Architecture=Two-layer ReLU MLP with 64 neurons, Best Learning Rate=3.52 × 10−42024.11 | — | — | — | — | — | — | — | 12.24 | |
| AdamWModel Architecture=Two-layer ReLU MLP with 64 neurons, Best Learning Rate=3.52 × 10−42024.11 | — | — | — | — | — | — | — | 12.23 | |
| CRONOSModel Architecture=Two-layer ReLU MLP with 64 neurons, Best Learning Rate=NA2024.11 | — | — | — | — | — | — | — | 85.8 | |
| CTMd_model=256, Backbone=HRF, Steps=200K, Seeding Protocol=single seed2026.05 | — | 92.8 | — | — | — | — | — | — | |
| SGDModel Architecture=Two-layer ReLU MLP with 64 neurons, Best Learning Rate=1.97 × 10−22024.11 | — | — | — | — | — | — | — | 6.63 | |
| ShampooModel Architecture=Two-layer ReLU MLP with 64 neurons, Best Learning Rate=2.00 × 10−22024.11 | — | — | — | — | — | — | — | 5.87 | |
| TIDEd_model=256, Backbone=HRF, Steps=50K, Seeding Protocol=multi-seeded2026.05 | — | 94.24 | 94.16 | 94.02 | 0.3 | 92.68 | 2.79 | — | |
| YogiModel Architecture=Two-layer ReLU MLP with 64 neurons, Best Learning Rate=3.52 × 10−42024.11 | — | — | — | — | — | — | — | 19.44 |