Image Classification on ImageNet-Sketch (val)
67.3Top-1 AccEVA CLIP-g
Evaluation Results
| Method | Links | |
|---|---|---|
| EVA CLIP-gzero-shot evaluation protocol=True2022.11 | 67.3 | |
| Open CLIP-Hzero-shot evaluation protocol=True2022.11 | 66.6 | |
| Open CLIP-gzero-shot evaluation protocol=True2022.11 | 65.2 | |
| ALIGNzero-shot evaluation protocol=True2022.11 | 64.8 | |
| PaLI-17BShot Setting=0-shot2022.09 | 63.83 | |
| PaLI-3BShot Setting=0-shot2022.09 | 61.11 | |
| PaLI-15BShot Setting=0-shot2022.09 | 61.03 | |
| OpenAI CLIP-Lzero-shot evaluation protocol=True2022.11 | 59.6 | |
| MAEModel=ViT-H, Resolution=448, Training Objective=MAE pre-training2021.11 | 50.9 | |
| MAEModel=ViT-H, Resolution=224, Training Objective=MAE pre-training2021.11 | 49.6 | |
| OTTEREvaluation Protocol=Linear Probing, Label Distribution Estimation=BBSE+OT2024.04 | 48.3 | |
| MAEModel=ViT-L, Resolution=224, Training Objective=MAE pre-training2021.11 | 45.3 | |
| Linear Probing baselineEvaluation Protocol=Linear Probing, Label Distribution Estimation=None2024.04 | 43.4 | |
| OTTEREvaluation Protocol=Zero-Shot, Label Distribution Estimation=BBSE+OT2024.04 | 40.4 | |
| Zero-Shot baselineEvaluation Protocol=Zero-Shot, Label Distribution Estimation=None2024.04 | 39.8 | |
| Supervised BaselineModel=ViT-H, Resolution=224, Training Objective=Supervised2021.11 | 38 | |
| BBSE+PMEvaluation Protocol=Linear Probing, Label Distribution Estimation=BBSE+PM2024.04 | 37.9 | |
| Supervised BaselineModel=ViT-L, Resolution=224, Training Objective=Supervised2021.11 | 37.5 | |
| MSNArchitecture=ViT-B/16, Pre-training method=MSN2022.04 | 36.3 | |
| LPTBackbone=ViT-B2022.10 | 36.22 | |
| Noisy StudentNotes=Previous Best2021.11 | 36 | |
| Supervised BaselineModel=ViT-B, Resolution=224, Training Objective=Supervised2021.11 | 35.6 | |
| MAEModel=ViT-B, Resolution=224, Training Objective=MAE pre-training2021.11 | 34.5 | |
| MAEArchitecture=ViT-B/16, Pre-training method=MAE2022.04 | 34.5 | |
| Full fine-tuneBackbone=ViT-B2022.10 | 32.25 | |
| Linear ProbeBackbone=ViT-B2022.10 | 31.55 | |
| LSNet-BFLOPS=1.3G2025.03 | 30.7 | |
| PVTv2-B1FLOPS=2.1G2025.03 | 28.9 | |
| EdgeNeXt-SFLOPS=1.3G2025.03 | 28.8 | |
| FastViT-T12FLOPS=1.4G2025.03 | 27.6 | |
| LSNet-SFLOPS=0.5G2025.03 | 27.5 | |
| FasterNet-T2FLOPS=1.9G2025.03 | 27.2 | |
| UniRepLKNet-AFLOPS=0.6G2025.03 | 26 | |
| LSNet-TFLOPS=0.3G2025.03 | 25.5 | |
| FastViT-T8FLOPS=0.7G2025.03 | 25.5 | |
| PoolFormer-S12FLOPS=1.8G2025.03 | 25.2 | |
| ResNet50Architecture=ResNet50, Pre-training method=Supervised2022.04 | 24.2 | |
| EfficientViT-M3FLOPS=0.3G2025.03 | 23.4 | |
| EdgeNeXt-XSFLOPS=0.5G2025.03 | 22 | |
| StarNet-S1FLOPS=0.4G2025.03 | 21.8 | |
| PVTv2-B0FLOPS=0.6G2025.03 | 21.5 | |
| PVT-TinyFLOPS=1.9G2025.03 | 21.5 | |
| MoCo-v2 + Patch-base NSalpha=2, k=163842021.10 | 19.49 | |
| MoCo-v2 + Patch-base NSalpha=3, k=655362021.10 | 18.79 | |
| MoCo-v2 + Patch-base NSalpha=2, k=327682021.10 | 18.7 | |
| MoCo-v2 + Patch-base NSalpha=2, k=655362021.10 | 18.58 | |
| EdgeNeXt-XXSFLOPS=0.3G2025.03 | 18.5 | |
| MoCo-v22021.10 | 17.47 | |
| MoCo-v2 + MoCHiN=512, K=1024, s=5122021.10 | 16.32 | |
| FasterNet-T0FLOPS=0.3G2025.03 | 16.3 | |
| MoCo-v2 + MoCHiN=128, K=1024, s=5122021.10 | 16 | |
| BBSE+PMEvaluation Protocol=Zero-Shot, Label Distribution Estimation=BBSE+PM2024.04 | 0.8 |