Classification on VOC 2007
95AccuracyDREAM
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DREAMWay=5, Shot=52026.03 | 95 | — | |
| BYOLArchitecture=ResNet-50, Pre-training Dataset=ImageNet, Evaluation Protocol=Fine-tuned2020.06 | 85.4 | — | |
| Supervised-INArchitecture=ResNet-50, Pre-training Dataset=ImageNet, Evaluation Protocol=Fine-tuned2020.06 | 85 | — | |
| SupervisedProtocol=Many-Shot classification, Pre-training=ImageNet2021.12 | 84.76 | — | |
| BASIC-LBackbone=BASIC-L2021.11 | 84.6 | — | |
| SimCLR (repro)Architecture=ResNet-50, Pre-training Dataset=ImageNet, Evaluation Protocol=Fine-tuned2020.06 | 84.5 | — | |
| CLIPBackbone=ViT-L/14-3362021.11 | 84.3 | — | |
| BASIC-MBackbone=BASIC-M2021.11 | 84.2 | — | |
| SimCLRArchitecture=ResNet-50, Pre-training Dataset=ImageNet, Evaluation Protocol=Fine-tuned2020.06 | 84.1 | — | |
| CLIPBackbone=ViT-B/162021.11 | 83.9 | — | |
| BASIC-SBackbone=BASIC-S2021.11 | 83.4 | — | |
| MMCLProtocol=Many-Shot classification, Pre-training=ImageNet2021.12 | 82.83 | — | |
| Supervised-INArchitecture=ResNet-50, Pre-training Dataset=ImageNet, Evaluation Protocol=Linear evaluation2020.06 | 82.8 | — | |
| BYOLArchitecture=ResNet-50, Pre-training Dataset=ImageNet, Evaluation Protocol=Linear evaluation2020.06 | 82.5 | — | |
| PCL-v2Protocol=Many-Shot classification, Pre-training=ImageNet2021.12 | 82.2 | — | |
| CLIPBackbone=ResNet-502021.11 | 82.1 | — | |
| PCL-v1Protocol=Many-Shot classification, Pre-training=ImageNet2021.12 | 82.08 | — | |
| SimCLR (repro)Architecture=ResNet-50, Pre-training Dataset=ImageNet, Evaluation Protocol=Linear evaluation2020.06 | 81.4 | — | |
| SimCLRArchitecture=ResNet-50, Pre-training Dataset=ImageNet, Evaluation Protocol=Linear evaluation2020.06 | 80.5 | — | |
| MoCHiProtocol=Many-Shot classification, Pre-training=ImageNet2021.12 | 79.73 | — | |
| CLIPWay=5, Shot=52026.03 | 79.6 | — | |
| MoCo-v2Protocol=Many-Shot classification, Pre-training=ImageNet2021.12 | 79.34 | — | |
| MoCoProtocol=Many-Shot classification, Pre-training=ImageNet2021.12 | 74.61 | — | |
| InsDisProtocol=Many-Shot classification, Pre-training=ImageNet2021.12 | 71.9 | — | |
| Random initArchitecture=ResNet-50, Pre-training Dataset=ImageNet, Evaluation Protocol=Fine-tuned2020.06 | 67.3 | — | |
| REPAWay=5, Shot=52026.03 | 61.2 | — | |
| xCLIPBackbone=ViT-B/16, Pre-trained on=IT35M, Mode=Zero-shot2022.10 | 52.2 | — | |
| nCLIPBackbone=ViT-B/16, Pre-trained on=IT35M, Mode=Zero-shot2022.10 | 51.4 | — | |
| FLUIDWay=5, Shot=52026.03 | 47.9 | — | |
| MARWay=5, Shot=52026.03 | 47.2 | — | |
| CLIPBackbone=ViT-B/16, Pre-trained on=IT35M, Mode=Zero-shot2022.10 | 46.4 | — | |
| A-CLIPImage Encoder=ViT-B/16, Protocol=Zero-shot2025.03 | — | 51.1 | |
| CLIPImage Encoder=ViT-B/16, Protocol=Zero-shot2025.03 | — | 50.5 | |
| CLIPPre-training Dataset=DataComp-1B, Backbone=ViT-B/16, Training Epochs=10 epochs, Zero-shot protocol=EVA-CLIP2026.03 | — | 37.3 | |
| CLIP-PGSImage Encoder=ViT-B/16, p=0.5, Protocol=Zero-shot2025.03 | — | 50 | |
| CLIP-PGSImage Encoder=ViT-B/16, p=0.3, Protocol=Zero-shot2025.03 | — | 55.1 | |
| CLIP-PGSImage Encoder=ViT-S/16, p=0.5, Protocol=Zero-shot2025.03 | — | 48 | |
| CLIP-PGSImage Encoder=ViT-S/16, p=0.3, Protocol=Zero-shot2025.03 | — | 47.1 | |
| CLIP-ResNet-50x64Backbone=ResNet-50x64, Evaluation Protocol=Zero-shot2021.11 | — | 83.8 | |
| CLIP-ResNet-50x64Evaluation Protocol=Linear Probing2021.11 | — | 88.9 | |
| CLIP-ViT-L/14Backbone=ViT-L/14, Resolution=336px, Evaluation Protocol=Zero-shot2021.11 | — | 84.3 | |
| CLIP-ViT-L/14Resolution=336px, Evaluation Protocol=Linear Probing2021.11 | — | 89.9 | |
| E-CLIPImage Encoder=ViT-B/16, Protocol=Zero-shot2025.03 | — | 43.6 | |
| EfficientNet-L2Resolution=800px, Evaluation Protocol=Linear Probing2021.11 | — | 89.4 | |
| FLAIRBackbone=ViT-B/16, Pre-training Dataset=CC3M-recap, Evaluation Protocol=Zero-shot2026.03 | — | 26.6 | |
| FLIPImage Encoder=ViT-B/16, Protocol=Zero-shot2025.03 | — | 46.6 | |
| Florence-CoSwin-HBackbone=CoSwin-H, Resolution=384px, Evaluation Protocol=Zero-shot2021.11 | — | 85.5 | |
| Florence-CoSwin-HResolution=384px, Backbone=CoSwin-H, Evaluation Protocol=Linear Probing2021.11 | — | 90.5 | |
| ITOPre-training Dataset=DataComp-1B, Backbone=ViT-B/16, Training Epochs=10 epochs, Zero-shot protocol=EVA-CLIP2026.03 | — | 38.3 | |
| ITO sub2Backbone=ViT-B/16, Pre-training Dataset=CC3M-recap, Evaluation Protocol=Zero-shot2026.03 | — | 27.1 | |
| ITO sub2Pre-training Dataset=DataComp-1B, Backbone=ViT-B/16, Training Epochs=10 epochs, Zero-shot protocol=EVA-CLIP2026.03 | — | 38.1 | |
| ITO sub3Backbone=ViT-B/16, Pre-training Dataset=CC3M-recap, Evaluation Protocol=Zero-shot2026.03 | — | 22.5 | |
| SimCLRv2Backbone=ResNet-152x3, Evaluation Protocol=Linear Probing2021.11 | — | 86.7 | |
| ViT-L/16Resolution=384px, Evaluation Protocol=Linear Probing2021.11 | — | 86.1 |