Image Classification on CLIP Classification Suite
97.2CIFAR-10 AccuracyFLIP
Evaluation Results
| Method | Links | ||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| FLIPBackbone=ViT-L/14, Pre-training Data=LAION-400M, Image Size=224, Evaluation Protocol=Zero-shot2022.12 | 97.2 | 89.3 | 84.1 | 63 | 73.1 | 90.7 | 29.1 | 83.1 | 60.4 | 92.6 | 93.8 | 75 | 80.3 | 98.5 | 53.5 | 70.8 | 41.4 | 34.8 | 23.1 | 50.3 | 74.1 | 55.8 | 22.7 | 54 | 58.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CLIPBackbone=ViT-L/14, Pre-training Data=WIT-400M, Image Size=224, Evaluation Protocol=Zero-shot2022.12 | 96.2 | 92.9 | 77.9 | 48.3 | 67.7 | 77.3 | 36.1 | 84.1 | 55.3 | 93.5 | 92.6 | 78.7 | 87.2 | 99.3 | 59.9 | 71.6 | 50.3 | 23.1 | 32.7 | 58.8 | 76.2 | 60.3 | 24.3 | 63.3 | 64 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CLIP (our repro.)Backbone=ViT-L/14, Pre-training Data=LAION-400M, Image Size=224, Evaluation Protocol=Zero-shot2022.12 | 96 | 88.1 | 81.3 | 60.5 | 72.3 | 89.1 | 25.8 | 81.1 | 59.3 | 93.2 | 93.2 | 74.6 | 69.1 | 96.5 | 50.7 | 69.2 | 50.2 | 29.4 | 21.4 | 53.1 | 71.5 | 53.5 | 18.5 | 53.3 | 57.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CLIP (our eval.)Backbone=ViT-L/14, Pre-training Data=WIT-400M, Image Size=224, Evaluation Protocol=Zero-shot2022.12 | 95.2 | 91 | 75.6 | 51.2 | 66.6 | 75 | 32.3 | 83.3 | 55 | 93.6 | 92.4 | 77.7 | 76 | 99.3 | 62 | 71.6 | 51.6 | 26.9 | 30.9 | 51.6 | 76.1 | 59.5 | 22.2 | 55.3 | 67.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| OpenCLIP (our eval.)Backbone=ViT-L/14, Pre-training Data=LAION-400M, Image Size=224, Evaluation Protocol=Zero-shot2022.12 | 94.1 | 87.4 | 77.1 | 61.3 | 70.7 | 86.2 | 21.8 | 83.5 | 54.9 | 90.8 | 94 | 72.1 | 71.5 | 98.2 | 53.3 | 67.7 | 47.3 | 29.3 | 21.6 | 51.1 | 71.3 | 50.5 | 22 | 55.3 | 57.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| B-cosified RN-50 CLIPEvaluation Protocol=Linear-Probe, Backbone=RN-50, Pre-training Dataset=ImageNet, Learning Scheduler=Cosine2024.11 | 91 | — | 74 | — | — | 71 | 36 | 83 | 73 | 86 | 92 | 91 | 97 | 98 | 95 | 89 | 82 | — | — | 82 | — | — | — | — | 56 | 60 | 72 | 66 | 76 | 0.52 | 68 | 57 | 55 | 58 | 14 | 39 | 48 | 0.48 | |
| B-cosified RN-50 CLIPEvaluation Protocol=Linear-Probe, Backbone=RN-50, Pre-training Dataset=ImageNet, Learning Scheduler=Cyclic2024.11 | 91 | — | 74 | — | — | 71 | 36 | 83 | 73 | 88 | 92 | 91 | 97 | 98 | 95 | 89 | 83 | — | — | 84 | — | — | — | — | 56 | 60 | 72 | 66 | 76 | 0.53 | 67 | 58 | 53 | 58 | 13 | 39 | 49 | 0.5 | |
| Standard CLIPEvaluation Protocol=Linear-Probe, Backbone=RN-502024.11 | 89 | — | 70 | — | — | 80 | 42 | 82 | 74 | 88 | 92 | 92 | 98 | 97 | 94 | 91 | 84 | — | — | 82 | — | — | — | — | 72 | 63 | 71 | 65 | 76 | 0.53 | 62 | 61 | 52 | 56 | 14 | 37 | 48 | 0.52 | |
| Text2Concept (T2C)Evaluation Protocol=Linear-Probe2024.11 | 89 | — | 70 | — | — | 33 | 23 | 82 | 66 | 89 | 88 | 73 | 97 | 96 | 95 | 84 | 69 | — | — | 84 | — | — | — | — | 51 | 48 | 73 | 60 | 75 | 0.53 | 53 | 49 | 43 | 50 | 14 | 33 | 44 | 0.47 | |
| B-cosified RN-50 CLIPEvaluation Protocol=Linear-Probe, Backbone=RN-50, Pre-training Dataset=CC3M, Learning Scheduler=Cyclic2024.11 | 89 | — | 72 | — | — | 69 | 34 | 81 | 70 | 85 | 86 | 89 | 97 | 97 | 95 | 88 | 82 | — | — | 81 | — | — | — | — | 60 | 61 | 68 | 65 | 76 | 0.55 | 64 | 63 | 54 | 59 | 14 | 40 | 47 | 0.47 | |
| B-cosified RN-50 CLIPEvaluation Protocol=Linear-Probe, Backbone=RN-50, Pre-training Dataset=CC3M, Learning Scheduler=Cosine2024.11 | 88 | — | 71 | — | — | 67 | 33 | 81 | 69 | 84 | 86 | 89 | 98 | 96 | 95 | 87 | 81 | — | — | 82 | — | — | — | — | 62 | 61 | 67 | 63 | 75 | 0.54 | 65 | 61 | 53 | 57 | 14 | 40 | 47 | 0.47 |