Image Classification on ImageNet-1K (val) (Accuracy based on N Images per Class)
24Accuracy (1 img/cls)RILS
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| RILSPT Dataset=L-20M, PT Epo.=25(~ 400), Vision Encoder=ViT-B/16, Evaluation Protocol=logistic regression accuracy (%)2023.01 | 24 | 34.6 | 45.7 | 51.8 | |
| MAE+CLIPPT Dataset=L-20M, PT Epo.=25(~ 400), Vision Encoder=ViT-B/16, Evaluation Protocol=logistic regression accuracy (%)2023.01 | 21.1 | 31.1 | 41.6 | 47.5 | |
| CLIPPT Dataset=L-20M, PT Epo.=25(~ 400), Vision Encoder=ViT-B/16, Evaluation Protocol=logistic regression accuracy (%)2023.01 | 19.4 | 29.2 | 39.8 | 46.3 | |
| SLIPPT Dataset=L-20M, PT Epo.=25(~ 400), Vision Encoder=ViT-B/16, Evaluation Protocol=logistic regression accuracy (%)2023.01 | 17.7 | 27.2 | 38.6 | 46.4 | |
| MAEPT Dataset=IN-1K(~ 1.3M), PT Epo.=1600, Vision Encoder=ViT-B/16, Evaluation Protocol=logistic regression accuracy (%)2023.01 | 4.3 | 10.6 | 22.4 | 31.6 | |
| MAEPT Dataset=L-20M, PT Epo.=25(~ 400), Vision Encoder=ViT-B/16, Evaluation Protocol=logistic regression accuracy (%)2023.01 | 3.4 | 5.2 | 10.1 | 14.8 | |
| BEiTPT Dataset=IN-1K(~ 1.3M), PT Epo.=800, Vision Encoder=ViT-B/16, Evaluation Protocol=logistic regression accuracy (%)2023.01 | 1.3 | 2.2 | 4.4 | 7.4 |