Deep Metric Learning on CUB-2011
88.5R@1VPTSP-G
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| VPTSP-GBackbone=ViT-B16, Feature Dimension=512, Batch Size=64/322024.02 | 88.5 | 92.8 | 95.1 | — | — | |
| VPTSP (CLIP vision)Architecture=Vit-L/14, Pre-training Set=Laion2b2024.02 | 86.7 | — | — | — | — | |
| VPTSP-GBackbone=ViT-S16, Feature Dimension=384, Batch Size=64/322024.02 | 86.6 | 91.7 | 94.8 | — | — | |
| Hyp-ViTBackbone=ViT-S16, Feature Dimension=384, Batch Size=882/9002024.02 | 85.6 | 91.4 | 94.8 | — | — | |
| VPTSP-MBackbone=ViT-S16, Feature Dimension=384, Batch Size=64/322024.02 | 85.4 | 91.2 | 94.6 | — | — | |
| VPT-BaseBackbone=ViT-S16, Feature Dimension=384, Batch Size=64/322024.02 | 85.1 | 91.1 | 94 | — | — | |
| VPTSP (MAE)Architecture=Vit-B/16, Pre-training Set=ImageNet 1K2024.02 | 85.1 | — | — | — | — | |
| VPTSP (DINO)Architecture=Vit-S/16, Pre-training Set=ImageNet 1K2024.02 | 82.1 | — | — | — | — | |
| VPTSP (DeiT)Architecture=Vit-S/16, Pre-training Set=ImageNet 1K2024.02 | 80.2 | — | — | — | — | |
| SoftTripleEmbedding dimension=642019.09 | 60.1 | 71.9 | 81.2 | 88.5 | 66.2 | |
| SoftMaxnormEmbedding dimension=642019.09 | 57.8 | 70 | 80.1 | 87.9 | 65.3 | |
| Npairs*Embedding dimension=642019.09 | 51 | 63.3 | 74.3 | 83.2 | 60.4 | |
| ProxyNCAEmbedding dimension=642019.09 | 49.2 | 61.9 | 67.9 | 72.4 | 59.5 | |
| ClusteringEmbedding dimension=642019.09 | 48.2 | 61.4 | 71.8 | 81.9 | 59.2 | |
| LiftedStructEmbedding dimension=642019.09 | 43.6 | 56.6 | 68.6 | 79.6 | 56.5 | |
| SemiHardEmbedding dimension=642019.09 | 42.6 | 55 | 66.4 | 77.2 | 55.4 |