Long-caption Cross-modal Retrieval on Urban-1k (test)
91.1T2I Recall@1HyFL-CLIP
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| HyFL-CLIPModel Type=Euclidean VLM with our hyperbolic fine-tuning, Backbone=ViT-B, Protocol=Zero-shot2026.07 | 91.1 | 91.8 | |
| UNCHA + OursModel Type=State-of-the-art hyperbolic VLM, Backbone=ViT-B, Protocol=Zero-shot, Framework=hyperbolic fine-tuning without distillation loss2026.07 | 77.1 | 76.8 | |
| HyCoCLIP + OursModel Type=State-of-the-art hyperbolic VLM, Backbone=ViT-B, Protocol=Zero-shot, Framework=hyperbolic fine-tuning without distillation loss2026.07 | 75.3 | 76.7 | |
| MERU + OursModel Type=State-of-the-art hyperbolic VLM, Backbone=ViT-B, Protocol=Zero-shot, Framework=hyperbolic fine-tuning without distillation loss2026.07 | 74.2 | 74.3 | |
| OpenCLIPModel Type=Euclidean VLM, Backbone=ViT-B, Protocol=Zero-shot2026.07 | 53.4 | 67.5 | |
| HyCoCLIPModel Type=State-of-the-art hyperbolic VLM, Backbone=ViT-B, Protocol=Zero-shot2026.07 | 47.7 | 56.6 | |
| MERUModel Type=State-of-the-art hyperbolic VLM, Backbone=ViT-B, Protocol=Zero-shot2026.07 | 45.7 | 51.9 | |
| UNCHAModel Type=State-of-the-art hyperbolic VLM, Backbone=ViT-B, Protocol=Zero-shot2026.07 | 39.9 | 43.8 |