Long-caption Cross-modal Retrieval on DCI (test)
71.79T2I Recall@1HyFL-CLIP
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| HyFL-CLIPModel Type=Euclidean VLM with our hyperbolic fine-tuning, Backbone=ViT-B, Protocol=Zero-shot2026.07 | 71.79 | 71.54 | |
| UNCHA + OursModel Type=State-of-the-art hyperbolic VLM, Backbone=ViT-B, Protocol=Zero-shot, Framework=hyperbolic fine-tuning without distillation loss2026.07 | 60.58 | 62.48 | |
| HyCoCLIP + OursModel Type=State-of-the-art hyperbolic VLM, Backbone=ViT-B, Protocol=Zero-shot, Framework=hyperbolic fine-tuning without distillation loss2026.07 | 58.88 | 60.68 | |
| MERU + OursModel Type=State-of-the-art hyperbolic VLM, Backbone=ViT-B, Protocol=Zero-shot, Framework=hyperbolic fine-tuning without distillation loss2026.07 | 55.03 | 58.93 | |
| HyCoCLIPModel Type=State-of-the-art hyperbolic VLM, Backbone=ViT-B, Protocol=Zero-shot2026.07 | 47.87 | 49.02 | |
| OpenCLIPModel Type=Euclidean VLM, Backbone=ViT-B, Protocol=Zero-shot2026.07 | 46.97 | 50.78 | |
| MERUModel Type=State-of-the-art hyperbolic VLM, Backbone=ViT-B, Protocol=Zero-shot2026.07 | 46.52 | 49.77 | |
| UNCHAModel Type=State-of-the-art hyperbolic VLM, Backbone=ViT-B, Protocol=Zero-shot2026.07 | 45.12 | 44.82 |