Text-to-Image Retrieval on BRSET (test)
9,994R@1Knowledge-Enhanced Multimodal Transformer
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Knowledge-Enhanced Multimodal TransformerEvaluation Protocol=Proposed2025.12 | 9,994 | 10,000 | 10,000 | |
| CLIPEvaluation Protocol=Fine-tuned2025.12 | 129 | 497 | 811 | |
| CLIPEvaluation Protocol=Zero-shot2025.12 | 0 | 25 | 74 |