Text-to-Image Retrieval on MS-COCO 5-fold 1K (test)
91R@1SEPS
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| SEPSVisual Encoder=ViT-Base-384, Textual Encoder=BERT-base, Image Resolution=384x384, Patch Size=24x24, Fine-Grained Alignment (FG)=true2025.11 | 91 | 99.5 | 99.8 | 576.1 | |
| SEPSVisual Encoder=ViT-Base-224, Textual Encoder=BERT-base, Image Resolution=224x224, Patch Size=14x14, Fine-Grained Alignment (FG)=true2025.11 | 88.5 | 99.3 | 99.8 | 569.5 | |
| SEPSVisual Encoder=Swin-Base-384, Textual Encoder=BERT-base, Image Resolution=384x384, Patch Size=12x12, Fine-Grained Alignment (FG)=true2025.11 | 87.1 | 99.2 | 99.9 | 571.2 | |
| SEPSVisual Encoder=Swin-Base-224, Textual Encoder=BERT-base, Image Resolution=224x224, Patch Size=7x7, Fine-Grained Alignment (FG)=true2025.11 | 84.7 | 99 | 99.8 | 563.9 | |
| CORAVisual Backbone=Faster R-CNN, Semantic Encoder=BERT, Cross Attention=false, Ensemble=true2024.06 | 67.3 | 92.4 | 96.9 | 535.6 | |
| HREMVisual Backbone=Faster R-CNN, Semantic Encoder=BERT, Cross Attention=false, Ensemble=true, Venue=CVPR’232024.06 | 67.1 | 92 | 96.6 | 534.5 | |
| CHANVisual Backbone=Faster R-CNN, Semantic Encoder=BERT, Cross Attention=true, Ensemble=false, Venue=CVPR’232024.06 | 66.5 | 92.1 | 96.7 | 532.5 | |
| CORAVisual Backbone=Faster R-CNN, Semantic Encoder=BERT, Cross Attention=false, Ensemble=false2024.06 | 66.2 | 91.9 | 96.6 | 532.7 | |
| CORAVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=false, Ensemble=true2024.06 | 66 | 92 | 96.7 | 532.1 | |
| CODERVisual Backbone=Faster R-CNN, Semantic Encoder=BERT, Cross Attention=true, Ensemble=false, Venue=ECCV’222024.06 | 65.5 | 91.5 | 96.2 | 530.7 | |
| GraDualVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=true, Ensemble=true, Venue=WACV’232024.06 | 65.3 | 91.9 | 96.4 | 525.6 | |
| CORAVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=false, Ensemble=false2024.06 | 64.9 | 91.3 | 96.4 | 528.5 | |
| MV-VSEVisual Backbone=Faster R-CNN, Semantic Encoder=BERT, Cross Attention=false, Ensemble=true, Venue=IJCAI’222024.06 | 64.9 | 91.2 | 96 | 528.1 | |
| VSE∞Visual Backbone=Faster R-CNN, Semantic Encoder=BERT, Cross Attention=false, Ensemble=false, Venue=CVPR’212024.06 | 64.8 | 91.4 | 96.3 | 527.5 | |
| SDEVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=false, Ensemble=true, Venue=CVPR’232024.06 | 64.7 | 91.4 | 96.2 | 528 | |
| NAAFVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=true, Ensemble=true, Venue=CVPR’232024.06 | 64.1 | 90.7 | 96.5 | 527.2 | |
| CHANVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=true, Ensemble=false, Venue=CVPR’232024.06 | 63.8 | 90.4 | 95.8 | 525.1 | |
| HREMVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=false, Ensemble=true, Venue=CVPR’232024.06 | 63.7 | 90.7 | 96 | 527 | |
| CORACross-Attention=false, Backbone=Faster R-CNN + Bi-GRU, Ensemble=true2024.06 | 63.4 | 90.9 | 96.2 | 526.6 | |
| SGARFVenue=AAAI’21, Cross-Attention=true, Backbone=Faster R-CNN + Bi-GRU, Ensemble=true2024.06 | 63.2 | 90.7 | 96.1 | 524.3 | |
| SGARFVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=true, Ensemble=true, Venue=AAAI’232024.06 | 63.2 | 90.7 | 96.1 | 524.3 | |
| CORACross-Attention=false, Backbone=Faster R-CNN + Bi-GRU, Ensemble=false2024.06 | 62.9 | 90.6 | 96 | 524.6 | |
| VSRNVenue=ICCV’19, Cross-Attention=false, Backbone=Faster R-CNN + Bi-GRU, Ensemble=false2024.06 | 62.8 | 89.7 | 95.1 | 516.8 | |
| VSRNVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=false, Ensemble=false, Venue=ICCV’192024.06 | 62.8 | 89.7 | 95.1 | 516.8 | |
| MV-VSEVenue=IJCAI’22, Cross-Attention=false, Backbone=Faster R-CNN + Bi-GRU, Ensemble=true2024.06 | 62.7 | 90.4 | 95.7 | 521.9 | |
| MV-VSEVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=false, Ensemble=true, Venue=IJCAI’222024.06 | 62.7 | 90.4 | 95.7 | 521.9 | |
| CODERVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=true, Ensemble=false, Venue=ECCV’222024.06 | 62.5 | 90.3 | 95.7 | 521.6 | |
| VSE∞Venue=CVPR’21, Cross-Attention=false, Backbone=Faster R-CNN + Bi-GRU, Ensemble=false2024.06 | 61.7 | 90.3 | 95.6 | 520.8 | |
| VSE∞Visual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=false, Ensemble=false, Venue=CVPR’212024.06 | 61.7 | 90.3 | 95.6 | 520.8 | |
| CAANVenue=CVPR’20, Cross-Attention=true, Backbone=Faster R-CNN + Bi-GRU, Ensemble=false2024.06 | 61.3 | 89.7 | 95.2 | 515.6 | |
| CAANVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=true, Ensemble=false, Venue=CVPR’202024.06 | 61.3 | 89.7 | 95.2 | 515.6 | |
| SCANVenue=ECCV’18, Cross-Attention=true, Backbone=Faster R-CNN + Bi-GRU, Ensemble=true2024.06 | 58.8 | 88.4 | 94.8 | 507.9 | |
| SCANVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=true, Ensemble=true, Venue=ECCV’182024.06 | 58.8 | 88.4 | 94.8 | 507.9 | |
| SGMVenue=WACV’20, Cross-Attention=true, Backbone=Faster R-CNN + Bi-GRU, Ensemble=false2024.06 | 57.5 | 87.3 | 94.3 | 504.1 | |
| SGMVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=true, Ensemble=false, Venue=WACV’202024.06 | 57.5 | 87.3 | 94.3 | 504.1 |