Image-to-Text Retrieval on MS-COCO 5-fold 1K (test)
90.9R@1SEPS
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| SEPSVisual Encoder=ViT-Base-384, Textual Encoder=BERT-base, Image Resolution=384x384, Patch Size=24x24, Fine-Grained Alignment (FG)=true2025.11 | 90.9 | 96.1 | 98.8 | 576.1 | |
| SEPSVisual Encoder=Swin-Base-384, Textual Encoder=BERT-base, Image Resolution=384x384, Patch Size=12x12, Fine-Grained Alignment (FG)=true2025.11 | 89.5 | 96.5 | 99 | 571.2 | |
| SEPSVisual Encoder=ViT-Base-224, Textual Encoder=BERT-base, Image Resolution=224x224, Patch Size=14x14, Fine-Grained Alignment (FG)=true2025.11 | 89 | 94.8 | 98 | 569.5 | |
| SEPSVisual Encoder=Swin-Base-224, Textual Encoder=BERT-base, Image Resolution=224x224, Patch Size=7x7, Fine-Grained Alignment (FG)=true2025.11 | 87.2 | 94.9 | 98.3 | 563.9 | |
| HREMVisual Backbone=Faster R-CNN, Semantic Encoder=BERT, Cross Attention=false, Ensemble=true, Venue=CVPR’232024.06 | 82.9 | 96.9 | 99 | 534.5 | |
| CORAVisual Backbone=Faster R-CNN, Semantic Encoder=BERT, Cross Attention=false, Ensemble=true2024.06 | 82.8 | 97.3 | 99 | 535.6 | |
| CORAVisual Backbone=Faster R-CNN, Semantic Encoder=BERT, Cross Attention=false, Ensemble=false2024.06 | 82.4 | 96.8 | 98.8 | 532.7 | |
| CODERVisual Backbone=Faster R-CNN, Semantic Encoder=BERT, Cross Attention=true, Ensemble=false, Venue=ECCV’222024.06 | 82.1 | 96.6 | 98.8 | 530.7 | |
| CORAVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=false, Ensemble=true2024.06 | 81.7 | 96.7 | 99 | 532.1 | |
| CHANVisual Backbone=Faster R-CNN, Semantic Encoder=BERT, Cross Attention=true, Ensemble=false, Venue=CVPR’232024.06 | 81.4 | 96.9 | 98.9 | 532.5 | |
| CORACross-Attention=false, Backbone=Faster R-CNN + Bi-GRU, Ensemble=true2024.06 | 81.2 | 96.2 | 98.7 | 526.6 | |
| HREMVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=false, Ensemble=true, Venue=CVPR’232024.06 | 81.2 | 96.5 | 98.9 | 527 | |
| CORAVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=false, Ensemble=false2024.06 | 80.9 | 96.3 | 98.8 | 528.5 | |
| SDEVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=false, Ensemble=true, Venue=CVPR’232024.06 | 80.6 | 96.3 | 98.8 | 528 | |
| CORACross-Attention=false, Backbone=Faster R-CNN + Bi-GRU, Ensemble=false2024.06 | 80.5 | 96 | 98.6 | 524.6 | |
| NAAFVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=true, Ensemble=true, Venue=CVPR’232024.06 | 80.5 | 96.5 | 98.8 | 527.2 | |
| MV-VSEVisual Backbone=Faster R-CNN, Semantic Encoder=BERT, Cross Attention=false, Ensemble=true, Venue=IJCAI’222024.06 | 80.4 | 96.6 | 99 | 528.1 | |
| CHANVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=true, Ensemble=false, Venue=CVPR’232024.06 | 79.7 | 96.7 | 98.7 | 525.1 | |
| VSE∞Visual Backbone=Faster R-CNN, Semantic Encoder=BERT, Cross Attention=false, Ensemble=false, Venue=CVPR’212024.06 | 79.7 | 96.4 | 98.9 | 527.5 | |
| SGARFVenue=AAAI’21, Cross-Attention=true, Backbone=Faster R-CNN + Bi-GRU, Ensemble=true2024.06 | 79.6 | 96.2 | 98.5 | 524.3 | |
| SGARFVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=true, Ensemble=true, Venue=AAAI’232024.06 | 79.6 | 96.2 | 98.5 | 524.3 | |
| CODERVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=true, Ensemble=false, Venue=ECCV’222024.06 | 78.9 | 95.6 | 98.6 | 521.6 | |
| MV-VSEVenue=IJCAI’22, Cross-Attention=false, Backbone=Faster R-CNN + Bi-GRU, Ensemble=true2024.06 | 78.7 | 95.7 | 98.7 | 521.9 | |
| MV-VSEVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=false, Ensemble=true, Venue=IJCAI’222024.06 | 78.7 | 95.7 | 98.7 | 521.9 | |
| VSE∞Venue=CVPR’21, Cross-Attention=false, Backbone=Faster R-CNN + Bi-GRU, Ensemble=false2024.06 | 78.5 | 96 | 98.7 | 520.8 | |
| VSE∞Visual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=false, Ensemble=false, Venue=CVPR’212024.06 | 78.5 | 96 | 98.7 | 520.8 | |
| GraDualVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=true, Ensemble=true, Venue=WACV’232024.06 | 77 | 96.4 | 98.6 | 525.6 | |
| VSRNVenue=ICCV’19, Cross-Attention=false, Backbone=Faster R-CNN + Bi-GRU, Ensemble=false2024.06 | 76.2 | 94.8 | 98.2 | 516.8 | |
| VSRNVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=false, Ensemble=false, Venue=ICCV’192024.06 | 76.2 | 94.8 | 98.2 | 516.8 | |
| CAANVenue=CVPR’20, Cross-Attention=true, Backbone=Faster R-CNN + Bi-GRU, Ensemble=false2024.06 | 75.5 | 95.4 | 98.5 | 515.6 | |
| CAANVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=true, Ensemble=false, Venue=CVPR’202024.06 | 75.5 | 95.4 | 98.5 | 515.6 | |
| SGMVenue=WACV’20, Cross-Attention=true, Backbone=Faster R-CNN + Bi-GRU, Ensemble=false2024.06 | 73.4 | 93.8 | 97.8 | 504.1 | |
| SGMVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=true, Ensemble=false, Venue=WACV’202024.06 | 73.4 | 93.8 | 97.8 | 504.1 | |
| SCANVenue=ECCV’18, Cross-Attention=true, Backbone=Faster R-CNN + Bi-GRU, Ensemble=true2024.06 | 72.7 | 94.8 | 98.4 | 507.9 | |
| SCANVisual Backbone=Faster R-CNN, Semantic Encoder=Bi-GRU, Cross Attention=true, Ensemble=true, Venue=ECCV’182024.06 | 72.7 | 94.8 | 98.4 | 507.9 |