Cross-modal Retrieval on RSICD (test)
21.13Image-to-Text R@1Full-FT GeoRSCLIP
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Full-FT GeoRSCLIPBackbone (image/text)=GeoRSCLIP(ViT-B-32-RET-2), Trainable Params=151M, Extra Data=true, Evaluation Protocol=Full Fine-tuned2024.04 | 21.13 | 41.72 | 55.63 | 15.59 | 41.19 | 57.99 | 38.87 | |
| HarMABackbone (image/text)=GeoRSCLIP(ViT-B-32-RET-2), Trainable Params=0.50M, Extra Data=false2024.04 | 20.52 | 41.37 | 54.66 | 15.84 | 41.92 | 59.39 | 38.95 | |
| Full-FT GeoRSCLIPBackbone (image/text)=GeoRSCLIP(ViT-B-32-RET-2), Trainable Params=151M, Evaluation Protocol=Full Fine-tuned2024.04 | 18.85 | 38.15 | 53.16 | 14.27 | 39.71 | 57.49 | 36.94 | |
| RemoteCLIPTraining Dataset=RET-3 + DET-10 + SEG-4, Training Samples=165,745, Date=Jun 2023, Image Backbone Name=ViT-L-14, Image Backbone Params=304, Text Backbone Name=Transformer, Text Backbone Params=124, Total Params=4282023.06 | 18.39 | 37.42 | 51.05 | 14.73 | 39.93 | 56.58 | 36.35 | |
| CLIP-CLTraining Dataset=RET-3, Training Samples=13,713, Image Backbone Name=ViT-B-32, Image Backbone Params=88, Text Backbone Name=Transformer, Text Backbone Params=64, Total Params=1512023.06 | 17.84 | 35.96 | 50.14 | 13.89 | 35.15 | 50.08 | 33.96 | |
| RemoteCLIPTraining Dataset=RET-3 + DET-10 + SEG-4, Training Samples=165,745, Date=Jun 2023, Image Backbone Name=ViT-B-32, Image Backbone Params=87, Text Backbone Name=Transformer, Text Backbone Params=63, Total Params=1512023.06 | 17.02 | 37.97 | 51.51 | 13.71 | 37.11 | 54.25 | 35.26 | |
| HarMABackbone (image/text)=CLIP(ViT-B-32), Trainable Params=0.50M, Extra Data=false2024.04 | 16.36 | 34.48 | 47.74 | 12.92 | 37.17 | 53.07 | 33.62 | |
| Full-FT CLIPBackbone (image/text)=CLIP(ViT-B-32), Trainable Params=151M, Evaluation Protocol=Full Fine-tuned2024.04 | 15.89 | 36.14 | 47.93 | 12.21 | 32.97 | 48.84 | 32.33 | |
| PE-RSITRBackbone (image/text)=CLIP(ViT-B-32), Trainable Params=0.16M, Evaluation Protocol=Parameter-efficient Fine-tuned2024.04 | 14.13 | 31.51 | 44.78 | 11.63 | 33.92 | 50.73 | 31.12 | |
| RemoteCLIPTraining Dataset=RET-3 + DET-10 + SEG-4, Training Samples=165,745, Date=Jun 2023, Image Backbone Name=ResNet-50, Image Backbone Params=38, Text Backbone Name=Transformer, Text Backbone Params=64, Total Params=1022023.06 | 13.36 | 32.94 | 44.83 | 10.76 | 32.83 | 48.75 | 30.58 | |
| FBCLMTraining Dataset=RSICD, Training Samples=10,921, Date=May 2022, Image Backbone Name=Resnet-18, Text Backbone Name=BERT2023.06 | 13.27 | 27.17 | 37.6 | 13.54 | 38.74 | 56.94 | 31.21 | |
| CLIP-CLTraining Dataset=RET-3, Training Samples=13,713, Image Backbone Name=ResNet-50, Image Backbone Params=38, Text Backbone Name=Transformer, Text Backbone Params=64, Total Params=1022023.06 | 12.99 | 26.35 | 36.32 | 8.56 | 25.6 | 39.16 | 24.83 | |
| UniAdapterBackbone (image/text)=CLIP(ViT-B-32), Trainable Params=0.55M, Evaluation Protocol=Parameter-efficient Fine-tuned2024.04 | 12.65 | 30.81 | 42.74 | 9.61 | 30.06 | 47.16 | 28.84 | |
| AdaptFormerBackbone (image/text)=CLIP(ViT-B-32), Trainable Params=0.17M, Evaluation Protocol=Parameter-efficient Fine-tuned2024.04 | 12.46 | 28.49 | 41.86 | 9.09 | 29.89 | 46.81 | 28.1 | |
| Cross-Modal AdapterBackbone (image/text)=CLIP(ViT-B-32), Trainable Params=0.16M, Evaluation Protocol=Parameter-efficient Fine-tuned2024.04 | 11.18 | 27.31 | 40.62 | 9.57 | 30.74 | 48.36 | 27.96 | |
| Rahhal et al.Training Dataset=RSICD, Training Samples=10,921, Date=Oct 2022, Image Backbone Name=ViT-B-32, Image Backbone Params=87, Text Backbone Name=Transformer, Text Backbone Params=63, Total Params=1512023.06 | 10.7 | 29.64 | 41.53 | 9.14 | 28.96 | 44.59 | 27.43 | |
| PIRTraining Dataset=RSICD, Training Samples=10,921, Date=Oct 2023, Image Backbone Name=SwinT + ResNet-50, Text Backbone Name=Bert2023.06 | 9.88 | 27.26 | 39.16 | 6.97 | 24.56 | 38.92 | 24.46 | |
| PIRBackbone (image/text)=Swin Transformer, Bert2024.04 | 9.88 | 27.26 | 39.16 | 6.97 | 24.56 | 38.92 | 24.46 | |
| AdapterBackbone (image/text)=CLIP(ViT-B-32), Trainable Params=0.17M, Evaluation Protocol=Parameter-efficient Fine-tuned2024.04 | 8.73 | 24.73 | 37.81 | 8.43 | 26.02 | 43.33 | 24.84 | |
| DOVETraining Dataset=RSICD, Training Samples=10,921, Date=Oct 2023, Image Backbone Name=ResNet-50, Text Backbone Name=GRU2023.06 | 8.66 | 22.35 | 34.95 | 6.04 | 23.95 | 40.35 | 22.72 | |
| HVSATraining Dataset=RSICD, Training Samples=10,921, Date=Sept 2023, Image Backbone Name=ResNet-18, Text Backbone Name=GRU2023.06 | 7.47 | 20.62 | 32.11 | 5.51 | 21.13 | 34.13 | 16.43 | |
| HyperMatchTraining Dataset=RSICD, Training Samples=10,921, Date=Dec 2022, Image Backbone Name=ResNet-18, Text Backbone Name=GRU2023.06 | 7.14 | 20.04 | 31.02 | 6.08 | 20.37 | 33.82 | 19.75 | |
| CLIP-AdapterBackbone (image/text)=CLIP(ViT-B-32), Trainable Params=0.52M, Evaluation Protocol=Parameter-efficient Fine-tuned2024.04 | 7.11 | 19.48 | 31.01 | 7.67 | 24.87 | 39.73 | 21.65 | |
| GaLRTraining Dataset=RSICD, Training Samples=10,921, Date=Apr 2022, Image Backbone Name=ResNet-18, Text Backbone Name=GRU2023.06 | 6.59 | 19.85 | 31.04 | 4.69 | 19.48 | 32.13 | 18.96 | |
| GaLR with MRBackbone (image/text)=ResNet18, biGRU, Trainable Params=46.89M2024.04 | 6.59 | 19.85 | 31.04 | 4.69 | 19.48 | 32.13 | 18.96 | |
| SCANTraining Dataset=RSICD, Training Samples=10,921, Date=Mar 2018, Image Backbone Name=ResNet-101, Text Backbone Name=GRU2023.06 | 5.85 | 12.89 | 19.84 | 3.71 | 16.4 | 26.73 | 14.23 | |
| KCRTraining Dataset=RSICD, Training Samples=10,921, Date=Sept 2022, Image Backbone Name=ResNet-101, Text Backbone Name=Transformer2023.06 | 5.84 | 22.31 | 36.12 | 4.76 | 18.59 | 27.2 | 19.14 | |
| CMFM-NetTraining Dataset=RSICD, Training Samples=10,921, Date=Dec 2022, Image Backbone Name=ResNet-18, Text Backbone Name=GRU2023.06 | 5.4 | 18.66 | 28.55 | 5.31 | 18.57 | 30.03 | 17.75 | |
| AMFMNTraining Dataset=RSICD, Training Samples=10,921, Date=Aug 2020, Image Backbone Name=ResNet-50, Text Backbone Name=GloVe fasText2023.06 | 5.39 | 15.08 | 23.4 | 4.9 | 18.28 | 31.44 | 16.42 | |
| MTFNTraining Dataset=RSICD, Training Samples=10,921, Date=Aug 2019, Image Backbone Name=ResNet, Text Backbone Name=GRU2023.06 | 5.02 | 12.52 | 19.74 | 4.9 | 17.17 | 29.49 | 14.81 | |
| LW-MRC-uTraining Dataset=RSICD, Training Samples=10,921, Date=Dec 2020, Image Backbone Name=Big Transfer, Text Backbone Name=Bi-LSTM2023.06 | 4.39 | 13.35 | 20.29 | 4.3 | 18.85 | 32.34 | 15.59 | |
| VSE++Training Dataset=RSICD, Training Samples=10,921, Date=Jul 2017, Image Backbone Name=VGG19, Text Backbone Name=GRU2023.06 | 3.38 | 9.51 | 17.46 | 2.82 | 11.32 | 18.1 | 10.43 |