Phrase Localization on Flickr30K Entities (test)
83.15AccuracySimVG-DB
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| SimVG-DBVisual Encoder=ViT-L/32, Time (ms)=1162024.09 | 83.15 | — | — | — | — | — | |
| SimVG-TBVisual Encoder=ViT-L/32, Time (ms)=1012024.09 | 82.61 | — | — | — | — | — | |
| SimVG-DBVisual Encoder=ViT-B/32, Time (ms)=522024.09 | 82.04 | — | — | — | — | — | |
| Dyn.MDETRVisual Encoder=ViT-B/162024.09 | 81.89 | — | — | — | — | — | |
| SimVG-TBVisual Encoder=ViT-B/32, Time (ms)=442024.09 | 81.59 | — | — | — | — | — | |
| SeqTRVisual Encoder=DN53, Time (ms)=502024.09 | 81.23 | — | — | — | — | — | |
| TransCPVisual Encoder=RN50, Time (ms)=74*2024.09 | 80.04 | — | — | — | — | — | |
| VLTVGVisual Encoder=RN50, Time (ms)=79*2024.09 | 79.18 | — | — | — | — | — | |
| TransVGVisual Encoder=RN101, Time (ms)=622024.09 | 79.1 | — | — | — | — | — | |
| Proposal upperboundFeature extractor=Other2015.11 | 77.9 | — | — | — | — | — | |
| Proposal upperboundFeature extractor=VGG-CLS2015.11 | 77.9 | — | — | — | — | — | |
| Proposal upperboundFeature extractor=VGG-DET2015.11 | 77.9 | — | — | — | — | — | |
| VGTRVisual Encoder=RN502024.09 | 75.44 | — | — | — | — | — | |
| ReSCLVisual Encoder=DN53, Time (ms)=362024.09 | 69.28 | — | — | — | — | — | |
| FAOAVisual Encoder=DN53, Time (ms)=392024.09 | 68.71 | — | — | — | — | — | |
| GroundeRSupervision level=Semi-supervised, Annotation percentage=100.0%, Feature extractor=VGG-DET2015.11 | 48.38 | — | — | — | — | — | |
| GroundeRSupervision level=Supervised, Feature extractor=VGG-DET2015.11 | 47.81 | — | — | — | — | — | |
| GroundeRSupervision level=Semi-supervised, Annotation percentage=50.0%, Feature extractor=VGG-DET2015.11 | 46.65 | — | — | — | — | — | |
| GroundeRSupervision level=Semi-supervised, Annotation percentage=25.0%, Feature extractor=VGG-DET2015.11 | 45.32 | — | — | — | — | — | |
| GroundeRSupervision level=Semi-supervised, Annotation percentage=12.5%, Feature extractor=VGG-DET2015.11 | 44.96 | — | — | — | — | — | |
| GroundeRSupervision level=Semi-supervised, Annotation percentage=6.25%, Feature extractor=VGG-DET2015.11 | 44.02 | — | — | — | — | — | |
| DSPESupervision level=Supervised, Feature extractor=VGG-DET2015.11 | 43.89 | — | — | — | — | — | |
| GroundeRSupervision level=Semi-supervised, Annotation percentage=100.0%, Feature extractor=VGG-CLS2015.11 | 42.43 | — | — | — | — | — | |
| GroundeRSupervision level=Semi-supervised, Annotation percentage=3.12%, Feature extractor=VGG-DET2015.11 | 42.32 | — | — | — | — | — | |
| GroundeRSupervision level=Supervised, Feature extractor=VGG-CLS2015.11 | 41.56 | — | — | — | — | — | |
| GroundeRSupervision level=Semi-supervised, Annotation percentage=50.0%, Feature extractor=VGG-CLS2015.11 | 40.72 | — | — | — | — | — | |
| GroundeRSupervision level=Semi-supervised, Annotation percentage=25.0%, Feature extractor=VGG-CLS2015.11 | 39.31 | — | — | — | — | — | |
| GroundeRSupervision level=Semi-supervised, Annotation percentage=12.5%, Feature extractor=VGG-CLS2015.11 | 38.67 | — | — | — | — | — | |
| GroundeRSupervision level=Semi-supervised, Annotation percentage=6.25%, Feature extractor=VGG-CLS2015.11 | 37.1 | — | — | — | — | — | |
| GroundeRSupervision level=Semi-supervised, Annotation percentage=3.12%, Feature extractor=VGG-CLS2015.11 | 33.02 | — | — | — | — | — | |
| GroundeRSupervision level=Unsupervised, Feature extractor=VGG-DET2015.11 | 28.94 | — | — | — | — | — | |
| SCRCSupervision level=Supervised, Feature extractor=VGG-CLS2015.11 | 27.8 | — | — | — | — | — | |
| CCASupervision level=Supervised, Feature extractor=VGG-CLS2015.11 | 27.42 | — | — | — | — | — | |
| GroundeRSupervision level=Unsupervised, Feature extractor=VGG-CLS2015.11 | 24.66 | — | — | — | — | — | |
| Deep FragmentsSupervision level=Unsupervised, Feature extractor=Other2015.11 | 21.78 | — | — | — | — | — | |
| BANDetector=Bottom-Up [2]2018.05 | — | 69.69 | 84.22 | 86.35 | — | 87.45 | |
| BAN2019.08 | — | 69.69 | 84.22 | 86.35 | — | 87.45 | |
| CCA baselineBackbone=Fast-RCNN, Proposals=100 EdgeBox2015.11 | — | 40.11 | 61.52 | 67.17 | 41.96 | — | |
| Deep Structure-Preserving Embedding (a)negative mining=false, λ1=2, λ2=0, λ3=0, Backbone=Fast-RCNN, Proposals=100 EdgeBox2015.11 | — | 35.83 | 60.51 | 66.7 | 40.5 | — | |
| Deep Structure-Preserving Embedding (a) Fine-tunednegative mining=true, Fine-tuning epochs=5, λ1=2, λ2=0, λ3=0, Backbone=Fast-RCNN, Proposals=100 EdgeBox2015.11 | — | 41.77 | 63.01 | 68.27 | 46.55 | — | |
| Deep Structure-Preserving Embedding (b)negative mining=false, λ1=2, λ2=0, λ3=0.1, Backbone=Fast-RCNN, Proposals=100 EdgeBox2015.11 | — | 36.59 | 60.44 | 66.92 | 40.85 | — | |
| Deep Structure-Preserving Embedding (b) Fine-tunednegative mining=true, Fine-tuning epochs=5, λ1=2, λ2=0, λ3=0.1, Backbone=Fast-RCNN, Proposals=100 EdgeBox2015.11 | — | 43.77 | 64.22 | 68.84 | 47.38 | — | |
| Deep Structure-Preserving Embedding (c)negative mining=false, λ1=2, λ2=0.1, λ3=0, Backbone=Fast-RCNN, Proposals=100 EdgeBox2015.11 | — | 36.74 | 60.35 | 66.73 | 41.22 | — | |
| Deep Structure-Preserving Embedding (c) Fine-tunednegative mining=true, Fine-tuning epochs=5, λ1=2, λ2=0.1, λ3=0, Backbone=Fast-RCNN, Proposals=100 EdgeBox2015.11 | — | 42.88 | 63.41 | 68.47 | 46.78 | — | |
| Deep Structure-Preserving Embedding (d)negative mining=false, λ1=2, λ2=0.1, λ3=0.1, Backbone=Fast-RCNN, Proposals=100 EdgeBox2015.11 | — | 36.72 | 61.14 | 67.21 | 41.13 | — | |
| Deep Structure-Preserving Embedding (d) Fine-tunednegative mining=true, Fine-tuning epochs=5, λ1=2, λ2=0.1, λ3=0.1, Backbone=Fast-RCNN, Proposals=100 EdgeBox2015.11 | — | 43.89 | 64.46 | 68.66 | 47.72 | — | |
| Fukui et al.Detector=Fast RCNN [7]2018.05 | — | 48.69 | — | — | — | — | |
| Hinami and SatohDetector=Query-Adaptive RCNN [10]2018.05 | — | 65.21 | — | — | — | — | |
| Hu et al.Detector=Edge Boxes [42]2018.05 | — | 27.8 | — | 62.9 | — | 76.9 | |
| Plummer et al.Detector=Fast RCNN [7], Additional Features=box size and color2018.05 | — | 50.89 | 71.09 | 75.73 | — | 85.12 | |
| Rohrbach et al.Detector=Fast RCNN [7]2018.05 | — | 42.43 | — | — | — | 77.9 | |
| Rohrbach et al.Detector=Fast RCNN [7]2018.05 | — | 48.38 | — | — | — | 77.9 | |
| VisualBERTEarly Fusion=false2019.08 | — | — | — | — | — | 87.45 | |
| VisualBERTCOCO Pre-training=false2019.08 | — | — | — | — | — | 87.45 | |
| VisualBERTFull Model=true2019.08 | — | 71.33 | 84.98 | 86.51 | — | — | |
| Wang et al.Detector=Fast RCNN [7]2018.05 | — | 42.08 | — | — | — | 76.91 | |
| Wang et al.Detector=Fast RCNN [7]2018.05 | — | 43.89 | 64.46 | 68.66 | — | 76.91 | |
| Yeh et al.Detector=YOLOv2 [24], Additional Features=semantic segmentation, object detection, pose-estimation2018.05 | — | 53.97 | — | — | — | — | |
| Zhang et al.Detector=MCG [3]2018.05 | — | 28.5 | 52.7 | 61.3 | — | — |