Referring Image Segmentation on G-Ref UMD (test)
66.22IoUVLT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| VLTVisual Backbone=Swin-B, Textual Encoder=BERT2022.10 | 66.22 | 78.03 | |
| MaILResolution=Fixed aspect ratio2021.11 | 62.87 | — | |
| MaILVisual Backbone=ViLT, Textual Encoder=BERT, large-scale pre-training=true2022.10 | 62.87 | — | |
| LAVTVisual Backbone=Swin-B, Textual Encoder=BERT2022.10 | 62.09 | — | |
| MaILResolution=416x4162021.11 | 61.39 | — | |
| CRISVisual Backbone=CLIP-R101, Textual Encoder=CLIP, large-scale pre-training=true2022.10 | 60.36 | — | |
| VLTVisual Backbone=Darknet53, Textual Encoder=bi-GRU2022.10 | 57.73 | 60.96 | |
| VLT2021.11 | 56.65 | — | |
| LTS2021.11 | 54.25 | — | |
| LTSVisual Backbone=Darknet53, Textual Encoder=bi-GRU2022.10 | 54.25 | — | |
| ISFPVisual Backbone=Darknet53, Textual Encoder=Bi-GRU2022.10 | 53 | — | |
| CGAN2021.11 | 51.69 | — | |
| CGANVisual Backbone=DeepLab-R101, Textual Encoder=bi-GRU2022.10 | 51.69 | — | |
| MCN2021.11 | 49.4 | — | |
| MCNVisual Backbone=Darknet53, Textual Encoder=bi-GRU2022.10 | 49.4 | — | |
| MAttNet2021.11 | 48.61 | — | |
| MAttNetVisual Backbone=MaskRCNN-R101, Textual Encoder=bi-LSTM2022.10 | 48.61 | — | |
| CAC2021.11 | 46.95 | — | |
| CACVisual Backbone=ResNet101, Textual Encoder=bi-LSTM2022.10 | 46.95 | — |