Referring Expression Segmentation on PhraseCut
53.7mIoUMDETR
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| MDETR2021.12 | 53.7 | — | — | |
| CLIPSeg (PC, D = 128)threshold (t)=0.3, decoder_dimension (D)=128, training_data=PhraseCut (text labels only)2021.12 | 48.2 | 56.5 | 78.2 | |
| CLIPSeg (PC)threshold (t)=0.3, training_data=PhraseCut (text labels only)2021.12 | 46.1 | 56.2 | 78.2 | |
| CLIPSeg (PC+)threshold (t)=0.3, training_data=PhraseCut (text and visual samples)2021.12 | 43.4 | 54.7 | 76.7 | |
| HulaNet2021.12 | 41.3 | 50.8 | — | |
| Mask-RCNN top2021.12 | 39.4 | 47.4 | — | |
| ViTSeg (PC)threshold (t)=0.3, backbone=ImageNet-trained ViT, training_data=PhraseCut (text labels only)2021.12 | 38.9 | 51.2 | 74.4 | |
| CLIP-Deconvthreshold (t)=0.32021.12 | 37.7 | 49.5 | 71.2 | |
| ViTSeg (PC+)threshold (t)=0.1, backbone=ImageNet-trained ViT, training_data=PhraseCut (text and visual samples)2021.12 | 28.4 | 35.4 | 58.3 | |
| RMI2021.12 | 21.1 | 42.5 | — |