Scene Graph Detection on Visual Genome (test)
35.8Recall@100SVRP
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| SVRPModel Input=Vision + Language2023.11 | 35.8 | — | 31.8 | — | 10.5 | 12.8 | |
| MOTIFModel Input=Vision2023.11 | 35.1 | 21.7 | 31 | — | 6.7 | 7.7 | |
| GPSNetModel Input=Vision2023.11 | 35 | 22.3 | 30.3 | — | 5.9 | 7.1 | |
| VCTreeModel Input=Vision2023.11 | 34.6 | 22 | 30.2 | — | 6.7 | 8 | |
| TransformerModel Input=Vision2023.11 | 34.3 | — | 30 | — | 7.4 | 8.8 | |
| VLPromptModel Input=Vision + Language2023.11 | 31.6 | 24 | 28.7 | 10.7 | 17.5 | 20.2 | |
| GPSNet + IETransModel Input=Vision2023.11 | 28.1 | — | 25.9 | — | 14.6 | 16.5 | |
| GPSNet + HiLoModel Input=Vision2023.11 | 27.9 | — | 25.6 | — | 15.8 | 18 | |
| GPS-Netbackbone=VGG-16, detector=Faster R-CNN2020.03 | 9.8 | — | — | — | — | — | |
| VCTREE-HLbackbone=VGG-16, detector=Faster R-CNN2020.03 | 8 | — | — | — | — | — | |
| KERNbackbone=VGG-16, detector=Faster R-CNN2020.03 | 7.3 | — | — | — | — | — | |
| FREQbackbone=VGG-16, detector=Faster R-CNN2020.03 | 7.1 | — | — | — | — | — | |
| MOTIFSbackbone=VGG-16, detector=Faster R-CNN2020.03 | 6.6 | — | — | — | — | — | |
| IMPbackbone=VGG-16, detector=Faster R-CNN2020.03 | 4.8 | — | — | — | — | — |