Video Human-Object Interaction Detection on VidHOI (test)
26.92Full Interaction APTUTOR
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| TUTORbackbone=ResNet-50, P (human poses)=false2022.06 | 26.92 | 23.49 | 37.12 | 32.21 | 21.28 | |
| TimeSformer w/ decoderbackbone=TimeSformer, P (human poses)=false2022.06 | 23.17 | 21.79 | 34.57 | 27.84 | 18.9 | |
| QPICbackbone=SlowFast, P (human poses)=false2022.06 | 22.92 | 21.64 | 33.43 | 28.41 | 13.47 | |
| HOTRbackbone=SlowFast, P (human poses)=false2022.06 | 22.84 | 21.15 | 32.86 | 27.12 | 13.29 | |
| QPICbackbone=ResNet-50, P (human poses)=false, image-based=true2022.06 | 21.4 | 20.56 | 32.9 | 28.87 | 9.74 | |
| HOTRbackbone=ResNet-50, P (human poses)=false, image-based=true2022.06 | 21.14 | 19.83 | 30.75 | 28.36 | 9.81 | |
| STIGPNbackbone=ResNet-50, P (human poses)=false2022.06 | 19.39 | 18.22 | 28.13 | 26.58 | 18.46 | |
| GPNNbackbone=ResNet-101, P (human poses)=false2022.06 | 18.47 | 16.41 | 24.5 | 26.41 | 16.06 | |
| ST-HOIbackbone=SlowFast, P (human poses)=false2022.06 | 17.6 | 17.3 | 27.2 | 25 | 14.4 | |
| PMFbackbone=SlowFast, P (human poses)=true2022.06 | 16.31 | 14.28 | 23.86 | 21.77 | 8.42 |