Human-Object Interaction Detection on HICO-DET (UO)
45.28mAP (Full)DA-HOI
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| DA-HOIDetector=Grounding-DINO2026.02 | 45.28 | 50.15 | 44.31 | |
| DA-HOIDetector=Yolo-World2026.02 | 44.82 | 47.95 | 44.19 | |
| DA-HOIDetector=ResNet50 DETR2026.02 | 43.6 | 48.67 | 42.58 | |
| SL-HOIBackbone=DINOv3-ViT-L/16, Object detection pre-training=false2026.03 | 42.49 | 40.53 | 42.99 | |
| BC-HOIBackbone=ResNet50 + BLIP-2-ViT-G/14, Object detection pre-training=true2026.03 | 40.99 | 42.31 | 40.67 | |
| VRDiffBackbone=ResNet50 + CLIP-ViT-L/14, Object detection pre-training=true2026.03 | 40.45 | 38.92 | 40.83 | |
| GRASP-HOISetting=Zero-shot, Protocol=Open-vocabulary2025.12 | 37.69 | 39.98 | 38.15 | |
| LAINBackbone=ViT-L2025.05 | 37.6 | 40.78 | 36.96 | |
| CMMPBackbone=ResNet50 + CLIP-ViT-L/14, Object detection pre-training=true2026.03 | 37.13 | 35.98 | 37.42 | |
| CMMPSetting=Zero-shot, Protocol=Open-vocabulary2025.12 | 36.74 | 39.67 | 36.15 | |
| EZ-HOIBackbone=ResNet50 + CLIP-ViT-L/14, Object detection pre-training=true2026.03 | 36.73 | 34.24 | 37.35 | |
| EZ-HOISetting=Zero-shot, Protocol=Open-vocabulary2025.12 | 36.38 | 38.17 | 36.02 | |
| EZ-HOI2026.02 | 36.38 | 38.17 | 36.02 | |
| LAINBackbone=ViT-B2025.05 | 34.27 | 37.88 | 33.55 | |
| LAIN2026.02 | 34.27 | 37.88 | 33.55 | |
| HOLaBackbone=ResNet50 + CLIP-ViT-B/16, Object detection pre-training=true2026.03 | 34.19 | 30.61 | 35.08 | |
| BC-HOISetting=Zero-shot, Protocol=Open-vocabulary2025.12 | 34.18 | 19.94 | 37.03 | |
| BC-HOIFeatures=BLIP22026.02 | 34.18 | 19.94 | 37.03 | |
| CLIP4HOIBackbone=ResNet50 + CLIP-ViT-B/16, Object detection pre-training=true2026.03 | 34.08 | 28.47 | 35.48 | |
| BCOMBackbone=ResNet50 + CLIP-ViT-L/14, Object detection pre-training=true2026.03 | 33.74 | 28.52 | 35.04 | |
| HOIGenSetting=Zero-shot, Protocol=Open-vocabulary2025.12 | 33.48 | 36.35 | 32.9 | |
| LOGICHOIBackbone=ResNet50 + CLIP-ViT-B/32, Object detection pre-training=true2026.03 | 33.17 | 25.97 | 34.93 | |
| HOICLIPBackbone=ResNet50 + CLIP-ViT-B/32, Object detection pre-training=true2026.03 | 32.99 | 25.53 | 34.85 | |
| CLIP4HOIBackbone=ViT-B2025.05 | 32.58 | 31.79 | 32.73 | |
| CLIP4HOI2026.02 | 32.58 | 31.79 | 32.73 | |
| UniHOIBackbone=ResNet50 + BLIP-2-ViT-G/14, Object detection pre-training=true2026.03 | 32.27 | 28.68 | 33.16 | |
| CMMPBackbone=ViT-B2025.05 | 31.59 | 33.76 | 31.15 | |
| CMMP2026.02 | 31.59 | 33.76 | 31.15 | |
| UniHOI (BLIP2)Setting=Zero-shot, Protocol=Open-vocabulary2025.12 | 31.56 | 19.72 | 34.76 | |
| UniHOIFeatures=BLIP22026.02 | 31.56 | 19.72 | 34.76 | |
| DA-HOITraining-free=true2026.02 | 31.5 | — | — | |
| GEN-VLKTBackbone=ResNet50 + CLIP-ViT-B/32, Object detection pre-training=true2026.03 | 30.56 | 21.36 | 32.91 | |
| HOICLIPBackbone=ViT-B2025.05 | 28.53 | 16.2 | 30.99 | |
| HOICLIPSetting=Zero-shot, Protocol=Open-vocabulary2025.12 | 28.53 | 16.2 | 30.99 | |
| HOICLIP2026.02 | 28.53 | 16.2 | 30.99 | |
| LOGICHOIBackbone=ViT-B2025.05 | 28.23 | 15.67 | 30.42 | |
| SGC-NetBackbone=CLIP-ViT-B/16, Object detection pre-training=false2026.03 | 27.22 | 23.27 | 28.34 | |
| GEN-VLKTBackbone=ViT-B2025.05 | 25.63 | 10.51 | 28.92 | |
| GEN-VLKTSetting=Zero-shot, Protocol=Open-vocabulary2025.12 | 25.63 | 10.51 | 28.92 | |
| GEN-VLKT2026.02 | 25.63 | 10.51 | 28.92 | |
| ADA-CMTraining-free=true2026.02 | 25.19 | — | — | |
| CLIPBackbone=ViT-B2025.05 | 23.36 | 28.66 | 22.29 | |
| INP-CCBackbone=CLIP-ViT-B/16, Object detection pre-training=false2026.03 | 23.13 | 17.38 | 24.74 | |
| THIDBackbone=CLIP-ViT-B/16, Object detection pre-training=false2026.03 | 22.96 | 15.53 | 24.32 | |
| CMD-SEBackbone=CLIP-ViT-B/16, Object detection pre-training=false2026.03 | 22.35 | 16.7 | 23.95 | |
| ATLBackbone=ViT-B2025.05 | 20.47 | 15.11 | 21.54 | |
| FCLBackbone=ViT-B2025.05 | 19.87 | 15.54 | 20.74 |