Action Prediction on ProcTHOR multi-object
94IID AccuracyCDE (ViT-CLIP)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| CDE (ViT-CLIP)Encoder Backbone=ViT-CLIP2025.08 | 94 | 48 | 0.46 | |
| CDE (ViT-DINO)Encoder Backbone=ViT-DINO2025.08 | 92 | 45 | 0.47 | |
| CDE (ViT-MAE)Encoder Backbone=ViT-MAE2025.08 | 91 | 30 | 0.61 | |
| Oracle-maskSupervision Type=Ground truth mask2025.08 | 90 | 42 | 0.48 | |
| ResNetEncoder Backbone=ResNet2025.08 | 83 | 30 | 0.53 | |
| Slot-matchFeature Aggregation Strategy=Slot-match2025.08 | 66 | 21 | 0.45 | |
| Slot-denseFeature Aggregation Strategy=Slot-dense2025.08 | 51 | 19 | 0.32 | |
| Slot-avgFeature Aggregation Strategy=Slot-avg2025.08 | 49 | 15 | 0.34 |