Image Classification on FACET
70.8Macro F1Baseline CLIP (ViT-L/14)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Baseline CLIP (ViT-L/14)Zero-shot=true, Backbone=ViT-L/142026.03 | 70.8 | 8.9 | 49.8 | |
| OursZero-shot=true, Backbone=ViT-L/14, Attribute-annotated data=Not Required, Training=No, Debiased Modalities=I&T2026.03 | 70.7 | 8.3 | 47.5 | |
| FairerCLIPZero-shot=true, Backbone=ViT-L/14, Attribute-annotated data=Required, Training=Yes, Debiased Modalities=I&T2026.03 | 69.8 | 9.2 | 50.1 | |
| RoboShotZero-shot=true, Backbone=ViT-L/14, Attribute-annotated data=Not Required, Training=No, Debiased Modalities=I&T2026.03 | 69.3 | 8.5 | 47.3 | |
| Orth-CaliZero-shot=true, Backbone=ViT-L/14, Attribute-annotated data=Not Required, Training=No, Debiased Modalities=T2026.03 | 69.3 | 9.2 | 50.4 | |
| SFIDZero-shot=true, Backbone=ViT-L/14, Attribute-annotated data=Required, Training=Yes, Debiased Modalities=I&T2026.03 | 69.2 | 9.3 | 49.8 | |
| PRISM-miniZero-shot=true, Backbone=ViT-L/14, Attribute-annotated data=Not Required, Training=Yes, Debiased Modalities=I&T2026.03 | 69.2 | 8.6 | 49.2 | |
| PRISMZero-shot=true, Backbone=ViT-L/14, Attribute-annotated data=Not Required, Training=Yes, Debiased Modalities=I&T2026.03 | 69 | 8.1 | 48.3 | |
| Orth-ProjZero-shot=true, Backbone=ViT-L/14, Attribute-annotated data=Not Required, Training=No, Debiased Modalities=T2026.03 | 68.6 | 9 | 50.2 | |
| Baseline BLIPProtocol=Zero-shot, Base Model=BLIP2026.03 | 68.4 | 8.8 | 48 | |
| Debias_VLMProtocol=Zero-shot, Base Model=BLIP, Modality=Image & Text, Technique=Training-free2026.03 | 68.4 | 7.3 | 44.7 | |
| FairerCLIPProtocol=Zero-shot, Base Model=BLIP, Modality=Image & Text, Technique=Training-based2026.03 | 61 | 8.9 | 46.5 | |
| SFIDProtocol=Zero-shot, Base Model=BLIP, Modality=Image & Text, Technique=Training-based2026.03 | 60.6 | 8.2 | 47.1 | |
| PRISMProtocol=Zero-shot, Base Model=BLIP, Modality=Image & Text, Technique=Training-based2026.03 | 60.3 | 6.8 | 46 | |
| RoboShotProtocol=Zero-shot, Base Model=BLIP, Modality=Image & Text, Technique=Training-based2026.03 | 60.3 | 7.6 | 44.4 | |
| PRISM-miniProtocol=Zero-shot, Base Model=BLIP, Modality=Image & Text, Technique=Training-free2026.03 | 60 | 7.4 | 46.1 | |
| Orth-CaliProtocol=Zero-shot, Base Model=BLIP, Modality=Text, Technique=Training-free2026.03 | 59.5 | 8 | 46 | |
| CLIPBackbone=ResNet50, Zero-shot=true2026.03 | 59.2 | 9.3 | 48.9 | |
| Debias_VLMBackbone=ResNet50, Modality=I&T, Zero-shot=true2026.03 | 58.6 | 8.5 | 45.8 | |
| Orth-ProjProtocol=Zero-shot, Base Model=BLIP, Modality=Text, Technique=Training-free2026.03 | 58.3 | 8.5 | 46.3 | |
| SFIDBackbone=ResNet50, Modality=I&T, Zero-shot=true2026.03 | 50.8 | 9.4 | 47.2 | |
| FairerCLIPBackbone=ResNet50, Modality=I&T, Zero-shot=true2026.03 | 50.7 | 8.9 | 47.2 | |
| RoboShotBackbone=ResNet50, Modality=I&T, Zero-shot=true2026.03 | 50.7 | 9 | 45.6 | |
| PRISMBackbone=ResNet50, Modality=I&T, Zero-shot=true2026.03 | 50.5 | 8.4 | 47.6 | |
| Orth-CaliBackbone=ResNet50, Modality=T, Zero-shot=true2026.03 | 49.6 | 8.9 | 47.6 | |
| PRISM-miniBackbone=ResNet50, Modality=I&T, Zero-shot=true2026.03 | 49.2 | 8.9 | 46.6 | |
| Orth-ProjBackbone=ResNet50, Modality=T, Zero-shot=true2026.03 | 47.8 | 8.8 | 47.9 |