Image Classification on MetaShift Animals
76.6Worst Case AccuracyDFR
Evaluation Results
| Method | Links | |
|---|---|---|
| DFRBackbone (attn x emb)=ViT-G (1.1B), Evaluation Protocol=Deep Feature Reweighting (requires group labels)2026.06 | 76.6 | |
| LR (full image)Backbone (attn x emb)=ViT-G (1.1B), Evaluation Protocol=Linear probe on full image (no crops)2026.06 | 75.9 | |
| DFRBackbone (attn x emb)=ViT-B (86M), Evaluation Protocol=Deep Feature Reweighting (requires group labels)2026.06 | 73.5 | |
| DFR† + A²Backbone (attn x emb)=ViT-S × ViT-S, Evaluation Protocol=Deep Feature Reweighting (requires group labels), uses group labels=true2026.06 | 72.9 | |
| DFRBackbone (attn x emb)=ViT-S (21M), Evaluation Protocol=Deep Feature Reweighting (requires group labels)2026.06 | 72 | |
| A²_LR cross-modelBackbone (attn x emb)=ViT-S × ViT-G, Evaluation Protocol=Attending on Attention (ours)2026.06 | 70.8 | |
| iFAM (K=4)Backbone (attn x emb)=ViT-B (86M), Evaluation Protocol=End-to-end attention learning2026.06 | 68.8 | |
| A²_LRBackbone (attn x emb)=ViT-S × ViT-S, Evaluation Protocol=Attending on Attention (ours)2026.06 | 67 | |
| iFAM (K=8)Backbone (attn x emb)=ViT-B (86M), Evaluation Protocol=End-to-end attention learning2026.06 | 63.5 | |
| LR (full image)Backbone (attn x emb)=ViT-S (21M), Evaluation Protocol=Linear probe on full image (no crops)2026.06 | 63.2 | |
| OpenCLIP ZS (full image)Backbone (attn x emb)=ViT-L-14, Evaluation Protocol=Zero-shot CLIP-based references (no crops)2026.06 | 58.8 | |
| A²_ZSBackbone (attn x emb)=ViT-S × OpenCLIP ViT-L-14, Evaluation Protocol=Attending on Attention (ours)2026.06 | 54.7 | |
| TTRBackbone (attn x emb)=OpenCLIP ViT-L-14, Evaluation Protocol=Zero-shot CLIP-based references (no crops)2026.06 | 49.1 |