Image Classification on MetaShift Cat vs. Dog
88.8WGAOpenCLIP ZS (full image)
Evaluation Results
| Method | Links | |
|---|---|---|
| OpenCLIP ZS (full image)Backbone (attn x emb)=ViT-L-14, Evaluation Protocol=Zero-shot CLIP-based references (no crops)2026.06 | 88.8 | |
| A²_ZSBackbone (attn x emb)=ViT-S × OpenCLIP ViT-L-14, Evaluation Protocol=Attending on Attention (ours)2026.06 | 88.8 | |
| TTRBackbone (attn x emb)=OpenCLIP ViT-L-14, Evaluation Protocol=Zero-shot CLIP-based references (no crops)2026.06 | 82.7 | |
| LR (full image)Backbone (attn x emb)=ViT-G (1.1B), Evaluation Protocol=Linear probe on full image (no crops)2026.06 | 76.8 | |
| A²_LR cross-modelBackbone (attn x emb)=ViT-S × ViT-G, Evaluation Protocol=Attending on Attention (ours)2026.06 | 76.6 | |
| DFRBackbone (attn x emb)=ViT-G (1.1B), Evaluation Protocol=Deep Feature Reweighting (requires group labels)2026.06 | 74.2 | |
| iFAM (K=8)Backbone (attn x emb)=ViT-B (86M), Evaluation Protocol=End-to-end attention learning2026.06 | 73.5 | |
| iFAM (K=4)Backbone (attn x emb)=ViT-B (86M), Evaluation Protocol=End-to-end attention learning2026.06 | 72.1 | |
| A²_LRBackbone (attn x emb)=ViT-S × ViT-S, Evaluation Protocol=Attending on Attention (ours)2026.06 | 68.3 | |
| DFR† + A²Backbone (attn x emb)=ViT-S × ViT-S, Evaluation Protocol=Deep Feature Reweighting (requires group labels), uses group labels=true2026.06 | 67.7 | |
| DFRBackbone (attn x emb)=ViT-B (86M), Evaluation Protocol=Deep Feature Reweighting (requires group labels)2026.06 | 65.9 | |
| LR (full image)Backbone (attn x emb)=ViT-S (21M), Evaluation Protocol=Linear probe on full image (no crops)2026.06 | 56 | |
| DFRBackbone (attn x emb)=ViT-S (21M), Evaluation Protocol=Deep Feature Reweighting (requires group labels)2026.06 | 54.5 |