Fine-grained Image Classification on UCF101
68.52AccuracyFair Context Learning (FCL)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Fair Context Learning (FCL)mode=Test-Time Adaptation2026.02 | 68.52 | — | |
| TPTmode=Test-Time Adaptation2026.02 | 68.41 | — | |
| TPTBackbone=CLIP-ViT-B/16, Epsilon (ε)=4.02025.04 | 67.9 | 0 | |
| MTAmode=Test-Time Adaptation2026.02 | 67.8 | — | |
| MTABackbone=CLIP-ViT-B/16, Epsilon (ε)=4.02025.04 | 67.5 | 27.5 | |
| MTABackbone=CLIP-ViT-B/16, Perturbation bound (epsilon)=4.02025.12 | 67.5 | 27.5 | |
| TPTBackbone=CLIP-ViT-B/16, Augmented views=15, Perturbation (epsilon)=4.02026.06 | 67.5 | 0 | |
| TPSmode=Test-Time Adaptation2026.02 | 67.46 | — | |
| R-TPTmode=Test-Time Adaptation2026.02 | 67.35 | — | |
| TTLmode=Test-Time Adaptation2026.02 | 67.3 | — | |
| R-TPTBackbone=CLIP-ViT-B/16, Epsilon (ε)=4.02025.04 | 67.2 | 43.2 | |
| R-TPTBackbone=CLIP-ViT-B/16, Perturbation bound (epsilon)=4.02025.12 | 67.2 | 43.2 | |
| RLCFmode=Test-Time Adaptation2026.02 | 66.96 | — | |
| ZEROmode=Test-Time Adaptation2026.02 | 66.53 | — | |
| SS-TPTBackbone=CLIP-ViT-B/16, Augmented views=15, Perturbation (epsilon)=4.02026.06 | 66 | 43.1 | |
| TAPTBackbone=CLIP-ViT-B/16, Augmented views=15, Perturbation (epsilon)=4.02026.06 | 65.9 | 7.1 | |
| TTCBackbone=CLIP-ViT-B/16, Perturbation bound (epsilon)=4.02025.12 | 65.8 | 1.6 | |
| C-TPTBackbone=CLIP-ViT-B/16, Epsilon (ε)=4.02025.04 | 65.5 | 0 | |
| R-TPTBackbone=CLIP-ViT-B/16, Augmented views=15, Perturbation (epsilon)=4.02026.06 | 65.5 | 40.7 | |
| CLIPmode=Zero-shot2026.02 | 65.24 | — | |
| CLIPBackbone=CLIP-ViT-B/16, Epsilon (ε)=4.02025.04 | 65.2 | 0 | |
| CLIPBackbone=CLIP-ViT-B/16, Perturbation bound (epsilon)=4.02025.12 | 65.2 | 0 | |
| CLIPBackbone=CLIP-ViT-B/16, Augmented views=15, Perturbation (epsilon)=4.02026.06 | 65.2 | 0 | |
| TTPBackbone=CLIP-ViT-B/16, Perturbation bound (epsilon)=4.02025.12 | 65 | 47.2 | |
| C-TPTmode=Test-Time Adaptation2026.02 | 64.97 | — | |
| TTCBackbone=CLIP-ViT-B/16, Augmented views=15, Perturbation (epsilon)=4.02026.06 | 64.5 | 1.8 | |
| DOCBackbone=CLIP-ViT-B/16, Augmented views=15, Perturbation (epsilon)=4.02026.06 | 64.1 | 0.8 | |
| MTABackbone=CLIP-ViT-B/32, Perturbation Bound (epsilon)=4.02025.12 | 63.3 | 39.1 | |
| EnsembleBackbone=CLIP-ViT-B/16, Epsilon (ε)=4.02025.04 | 63 | 30.6 | |
| EnsembleBackbone=CLIP-ViT-B/16, Perturbation bound (epsilon)=4.02025.12 | 63 | 30.6 | |
| EnsembleBackbone=CLIP-ViT-B/16, Augmented views=15, Perturbation (epsilon)=4.02026.06 | 63 | 19.5 | |
| R-TPTBackbone=CLIP-ViT-B/32, Perturbation Bound (epsilon)=4.02025.12 | 62.8 | 41 | |
| TTCBackbone=CLIP-ViT-B/32, Perturbation Bound (epsilon)=4.02025.12 | 62.6 | 6.1 | |
| CLIPBackbone=CLIP-ViT-B/32, Perturbation Bound (epsilon)=4.02025.12 | 61.6 | 0 | |
| TTPBackbone=CLIP-ViT-B/32, Perturbation Bound (epsilon)=4.02025.12 | 61.3 | 46.6 | |
| TPTBackbone=CLIP-ResNet50, Epsilon=1.02025.04 | 60.6 | 0.3 | |
| MTABackbone=CLIP-ResNet50, Epsilon=1.02025.04 | 60.6 | 31.3 | |
| C-TPTBackbone=CLIP-ResNet50, Epsilon=1.02025.04 | 60.1 | 0.1 | |
| R-TPTBackbone=CLIP-ResNet50, Epsilon=1.02025.04 | 59.7 | 50.9 | |
| CLIPBackbone=CLIP-ResNet50, Epsilon=1.02025.04 | 58.9 | 0 | |
| EnsembleBackbone=CLIP-ViT-B/32, Perturbation Bound (epsilon)=4.02025.12 | 54.9 | 36.9 | |
| EnsembleBackbone=CLIP-ResNet50, Epsilon=1.02025.04 | 53.9 | 43 | |
| TPTBackbone=CLIP-ViT-B/32, Pre-trained=TeCoA, Augmented views=15, Epsilon=4.02026.06 | 34.9 | 9.9 | |
| CLIPBackbone=CLIP-ViT-B/32, Pre-trained=TeCoA, Augmented views=15, Epsilon=4.02026.06 | 34.6 | 7.2 | |
| SS-TPTBackbone=CLIP-ViT-B/32, Pre-trained=TeCoA, Augmented views=15, Epsilon=4.02026.06 | 28.3 | 14.9 | |
| R-TPTBackbone=CLIP-ViT-B/32, Pre-trained=TeCoA, Augmented views=15, Epsilon=4.02026.06 | 28.2 | 14.5 | |
| EnsembleBackbone=CLIP-ViT-B/32, Pre-trained=TeCoA, Augmented views=15, Epsilon=4.02026.06 | 27 | 14.1 |