Image Classification on Pets
99.75AccuracyViT-22B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| ViT-22BEvaluation Protocol=Finetuning, Resolution=224, Seeds=32023.02 | 99.75 | — | — | — | |
| ViT-22BEvaluation Protocol=Linear, Resolution=224, Seeds=32023.02 | 98.15 | — | — | — | |
| ViT-22BEvaluation Protocol=H2T (Head2Toe), Resolution=224, Seeds=32023.02 | 97.46 | — | — | — | |
| StatABackbone=ViT-L/14, Keff=Very Low2025.01 | 97.1 | — | — | — | |
| DFN5B-CLIP-H/14+Zero-shot=true2024.02 | 96.8 | — | — | — | |
| HIVEFoundation Model=SigLIP2026.03 | 96.78 | — | — | — | |
| BaseFoundation Model=SigLIP2026.03 | 96.76 | — | — | — | |
| SAFoundation Model=SigLIP2026.03 | 96.76 | — | — | — | |
| DFN5B-CLIP-H/14Zero-shot=true2024.02 | 96.5 | — | — | — | |
| EVA-CLIP-8BZero-shot=true2024.02 | 96.4 | — | — | — | |
| InternVL-CZero-shot=true2024.02 | 96.3 | — | — | — | |
| StatABackbone=ViT-L/14, Keff=Low2025.01 | 96.3 | — | — | — | |
| EVA-CLIP-18BZero-shot=true2024.02 | 96.1 | — | — | — | |
| BaseFoundation Model=CLIP2026.03 | 96.02 | — | — | — | |
| EVA-02-CLIP-E/14+Zero-shot=true2024.02 | 96 | — | — | — | |
| SAFoundation Model=CLIP2026.03 | 95.98 | — | — | — | |
| HIVEFoundation Model=CLIP2026.03 | 95.92 | — | — | — | |
| CLIP-1+DINO(v2+v3)Setting=Inductive, Test-Time Adaptation (TTA)=false2025.06 | 95.4 | — | — | — | |
| OpenCLIP-G/14Zero-shot=true2024.02 | 95.3 | — | — | — | |
| StatABackbone=ResNet-101, Keff=Very Low2025.01 | 95.2 | — | — | — | |
| GalLoP+OursShots=42026.05 | 95.2 | — | — | — | |
| CLIP-1+DINOv3Setting=Inductive, Test-Time Adaptation (TTA)=false2025.06 | 95.1 | — | — | — | |
| GPUASamples per class=full training set2026.06 | 95 | — | — | — | |
| EVA-01-CLIP-g/14+Zero-shot=true2024.02 | 94.9 | — | — | — | |
| StatABackbone=ViT-L/14, K_eff=Medium2025.01 | 94.8 | — | — | — | |
| PromptSRC+OursShots=42026.05 | 94.7 | — | — | — | |
| StatAEncoder=ViT-L/14, K_eff=All2025.01 | 94.6 | — | — | — | |
| GPUA*Samples per class=162026.06 | 94.5 | — | — | — | |
| StatAScenario=High, Backbone=ResNet-101, Batch Size=1282025.01 | 94.4 | — | — | — | |
| StatABackbone=ResNet-101, Batch size=128, Scenario=High2025.01 | 94.4 | — | — | — | |
| CLIP-1+DINOv2Setting=Inductive, Test-Time Adaptation (TTA)=false2025.06 | 94.4 | — | — | — | |
| StatABackbone=ViT-B/32, Keff=Very Low2025.01 | 94.3 | — | — | — | |
| StatABackbone=ViT-B/32, Batch Size=64, Evaluation Protocol=Batch test-time adaptation, K_eff=Very Low2025.01 | 94.3 | — | — | — | |
| VeRAParameter Count=0.10M2025.12 | 94.23 | — | — | — | |
| EVA-01-CLIP-g/14Zero-shot=true2024.02 | 94.2 | — | — | — | |
| StatAScenario=Separate, Backbone=ResNet-101, Batch Size=1282025.01 | 94.2 | — | — | — | |
| StatABackbone=ResNet-101, Batch size=128, Scenario=Separate2025.01 | 94.2 | — | — | — | |
| COSMIC*Setting=Inductive, Test-Time Adaptation (TTA)=true2025.06 | 94.2 | — | — | — | |
| COSMICInference-time adaptation=true2026.06 | 94.2 | — | — | — | |
| StatAScenario=Medium, Backbone=ResNet-101, Batch Size=1282025.01 | 94.1 | — | — | — | |
| StatABackbone=ResNet-101, Batch size=128, Scenario=Medium2025.01 | 94.1 | — | — | — | |
| GalLoPShots=42026.05 | 94.1 | — | — | — | |
| LoRA+Parameter Count=0.67M2025.12 | 94.07 | — | — | — | |
| LoRAParameter Count=0.67M2025.12 | 94.06 | — | — | — | |
| AdaLoRAParameter Count=0.67M2025.12 | 93.91 | — | — | — | |
| StatABackbone=ViT-L/14, Keff=Medium2025.01 | 93.9 | — | — | — | |
| Partial-LoRAParameter Count=0.14M2025.12 | 93.86 | — | — | — | |
| TransCLIP-FSShots=1, Backbone=ViT-B/162024.06 | 93.8 | — | — | — | |
| MLCDPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot classification2024.07 | 93.7 | — | — | — | |
| PromptSRCShots=42026.05 | 93.7 | — | — | — | |
| CLIPPre-training Data=WIT-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot classification2024.07 | 93.6 | — | — | — | |
| PLOT++Shot=162026.05 | 93.6 | — | — | — | |
| PLOTShots=42026.05 | 93.6 | — | — | — | |
| DoRAParameter Count=0.77M2025.12 | 93.55 | — | — | — | |
| CLIP (reported)Pre-training Data=WIT-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot classification2024.07 | 93.5 | — | — | — | |
| CLIPBackbone=ViT-L/142025.01 | 93.5 | — | — | — | |
| CLIPEncoder=ViT-L/14, K_eff=All2025.01 | 93.5 | — | — | — | |
| ProDAShots=42026.05 | 93.5 | — | — | — | |
| Partial-AdaLoRAParameter Count=0.12M2025.12 | 93.38 | — | — | — | |
| CoCoOpShots=42026.05 | 93.3 | — | — | — | |
| InMaPBackbone=ViT-B/16, Evaluation Protocol=Transductive Zero-Shot2024.04 | 93.2 | — | — | — | |
| CoCoOpShot=162026.05 | 93.2 | — | — | — | |
| KgCoOpShot=162026.05 | 93.2 | — | — | — | |
| TransCLIPBatch Size=1000, Backbone=ViT-B/16, Keff (Effective classes per task)=Medium (5-25)2025.01 | 93.1 | — | — | — | |
| StatABackbone=ResNet-50, Keff=Very Low2025.01 | 93.1 | — | — | — | |
| StatABatch Size=1000, Backbone=ViT-B/16, Keff=Medium (5-25), Protocol=Batch test-time adaptation2025.01 | 93.1 | — | — | — | |
| +DP-FMShot=12026.05 | 93.1 | — | — | — | |
| StatABackbone=ResNet-101, Keff=Low2025.01 | 92.9 | — | — | — | |
| InMaP + ZLaPBackbone=ViT-B/16, Evaluation Protocol=Transductive Zero-Shot2024.04 | 92.8 | — | — | -0.4 | |
| ProGradShot=162026.05 | 92.8 | — | — | — | |
| +DP-FMShot=162026.05 | 92.8 | — | — | — | |
| MaPLeShots=42026.05 | 92.8 | — | — | — | |
| CoCoOpShot=42026.05 | 92.7 | — | — | — | |
| PLOT++Shot=42026.05 | 92.7 | — | — | — | |
| FLIPPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot classification2024.07 | 92.6 | — | — | — | |
| KgCoOpShot=42026.05 | 92.6 | — | — | — | |
| TIP-AdapterShot=162026.05 | 92.6 | — | — | — | |
| StatABackbone=ViT-B/32, Keff=Low2025.01 | 92.5 | — | — | — | |
| StatABackbone=ViT-B/32, Batch Size=64, Evaluation Protocol=Batch test-time adaptation, K_eff=Low2025.01 | 92.5 | — | — | — | |
| CoOpShot=42026.05 | 92.5 | — | — | — | |
| SupervisedProtocol=Many-Shot classification, Pre-training=ImageNet2021.12 | 92.42 | — | — | — | |
| TransCLIP-FSShots=16, Backbone=ViT-B/162024.06 | 92.4 | — | — | — | |
| LoCoOpShots=42026.05 | 92.4 | — | — | — | |
| TCP+OursShots=42026.05 | 92.4 | — | — | — | |
| StatA2026.06 | 92.4 | — | — | — | |
| CLIP-AdapterShot=162026.05 | 92.3 | — | — | — | |
| +DP-FMShot=42026.05 | 92.2 | — | — | — | |
| KgCoOpShot=12026.05 | 92.1 | — | — | — | |
| ProGradShot=42026.05 | 92.1 | — | — | — | |
| DMN*Setting=Inductive, Test-Time Adaptation (TTA)=true2025.06 | 92 | — | — | — | |
| DMN2026.06 | 92 | — | — | — | |
| ZLaPBatch Size=1000, Backbone=ViT-B/16, Keff (Effective classes per task)=Very High (50-100)2025.01 | 91.9 | — | — | — | |
| TransCLIPBatch Size=1000, Backbone=ViT-B/16, Keff=Very High (50-100), Protocol=Batch test-time adaptation2025.01 | 91.9 | — | — | — | |
| CoCoOpShot=12026.05 | 91.9 | — | — | — | |
| PLOT++Shot=12026.05 | 91.9 | — | — | — | |
| CLIP-LoRAShot=12026.05 | 91.9 | — | — | — | |
| TIP-AdapterShot=42026.05 | 91.9 | — | — | — | |
| CoOpShots=42026.05 | 91.9 | — | — | — | |
| TCP*Shots=42026.05 | 91.8 | — | — | — | |
| TransCLIP-FSShots=4, Backbone=ViT-B/162024.06 | 91.6 | — | — | — |