Image Classification on Caltech
99.93AccuracyMiPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MiPOBase Model=SD 2.0, Zero-shot=true2026.07 | 99.93 | — | |
| SD 2.0Zero-shot=true2026.07 | 99.65 | — | |
| MiPOBase Model=SD 1.5, Zero-shot=true2026.07 | 99.51 | — | |
| GPUA*Samples per class=162026.06 | 98.1 | — | |
| SD 1.5Zero-shot=true2026.07 | 97.81 | — | |
| IsoCLIPIntra-modal=✓, Classifier=NCM, Backbone=ViT-B/16-open2026.03 | 97.2 | — | |
| ProDiaLModel Size=Vim-small, Parameters=1.32M, Projection Type=Both-Proj2024.11 | 97.16 | — | |
| LoRAModel Size=Vim-small, Parameters=0.88M, Projection Type=In-Proj2024.11 | 97.16 | — | |
| DoRAModel Size=Vim-small, Parameters=0.91M, Projection Type=In-Proj2024.11 | 97.16 | — | |
| ProDiaLModel Size=Vim-small, Parameters=0.82M, Projection Type=In-Proj2024.11 | 97.09 | — | |
| Image-TextIntra-modal=✗, Classifier=Zero-Shot, Backbone=ViT-B/16-open2026.03 | 96.9 | — | |
| LoRAModel Size=Vim-small, Parameters=1.22M, Projection Type=Both-Proj2024.11 | 96.85 | — | |
| DoRAModel Size=Vim-small, Parameters=1.27M, Projection Type=Both-Proj2024.11 | 96.85 | — | |
| ProDiaLModel Size=Vim-small, Parameters=0.51M, Projection Type=Out-Proj2024.11 | 96.85 | — | |
| COSMICInference-time adaptation=true2026.06 | 96.8 | — | |
| DoRAModel Size=Vim-small, Parameters=0.59M, Projection Type=Out-Proj2024.11 | 96.78 | — | |
| LoRAModel Size=Vim-small, Parameters=0.48M, Projection Type=Out-Proj2024.11 | 96.7 | — | |
| Image-ImageIntra-modal=✓, Classifier=NCM, Backbone=ViT-B/16-open2026.03 | 96.7 | — | |
| BitFitModel Size=Vim-small, Parameters=0.11M, Projection Type=Base2024.11 | 96.62 | — | |
| +DP-FMShot=162026.05 | 96.6 | — | |
| StrongModel Size=Vim-small, Parameters=1.85M, Projection Type=Base2024.11 | 96.47 | — | |
| ProDiaLProjection Type=Both-Proj, Number of Parameters=0.65M2024.11 | 96.24 | — | |
| CLIP-LoRAShot=162026.05 | 96.1 | — | |
| DoRAProjection Type=Both-Proj, Number of Parameters=0.65M2024.11 | 96.09 | — | |
| LoRAProjection Type=Both-Proj, Number of Parameters=0.61M2024.11 | 96.01 | — | |
| +DP-FMShot=42026.05 | 96 | — | |
| PLOT++Shot=162026.05 | 96 | — | |
| ProDiaLProjection Type=In-Proj, Number of Parameters=0.41M2024.11 | 95.93 | — | |
| ProGradShot=162026.05 | 95.9 | — | |
| FTModel Size=Vim-small, Parameters=7.12M, Projection Type=Out-Proj2024.11 | 95.86 | — | |
| LoRAProjection Type=In-Proj, Number of Parameters=0.39M2024.11 | 95.78 | — | |
| StrongProjection Type=Baseline, Number of Parameters=0.94M2024.11 | 95.7 | — | |
| TIP-AdapterShot=162026.05 | 95.7 | — | |
| FTProjection Type=Out-Proj, Number of Parameters=1.77M2024.11 | 95.63 | — | |
| DoRAProjection Type=In-Proj, Number of Parameters=0.41M2024.11 | 95.55 | — | |
| ProDiaLProjection Type=Out-Proj, Number of Parameters=0.25M2024.11 | 95.55 | — | |
| CoOpShot=162026.05 | 95.5 | — | |
| DoRAProjection Type=Out-Proj, Number of Parameters=0.26M2024.11 | 95.47 | — | |
| LoRAProjection Type=Out-Proj, Number of Parameters=0.24M2024.11 | 95.4 | — | |
| DMN2026.06 | 95.4 | — | |
| GPUASamples per class=full training set2026.06 | 95.3 | — | |
| FTProjection Type=In-Proj, Number of Parameters=3.56M2024.11 | 95.24 | — | |
| CoCoOpShot=162026.05 | 95.2 | — | |
| KgCoOpShot=162026.05 | 95.2 | — | |
| FTModel Size=Vim-small, Parameters=14.20M, Projection Type=In-Proj2024.11 | 95.17 | — | |
| PLOT++Shot=42026.05 | 95.1 | — | |
| FTProjection Type=Both-Proj, Number of Parameters=5.33M2024.11 | 95.01 | — | |
| KgCoOpShot=42026.05 | 95 | — | |
| CLIP-LoRAShot=42026.05 | 95 | — | |
| FTModel Size=Vim-small, Parameters=21.27M, Projection Type=Both-Proj2024.11 | 94.94 | — | |
| CLIP-AdapterShot=162026.05 | 94.9 | — | |
| CoCoOpShot=42026.05 | 94.8 | — | |
| TIP-AdapterShot=42026.05 | 94.8 | — | |
| DPE2026.06 | 94.8 | — | |
| CoOpShot=42026.05 | 94.5 | — | |
| +DP-FMShot=12026.05 | 94.4 | — | |
| ProGradShot=42026.05 | 94.4 | — | |
| ZERO2026.06 | 94.4 | — | |
| PLOT++Shot=12026.05 | 94.3 | — | |
| PyramidCLIPTask=Linear Probe, Pretrain Dataset=143M, Image Encoder=ResNet502022.04 | 94.2 | — | |
| TPTBackbone=ViT-B/16, Evaluation Protocol=Transductive Zero-Shot2024.04 | 94.2 | — | |
| KgCoOpShot=12026.05 | 94.2 | — | |
| MTA2026.06 | 94.2 | — | |
| TDA2026.06 | 94.2 | — | |
| StatA2026.06 | 94.2 | — | |
| CoCoOpShot=12026.05 | 94.1 | — | |
| Full-FTModel Size=Vim-small, Parameters=25.45M, Projection Type=Base2024.11 | 94.09 | — | |
| TransCLIP-FSShots=4, Backbone=ViT-B/162024.06 | 94 | — | |
| TransCLIP-FSShots=16, Backbone=ViT-B/162024.06 | 94 | — | |
| TIP-AdapterShot=12026.05 | 94 | — | |
| CLIP-AdapterShot=42026.05 | 94 | — | |
| DeCLIPTask=Linear Probe, Pretrain Dataset=88M, Image Encoder=ResNet502022.04 | 93.9 | — | |
| CLIPBackbone=ViT-B/16, Evaluation Protocol=Transductive Zero-Shot2024.04 | 93.9 | — | |
| TIPPLEInference-time adaptation=true2026.06 | 93.9 | — | |
| CLIP* L/14MAP head=true, Grain=Coarse, Frozen representation=true, backbone=ViT-L/142023.06 | 93.8 | — | |
| CLIP-LoRAShot=12026.05 | 93.8 | — | |
| BitFitProjection Type=Baseline, Number of Parameters=0.06M2024.11 | 93.71 | — | |
| CLIP-DNBackbone=ViT-B/16, Evaluation Protocol=Transductive Zero-Shot2024.04 | 93.6 | — | |
| CLIP* (8k)MAP head=true, Grain=Coarse, Frozen representation=true2023.06 | 93.5 | — | |
| ProGradShot=12026.05 | 93.5 | — | |
| CLIP-ViT-B/16Shots=0, Backbone=ViT-B/162024.06 | 93.2 | — | |
| CapPa L/14MAP head=true, Grain=Coarse, Frozen representation=true, backbone=ViT-L/142023.06 | 93.2 | — | |
| CLIP2026.06 | 93.2 | — | |
| TransCLIP-FSShots=1, Backbone=ViT-B/162024.06 | 93.1 | — | |
| CLIP* (16k)MAP head=true, Grain=Coarse, Frozen representation=true2023.06 | 93.1 | — | |
| ZLaP2026.06 | 93.1 | — | |
| CLIPShot=02026.05 | 92.9 | — | |
| Full-FTProjection Type=Baseline, Number of Parameters=7.00M2024.11 | 92.86 | — | |
| DINOv2-ARLN=20M, Zero-shot=true2024.09 | 92.8 | — | |
| DINOv2-ARLN=20M, Zero-shot=true2024.09 | 92.8 | — | |
| OpenAI-CLIP ViT-LN=400M, Zero-shot=true2024.09 | 92.6 | — | |
| OpenAI-CLIP VIT-LN=400M, Zero-shot=true2024.09 | 92.6 | — | |
| LAION-CLIP ViT-LN=400M, Zero-shot=true2024.09 | 92.5 | — | |
| LAION-CLIP VIT-LN=400M, Zero-shot=true2024.09 | 92.5 | — | |
| CoOpShot=12026.05 | 92.5 | — | |
| CLIP-AdapterShot=12026.05 | 92 | — | |
| CLIP + ZLaPBackbone=ViT-B/16, Evaluation Protocol=Transductive Zero-Shot2024.04 | 91.8 | -2.1 | |
| DINOv2-MpNetN=20M, Zero-shot=true2024.09 | 91.8 | — | |
| DINOv2-MpNetN=20M, Zero-shot=true2024.09 | 91.8 | — | |
| CLIPTask=Linear Probe, Pretrain Dataset=143M, Image Encoder=ResNet502022.04 | 91.5 | — |