Image Classification on SUN397
98.4AccuracyLaCLIP
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| LaCLIPArchitecture=ViT-B/16, Pre-training Data=LAION-400M, Protocol=5-way 5-shot2023.05 | 98.4 | — | — | |
| LaCLIPArchitecture=ViT-B/32, Pre-training Data=LAION-400M, Protocol=5-way 5-shot2023.05 | 98.2 | — | — | |
| CLIPArchitecture=ViT-B/16, Pre-training Data=LAION-400M, Protocol=5-way 5-shot2023.05 | 98.2 | — | — | |
| LaSLIPArchitecture=ViT-B/16, Pre-training Data=CC12M, Protocol=5-way 5-shot2023.05 | 98.1 | — | — | |
| CLIPArchitecture=ViT-B/32, Pre-training Data=LAION-400M, Protocol=5-way 5-shot2023.05 | 98 | — | — | |
| LaCLIPArchitecture=ViT-B/16, Pre-training Data=CC12M, Protocol=5-way 5-shot2023.05 | 98 | — | — | |
| SLIPArchitecture=ViT-B/16, Pre-training Data=CC12M, Protocol=5-way 5-shot2023.05 | 97.5 | — | — | |
| CLIPArchitecture=ViT-B/16, Pre-training Data=CC12M, Protocol=5-way 5-shot2023.05 | 97.2 | — | — | |
| LaCLIPArchitecture=ViT-B/16, Pre-training Data=RedCaps, Protocol=5-way 5-shot2023.05 | 96.2 | — | — | |
| CLIPArchitecture=ViT-B/16, Pre-training Data=RedCaps, Protocol=5-way 5-shot2023.05 | 95.9 | — | — | |
| LaCLIPArchitecture=ViT-B/16, Pre-training Data=CC3M, Protocol=5-way 5-shot2023.05 | 95.2 | — | — | |
| CLIPArchitecture=ViT-B/16, Pre-training Data=CC3M, Protocol=5-way 5-shot2023.05 | 94.1 | — | — | |
| TPTEvaluation Protocol=Cross-dataset transfer evaluation2024.02 | 84.67 | — | — | |
| X-FM_baseLinear evaluation=true, Model size=Base, Patch size=16*16, Resolution=224*2242023.01 | 83.9 | — | — | |
| fine-tunedBackbone=ViT-L/142026.02 | 82.8 | — | — | |
| Multiple modelsBackbone=ViT-L/142022.08 | 82.4 | — | — | |
| FLAVALinear evaluation=true, Model size=Base, Patch size=16*16, Resolution=224*2242023.01 | 82.1 | — | — | |
| MLCDPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 82.1 | — | — | |
| CLIPPre-training Data=WIT-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 81.8 | — | — | |
| CLIP+Pre-training Data=WIT-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 81.5 | — | — | |
| ISO-CTS/LARVBackbone=ViT-L/142026.02 | 81.4 | — | 0.3 | |
| OpenCLIPPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 81 | — | — | |
| Joint patchingBackbone=ViT-L/142022.08 | 80.9 | — | — | |
| ISO-C/LARVBackbone=ViT-L/142026.02 | 80.8 | — | 0.1 | |
| 2SFS LoRABackbone=ViT-L/14, Shots (k=16)=16, Evaluation Protocol=all-to-all2025.03 | 80.7 | — | — | |
| Seq. patchingBackbone=ViT-L/14, Seed=02022.08 | 80 | — | — | |
| UNICOMPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 80 | — | — | |
| PyramidCLIPImage Encoder=ViT-B/16, Pretrain Dataset Size=143M, Evaluation Protocol=Linear probe2022.04 | 79.9 | — | — | |
| TSV-M/LARVBackbone=ViT-L/142026.02 | 79.8 | — | 1.6 | |
| Multiple modelsBackbone=ViT-B/162022.08 | 79 | — | — | |
| IndividualBackbone=ViT-B/16, Pre-trained=ImageNet-21k, Number of Merged Models=12024.05 | 78.98 | — | — | |
| fine-tunedBackbone=ViT-B/162026.02 | 78.9 | — | — | |
| Seq. patchingBackbone=ViT-L/14, Seed=12022.08 | 78.8 | — | — | |
| CLIP*Image Encoder=ViT-B/16, Pretrain Dataset Size=400M, Evaluation Protocol=Linear probe2022.04 | 78.4 | — | — | |
| Seq. patchingBackbone=ViT-L/14, Seed=22022.08 | 78.4 | — | — | |
| CLIPLinear evaluation=true, Model size=Base, Patch size=16*16, Resolution=224*2242023.01 | 78.4 | — | — | |
| DaVinciLinear evaluation=true, Model size=Base, Patch size=16*16, Resolution=224*2242023.01 | 78 | — | — | |
| MMRLShots=16 shots2025.03 | 77.7 | — | — | |
| AWTShots=162024.07 | 77.57 | — | — | |
| PyramidCLIPImage Encoder=ViT-B/32, Pretrain Dataset Size=143M, Evaluation Protocol=Linear probe2022.04 | 77.5 | — | — | |
| TIES/LARVBackbone=ViT-L/142026.02 | 77.3 | — | 2.6 | |
| ProSRCShot=16, Time (s)=14403, Mem. (M)=33732024.12 | 77.23 | — | — | |
| PromptSRCShots=16 shots2025.03 | 77.23 | — | — | |
| 2SFS LoRABackbone=ViT-B/16, Shots (k=16)=16, Evaluation Protocol=all-to-all2025.03 | 76.9 | — | — | |
| Skip TuningShot=16, Time (s)=962, Mem. (M)=5282024.12 | 76.8 | — | — | |
| ISO-C/LARVBackbone=ViT-B/162026.02 | 76.7 | — | 0.9 | |
| CLIP*Image Encoder=ViT-B/32, Pretrain Dataset Size=400M, Evaluation Protocol=Linear probe2022.04 | 76.6 | — | — | |
| BASIC-LBackbone=BASIC-L2021.11 | 76.2 | — | — | |
| EMR-MERGINGBackbone=ViT-B/16, Pre-trained=ImageNet-21k, Number of Merged Models=302024.05 | 76.19 | — | — | |
| PyramidCLIPTask=Linear Probe, Pretrain Dataset=143M, Image Encoder=ResNet502022.04 | 76.1 | — | — | |
| MMRLShots=8 shots2025.03 | 76 | — | — | |
| Joint patchingBackbone=ViT-B/162022.08 | 75.9 | — | — | |
| AdaMergingBackbone=ViT-L/142026.02 | 75.9 | — | — | |
| ProSRCShot=8, Time (s)=7636, Mem. (M)=33732024.12 | 75.73 | — | — | |
| PromptSRCShots=8 shots2025.03 | 75.73 | — | — | |
| MaPLeShot=16, Time (s)=14070, Mem. (M)=30612024.12 | 75.53 | — | — | |
| MaPLeShots=16 shots2025.03 | 75.53 | — | — | |
| ISO-CTS/LARVBackbone=ViT-B/162026.02 | 75.5 | — | -0.6 | |
| Multiple modelsBackbone=ViT-B/322022.08 | 75.1 | — | — | |
| fine-tunedBackbone=ViT-B/322026.02 | 74.9 | — | — | |
| Skip TuningShot=8, Time (s)=508, Mem. (M)=5282024.12 | 74.77 | — | — | |
| CLIPTask=Linear Probe, Pretrain Dataset=143M, Image Encoder=ResNet502022.04 | 74.7 | — | — | |
| CoOpShot=16, Time (s)=3879, Mem. (M)=45512024.12 | 74.67 | — | — | |
| CoOpShots=16 shots2025.03 | 74.67 | — | — | |
| MMAShots=16 shots2025.03 | 74.63 | — | — | |
| 2SFS LoRABackbone=ViT-B/32, Shots (k=16)=16, Evaluation Protocol=all-to-all2025.03 | 74.5 | — | — | |
| TA/LARVBackbone=ViT-L/142026.02 | 74.5 | — | 2.5 | |
| ProSRCShot=4, Time (s)=3841, Mem. (M)=33732024.12 | 74 | — | — | |
| PromptSRCShots=4 shots2025.03 | 74 | — | — | |
| TSV-M/LARVBackbone=ViT-B/162026.02 | 74 | — | 0.9 | |
| MMRLShots=4 shots2025.03 | 73.93 | — | — | |
| Seq. patchingBackbone=ViT-B/16, Seed=02022.08 | 73.4 | — | — | |
| Parallel patchingBackbone=ViT-L/142022.08 | 73.4 | — | — | |
| CLIPTask=Linear Probe, Pretrain Dataset=400M, Image Encoder=ResNet502022.04 | 73.3 | — | — | |
| CLIPBackbone=ResNet-50, Evaluation Protocol=Linear probe2021.06 | 73.3 | — | — | |
| Linear probe CLIPShots=16 shots2025.03 | 73.28 | — | — | |
| MaPLeShot=8, Time (s)=7060, Mem. (M)=30612024.12 | 73.23 | — | — | |
| MaPLeShots=8 shots2025.03 | 73.23 | — | — | |
| FLIPPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot classification2024.07 | 73.1 | — | — | |
| Skip TuningShot=4, Time (s)=280, Mem. (M)=5282024.12 | 73.07 | — | — | |
| BASIC-MBackbone=BASIC-M2021.11 | 72.9 | — | — | |
| ISO-CTS/LARVBackbone=ViT-B/322026.02 | 72.9 | — | 1.8 | |
| DeCLIPTask=Linear Probe, Pretrain Dataset=88M, Image Encoder=ResNet502022.04 | 72.8 | — | — | |
| StatAEncoder=ViT-L/14, K_eff=All2025.01 | 72.8 | — | — | |
| AWTShots=12024.07 | 72.65 | — | — | |
| Joint patchingBackbone=ViT-B/322022.08 | 72.6 | — | — | |
| Simple AveragingBackbone=ViT-L/142026.02 | 72.5 | — | — | |
| Seq. patchingBackbone=ViT-B/16, Seed=12022.08 | 72.4 | — | — | |
| MMAShots=8 shots2025.03 | 72.3 | — | — | |
| TIES/LARVBackbone=ViT-B/162026.02 | 72.2 | — | 1.5 | |
| CoCoOpShot=16, Time (s)=10694, Mem. (M)=28882024.12 | 72.15 | — | — | |
| CoCoOpShots=16 shots2025.03 | 72.15 | — | — | |
| MLCDPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot classification2024.07 | 72.1 | — | — | |
| StatABackbone=ViT-L/14, K_eff=Very High, Batch Size=1000, Evaluation Protocol=batch test-time adaptation2025.01 | 71.6 | — | — | |
| StatABackbone=ViT-L/14, Keff=Low2025.01 | 71.6 | — | — | |
| ProSRCShot=2, Time (s)=1947, Mem. (M)=33732024.12 | 71.6 | — | — | |
| PromptSRCShots=2 shots2025.03 | 71.6 | — | — | |
| CoOpShot=8, Time (s)=1988, Mem. (M)=45512024.12 | 71.53 | — | — | |
| CoOpShots=8 shots2025.03 | 71.53 | — | — | |
| MMRLShots=2 shots2025.03 | 71.53 | — | — |