Image Classification on Food101
95.3AccuracyCLIP+
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| CLIP+Pre-training Data=WIT-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 95.3 | — | — | — | — | — | — | |
| CLIPPre-training Data=WIT-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 95.2 | — | — | — | — | — | — | |
| MLCDPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 94.3 | — | — | — | — | — | — | |
| StatABackbone=ViT-L/14, Keff=Very Low2025.01 | 94.2 | — | — | — | — | — | — | |
| StatABackbone=ViT-L/14, Scenario=High, Batch Size=1282025.01 | 93.5 | — | — | — | — | — | — | |
| StatABackbone=ViT-L/14, Scenario=Separate, Batch Size=1282025.01 | 93.5 | — | — | — | — | — | — | |
| StatABackbone=ViT-L/14, Batch Size=128, Online Test-Time Adaptation=true, Scenario=High2025.01 | 93.5 | — | — | — | — | — | — | |
| StatABackbone=ViT-L/14, Batch Size=128, Online Test-Time Adaptation=true, Scenario=Separate2025.01 | 93.5 | — | — | — | — | — | — | |
| UNICOMPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 93.4 | — | — | — | — | — | — | |
| OpenCLIPPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 93.3 | — | — | — | — | — | — | |
| StatABackbone=ViT-L/14, Scenario=Medium, Batch Size=1282025.01 | 93.2 | — | — | — | — | — | — | |
| StatABackbone=ViT-L/14, Batch Size=128, Online Test-Time Adaptation=true, Scenario=Medium2025.01 | 93.2 | — | — | — | — | — | — | |
| StatABackbone=ViT-L/14, Keff=Low2025.01 | 93.1 | — | — | — | — | — | — | |
| CLIP (reported)Pre-training Data=WIT-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot classification2024.07 | 92.9 | — | — | — | — | — | — | |
| CLIP-ViT-B/16Img-text pairs=400M, Evaluation protocol=Linear probing2022.09 | 92.8 | — | — | — | — | — | — | |
| StatABackbone=ViT-L/14, K_eff=High2025.01 | 92.3 | — | — | — | — | — | — | |
| StatABackbone=ViT-L/14, Keff=Medium2025.01 | 92.1 | — | — | — | — | — | — | |
| StatABackbone=ViT-L/14, Scenario=Low, Batch Size=1282025.01 | 92 | — | — | — | — | — | — | |
| StatABackbone=ViT-L/14, Batch Size=128, Online Test-Time Adaptation=true, Scenario=Low2025.01 | 92 | — | — | — | — | — | — | |
| StatAEncoder=ViT-L/14, K_eff=All2025.01 | 91.8 | — | — | — | — | — | — | |
| CLIPPre-training Data=WIT-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot classification2024.07 | 91 | — | — | — | — | — | — | |
| CLIPBackbone=ViT-L/142025.01 | 90.9 | — | — | — | — | — | — | |
| CLIPEncoder=ViT-L/14, K_eff=All2025.01 | 90.9 | — | — | — | — | — | — | |
| CLIPBackbone=ViT-L/14, Setting=Zero-shot, Batch Size=1282025.01 | 90.9 | — | — | — | — | — | — | |
| CLIPBackbone=ViT-L/14, Batch Size=128, Online Test-Time Adaptation=false2025.01 | 90.9 | — | — | — | — | — | — | |
| CCMAAL iteration (t)=20, Feature extractor=DINOv2 ViT-g14, Initialization=Random2026.03 | 90.8 | — | — | — | — | — | — | |
| CoFTmode=unsupervised fine-tuning2026.02 | 90.58 | — | — | — | — | — | — | |
| CLIP-MoE + MoE-GRPOTraining shots=16-shot, Source dataset=ImageNet, Training epochs=32026.03 | 90.5 | — | — | — | — | — | — | |
| CoFT+mode=unsupervised fine-tuning2026.02 | 90.45 | — | — | — | — | — | — | |
| MLCDPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot classification2024.07 | 90.3 | — | — | — | — | — | — | |
| CLIP-1+DINO(v2+v3)Setting=Inductive, Test-Time Adaptation (TTA)=false2025.06 | 90.2 | — | — | — | — | — | — | |
| BADGEAL iteration (t)=20, Feature extractor=DINOv2 ViT-g14, Initialization=Random2026.03 | 90.2 | — | — | — | — | — | — | |
| Alfa-mixAL iteration (t)=20, Feature extractor=DINOv2 ViT-g14, Initialization=Random2026.03 | 90.2 | — | — | — | — | — | — | |
| pBALDAL iteration (t)=20, Feature extractor=DINOv2 ViT-g14, Initialization=Random2026.03 | 90.1 | — | — | — | — | — | — | |
| CCMAAL iteration (t)=16, Feature extractor=DINOv2 ViT-g14, Initialization=Random2026.03 | 90.1 | — | — | — | — | — | — | |
| CLIP-1+DINOv3Setting=Inductive, Test-Time Adaptation (TTA)=false2025.06 | 90 | — | — | — | — | — | — | |
| MarginsAL iteration (t)=20, Feature extractor=DINOv2 ViT-g14, Initialization=Random2026.03 | 90 | — | — | — | — | — | — | |
| StatABackbone=ResNet-101, Keff=Very Low2025.01 | 89.5 | — | — | — | — | — | — | |
| FLIPPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot classification2024.07 | 89.3 | — | — | — | — | — | — | |
| CLIP-Adapter (best α)Backbone=ViT-B/16, Shots=16, Blending ratio (alpha) selection strategy=Best residual ratio per dataset (oracle)2026.03 | 89.3 | — | — | — | — | — | — | |
| CLIP-AdapterBackbone=ViT-B/16, Shots=16, Blending ratio (alpha)=best2026.03 | 89.3 | — | — | — | — | — | — | |
| pBALDAL iteration (t)=16, Feature extractor=DINOv2 ViT-g14, Initialization=Random2026.03 | 89.3 | — | — | — | — | — | — | |
| BADGEAL iteration (t)=16, Feature extractor=DINOv2 ViT-g14, Initialization=Random2026.03 | 89.3 | — | — | — | — | — | — | |
| MarginsAL iteration (t)=16, Feature extractor=DINOv2 ViT-g14, Initialization=Random2026.03 | 89.2 | — | — | — | — | — | — | |
| GRIPmode=unsupervised fine-tuning2026.02 | 89.16 | — | — | — | — | — | — | |
| CLIP-Adapter (α = 0.2)Backbone=ViT-B/16, Shots=16, Blending ratio (alpha) selection strategy=Fixed residual ratio (α = 0.2), selected on ImageNet2026.03 | 89.1 | — | — | — | — | — | — | |
| CLIP-AdapterBackbone=ViT-B/16, Shots=16, Blending ratio (alpha)=0.22026.03 | 89.1 | — | — | — | — | — | — | |
| Alfa-mixAL iteration (t)=16, Feature extractor=DINOv2 ViT-g14, Initialization=Random2026.03 | 89.1 | — | — | — | — | — | — | |
| CLIP-1+DINOv2Setting=Inductive, Test-Time Adaptation (TTA)=false2025.06 | 89 | — | — | — | — | — | — | |
| HOSO-AdapterBackbone=ViT-B/16, Shots=16, Blending ratio (alpha) selection strategy=Optimised on hold-one-shot-out cache2026.03 | 88.97 | — | — | — | — | — | — | |
| HOSO-AdapterBackbone=ViT-B/16, Shots=16, Number of runs=32026.03 | 88.97 | — | — | — | — | — | — | |
| PromptKDrole=Target2026.03 | 88.84 | — | — | — | — | — | — | |
| VLM BaselineIR=1, Backbone=VLM2026.01 | 88.8 | — | — | — | — | — | — | |
| CAPTrole=Target2026.03 | 88.75 | — | — | — | — | — | — | |
| StatAScenario=High, Backbone=ResNet-101, Batch Size=1282025.01 | 88.7 | — | — | — | — | — | — | |
| StatABackbone=ResNet-101, Batch size=128, Scenario=High2025.01 | 88.7 | — | — | — | — | — | — | |
| ProbCoverAL iteration (t)=20, Feature extractor=DINOv2 ViT-g14, Initialization=Random2026.03 | 88.7 | — | — | — | — | — | — | |
| CLIP-MoETraining shots=16-shot, Source dataset=ImageNet, Training epochs=32026.03 | 88.7 | — | — | — | — | — | — | |
| CPLmode=unsupervised fine-tuning2026.02 | 88.64 | — | — | — | — | — | — | |
| StatAScenario=Separate, Backbone=ResNet-101, Batch Size=1282025.01 | 88.5 | — | — | — | — | — | — | |
| StatABackbone=ResNet-101, Batch size=128, Scenario=Separate2025.01 | 88.5 | — | — | — | — | — | — | |
| FLAVAImg-text pairs=70M, Evaluation protocol=Linear probing2022.09 | 88.5 | — | — | — | — | — | — | |
| ProbCoverAL iteration (t)=16, Feature extractor=DINOv2 ViT-g14, Initialization=Random2026.03 | 88.4 | — | — | — | — | — | — | |
| CLIP-MoE + Det-FTTraining shots=16-shot, Source dataset=ImageNet, Training epochs=32026.03 | 88.3 | — | — | — | — | — | — | |
| RandomAL iteration (t)=20, Feature extractor=DINOv2 ViT-g14, Initialization=Random2026.03 | 88.2 | — | — | — | — | — | — | |
| StatABackbone=ResNet-101, Keff=Low2025.01 | 88.1 | — | — | — | — | — | — | |
| StatAScenario=Medium, Backbone=ResNet-101, Batch Size=1282025.01 | 88.1 | — | — | — | — | — | — | |
| StatABackbone=ResNet-101, Batch size=128, Scenario=Medium2025.01 | 88.1 | — | — | — | — | — | — | |
| ZLaPBatch Size=1000, Backbone=ViT-B/16, Keff (Effective classes per task)=High (25-50)2025.01 | 88 | — | — | — | — | — | — | |
| TransCLIPBatch Size=1000, Backbone=ViT-B/16, Keff=High (25-50), Protocol=Batch test-time adaptation2025.01 | 88 | — | — | — | — | — | — | |
| StatABatch Size=1000, Backbone=ViT-B/16, Keff=High (25-50), Protocol=Batch test-time adaptation2025.01 | 88 | — | — | — | — | — | — | |
| StatABatch Size=1000, Backbone=ViT-B/16, Keff (Effective classes per task)=High (25-50)2025.01 | 87.7 | — | — | — | — | — | — | |
| StatABatch Size=1000, Backbone=ViT-B/16, Keff (Effective classes per task)=Very High (50-100)2025.01 | 87.7 | — | — | — | — | — | — | |
| MMRL + MetaTPTadaptation=MetaTPT2025.12 | 87.61 | — | — | — | — | — | — | |
| TransCLIPBatch Size=1000, Backbone=ViT-B/16, Keff (Effective classes per task)=Medium (5-25)2025.01 | 87.5 | — | — | — | — | — | — | |
| StatABatch Size=1000, Backbone=ViT-B/16, Keff=Medium (5-25), Protocol=Batch test-time adaptation2025.01 | 87.5 | — | — | — | — | — | — | |
| PromptSRCShots=16 shots2025.03 | 87.5 | — | — | — | — | — | — | |
| PromptSRCshots=162025.05 | 87.5 | — | — | — | — | — | — | |
| TOGAShots=16, Venue=-2026.03 | 87.5 | — | — | — | — | — | — | |
| Image-TextIntra-modal=✗, Classifier=Zero-Shot, Backbone=ViT-B/16-open2026.03 | 87.5 | — | — | — | — | — | — | |
| OpenCLIPPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot classification2024.07 | 87.4 | — | — | — | — | — | — | |
| OmniVLImg-text pairs=14M*, Evaluation protocol=Linear probing, Note=* denotes extra video data is used2022.09 | 87.4 | — | — | — | — | — | — | |
| CoCoOpShots=162025.10 | 87.4 | — | — | — | — | — | — | |
| RandomAL iteration (t)=16, Feature extractor=DINOv2 ViT-g14, Initialization=Random2026.03 | 87.4 | — | — | — | — | — | — | |
| MaPLe + MetaTPTadaptation=MetaTPT2025.12 | 87.28 | — | — | — | — | — | — | |
| CoCoOpShots=16 shots2025.03 | 87.25 | — | — | — | — | — | — | |
| CoCoOpshots=162025.05 | 87.25 | — | — | — | — | — | — | |
| MTABatch Size=1000, Backbone=ViT-B/162025.01 | 87.2 | — | — | — | — | — | — | |
| MTABatch Size=1000, Backbone=ViT-B/16, Protocol=Batch test-time adaptation2025.01 | 87.2 | — | — | — | — | — | — | |
| KgCoOpShots=162025.10 | 87.2 | — | — | — | — | — | — | |
| TransCLIPBatch Size=1000, Backbone=ViT-B/16, Keff (Effective classes per task)=Very High (50-100)2025.01 | 87.1 | — | — | — | — | — | — | |
| StatABackbone=ResNet-50, Keff=Very Low2025.01 | 87.1 | — | — | — | — | — | — | |
| StatABatch Size=1000, Backbone=ViT-B/16, Keff=Very High (50-100), Protocol=Batch test-time adaptation2025.01 | 87.1 | — | — | — | — | — | — | |
| CLIP-AdapterShots=162025.10 | 87.1 | — | — | — | — | — | — | |
| PLOT++Shots=162025.10 | 87.1 | — | — | — | — | — | — | |
| IsoCLIPIntra-modal=✓, Classifier=NCM, Backbone=ViT-B/16-open2026.03 | 87.1 | — | — | — | — | — | — | |
| MMRLShots=16 shots2025.03 | 87.03 | — | — | — | — | — | — | |
| MMRLshots=162025.05 | 87.03 | — | — | — | — | — | — | |
| CoCoOpShots=8 shots2025.03 | 86.97 | — | — | — | — | — | — | |
| CoCoOpshots=82025.05 | 86.97 | — | — | — | — | — | — |