Image Classification on SUN397 (Accuracy)
97.8AccuracySigLIP2-g-opt
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SigLIP2-g-optSize=g2026.02 | 97.8 | — | |
| StableRepPre-training Data Source=Synthetic, Evaluation Protocol=5-way, 5-shot2023.06 | 97.3 | — | |
| StableRepPre-training dataset=CC12M, Pre-training data source=Syn, Evaluation protocol=Few-shot2023.06 | 97.3 | — | |
| CLIPPre-training Data Source=Real, Evaluation Protocol=5-way, 5-shot2023.06 | 97.2 | — | |
| CLIPPre-training dataset=CC12M, Pre-training data source=Real, Evaluation protocol=Few-shot2023.06 | 97.2 | — | |
| PEcore GSize=G2026.02 | 96.9 | — | |
| SigLIP-L/16Backbone=SigLIP-L, Patch Size=162026.02 | 96.8 | — | |
| DFN-H+2026.02 | 96.8 | — | |
| StableRepPre-training dataset=RedCaps, Pre-training data source=Syn, Evaluation protocol=Few-shot2023.06 | 96.8 | — | |
| CLIPPre-training Data Source=Synthetic, Evaluation Protocol=5-way, 5-shot2023.06 | 96.7 | — | |
| CLIPPre-training dataset=CC12M, Pre-training data source=Syn, Evaluation protocol=Few-shot2023.06 | 96.7 | — | |
| SigLIP2-L/16Backbone=SigLIP2-L, Patch Size=162026.02 | 96.4 | — | |
| PEcore LSize=L2026.02 | 96.4 | — | |
| PEcore G (image only)Size=G, Input=image only2026.02 | 96.4 | — | |
| InternVL-C2026.02 | 96.3 | — | |
| CLIPPre-training dataset=RedCaps, Pre-training data source=Syn, Evaluation protocol=Few-shot2023.06 | 96.3 | — | |
| EVA 18BSize=18B2026.02 | 96.1 | — | |
| CLIPPre-training dataset=RedCaps, Pre-training data source=Real, Evaluation protocol=Few-shot2023.06 | 95.9 | — | |
| SimCLRPre-training Data Source=Real, Evaluation Protocol=5-way, 5-shot2023.06 | 94 | — | |
| SimCLRPre-training dataset=CC12M, Pre-training data source=Real, Evaluation protocol=Few-shot2023.06 | 94 | — | |
| SimCLRPre-training Data Source=Synthetic, Evaluation Protocol=5-way, 5-shot2023.06 | 92.9 | — | |
| SimCLRPre-training dataset=CC12M, Pre-training data source=Syn, Evaluation protocol=Few-shot2023.06 | 92.9 | — | |
| SimCLRPre-training dataset=RedCaps, Pre-training data source=Syn, Evaluation protocol=Few-shot2023.06 | 92 | — | |
| SimCLRPre-training dataset=RedCaps, Pre-training data source=Real, Evaluation protocol=Few-shot2023.06 | 91.8 | — | |
| CLIP* L/14Backbone Scale=L/14, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 85.7 | — | |
| CLIP* L/14MAP head=true, Grain=Coarse, Frozen representation=true, backbone=ViT-L/142023.06 | 85.7 | — | |
| CapPa L/14MAP head=true, Grain=Coarse, Frozen representation=true, backbone=ViT-L/142023.06 | 85.2 | — | |
| CapPa L/14Backbone Scale=L/14, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 84.9 | — | |
| CLIP L/14Backbone Scale=L/14, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 84.8 | — | |
| CLIP* (16k)Backbone Scale=B/16, Pre-training Batch Size=16k, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 83.3 | — | |
| CLIP* (8k)Backbone Scale=B/16, Pre-training Batch Size=8k, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 83.2 | — | |
| theta_B fine-tuneModel=ViT-L/14, Mode=fine-tune2026.02 | 82.94 | — | |
| CLIP* (16k)MAP head=true, Grain=Coarse, Frozen representation=true2023.06 | 82.9 | — | |
| CLIPBackbone Scale=B/16, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 82.5 | — | |
| CapPaBackbone Scale=B/16, Pre-training Batch Size=8k, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 82.4 | — | |
| PaRaMSModel Status=Protected Standalone (θ̂_def)2025.11 | 82.32 | — | |
| CapBackbone Scale=B/16, Pre-training Batch Size=8k, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 82.3 | — | |
| CapPaMAP head=true, Grain=Coarse, Frozen representation=true2023.06 | 82.1 | — | |
| CapMAP head=true, Grain=Coarse, Frozen representation=true2023.06 | 82 | — | |
| CLIP* (8k)MAP head=true, Grain=Coarse, Frozen representation=true2023.06 | 81.8 | — | |
| MergeGuardModel Status=Protected Standalone (θ̂_def)2025.11 | 81.52 | — | |
| SynCLR + GMAILEvaluation Protocol=Zero-shot2026.02 | 81.25 | — | |
| θBmode=fine-tune2026.02 | 79.91 | — | |
| θBMode=fine-tune2026.02 | 79.91 | — | |
| theta_A fine-tuneModel=ViT-B/16, Mode=fine-tune2026.02 | 79.91 | — | |
| theta_A fine-tuneModel=laion2b, Protocol=fine-tune2026.02 | 79.86 | — | |
| ADAProtection Status=None (theta_merge)2025.11 | 79.4 | — | |
| SOT-GLPShots=16, Backbone=ViT-B/16, Pre-training=CLIP2026.03 | 78.2 | — | |
| Fine-TunedBackbone=DINOv3 ViT-B/162025.12 | 78.1 | — | |
| MMRLshots=162025.05 | 77.7 | — | |
| MMRL++shots=162025.05 | 77.5 | — | |
| KD 2B to ViT-H, M+V+L4Backbone=ViT-H, Distillation=KD 2B, Target=M+V+L42026.02 | 77.35 | — | |
| PromptSRCshots=162025.05 | 77.23 | — | |
| Linear ProbeBackbone=DINOv3 ViT-B/162025.12 | 77.2 | — | |
| CLIP-LoRA + FMAShots=162025.10 | 77.2 | — | |
| PromptSRCShots=16, Backbone=ViT-B/16, Pre-training=CLIP2026.03 | 77.2 | — | |
| GalLoPShots=16, Backbone=ViT-B/16, Pre-training=CLIP2026.03 | 77.2 | — | |
| THESEUSModel=ViT-L/14, B (alignment batches)=102026.02 | 77.16 | — | |
| theta_B fine-tuneModel=laion400m, Protocol=fine-tune2026.02 | 76.76 | — | |
| Task MatrixBackbone=DINOv3 ViT-B/16, best layer=112025.12 | 76.7 | — | |
| TIESProtection Status=None (theta_merge)2025.11 | 76.5 | — | |
| theta_B zero-shotModel=ViT-L/14, Mode=zero-shot2026.02 | 76.34 | — | |
| SynCLREvaluation Protocol=Zero-shot2026.02 | 76.2 | — | |
| CS-AlignerBackbone=ViT-B/16, Number of Shots=16, Evaluation Protocol=fine-tuned2025.02 | 76.2 | — | |
| TOGAShots=16, Venue=-2026.03 | 76.2 | — | |
| M-TIESBackbone=ViT-L/142026.02 | 76.13 | — | |
| MMRLshots=82025.05 | 76 | — | |
| TIP-AdapterShots=162025.10 | 76 | — | |
| PLOT++Shots=162025.10 | 76 | — | |
| CLIP-LoRAShots=162025.10 | 76 | — | |
| PLOTShots=16, Backbone=ViT-B/16, Pre-training=CLIP2026.03 | 76 | — | |
| PromptSRCshots=82025.05 | 75.73 | — | |
| ProDAShots=16, Backbone=ViT-B/16, Pre-training=CLIP2026.03 | 75.7 | — | |
| TIESBackbone=ViT-L/142026.02 | 75.65 | — | |
| CLIP-Adapter (best α)Backbone=ViT-B/16, Shots=16, Blending ratio (alpha) selection strategy=Best residual ratio per dataset (oracle)2026.03 | 75.6 | — | |
| CLIP-AdapterBackbone=ViT-B/16, Shots=16, Blending ratio (alpha)=best2026.03 | 75.6 | — | |
| MaPLeshots=162025.05 | 75.53 | — | |
| MMRL++shots=82025.05 | 75.53 | — | |
| CLIP + GMAILEvaluation Protocol=Zero-shot2026.02 | 75.53 | — | |
| MaPLeShots=16, Backbone=ViT-B/16, Pre-training=CLIP2026.03 | 75.5 | — | |
| DAREBackbone=ViT-L/142026.02 | 75.37 | — | |
| Fine-Tuned*Note=Upper-bound2025.12 | 75.3 | — | |
| KD 2B to ViT-H, ViseBackbone=ViT-H, Distillation=KD 2B, Target=Vise2026.02 | 75.12 | — | |
| ProGradShots=162025.10 | 75.1 | — | |
| TOGAShots=8, Venue=-2026.03 | 75.1 | — | |
| CLIP-AdapterBackbone=ViT-B/16, Number of Shots=16, Evaluation Protocol=fine-tuned2025.02 | 75 | — | |
| CoOpShots=162025.10 | 74.9 | — | |
| Image-ImageIntra-modal=✓, Classifier=NCM, Backbone=ViT-B/16-open2026.03 | 74.9 | — | |
| ProposedShots=16, Venue=-2025.12 | 74.8 | — | |
| Task MatrixBackbone=CLIP ViT-B/32, Best Layer=11, Classes=3972025.12 | 74.8 | — | |
| EnsemblingBackbone=ViT-L/142026.02 | 74.76 | — | |
| CoOpShots=16, Backbone=ViT-B/16, Pre-training=CLIP2026.03 | 74.7 | — | |
| CoOpshots=162025.05 | 74.67 | — | |
| HOSO-AdapterBackbone=ViT-B/16, Shots=16, Blending ratio (alpha) selection strategy=Optimised on hold-one-shot-out cache2026.03 | 74.67 | — | |
| HOSO-AdapterBackbone=ViT-B/16, Shots=16, Number of runs=32026.03 | 74.67 | — | |
| MMAshots=162025.05 | 74.63 | — | |
| MetaCLIP+ViSE2026.02 | 74.619 | — | |
| Only ViSE2026.02 | 74.619 | — | |
| Fine-TunedBackbone=CLIP ViT-B/32, Classes=3972025.12 | 74.5 | — | |
| ISO-C/LARVBackbone=ViT-B/322026.02 | 74.4 | — |