Image Classification on FER 2013
0.998Top-1 AccCLIP
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| CLIPPre-training Data=WIT-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 0.998 | — | — | |
| CLIP+Pre-training Data=WIT-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 0.998 | — | — | |
| MLCDPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 0.993 | — | — | |
| OpenCLIPPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 0.992 | — | — | |
| UNICOMPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 0.985 | — | — | |
| X-FM_baseLinear evaluation=true, Model size=Base, Patch size=16*16, Resolution=224*2242023.01 | 0.726 | — | — | |
| CLIPLinear evaluation=true, Model size=Base, Patch size=16*16, Resolution=224*2242023.01 | 0.695 | — | — | |
| DaVinciLinear evaluation=true, Model size=Base, Patch size=16*16, Resolution=224*2242023.01 | 0.684 | — | — | |
| WUDI-MergingMerging Strategy=WUDI-Merging, Quantization Method=None, Model Backbone=CLIP-ViT-B/32, Bit-width=Full2026.05 | 0.642 | — | — | |
| RN50 BaselineIR=16, Backbone=ResNet-502026.01 | 0.63 | — | — | |
| FLAVALinear evaluation=true, Model size=Base, Patch size=16*16, Resolution=224*2242023.01 | 0.611 | — | — | |
| OpenCLIP-G/14Zero-shot=true2024.02 | 0.596 | — | — | |
| LoRA+Parameter Count=0.67M2025.12 | 0.5951 | — | — | |
| AdaLoRAParameter Count=0.67M2025.12 | 0.5932 | — | — | |
| EVA-CLIP-18BZero-shot=true2024.02 | 0.593 | — | — | |
| DoRAParameter Count=0.77M2025.12 | 0.591 | — | — | |
| EVA-02-CLIP-E/14+Zero-shot=true2024.02 | 0.59 | — | — | |
| VeRAParameter Count=0.10M2025.12 | 0.5881 | — | — | |
| Partial-AdaLoRAParameter Count=0.21M2025.12 | 0.5877 | — | — | |
| Partial-LoRAParameter Count=0.21M2025.12 | 0.5861 | — | — | |
| SmartCLIPBackbone=ViT-L/14, Zero-shot=true2025.07 | 0.586 | — | — | |
| xCLIPLinear probing=true, Backbone=ViT-B/16, Pre-train=IT35M2022.10 | 0.585 | — | — | |
| WUDI-MergingMerging Strategy=WUDI-Merging, Quantization Method=Full-precision, Bit-width=Full, Backbone=CLIP-ViT-B/322026.05 | 0.585 | — | — | |
| LoRAParameter Count=0.67M2025.12 | 0.5796 | — | — | |
| LongCLIPBackbone=ViT-L/14, Zero-shot=true2025.07 | 0.578 | — | — | |
| CLIPmode=zero-shot2022.08 | 0.577 | — | — | |
| CLIPLinear probing=true, Backbone=ViT-B/16, Pre-train=IT35M2022.10 | 0.565 | — | — | |
| InternVL-CZero-shot=true2024.02 | 0.562 | — | — | |
| DFN5B-CLIP-H/14+Zero-shot=true2024.02 | 0.561 | — | — | |
| EVA-CLIP-8BZero-shot=true2024.02 | 0.561 | — | — | |
| EVA-01-CLIP-g/14+Zero-shot=true2024.02 | 0.56 | — | — | |
| TIES-MergingMerging Strategy=TIES-Merging, Quantization Method=None, Model Backbone=CLIP-ViT-B/32, Bit-width=Full2026.05 | 0.549 | — | — | |
| nCLIPLinear probing=true, Backbone=ViT-B/16, Pre-train=IT35M2022.10 | 0.547 | — | — | |
| DFN5B-CLIP-H/14Zero-shot=true2024.02 | 0.547 | — | — | |
| LOUPEmode=zero-shot2022.08 | 0.533 | — | — | |
| EVA-01-CLIP-g/14Zero-shot=true2024.02 | 0.522 | — | — | |
| Simple AveragingMerging Strategy=Simple Averaging, Quantization Method=None, Model Backbone=CLIP-ViT-B/32, Bit-width=Full2026.05 | 0.516 | — | — | |
| Simple AveragingMerging Strategy=Simple Averaging, Quantization Method=Full-precision, Bit-width=Full, Backbone=CLIP-ViT-B/322026.05 | 0.502 | — | — | |
| CLIPBackbone=ViT-L/14, Zero-shot=true2025.07 | 0.49 | — | — | |
| TIES-MergingMerging Strategy=TIES-Merging, Quantization Method=Full-precision, Bit-width=Full, Backbone=CLIP-ViT-B/322026.05 | 0.473 | — | — | |
| Noun SubmanifoldBackbone=CLIP ViT-L-14, zero-shot=true2023.05 | 0.464 | — | — | |
| Task ArithmeticMerging Strategy=Task Arithmetic, Quantization Method=None, Model Backbone=CLIP-ViT-B/32, Bit-width=Full2026.05 | 0.461 | — | — | |
| Noun SubmanifoldBackbone=CLIP ViT-B-32, Evaluation Protocol=Zero-shot2023.05 | 0.447 | — | — | |
| SLIPZero-shot=true, Backbone=ViT-B/16, Pre-training Dataset=Laion100M, Epochs=302026.03 | 0.429 | — | — | |
| VLM BaselineIR=16, Backbone=VLM2026.01 | 0.414 | — | — | |
| Simple Averaging w/ AWQMerging Strategy=Simple Averaging, Quantization Method=AWQ, Model Backbone=CLIP-ViT-B/32, Bit-width=4-bit2026.05 | 0.4121 | — | — | |
| ITOZero-shot=true, Backbone=ViT-B/16, Pre-training Dataset=Laion100M, Epochs=302026.03 | 0.406 | — | — | |
| Simple Averaging + E-PMQMerging Strategy=Simple Averaging, Quantization Method=E-PMQ, Bit-width=4-bit, Backbone=CLIP-ViT-B/322026.05 | 0.4039 | — | — | |
| Simple Averaging + AWQMerging Strategy=Simple Averaging, Quantization Method=AWQ, Bit-width=4-bit, Backbone=CLIP-ViT-B/322026.05 | 0.4029 | — | — | |
| Simple Averaging w/ RTNMerging Strategy=Simple Averaging, Quantization Method=RTN, Model Backbone=CLIP-ViT-B/32, Bit-width=4-bit2026.05 | 0.3991 | — | — | |
| Simple Averaging + RTNMerging Strategy=Simple Averaging, Quantization Method=RTN, Bit-width=4-bit, Backbone=CLIP-ViT-B/322026.05 | 0.3982 | — | — | |
| TIES-Merging + E-PMQMerging Strategy=TIES-Merging, Quantization Method=E-PMQ, Bit-width=4-bit, Backbone=CLIP-ViT-B/322026.05 | 0.3951 | — | — | |
| Simple Averaging w/ E-PMQMerging Strategy=Simple Averaging, Quantization Method=E-PMQ, Model Backbone=CLIP-ViT-B/32, Bit-width=4-bit2026.05 | 0.3951 | — | — | |
| WUDI-Merging w/ E-PMQMerging Strategy=WUDI-Merging, Quantization Method=E-PMQ, Model Backbone=CLIP-ViT-B/32, Bit-width=4-bit2026.05 | 0.3924 | — | — | |
| WUDI-Merging + E-PMQMerging Strategy=WUDI-Merging, Quantization Method=E-PMQ, Bit-width=4-bit, Backbone=CLIP-ViT-B/322026.05 | 0.3898 | — | — | |
| TIES-Merging w/ E-PMQMerging Strategy=TIES-Merging, Quantization Method=E-PMQ, Model Backbone=CLIP-ViT-B/32, Bit-width=4-bit2026.05 | 0.388 | — | — | |
| Task Arithmetic w/ E-PMQMerging Strategy=Task Arithmetic, Quantization Method=E-PMQ, Model Backbone=CLIP-ViT-B/32, Bit-width=4-bit2026.05 | 0.3842 | — | — | |
| WUDI-Merging w/ RTNMerging Strategy=WUDI-Merging, Quantization Method=RTN, Model Backbone=CLIP-ViT-B/32, Bit-width=4-bit2026.05 | 0.3831 | — | — | |
| Simple Averaging + GPTQMerging Strategy=Simple Averaging, Quantization Method=GPTQ, Bit-width=4-bit, Backbone=CLIP-ViT-B/322026.05 | 0.3809 | — | — | |
| Simple Averaging w/ GPTQMerging Strategy=Simple Averaging, Quantization Method=GPTQ, Model Backbone=CLIP-ViT-B/32, Bit-width=4-bit2026.05 | 0.3802 | — | — | |
| POS PGABackbone=CLIP ViT-B-32, Evaluation Protocol=Zero-shot2023.05 | 0.38 | — | — | |
| Task Arithmetic + E-PMQMerging Strategy=Task Arithmetic, Quantization Method=E-PMQ, Bit-width=4-bit, Backbone=CLIP-ViT-B/322026.05 | 0.3792 | — | — | |
| WUDI-Merging w/ AWQMerging Strategy=WUDI-Merging, Quantization Method=AWQ, Model Backbone=CLIP-ViT-B/32, Bit-width=4-bit2026.05 | 0.3781 | — | — | |
| CLIPBackbone=CLIP ViT-B-32, Evaluation Protocol=Zero-shot2023.05 | 0.378 | — | — | |
| TIES-Merging w/ AWQMerging Strategy=TIES-Merging, Quantization Method=AWQ, Model Backbone=CLIP-ViT-B/32, Bit-width=4-bit2026.05 | 0.3773 | — | — | |
| POS PCABackbone=CLIP ViT-B-32, Evaluation Protocol=Zero-shot2023.05 | 0.376 | — | — | |
| WUDI-Merging w/ GPTQMerging Strategy=WUDI-Merging, Quantization Method=GPTQ, Model Backbone=CLIP-ViT-B/32, Bit-width=4-bit2026.05 | 0.3752 | — | — | |
| WUDI-Merging + RTNMerging Strategy=WUDI-Merging, Quantization Method=RTN, Bit-width=4-bit, Backbone=CLIP-ViT-B/322026.05 | 0.3667 | — | — | |
| CLIPZero-shot=true, Backbone=ViT-B/16, Pre-training Dataset=Laion100M, Epochs=302026.03 | 0.364 | — | — | |
| WUDI-Merging + AWQMerging Strategy=WUDI-Merging, Quantization Method=AWQ, Bit-width=4-bit, Backbone=CLIP-ViT-B/322026.05 | 0.3628 | — | — | |
| WUDI-Merging + GPTQMerging Strategy=WUDI-Merging, Quantization Method=GPTQ, Bit-width=4-bit, Backbone=CLIP-ViT-B/322026.05 | 0.3626 | — | — | |
| POS PGABackbone=CLIP ViT-L-14, zero-shot=true2023.05 | 0.358 | — | — | |
| CLIPBackbone=CLIP ViT-L-14, zero-shot=true2023.05 | 0.357 | — | — | |
| POS PCABackbone=CLIP ViT-L-14, zero-shot=true2023.05 | 0.357 | — | — | |
| TIES-Merging w/ RTNMerging Strategy=TIES-Merging, Quantization Method=RTN, Model Backbone=CLIP-ViT-B/32, Bit-width=4-bit2026.05 | 0.3566 | — | — | |
| Standard CLIPEvaluation Protocol=zero-shot2024.11 | 0.35 | — | — | |
| Task ArithmeticMerging Strategy=Task Arithmetic, Quantization Method=Full-precision, Bit-width=Full, Backbone=CLIP-ViT-B/322026.05 | 0.343 | — | — | |
| ITO sub2Backbone=ViT-L/16, Pre-training Dataset=DataComp-1B, Training Epochs=1, Evaluation Protocol=Zero-shot2026.03 | 0.337 | — | — | |
| CLIPBackbone=ViT-L/16, Pre-training Dataset=DataComp-1B, Training Epochs=1, Evaluation Protocol=Zero-shot2026.03 | 0.336 | — | — | |
| TIES-Merging + RTNMerging Strategy=TIES-Merging, Quantization Method=RTN, Bit-width=4-bit, Backbone=CLIP-ViT-B/322026.05 | 0.3282 | — | — | |
| SLIPPre-training Dataset=CC12M, Backbone=ViT-B/16, Epochs=30, Evaluation protocol=Zero-shot2026.03 | 0.317 | — | — | |
| Task Arithmetic w/ AWQMerging Strategy=Task Arithmetic, Quantization Method=AWQ, Model Backbone=CLIP-ViT-B/32, Bit-width=4-bit2026.05 | 0.3153 | — | — | |
| ITOBackbone=ViT-L/16, Pre-training Dataset=DataComp-1B, Training Epochs=1, Evaluation Protocol=Zero-shot2026.03 | 0.314 | — | — | |
| TIES-Merging w/ GPTQMerging Strategy=TIES-Merging, Quantization Method=GPTQ, Model Backbone=CLIP-ViT-B/32, Bit-width=4-bit2026.05 | 0.3136 | — | — | |
| Task Arithmetic w/ RTNMerging Strategy=Task Arithmetic, Quantization Method=RTN, Model Backbone=CLIP-ViT-B/32, Bit-width=4-bit2026.05 | 0.3132 | — | — | |
| TIES-Merging + AWQMerging Strategy=TIES-Merging, Quantization Method=AWQ, Bit-width=4-bit, Backbone=CLIP-ViT-B/322026.05 | 0.2983 | — | — | |
| CLIPPre-training Dataset=CC12M, Backbone=ViT-B/16, Epochs=30, Evaluation protocol=Zero-shot2026.03 | 0.296 | — | — | |
| Task Arithmetic w/ GPTQMerging Strategy=Task Arithmetic, Quantization Method=GPTQ, Model Backbone=CLIP-ViT-B/32, Bit-width=4-bit2026.05 | 0.2871 | — | — | |
| CLIPPre-train dataset=CC3M, Backbone=ViT-B/16, Training Epochs=30, Evaluation Protocol=Zero-shot2026.03 | 0.273 | — | — | |
| ITO sub2Pre-train dataset=CC3M, Backbone=ViT-B/16, Training Epochs=30, Evaluation Protocol=Zero-shot2026.03 | 0.269 | — | — | |
| Task Arithmetic + RTNMerging Strategy=Task Arithmetic, Quantization Method=RTN, Bit-width=4-bit, Backbone=CLIP-ViT-B/322026.05 | 0.2685 | — | — | |
| FLAIRPre-train dataset=CC3M, Backbone=ViT-B/16, Training Epochs=30, Evaluation Protocol=Zero-shot2026.03 | 0.268 | — | — | |
| TIES-Merging + GPTQMerging Strategy=TIES-Merging, Quantization Method=GPTQ, Bit-width=4-bit, Backbone=CLIP-ViT-B/322026.05 | 0.2523 | — | — | |
| Task Arithmetic + AWQMerging Strategy=Task Arithmetic, Quantization Method=AWQ, Bit-width=4-bit, Backbone=CLIP-ViT-B/322026.05 | 0.2492 | — | — | |
| Task Arithmetic + GPTQMerging Strategy=Task Arithmetic, Quantization Method=GPTQ, Bit-width=4-bit, Backbone=CLIP-ViT-B/322026.05 | 0.2304 | — | — | |
| B-cosified RN-50 CLIPPre-training Dataset=CC3M [47], Learning Scheduler=Cyclic, Evaluation Protocol=zero-shot2024.11 | 0.23 | — | — | |
| FLAIRPre-training Dataset=CC12M, Backbone=ViT-B/16, Epochs=30, Evaluation protocol=Zero-shot2026.03 | 0.228 | — | — | |
| ITOPre-train dataset=CC3M, Backbone=ViT-B/16, Training Epochs=30, Evaluation Protocol=Zero-shot2026.03 | 0.214 | — | — | |
| B-cosified RN-50 CLIPPre-training Dataset=ImageNet [16], Learning Scheduler=Cosine, Evaluation Protocol=zero-shot2024.11 | 0.21 | — | — | |
| SigLIPPre-training Dataset=CC12M, Backbone=ViT-B/16, Epochs=30, Evaluation protocol=Zero-shot2026.03 | 0.209 | — | — |