Image Classification on ImageNet-1K (Top-1 Accuracy)
97.1Top-1 AccAutoV
Evaluation Results
| Method | Links | |
|---|---|---|
| AutoVModel=Qwen2.5-VL 7B2025.06 | 97.1 | |
| APIModel=Qwen2.5-VL 7B2025.06 | 95.4 | |
| AutoVModel=LLaVA-OneVision 7B2025.06 | 95 | |
| BaseModel=Qwen2.5-VL 7B2025.06 | 94.7 | |
| APIModel=LLaVA-OneVision 7B2025.06 | 92.6 | |
| BaseModel=LLaVA-OneVision 7B2025.06 | 92.3 | |
| Absolute SOTAReference=[82]2023.05 | 91 | |
| RevColV2-L↑Size=512^2, Target=Pixel, Params=327M, FLOPs=417G, Training Protocol=ImageNet-1K pre-train + 22K intermediate fine-tune2023.09 | 88.4 | |
| RevColV2-L↑Size=384^2, Target=Pixel, Params=327M, FLOPs=215G, Training Protocol=ImageNet-1K pre-train + 22K intermediate fine-tune2023.09 | 88.3 | |
| ConvNeXt V2-L↑Size=384^2, Target=Pixel, Params=198M, FLOPs=103G, Training Protocol=ImageNet-1K pre-train + 22K intermediate fine-tune2023.09 | 88.2 | |
| Siglip2Type=Discriminative Pretraining, Pretrained on=Images, Arch.=ViT-G, Param.=1.2B, Attentive probing=true2025.12 | 88 | |
| SynCLR + GMAILEvaluation Protocol=Fine-tuning2026.02 | 87.95 | |
| SwinV2-L↑Size=384^2, Target=Label, Params=197M, FLOPs=115G, Training Protocol=ImageNet-1K pre-train + 22K intermediate fine-tune2023.09 | 87.6 | |
| RevCol-L↑Size=384^2, Target=Label, Params=273M, FLOPs=116G, Training Protocol=ImageNet-1K pre-train + 22K intermediate fine-tune2023.09 | 87.6 | |
| RevColV2-B↑Size=512^2, Target=Pixel, Params=88M, FLOPs=130G, Training Protocol=ImageNet-1K pre-train + 22K intermediate fine-tune2023.09 | 87.5 | |
| RevColV2-LSize=224^2, Target=Pixel, Params=327M, FLOPs=67G, Training Protocol=ImageNet-1K pre-train + 22K intermediate fine-tune2023.09 | 87.4 | |
| RevColV2-B↑Size=384^2, Target=Pixel, Params=88M, FLOPs=64G, Training Protocol=ImageNet-1K pre-train + 22K intermediate fine-tune2023.09 | 87.3 | |
| ConvNeXt V2-LSize=224^2, Target=Pixel, Params=198M, FLOPs=34G, Training Protocol=ImageNet-1K pre-train + 22K intermediate fine-tune2023.09 | 87.3 | |
| VIC-MAEArch.=ViT-L, Pre-training Data=K710 + MiT + IN1K, Evaluation Protocol=Self-Supervised2023.03 | 87.1 | |
| DeiT III-LSize=224^2, Target=Label, Params=304M, FLOPs=62G, Training Protocol=ImageNet-1K pre-train + 22K intermediate fine-tune2023.09 | 87 | |
| ViT-Hpretrain=MAE, FLOPs (G)=167, Param=632M2023.06 | 86.9 | |
| Hiera-Hpretrain=MAE, FLOPs (G)=125, Param=673M2023.06 | 86.9 | |
| MAEArch.=ViT-H, Pretrain Data=IN1K2022.06 | 86.9 | |
| SwinV2-LSize=256^2, Target=Label, Params=197M, FLOPs=48G, Training Protocol=ImageNet-1K pre-train + 22K intermediate fine-tune2023.09 | 86.9 | |
| ConvNeXt-BSize=384^2, Target=Pixel, Params=89M, FLOPs=45G, Training Protocol=ImageNet-1K pre-train + 22K intermediate fine-tune2023.09 | 86.8 | |
| MOAT-3Size=224^2, Target=Label, Params=190M, FLOPs=45G, Training Protocol=ImageNet-1K pre-train + 22K intermediate fine-tune2023.09 | 86.8 | |
| RevCol-B↑Size=384^2, Target=Label, Params=138M, FLOPs=49G, Training Protocol=ImageNet-1K pre-train + 22K intermediate fine-tune2023.09 | 86.7 | |
| OmniMAEArch.=ViT-H, Pretrain Data=IN1K + SSv22022.06 | 86.6 | |
| RevCol-LSize=224^2, Target=Label, Params=273M, FLOPs=39G, Training Protocol=ImageNet-1K pre-train + 22K intermediate fine-tune2023.09 | 86.6 | |
| COVERArch.=TimeSFormer-SR, Pre-training Data=JFT-3B+ K400+ MiT + IN1K, Evaluation Protocol=Supervised2023.03 | 86.6 | |
| C-RADIOv4Variant=H, Params=631M, Evaluation Protocol=kNN2026.01 | 86.59 | |
| SigLIP2Variant=g, Params=1,164M, Evaluation Protocol=kNN2026.01 | 86.39 | |
| RevColV2-LSize=224^2, Target=Pixel, Params=327M, FLOPs=67G, Training Protocol=ImageNet-1K pre-train2023.09 | 86.3 | |
| CoCaPPT=2B, TPU-DAYS=10k, Zero-shot=true2023.05 | 86.3 | |
| C-RADIOv3Variant=H, Params=631M, Evaluation Protocol=kNN2026.01 | 86.23 | |
| MCMAE-Lpretrain=MCMAE, FLOPs (G)=94, Param=323M2023.06 | 86.2 | |
| CAE-LSize=224^2, Target=DALL-E, Params=307M, FLOPs=62G, Training Protocol=ImageNet-1K pre-train2023.09 | 86.2 | |
| RevColV2-BSize=224^2, Target=Pixel, Params=88M, FLOPs=19G, Training Protocol=ImageNet-1K pre-train + 22K intermediate fine-tune2023.09 | 86.2 | |
| C-JEPAPretrain Epochs=600, Backbone=ViT-L/16, Evaluation Protocol=fine-tuning2024.10 | 86.2 | |
| DINOv2Type=Discriminative Pretraining, Pretrained on=Images, Arch.=ViT-G, Param.=1.1B, Attentive probing=true2025.12 | 86.2 | |
| Hiera-Lpretrain=MAE, FLOPs (G)=40, Param=214M2023.06 | 86.1 | |
| HIVEFoundation Model=SigLIP2026.03 | 86.06 | |
| SAFoundation Model=SigLIP2026.03 | 86.04 | |
| OMNIVOREArch.=ViT-L, Pre-training Data=IN1K + K400 + SUN RGB-D, Evaluation Protocol=Supervised2023.03 | 86 | |
| VIC-MAEArch.=ViT-L, Pre-training Data=K400 + IN1K, Evaluation Protocol=Self-Supervised2023.03 | 86 | |
| BaseFoundation Model=SigLIP2026.03 | 85.99 | |
| ViT-Lpretrain=MAE, FLOPs (G)=62, Param=304M2023.06 | 85.9 | |
| MAEArch.=ViT-L, Pretrain Data=IN1K2022.06 | 85.9 | |
| MAE-LSize=224^2, Target=Pixel, Params=307M, FLOPs=62G, Training Protocol=ImageNet-1K pre-train2023.09 | 85.9 | |
| ImageMAEStudent=ViT-L, Teacher=null2023.11 | 85.9 | |
| RADIOv2.5Variant=H, Params=631M, Evaluation Protocol=kNN2026.01 | 85.81 | |
| ConvNeXt V2-LSize=224^2, Target=Pixel, Params=198M, FLOPs=34G, Training Protocol=ImageNet-1K pre-train2023.09 | 85.8 | |
| ConvNeXt-BSize=224^2, Target=Pixel, Params=89M, FLOPs=15G, Training Protocol=ImageNet-1K pre-train + 22K intermediate fine-tune2023.09 | 85.8 | |
| InternVideo2Type=Discriminative Pretraining, Pretrained on=Videos, Arch.=ViT-G, Param.=1B, Attentive probing=true2025.12 | 85.8 | |
| SynCLREvaluation Protocol=Fine-tuning2026.02 | 85.8 | |
| DINOv3Variant=H+, Params=841M, Evaluation Protocol=kNN2026.01 | 85.77 | |
| SigLIP2Variant=SO400M, Params=412M, Evaluation Protocol=kNN2026.01 | 85.76 | |
| C-RADIOv4Variant=SO400M, Params=412M, Evaluation Protocol=kNN2026.01 | 85.76 | |
| ViT-Lpretrain=MaskFeat, FLOPs (G)=62, Param=304M2023.06 | 85.7 | |
| MaskFeat-LSize=224^2, Target=HOG, Params=307M, FLOPs=62G, Training Protocol=ImageNet-1K pre-train2023.09 | 85.7 | |
| DeiT III-BSize=224^2, Target=Label, Params=87M, FLOPs=18G, Training Protocol=ImageNet-1K pre-train + 22K intermediate fine-tune2023.09 | 85.7 | |
| VanillaBackbone=DeiTIII, Resolution=2242024.05 | 85.66 | |
| FlexiViTBackbone=DeiTIII, Resolution=2242024.05 | 85.66 | |
| MSPEBackbone=DeiTIII, Resolution=224, Training Epochs=52024.05 | 85.66 | |
| RevCol-BSize=224^2, Target=Label, Params=138M, FLOPs=17G, Training Protocol=ImageNet-1K pre-train + 22K intermediate fine-tune2023.09 | 85.6 | |
| MAE-LModel Scale=ViT-L, Pretrain Task=masked pixel pred, Pretrain Framework=masked autoencoder, Decoder=transformer decoder, # FWD / step=1, Epochs=1600, Causal attention during fine-tuning=false, Implementation=Authors Implementation2025.12 | 85.6 | |
| ScaleKDBackbone=ViT-B/16, Pre-training=Ours2024.11 | 85.53 | |
| FlexiViTBackbone=DeiTIII, Resolution=4482024.05 | 85.53 | |
| MSPEBackbone=DeiTIII, Resolution=448, Training Epochs=52024.05 | 85.53 | |
| MAEArch.=ViT-L, Pre-training Data=IN1K, Evaluation Protocol=Self-Supervised2023.03 | 85.5 | |
| SegMAN-L EncoderParams (M)=81, GFLOPs=16.8, Resolution=224x2242024.12 | 85.5 | |
| DINOv3Variant=7B, Params=6,716M, Evaluation Protocol=kNN2026.01 | 85.42 | |
| Swin-Lpretrain=SimMIM, FLOPs (G)=36, Param=197M2023.06 | 85.4 | |
| SimMIM-LSize=224^2, Target=Pixel, Params=197M, FLOPs=35G, Training Protocol=ImageNet-1K pre-train2023.09 | 85.4 | |
| MViTv2-LFLOPs (G)=42, Param=218M2023.06 | 85.3 | |
| VIC-MAEArch.=ViT-L, Pre-training Data=MiT, Evaluation Protocol=Self-Supervised2023.03 | 85.3 | |
| I-JEPAPretrain Epochs=600, Backbone=ViT-L/16, Evaluation Protocol=fine-tuning2024.10 | 85.3 | |
| OpenCLIPType=Discriminative Pretraining, Pretrained on=Images, Arch.=ViT-G, Param.=1.8B, Attentive probing=true2025.12 | 85.3 | |
| NEPA-LModel Scale=ViT-L, Pretrain Task=autoreg. embed pred, Pretrain Framework=autoregression, Decoder=none, # FWD / step=1, Epochs=800, Causal attention during fine-tuning=false2025.12 | 85.3 | |
| Hiera-B+pretrain=MAE, FLOPs (G)=13, Param=70M2023.06 | 85.2 | |
| ViT-Lpretrain=BEiT, DALLE, FLOPs (G)=62, Param=304M2023.06 | 85.2 | |
| BEITArch.=ViT-L, Pretrain Data=IN1K2022.06 | 85.2 | |
| OmniMAEArch.=ViT-L, Pretrain Data=IN1K + SSv22022.06 | 85.2 | |
| BEIT-LSize=224^2, Target=Pixel, Params=307M, FLOPs=62G, Training Protocol=ImageNet-1K pre-train2023.09 | 85.2 | |
| ViT-LSize=384^2, Target=Label, Params=307M, FLOPs=191G, Training Protocol=ImageNet-1K pre-train + 22K intermediate fine-tune2023.09 | 85.2 | |
| BEiT-LModel Scale=ViT-L, Pretrain Task=masked token pred, Pretrain Framework=masked modeling, Decoder=linear pred. head, # FWD / step=1, Epochs=800, Causal attention during fine-tuning=false2025.12 | 85.2 | |
| JEPA-LModel Scale=ViT-L, Pretrain Task=masked embed pred, Pretrain Framework=siamese & masked modeling, Decoder=transformer predictor, # FWD / step=2, Epochs=300, Causal attention during fine-tuning=false, Implementation=Authors Implementation2025.12 | 85.2 | |
| SigLIP-L/14Params (M)=428, Protocol=k-NN2025.08 | 85.16 | |
| SigLIP-L/14Params (M)=428, Evaluation Protocol=k-NN2025.08 | 85.16 | |
| FlexiViTBackbone=ViT, Resolution=4482024.05 | 85.11 | |
| MSPEBackbone=ViT, Resolution=448, Training Epochs=52024.05 | 85.11 | |
| SegMAN-B EncoderParams (M)=45, GFLOPs=9.9, Resolution=224x2242024.12 | 85.1 | |
| VanillaBackbone=ViT, Resolution=2242024.05 | 85.1 | |
| FlexiViTBackbone=ViT, Resolution=2242024.05 | 85.1 | |
| MSPEBackbone=ViT, Resolution=224, Training Epochs=52024.05 | 85.1 | |
| MEDiCBackbone=ViT-B, Epochs=300, Evaluation Protocol=Fine-tuning2026.03 | 85.07 | |
| MCMAE-Bpretrain=MCMAE, FLOPs (G)=28, Param=88M2023.06 | 85 | |
| VIC-MAEArch.=ViT-L, Pre-training Data=K400, Evaluation Protocol=Self-Supervised2023.03 | 85 | |
| SoftmaxBackbone=DINOv2-B2026.03 | 84.93 | |
| ConvNeXt V2-BSize=224^2, Target=Pixel, Params=89M, FLOPs=15G, Training Protocol=ImageNet-1K pre-train2023.09 | 84.9 |