Image Classification on Oxford-IIIT Pets (test)
97.56Mean AccuracyVision Transformer (ViT-H/14)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Vision Transformer (ViT-H/14)Pre-training Dataset=JFT-300M, Architecture=ViT-H/142020.10 | 97.56 | — | |
| Vision Transformer (ViT-L/16)Pre-training Dataset=JFT-300M, Architecture=ViT-L/162020.10 | 97.32 | — | |
| SAM-finalBackbone=EfficientNet-L2, Fine-tuning strategy=With SAM optimization2021.02 | 97.1 | — | |
| SAM-baselineBackbone=EfficientNet-L2, Fine-tuning strategy=Without SAM optimization2021.02 | 96.92 | — | |
| JFT - Adaptive TransferBackbone=AmoebaNet-B, Input Resolution=331 x 3312018.11 | 96.8 | — | |
| BiT-LArchitecture=ResNet152x42020.10 | 96.62 | — | |
| BiT-LBackbone=ResNet152 x 42021.02 | 96.62 | — | |
| JFT - AnimalBackbone=AmoebaNet-B, Input Resolution=331 x 3312018.11 | 96.4 | — | |
| ALIGNBackbone=EfficientNet-L2, Resolution=289/3602021.02 | 96.19 | — | |
| Best Published Result [14]Backbone=AmoebaNet-B, Input Resolution=480 x 4802018.11 | 95.9 | — | |
| GPipe (AmoebaNet-B (18, 512))Architecture=AmoebaNet-B (18, 512), Resolution=480x480, Protocol=Fine-tuned, Number of fine-tuning runs=5, Crop=Single-crop2018.11 | 95.9 | — | |
| 2SFSBackbone=ViT-L/14, Shots=162025.03 | 95.5 | — | |
| ImageNet - Adaptive TransferBackbone=AmoebaNet-B, Input Resolution=331 x 3312018.11 | 95.1 | — | |
| TNT-BParams (M)=65.6, Resolution=384x384, Pre-training=ImageNet2021.02 | 95 | — | |
| LowFormer-B3GPU Throughput (images/sec)=424, Resolution=384x384, Fine-tuned=true2026.03 | 95 | — | |
| CeiT-SGPU Throughput (images/sec)=260, Resolution=384x384, Fine-tuned=true2026.03 | 94.9 | — | |
| ViT-H/16Param (M)=632, Pre-trained=ImageNet-22k2021.03 | 94.82 | — | |
| ViT-L/16Param (M)=307, Pre-trained=ImageNet-22k2021.03 | 94.73 | — | |
| CvT-W24Param (M)=277, Pre-trained=ImageNet-22k2021.03 | 94.73 | — | |
| TNT-SParams (M)=23.8, Resolution=384x384, Pre-training=ImageNet2021.02 | 94.7 | — | |
| TNT-SGPU Throughput (images/sec)=141, Resolution=384x384, Fine-tuned=true2026.03 | 94.7 | — | |
| Vision Transformer (ViT-L/16)Pre-training Dataset=ImageNet-21k, Architecture=ViT-L/162020.10 | 94.67 | — | |
| Entire JFT DatasetBackbone=AmoebaNet-B, Input Resolution=331 x 3312018.11 | 94.5 | — | |
| ImageNet - Entire DatasetBackbone=AmoebaNet-B, Input Resolution=331 x 3312018.11 | 94.5 | — | |
| BiT-MParam (M)=928, Pre-trained=ImageNet-22k2021.03 | 94.46 | — | |
| ViT-B/16Param (M)=86, Pre-trained=ImageNet-22k2021.03 | 94.43 | — | |
| ViTAE-SParams (M)=23.6, Fine-tuning=true2021.06 | 94.2 | — | |
| MoCo v3Backbone=ViT-H, Pre-training=ImageNet-1k Self-Supervised, Evaluation Protocol=End-to-end fine-tuning2021.04 | 94.2 | — | |
| CvT-21Param (M)=32, Pre-trained=ImageNet-22k2021.03 | 94.03 | — | |
| FTMod.=-, Params=86M2024.07 | 93.9 | — | |
| Method [29]2018.11 | 93.8 | — | |
| ViT-B/16Params (M)=86.4, Resolution=384x384, Pre-training=ImageNet2021.02 | 93.8 | — | |
| ViT-B/16Params (M)=86.5, Fine-tuning=true2021.06 | 93.8 | — | |
| ImNet supervisedBackbone=ViT-B, Pre-training=ImageNet-1k Supervised, Evaluation Protocol=End-to-end fine-tuning2021.04 | 93.8 | — | |
| ViT-B/16GPU Throughput (images/sec)=117, Resolution=384x384, Fine-tuned=true2026.03 | 93.8 | — | |
| 2SFSBackbone=ViT-B/16, Shots=162025.03 | 93.7 | — | |
| MoCo v3Backbone=ViT-L, Pre-training=ImageNet-1k Self-Supervised, Evaluation Protocol=End-to-end fine-tuning2021.04 | 93.7 | — | |
| ViT-L/16Params (M)=304.3, Fine-tuning=true2021.06 | 93.6 | — | |
| ImNet supervisedBackbone=ViT-L, Pre-training=ImageNet-1k Supervised, Evaluation Protocol=End-to-end fine-tuning2021.04 | 93.6 | — | |
| ViT-L/16GPU Throughput (images/sec)=36, Resolution=384x384, Fine-tuned=true2026.03 | 93.6 | — | |
| ResNet-152-SAMBackbone=ResNet-152, SAM Pre-training=true2021.06 | 93.3 | — | |
| CvT-13Param (M)=20, Pre-trained=ImageNet-22k2021.03 | 93.25 | — | |
| MoCo v3Backbone=ViT-B, Pre-training=ImageNet-1k Self-Supervised, Evaluation Protocol=End-to-end fine-tuning2021.04 | 93.2 | — | |
| VPTMod.=V, Params=73K2024.07 | 93.2 | — | |
| Full TrainingBackbone=ViT2025.10 | 93.15 | — | |
| ViT-B/16-SAMBackbone=ViT-B/16, SAM Pre-training=true2021.06 | 93.1 | — | |
| CoCoOpMod.=L, Params=35K2024.07 | 93 | — | |
| ViT-S/16-SAMBackbone=ViT-S/16, SAM Pre-training=true2021.06 | 92.9 | — | |
| MaPLeMod.=L & V, Params=1.2M2024.07 | 92.9 | — | |
| ViTAE-TParams (M)=4.8, Fine-tuning=true2021.06 | 92.6 | — | |
| Mixer-B/16-SAMBackbone=Mixer-B/16, SAM Pre-training=true2021.06 | 92.5 | — | |
| ReLICv2Evaluation protocol=Linear evaluation, Backbone=ResNet-50, Pre-trained=ImageNet2022.01 | 92.4 | — | |
| AdaBetBackbone=ViT2025.10 | 92.23 | — | |
| ReLICv2Evaluation protocol=Fine-tuned, Backbone=ResNet-50, Pre-trained=ImageNet2022.01 | 92.2 | — | |
| Supervised-INEvaluation protocol=Fine-tuned, Backbone=ResNet-50, Pre-trained=ImageNet2022.01 | 92.1 | — | |
| Elastic TrainerBackbone=ViT2025.10 | 91.96 | — | |
| ViT-B/16Backbone=ViT-B/16, SAM Pre-training=false2021.06 | 91.9 | — | |
| NNCLREvaluation protocol=Linear evaluation, Backbone=ResNet-50, Pre-trained=ImageNet2022.01 | 91.8 | — | |
| NNCLRProtocol=Linear classification2022.12 | 91.8 | — | |
| BYOLEvaluation protocol=Fine-tuned, Backbone=ResNet-50, Pre-trained=ImageNet2022.01 | 91.7 | — | |
| CoOpMod.=L, Params=9K2024.07 | 91.7 | — | |
| ResNet-50-SAMBackbone=ResNet-50, SAM Pre-training=true2021.06 | 91.6 | — | |
| Supervised-INEvaluation protocol=Linear evaluation, Backbone=ResNet-50, Pre-trained=ImageNet2022.01 | 91.5 | — | |
| SupervisedProtocol=Linear classification2022.12 | 91.5 | — | |
| BlackVIPMod.=V, Params=150K2024.07 | 91.4 | — | |
| LoRAMod.=V, Params=150K2024.07 | 91 | — | |
| SCEProtocol=Linear classification2022.12 | 90.9 | — | |
| BlackVIPMod.=V, Params=9K2024.07 | 90.8 | — | |
| PruneTrainBackbone=ViT2025.10 | 90.57 | — | |
| BlackVIPMod.=V, Params=68K2024.07 | 90.5 | — | |
| JFT - BirdBackbone=AmoebaNet-B, Input Resolution=331 x 3312018.11 | 90.4 | — | |
| ViT-S/16Backbone=ViT-S/16, SAM Pre-training=false2021.06 | 90.4 | — | |
| BYOLEvaluation protocol=Linear evaluation, Backbone=ResNet-50, Pre-trained=ImageNet2022.01 | 90.4 | — | |
| BYOLProtocol=Linear classification2022.12 | 90.4 | — | |
| 2SFSBackbone=ViT-B/32, Shots=162025.03 | 90.3 | — | |
| Transfer LearningBackbone=ViT2025.10 | 90.21 | — | |
| VPMod.=V, Params=69K2024.07 | 90.2 | — | |
| ProDAshots=162022.05 | 90 | — | |
| BlackVIP-SEMod.=V, Params=1K2024.07 | 90 | — | |
| AdaBetBackbone=MobileNetV22025.10 | 89.93 | — | |
| LoRAMod.=L, Params=150K2024.07 | 89.7 | — | |
| ProDAshots=82022.05 | 89.4 | — | |
| JFT - FoodBackbone=AmoebaNet-B, Input Resolution=331 x 3312018.11 | 89.2 | — | |
| SimCLREvaluation protocol=Fine-tuned, Backbone=ResNet-50, Pre-trained=ImageNet2022.01 | 89.2 | — | |
| JFT - TransportBackbone=AmoebaNet-B, Input Resolution=331 x 3312018.11 | 89.1 | — | |
| HisTPTBackbone=ViT-B/16, Protocol=Continuous Test-time Prompt Tuning2024.10 | 89.1 | — | |
| ZSMod.=-, Params=-2024.07 | 89.1 | — | |
| ProDAshots=42022.05 | 89 | — | |
| LoRAMod.=L & V, Params=150K2024.07 | 89 | — | |
| Last-K LayersBackbone=ViT2025.10 | 88.97 | — | |
| JFT - CarBackbone=AmoebaNet-B, Input Resolution=331 x 3312018.11 | 88.9 | — | |
| CLIPDataset size=400M (private), Visual encoder=ViT-B/16, Zero-shot=true2022.10 | 88.9 | — | |
| AdaBetBackbone=ResNet502025.10 | 88.74 | — | |
| Mixer-S/16-SAMBackbone=Mixer-S/16, SAM Pre-training=true2021.06 | 88.7 | — | |
| JFT - VehicleBackbone=AmoebaNet-B, Input Resolution=331 x 3312018.11 | 88.6 | — | |
| ViTModel variant=S, # parameters=22.05M, Pre-training=ImageNet, Evaluation Protocol=Fine-tuning2023.06 | 88.6 | — | |
| ViTModel variant=T, # parameters=5.72M, Pre-training=ImageNet, Evaluation Protocol=Fine-tuning2023.06 | 88.5 | — | |
| CoOpshots=162022.05 | 88.4 | — | |
| ProDAshots=22022.05 | 88.4 | — | |
| ProDAshots=12022.05 | 88.2 | — |