Image Classification on ImageNet (Accuracy and Time)
79.1AccuracyDINOv2-L
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DINOv2-LBackbone=ViT-B, Protocol=Knowledge distillation (online), Compute Setting=full training; ≈ 150% compute, Epochs=1002026.05 | 79.1 | 47.8 | |
| TextTeacher (sample)Backbone=ViT-B, Protocol=Text Guided (offline), Epochs=1002026.05 | 79.1 | 32.1 | |
| BorLan (distribution)Backbone=ViT-B, Protocol=Text Guided (offline), Epochs=1002026.05 | 79 | 32.9 | |
| DINOv2-LBackbone=ViT-B, Protocol=Knowledge distillation (online), Compute Setting=compute matched, Epochs=1002026.05 | 78.9 | 47.8 | |
| CLIP-ViT-LBackbone=ViT-B, Protocol=Vision Guided (offline), Epochs=1002026.05 | 78.8 | 32.8 | |
| CoCaBackbone=ViT-B, Protocol=Knowledge distillation (online), Compute Setting=full training; ≈ 150% compute, Epochs=1002026.05 | 78.7 | 50.3 | |
| CLIP-ViT-LBackbone=ViT-B, Protocol=Knowledge distillation (online), Compute Setting=full training; ≈ 150% compute, Epochs=1002026.05 | 78.6 | 48.2 | |
| CLIP-ViT-BBackbone=ViT-B, Protocol=Vision Guided (offline), Epochs=1002026.05 | 78.6 | 33.1 | |
| DINOv2-LBackbone=ViT-B, Protocol=Vision Guided (offline), Epochs=1002026.05 | 78.6 | 31.2 | |
| CoCaBackbone=ViT-B, Protocol=Vision Guided (offline), Epochs=1002026.05 | 78.5 | 31.9 | |
| DINOv2-LBackbone=ViT-S, Protocol=Knowledge distillation (online), Compute Setting=full training; ≈ 150% compute, Epochs=1002026.05 | 78.4 | — | |
| DINOv2-BBackbone=ViT-B, Protocol=Vision Guided (offline), Epochs=1002026.05 | 78.4 | 30.6 | |
| TextTeacher (sample)Backbone=ViT-S, Protocol=Text Guided (offline), Epochs=1002026.05 | 78.4 | — | |
| BorLan + adaptive weightBackbone=ViT-B, Protocol=Text Guided (offline), Epochs=1002026.05 | 78.2 | 31.1 | |
| CLIP-ViT-LBackbone=ViT-B, Protocol=Knowledge distillation (online), Compute Setting=compute matched, Epochs=1002026.05 | 78.1 | 48.2 | |
| CoCaBackbone=ViT-S, Protocol=Vision Guided (offline), Epochs=1002026.05 | 78 | — | |
| CLIP-ViT-BBackbone=ViT-S, Protocol=Vision Guided (offline), Epochs=1002026.05 | 77.9 | — | |
| DINOv2-BBackbone=ViT-S, Protocol=Vision Guided (offline), Epochs=1002026.05 | 77.9 | — | |
| CoCaBackbone=ViT-S, Protocol=Knowledge distillation (online), Compute Setting=full training; ≈ 150% compute, Epochs=1002026.05 | 77.8 | — | |
| CLIP-ViT-LBackbone=ViT-S, Protocol=Vision Guided (offline), Epochs=1002026.05 | 77.8 | — | |
| DINOv2-LBackbone=ViT-S, Protocol=Vision Guided (offline), Epochs=1002026.05 | 77.8 | — | |
| BorLan + adaptive weightBackbone=ViT-S, Protocol=Text Guided (offline), Epochs=1002026.05 | 77.7 | — | |
| BaselineBackbone=ViT-S, Protocol=Classification Only, Epochs=1002026.05 | 77.6 | — | |
| CLIP-ViT-LBackbone=ViT-S, Protocol=Knowledge distillation (online), Compute Setting=full training; ≈ 150% compute, Epochs=1002026.05 | 77.6 | — | |
| CoCaBackbone=ViT-B, Protocol=Knowledge distillation (online), Compute Setting=compute matched, Epochs=1002026.05 | 77.6 | 50.3 | |
| VL2Lite (text only)Backbone=ViT-S, Protocol=Text Guided (offline), Epochs=1002026.05 | 77 | — | |
| BorLan (distribution)Backbone=ViT-S, Protocol=Text Guided (offline), Epochs=1002026.05 | 76.7 | — | |
| BaselineBackbone=ViT-B, Protocol=Classification Only, Epochs=1002026.05 | 76.5 | 30.7 | |
| VL2LiteBackbone=ViT-B, Protocol=Knowledge distillation (online), Compute Setting=full training; ≈ 150% compute, Epochs=1002026.05 | 76.2 | 47.8 | |
| FedMAR(Ours)Architecture=Tiny-ViT, Params=11.60M, GFLOPS=0.88, Pre-trained Dataset=Mini-ImageNet2026.07 | 75.99 | — | |
| VL2LiteBackbone=ViT-S, Protocol=Knowledge distillation (online), Compute Setting=full training; ≈ 150% compute, Epochs=1002026.05 | 75.7 | — | |
| VL2Lite (text only)Backbone=ViT-B, Protocol=Text Guided (offline), Epochs=1002026.05 | 75.5 | 31.4 | |
| VL2LiteBackbone=ViT-B, Protocol=Knowledge distillation (online), Compute Setting=compute matched, Epochs=1002026.05 | 75.2 | 47.8 | |
| DINOv2-LBackbone=ViT-S, Protocol=Knowledge distillation (online), Compute Setting=compute matched, Epochs=1002026.05 | 74.3 | — | |
| CLIP-ViT-LBackbone=ViT-S, Protocol=Knowledge distillation (online), Compute Setting=compute matched, Epochs=1002026.05 | 73.8 | — | |
| CoCaBackbone=ViT-S, Protocol=Knowledge distillation (online), Compute Setting=compute matched, Epochs=1002026.05 | 73.5 | — | |
| VL2LiteBackbone=ViT-S, Protocol=Knowledge distillation (online), Compute Setting=compute matched, Epochs=1002026.05 | 68.4 | — | |
| FeatARCArchitecture=ResNet-18, Params=11.70M, GFLOPS=1.83, Pre-trained Dataset=Mini-ImageNet2026.07 | 68.17 | — | |
| FedMAR(Ours)Architecture=ResNet-18, Params=22.50M, GFLOPS=3.64, Pre-trained Dataset=Mini-ImageNet2026.07 | 65.36 | — | |
| FedEMAArchitecture=ResNet-18, Params=38.47M, GFLOPS=7.40, Pre-trained Dataset=Mini-ImageNet2026.07 | 65.24 | — | |
| FedUArchitecture=ResNet-18, Params=38.47M, GFLOPS=7.40, Pre-trained Dataset=Mini-ImageNet2026.07 | 65.1 | — | |
| OrchestraArchitecture=ResNet-18, Params=11.84M, GFLOPS=7.31, Pre-trained Dataset=Mini-ImageNet2026.07 | 65.02 | — | |
| LDAWAArchitecture=ResNet-18, Params=15.39M, GFLOPS=1.83, Pre-trained Dataset=Mini-ImageNet2026.07 | 51.43 | — | |
| FedU2Architecture=ResNet-18, Params=15.39M, GFLOPS=1.83, Pre-trained Dataset=Mini-ImageNet2026.07 | 45.27 | — |