Image Classification on CIFAR-100 (train)
0.004Training LossAdamW
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| AdamWNetwork=ViT-7/8/12-7682025.05 | 0.004 | — | — | |
| AdamWNetwork=ViT-7/8/8-3842025.05 | 0.0042 | — | — | |
| AdamWNetwork=VGG-16/BN2025.05 | 0.0092 | — | — | |
| SAMNetwork=VGG-16/BN2025.05 | 0.0139 | — | — | |
| AdamWNetwork=ResNet-1102025.05 | 0.0149 | — | — | |
| Friendly-SAMNetwork=ViT-7/8/8-3842025.05 | 0.0228 | — | — | |
| SAMNetwork=ViT-7/8/12-7682025.05 | 0.0234 | — | — | |
| Friendly-SAMNetwork=ViT-7/8/12-7682025.05 | 0.0246 | — | — | |
| SAMNetwork=ViT-7/8/8-3842025.05 | 0.0255 | — | — | |
| ASAMNetwork=ViT-7/8/12-7682025.05 | 0.0347 | — | — | |
| ASAMNetwork=ViT-7/8/8-3842025.05 | 0.0349 | — | — | |
| ASAMNetwork=VGG-16/BN2025.05 | 0.0375 | — | — | |
| ZSharpNetwork=VGG-16/BN2025.05 | 0.0375 | — | — | |
| AdamWNetwork=ResNet-562025.05 | 0.0452 | — | — | |
| Friendly-SAMNetwork=VGG-16/BN2025.05 | 0.0495 | — | — | |
| Friendly-SAMNetwork=ResNet-1102025.05 | 0.0524 | — | — | |
| SAMNetwork=ResNet-1102025.05 | 0.0531 | — | — | |
| ZSharpNetwork=ViT-7/8/12-7682025.05 | 0.0709 | — | — | |
| ZSharpNetwork=ViT-7/8/8-3842025.05 | 0.073 | — | — | |
| ASAMNetwork=ResNet-1102025.05 | 0.0915 | — | — | |
| Friendly-SAMNetwork=ResNet-562025.05 | 0.1051 | — | — | |
| SAMNetwork=ResNet-562025.05 | 0.1097 | — | — | |
| ZSharpNetwork=ResNet-1102025.05 | 0.1656 | — | — | |
| ASAMNetwork=ResNet-562025.05 | 0.1952 | — | — | |
| ZSharpNetwork=ResNet-562025.05 | 0.251 | — | — | |
| DA-MuonArchitecture=ResNet-32, Mean base η=0.05002026.05 | 1.6241 | — | — | |
| DF-MuonArchitecture=ResNet-32, Mean base η=0.05002026.05 | 1.6271 | — | — | |
| SC-MuonArchitecture=ResNet-32, Mean base η=0.03502026.05 | 1.6321 | — | — | |
| Best fixed MuonArchitecture=ResNet-32, Mean base η=0.05002026.05 | 1.6682 | — | — | |
| AdamWArchitecture=ResNet-32, Mean base η=0.00302026.05 | 2.5096 | — | — | |
| Avg Pooling2013.01 | — | 11.2 | — | |
| Language Bias Bridge TrainingEvaluation Protocol=linear-probe, Pre-training=ImageNet-100, Backbone=GPT-2 baseline2026.04 | — | — | 21.8 | |
| MAEEvaluation Protocol=linear-probe, Pre-training=ImageNet-100, Backbone=GPT-2 baseline2026.04 | — | — | 24.5 | |
| Max Pooling2013.01 | — | 0.17 | — | |
| Rotation ImageEvaluation Protocol=linear-probe, Pre-training=ImageNet-100, Backbone=GPT-2 baseline2026.04 | — | — | 25.1 | |
| SimCLREvaluation Protocol=linear-probe, Pre-training=ImageNet-100, Backbone=GPT-2 baseline2026.04 | — | — | 16.2 | |
| Stochastic Pooling2013.01 | — | 21.22 | — | |
| VGG-10-KDBackbone=VGG-10, Augmentation=Knowledge Distillation (KD)2025.08 | — | — | 77.09 | |
| VGG-11-LEBackbone=VGG-11, Augmentation=Local Error (LE)2025.08 | — | — | 84.59 | |
| VGG-12-KDBackbone=VGG-12, Augmentation=Knowledge Distillation (KD)2025.08 | — | — | 79.79 | |
| VGG-13-LEBackbone=VGG-13, Augmentation=Local Error (LE)2025.08 | — | — | 75.56 |