Language Modeling on C4 LLaMA-130M (val)
18.504PerplexityALIAS Adam version
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| ALIAS Adam versionweight decay=wd2025.06 | 18.504 | — | 2.918 | — | |
| MOMOoptimizer=with Adam2025.06 | 18.634 | — | 2.925 | — | |
| D-ADAPTATIONoptimizer=with Adam2025.06 | 18.672 | — | 2.927 | — | |
| PRODIGY2025.06 | 18.727 | — | 2.93 | — | |
| DOG2025.06 | 18.897 | — | 2.939 | — | |
| SIGN-SGDLearning Rate (lr)=tuned, Weight Decay (wd)=active, Learning Rate Scheduler (cosine sc)=cosine2025.06 | 19.693 | — | 2.98 | — | |
| SIGN-SGDLearning Rate (lr)=tuned, Weight Decay (wd)=none, Learning Rate Scheduler (cosine sc)=cosine2025.06 | 19.923 | — | 2.992 | — | |
| NORMALIZED SGDLearning Rate (lr)=tuned, Weight Decay (wd)=active, Learning Rate Scheduler (cosine sc)=cosine2025.06 | 20.169 | — | 3.006 | — | |
| ALIASLearning Rate (lr)=parameter-free, Weight Decay (wd)=active, Learning Rate Scheduler (cosine sc)=none2025.06 | 20.169 | — | 3.006 | — | |
| ALIASLearning Rate (lr)=parameter-free, Weight Decay (wd)=none, Learning Rate Scheduler (cosine sc)=none2025.06 | 20.422 | — | 3.017 | — | |
| STEEPEST DESCENTLearning Rate (lr)=tuned, Weight Decay (wd)=active, Learning Rate Scheduler (cosine sc)=cosine2025.06 | 20.537 | — | 3.022 | — | |
| STEEPEST DESCENTLearning Rate (lr)=tuned, Weight Decay (wd)=none, Learning Rate Scheduler (cosine sc)=cosine2025.06 | 20.791 | — | 3.035 | — | |
| SIGN-SGDLearning Rate (lr)=tuned, Weight Decay (wd)=none, Learning Rate Scheduler (cosine sc)=constant2025.06 | 20.923 | — | 3.041 | — | |
| SIGN-SGDLearning Rate (lr)=tuned, Weight Decay (wd)=active, Learning Rate Scheduler (cosine sc)=constant2025.06 | 20.923 | — | 3.041 | — | |
| FOAM-2Training Tokens=2.6B, Precision=BF16, Folding Level=22025.12 | 22.51 | 0.54 | — | — | |
| FOAM-3Training Tokens=2.6B, Precision=BF16, Folding Level=32025.12 | 22.58 | 0.5 | — | — | |
| MUONTraining Tokens=2.6B, Precision=BF162025.12 | 22.75 | 0.63 | — | — | |
| Full-AdamTraining Tokens=2.6B, Precision=BF162025.12 | 22.86 | 0.8 | — | — | |
| Apollo2025.09 | 22.94 | 0.79 | — | 134 | |
| NORMALIZED SGDLearning Rate (lr)=tuned, Weight Decay (wd)=none, Learning Rate Scheduler (cosine sc)=cosine2025.06 | 22.982 | — | 3.135 | — | |
| FOAM-MiniTraining Tokens=2.6B, Precision=BF162025.12 | 23.1 | 0.46 | — | — | |
| APOLLO-1/4Training Tokens=2.6B, Precision=BF16, Projection Rank=d_model/42025.12 | 23.35 | 0.57 | — | — | |
| Adam-MiniTraining Tokens=2.6B, Precision=BF162025.12 | 23.73 | 0.53 | — | — | |
| APOLLO-1/8Training Tokens=2.6B, Precision=BF16, Projection Rank=d_model/82025.12 | 23.74 | 0.52 | — | — | |
| CR-NetAlignment Target=Memory overhead2025.09 | 23.74 | 0.79 | — | 106 | |
| APOLLO-MiniTraining Tokens=2.6B, Precision=BF162025.12 | 23.83 | 0.46 | — | — | |
| GWT-MiniTraining Tokens=2.6B, Precision=BF162025.12 | 23.84 | 0.46 | — | — | |
| CR-NetAlignment Target=Parameter complexity2025.09 | 24.31 | 0.67 | — | 90 | |
| Full-rank2025.09 | 24.36 | 1 | — | 134 | |
| CoLA2025.09 | 24.48 | 0.7 | — | 94 | |
| LORO2025.09 | 24.59 | 0.7 | — | 94 | |
| FLoRA2025.09 | 25.29 | 0.86 | — | 134 | |
| RSO2025.09 | 25.34 | 0.79 | — | 134 | |
| GaLore2025.09 | 25.36 | 0.79 | — | 134 | |
| VeLoRA2025.09 | 25.88 | 0.86 | — | 134 | |
| SLTrain2025.09 | 26.04 | 0.72 | — | 97 | |
| GaLore-1/4Training Tokens=2.6B, Precision=BF16, Projection Rank=d_model/42025.12 | 26.47 | 0.57 | — | — | |
| ReLoRA2025.09 | 29.37 | 0.86 | — | 134 | |
| GaLore-1/8Training Tokens=2.6B, Precision=BF16, Projection Rank=d_model/82025.12 | 30.02 | 0.52 | — | — | |
| LoRA2025.09 | 33.92 | 0.86 | — | 134 |