Language Modeling on Pre-training corpus
1.577LossMuon
Evaluation Results
| Method | Links | |
|---|---|---|
| MuonOptimizer=Muon, Model Size=3B2026.04 | 1.577 | |
| Adam+NexusOptimizer=Adam+Nexus, Model Size=3B2026.04 | 1.602 | |
| AdamWOptimizer=AdamW, Model Size=3B2026.04 | 1.606 | |
| Panda-3Bd_model=4096, f_size=4096, n_layers=28, GQA=3, d_model / sqrt(N)=0.077, r=12025.10 | 2.619 | |
| Surefire-3Bd_model=4096, f_size=4096, n_layers=28, GQA=7, d_model / sqrt(N)=0.077, r=12025.10 | 2.62 | |
| LLaMA-3.2-3Bd_model=3072, f_size=8192, n_layers=28, GQA=3, d_model / sqrt(N)=0.058, r=4.802025.10 | 2.625 | |
| Panda-1Bd_model=2560, f_size=4096, n_layers=16, GQA=4, d_model / sqrt(N)=0.082, r=1.072025.10 | 2.782 | |
| LLaMA-3.2-1Bd_model=2048, f_size=8192, n_layers=16, GQA=4, d_model / sqrt(N)=0.066, r=4.802025.10 | 2.803 | |
| Surefire-1Bd_model=2560, f_size=6144, n_layers=16, GQA=9, d_model / sqrt(N)=0.082, r=3.62025.10 | 2.804 |