Language Modeling on Code 24B tokens
0.6994Cross-Entropy LossMuSe†
Evaluation Results
| Method | Links | |
|---|---|---|
| MuSe†Model Size=1B, Train Attention=MuSe, Fine-tuned=CUDNN attention for 0.1% tokens, Test Attention=CUDNN2025.09 | 0.6994 | |
| MuSeModel Size=1B, Train Attention=MuSe, Test Attention=MuSe2025.09 | 0.7001 | |
| MuSe†Model Size=1B, Train Attention=MuSe, Fine-tuned=CUDNN attention for 0.1% tokens, Test Attention=MuSe2025.09 | 0.702 | |
| CUDNNModel Size=1B, Train Attention=CUDNN, Test Attention=CUDNN2025.09 | 0.7026 | |
| MuSeModel Size=1B, Train Attention=MuSe, Test Attention=CUDNN2025.09 | 0.7108 |