ResearchBenchmarksLanguage Modeling on Scientific PDFs (24B tokens)Follow1.6166Cross-Entropy LossMuSe1.616461.6174051.618351.619295Sep 12, 2025Evaluation ResultsMethodMethodLinksCross-Entropy LossMuSeModel Size=1B, Train A...Model Size=1B, Train Attention=MuSe, Test Attention=CUDNN2025.091.6166MuSeModel Size=1B, Train A...Model Size=1B, Train Attention=MuSe, Test Attention=MuSe2025.091.6188CUDNNModel Size=1B, Train A...Model Size=1B, Train Attention=CUDNN, Test Attention=CUDNN2025.091.6201